02-23-2009
As an idea .. I was thinking of maybe counting the total number of words in the text file and then running a for loop to that number to check for n-grams.
I haven't yet tried this idea. Right now I am working out of text files. Creating 2 text files for bigrams is easy .. but in real time if I have to create n files when checking for n-grams ... how feasible is that in terms of memory and CPU cycles both ??
This is why I am asking here, if any of you can help me please
Thanks
SG
10 More Discussions You Might Find Interesting
1. Shell Programming and Scripting
Dear All,
To find the length of the longest line from a file i have used wc -L which is giving the proper output...
But the problem is AIX os does not support wc -L command.
so is there any other way 2 to find out the length of the longest line using awk or sed ?
Regards,
Pankaj (1 Reply)
Discussion started by: panknil
1 Replies
2. Shell Programming and Scripting
Okie here is my problem,
1. I have a directory with a ton of files.
2. I want to first get an input on how many days ago the files were created.
3. I will take those files and put it into another file
4. Then I will take the last # from each line and subtract by 1 then diff the line from the... (1 Reply)
Discussion started by: bigboizvince
1 Replies
3. Shell Programming and Scripting
I've got a script which finds *.txt files in directories and subdirectories after providing the path by the user and then searches in the files for phrase given by the user
How to write script in such way that the paths to the found *.txt files and the phrase given by the user were both... (2 Replies)
Discussion started by: patrykxes
2 Replies
4. Shell Programming and Scripting
Return the position of matched string from right, awk match can do from left only.
e.g return pos 7 for search string "service" from "AA-service"
or return the matched string "service", then caculate the string length.
Thanks!. (3 Replies)
Discussion started by: honglus
3 Replies
5. Shell Programming and Scripting
Hello everyone...
I need to find out, how to find longest line or possibly lines in several files which are arguments for script. The thing is, that I tried some possibilities before, but nothing worked correctly.
Example
when i use:
awk ' { if ( length > L ) { L=length ;s=$0 } }END{ print... (23 Replies)
Discussion started by: 1tempus1
23 Replies
6. Shell Programming and Scripting
Hello all,
I need to find the longest string in a select field and print that field.
I have tried a few different methods and I always end up one step from where I need to be.
Methods thus far:
nawk '{if (length($1) > long) long=length($1); if(length($1)==long) print $1}'
The above... (6 Replies)
Discussion started by: SEinT
6 Replies
7. Shell Programming and Scripting
Hello
My question is: How to find out the shell of the shell script which we are running? I am writing a script, say f1.sh, as below:
#!/bin/ksh
echo "Sample script"
From the first line, we can say this script will run in ksh. But, how can we prove it? Can we print anything inside... (6 Replies)
Discussion started by: guruprasadpr
6 Replies
8. Shell Programming and Scripting
I want to make a script which takes the number of argument, add those argument and gives output to the user, but I am not getting through...
Script that i am using is below :
#!/bin/bash
sum=0
for i in $@
do
sum=$sum+$1
echo $sum
shift
done
I am executing the script as... (3 Replies)
Discussion started by: mukulverma2408
3 Replies
9. Shell Programming and Scripting
I want to burst a report by using the page number value in the report header. Each section starts with *PAGE NO:* 1 Each section might have several pages, but the next section always starts back at 1.
So I want to find the "*PAGE NO:* 1" value and pull all lines that follow until "*PAGE NO:* 1"... (4 Replies)
Discussion started by: Scottie1954
4 Replies
10. Shell Programming and Scripting
Hello,
I am looking for a shell script that can
1- take as input a variable, like "server.cpu"
2- do a search for that variable in a directory that contains subdirectories.
The search will start at the last subdirectory working up to the top level if I can not find the file
3-... (7 Replies)
Discussion started by: georg2014
7 Replies
wc(1) General Commands Manual wc(1)
NAME
wc - Counts the lines, words, characters, and bytes in a file
SYNOPSIS
wc [-c | -m] [-lw] [file...]
The wc command counts the lines, words, and characters or bytes in a file, or in the standard input if you do not specify any files, and
writes the results to standard output. It also keeps a total count for all named files.
STANDARDS
Interfaces documented on this reference page conform to industry standards as follows:
wc: XCU5.0
Refer to the standards(5) reference page for more information about industry standards and associated tags.
OPTIONS
Counts bytes in the input. Counts lines in the input. Counts characters in the input. Counts words in the input.
OPERANDS
Specifies the pathname of the input file. If this operand is omitted, standard input is used.
DESCRIPTION
A word is defined as a string of characters delimited by white space as defined in the X/Open Base Definitions for XCU4.
The wc command counts lines, words, and bytes by default. Use the appropriate options to limit wc output. Specifying wc without options
is the equivalent of specifying wc -lwc. If any options are specified, only the requested information is output.
The order in which counts appear in the output line is lines, words, bytes. If an option is omitted, then the corresponding field in the
output is omitted. If the -m option is used, then character counts replace byte counts.
When you specify one or more files, wc displays the names of the files along with the counts. If standard input is used, then no file name
is displayed.
EXIT STATUS
The following exit values are returned: Successful completion. An error occurred.
EXAMPLES
To display the number of lines, words, and bytes in the file text, enter: wc text
This results in the following output: 27 185 722 text
The numbers 27, 185, and 722 are the number of lines, words, and bytes, respectively, in the file text. To display only one or two
of the three counts include the appropriate options. For example, the following command displays only line and byte counts: wc -cl
text
27 722 text To count lines, words, and bytes in more than one file, use wc with more than one input file or with a file name pat-
tern. For example, the following command can be issued in a directory containing the files text, text1, and text2: wc -l text*
27 text 112 text1 5 text2 144 total
The numbers 27, 112, and 5 are the numbers of lines in the files text, text1, and text2, respectively, and 144 is the total number
of lines in the three files. The file name is always appended to the output. To obtain a pure number for things like reporting
purposes, pipe all input to the wc command using cat. For example, the following command will report the total count of characters
in all files in a directory. echo There are `cat *.c | wc -c` characters in *.c files
There are 1869 characters in *.c files
ENVIRONMENT VARIABLES
The following environment variables affect the execution of wc: Provides a default value for the internationalization variables that are
unset or null. If LANG is unset or null, the corresponding value from the default locale is used. If any of the internationalization vari-
ables contain an invalid setting, the utility behaves as if none of the variables had been defined. If set to a non-empty string value,
overrides the values of all the other internationalization variables. Determines the locale for the interpretation of sequences of bytes
of text data as characters (for example, single-byte as opposed to multibyte characters in arguments and input files) and which characters
are defined as white space characters. Determines the locale for the format and contents of diagnostic messages written to standard error
and informative messages written to standard output. Determines the location of message catalogues for the processing of LC_MESSAGES.
SEE ALSO
Commands: cksum(1), ls(1)
Standards: standards(5)
wc(1)