how to get distinct counts in a column of a file Post: 302436410

9 More Discussions You Might Find Interesting

1. Solaris

file size counts??

Hello experts, I do - $ ls -lhtr logs2007* Is it possible that i can get the results of- totals size in MB/KB for ALL "logs2007*" note: in the same directory I have "logs2006*" & "logs2007*" files.

2. Shell Programming and Scripting

counts the number of distinct words

I'm looking to write a sample shell script that counts the number of distinct words in a text file given as Argument. Remark: White space characters are spaces, tabs, form feeds, and new lines. JUST with this commands tr, sort, grep. wc. Thanks.

3. Shell Programming and Scripting

have to retrieve the distinct values (not duplicate) from 2nd column and display

I have a text file names test2 with 3 columns as below . We have to retrieve the distinct values (not duplicate) from 2nd column and display. I have used the below command but giving some error. NS3303 NS CRAFT LTD NS3303 NS CHIRON VACCINES LTD NS3303 NS ALLIED MEDICARE LTD NS3303 NS...

4. Shell Programming and Scripting

Filtering lines for column elements based on corresponding counts in another column

Hi, I have a file like this ACC 2 2 21 aaa AC 443 3 22 aaa GCT 76 1 33 xxx TCG 34 2 33 aaa ACGT 33 1 22 ggg TTC 99 3 44 wee CCA 33 2 33 ggg AAC 1 3 55 ddd TTG 10 1 22 ddd TTGC 98 3 22 ddd GCT 23 1 21 sds GTC 23 4 32 sds ACGT 32 2 33 vvv CGT 11 2 33 eee CCC 87 2 44...

5. UNIX for Dummies Questions & Answers

count number of distinct values in each column with awk

Hi ! input: A|B|C|D A|F|C|E A|B|I|C A|T|I|B As the title of the thread says, I would need to get: 1|3|2|4 I tried different variants of this command, but I don't manage to obtain what I need: gawk 'BEGIN{FS=OFS="|"}{for(i=1; i<=NF; i++) a++} END {for (b in a) print b}' input ...

6. Shell Programming and Scripting

Select distinct rows in a file by last column

Hi, I have the following file: LOG:015608::ERR:2310:map_spsrec:Invalid parameter LOG:015608::ERR:2471:map_dgdrec:Invalid parameter LOG:015608::ERR:2487:map_nnmrec:Invalid number LOG:015608::ERR:2310:map_nmrec:Invalid number LOG:015608::ERR:2438:map_nmrec:Invalid number As a delimiter I...

7. UNIX for Dummies Questions & Answers

awk adding counts together from column

Hello Im new treat me nicely, I have a headache :) I have a script that seemed to work now it doesnt anyway, the last part is adding counts of unique items in a csv file eg 05492U34 38 05492U34 47 two columns, (many different values like this in file) i want...

8. Shell Programming and Scripting

Splitting the numeric vs alpha values in a column to distinct columns

How could i take an input file and split the numeric values from the alpha values (123 vs abc) to distinc columns, and if the source is blank to keep it blank (null) in both of the new columns: So if the source file had a column like: Value: |1 | |2.3| | | |No| I would...

9. Shell Programming and Scripting

Counts not matching in file

I can not figure out why there are 56,548 unique entries in test.bed. However, perl and awk see only 56,543 and that # is what my analysis see's as well. What happened to the 5 missing? Thank you :). The file is attached as well. cmccabe@DTV-A5211QLM:~/Desktop/NGS/bed/bedtools$wc -l...

LEARN ABOUT NETBSD

uniq

UNIQ(1) 						    BSD General Commands Manual 						   UNIQ(1)

NAME

     uniq -- report or filter out repeated lines in a file

SYNOPSIS

     uniq [-cdu] [-f fields] [-s chars] [input_file [output_file]]

DESCRIPTION

     The uniq utility reads the standard input comparing adjacent lines, and writes a copy of each unique input line to the standard output.  The
     second and succeeding copies of identical adjacent input lines are not written.  Repeated lines in the input will not be detected if they are
     not adjacent, so it may be necessary to sort the files first.

     The following options are available:

     -c      Precede each output line with the count of the number of times the line occurred in the input, followed by a single space.

     -d      Don't output lines that are not repeated in the input.

     -f fields
	     Ignore the first fields in each input line when doing comparisons.  A field is a string of non-blank characters separated from adja-
	     cent fields by blanks.  Field numbers are one based, i.e. the first field is field one.

     -s chars
	     Ignore the first chars characters in each input line when doing comparisons.  If specified in conjunction with the -f option, the
	     first chars characters after the first fields fields will be ignored.  Character numbers are one based, i.e. the first character is
	     character one.

     -u      Don't output lines that are repeated in the input.

     If additional arguments are specified on the command line, the first such argument is used as the name of an input file, the second is used
     as the name of an output file.

     The uniq utility exits 0 on success, and >0 if an error occurs.

COMPATIBILITY

     The historic +number and -number options have been deprecated but are still supported in this implementation.

SEE ALSO

     sort(1)

STANDARDS

     The uniq utility is expected to be IEEE Std 1003.2 (``POSIX.2'') compatible.

BSD
								  January 6, 2007							       BSD