Filter file to remove duplicate values in first column Post: 302977895

Sponsored Content

Top Forums Shell Programming and Scripting Filter file to remove duplicate values in first column Post 302977895 by LMHmedchem on Saturday 23rd of July 2016 02:15:21 AM

07-23-2016

Registered User

Filter file to remove duplicate values in first column

Hello,

I have a script that is generating a tab delimited output file.

Code:

num     Name            PCA_A1     PCA_A2       PCA_A3
0       compound_00     -3.5054     -1.1207     -2.4372
1       compound_01     -2.2641     0.4287      -1.6120
3       compound_03     -1.3053     1.8495      -1.0224
0       compound_00     -3.5054     -1.1207     -2.4372
4       compound_04     -1.1845     -0.3377     -2.9453
7       compound_07     -0.2988     1.3539      -1.6114
8       compound_08     2.6872     -1.3726      -5.9732
9       compound_09     -1.4546     -0.8284     -3.5016
4       compound_04     -1.1845     -0.3377     -2.9453
7       compound_07     -0.2988     1.3539      -1.6114
8       compound_08     2.6872     -1.3726      -5.9732

I need to trim this down so that there a no duplicates in the first column. Actually, the entire row would be a duplicate, but I don't see any reason to look at anything other than the index value. There is no particular rational to the order and there could be any number of duplicates of a given row.

The final results should look like this,

Code:

num     Name            PCA_A1     PCA_A2       PCA_A3
0       compound_00     -3.5054     -1.1207     -2.4372
1       compound_01     -2.2641     0.4287      -1.6120
3       compound_03     -1.3053     1.8495      -1.0224
4       compound_04     -1.1845     -0.3377     -2.9453
7       compound_07     -0.2988     1.3539      -1.6114
8       compound_08     2.6872     -1.3726      -5.9732
9       compound_09     -1.4546     -0.8284     -3.5016

I need one, and only one, instance of each index value ("num" column value) in the file, not just the lines with num values that appear only once. There always seems to be some confusion about that with discussions of "unique" lines.

The only thing I could think of was to sort the rows on the num column value and then loop through checking if the num value was equal to the previous line. If it is not equal, copy it to a new array, etc.

Any suggestions? There always seems to be some simple one line solution that I don't know about.

LMHmedchem

LMHmedchem

View Public Profile for LMHmedchem

Find all posts by LMHmedchem

9 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Filter/remove duplicate .dat file with certain criteria

I am a beginner in Unix. Though have been asked to write a script to filter(remove duplicates) data from a .dat file. File is very huge containig billions of records. contents of file looks like 30002157,40342424,OTC,mart_rec,100, ,0 30002157,40343369,OTC,mart_rec,95, ,0...

2. UNIX for Dummies Questions & Answers

[SOLVED] remove lines that have duplicate values in column two

Hi, I've got a file that I'd like to uniquely sort based on column 2 (values in column 2 begin with "comp"). I tried sort -t -nuk2,3 file.txtBut got: sort: multi-character tab `-nuk2,3' "man sort" did not help me out Any pointers? Input: Output:

3. Shell Programming and Scripting

Check to identify duplicate values at first column in csv file

Hello experts, I have a requirement where I have to implement two checks on a csv file: 1. Check to see if the value in first column is duplicate, if any value is duplicate script should exit. 2. Check to verify if the value at second column is between "yes" or "no", if it is anything else...

4. Linux

Filter a .CSV file based on the 5th column values

I have a .CSV file with the below format: "column 1","column 2","column 3","column 4","column 5","column 6","column 7","column 8","column 9","column 10 "12310","42324564756","a simple string with a , comma","string with or, without commas","string 1","USD","12","70%","08/01/2013",""...

5. Shell Programming and Scripting

Identify duplicate values at first column in csv file

Input 1,ABCD,no 2,system,yes 3,ABCD,yes 4,XYZ,no 5,XYZ,yes 6,pc,noCode used to find duplicate with regard to 2nd column awk 'NR == 1 {p=$2; next} p == $2 { print "Line" NR "$2 is duplicated"} {p=$2}' FS="," ./input.csv Now is there a wise way to de-duplicate the entire line (remove...

6. Shell Programming and Scripting

Remove duplicate values in a column(not in the file)

Hi Gurus, I have a file(weblog) as below abc|xyz|123|agentcode=sample code abcdeeess,agentcode=sample code abcdeeess,agentcode=sample code abcdeeess|agentadd=abcd stereet 23343,agentadd=abcd stereet 23343 sss|wwq|999|agentcode=sample1 code wqwdeeess,gentcode=sample1 code...

7. Shell Programming and Scripting

Find duplicate values in specific column and delete all the duplicate values

Dear folks I have a map file of around 54K lines and some of the values in the second column have the same value and I want to find them and delete all of the same values. I looked over duplicate commands but my case is not to keep one of the duplicate values. I want to remove all of the same...

8. Shell Programming and Scripting

Filter duplicate records from csv file with condition on one column

I have csv file with 30, 40 columns Pasting just three column for problem description I want to filter record if column 1 matches CN or DN then, check for values in column 2 if column contain 1235, 1235 then in column 3 values must be sequence of 2345, 2345 and if column 2 contains 6789, 6789...

9. Shell Programming and Scripting

CSV File:Filter duplicate records from column1 & another column having unique record

Hi Experts, I have csv file with 30, 40 columns Pasting just 2 column for problem description. Need to print error if below combination is not present in file check for column-1 (DocumentNumber) and filter columns where value in DocumentNumber field is same. For all such rows, the field...

LEARN ABOUT CENTOS

tcttest

TCTTEST(1)							   Tokyo Cabinet							TCTTEST(1)

NAME

       tcttest - test cases of the table database API

DESCRIPTION

       The command `tcttest' is a utility for facility test and performance test.  This command is used in the following format.  `path' specifies
       the path of a database file.  `rnum' specifies the number of iterations.  `bnum' specifies the number of  buckets.   `apow'  specifies  the
       power of the alignment.	`fpow' specifies the power of the free block pool.

	      tcttest  write  [-mt]  [-tl] [-td|-tb|-tt|-tx] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-ip] [-is] [-in] [-it] [-if] [-ix]
	      [-nl|-nb] [-rnd] path rnum [bnum [apow [fpow]]]
		     Store records with columns "str", "num", "type", and "flag".
	      tcttest read [-mt] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-nl|-nb] [-rnd] path
		     Retrieve all records of the database above.
	      tcttest remove [-mt] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-nl|-nb] [-rnd] path
		     Remove all records of the database above.
	      tcttest rcat [-mt] [-tl] [-td|-tb|-tt|-tx] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-ip] [-is]	[-in]  [-it]  [-if]  [-ix]
	      [-nl|-nb] [-pn num] [-dai|-dad|-rl|-ru] path rnum [bnum [apow [fpow]]]
		     Store records with partway duplicated keys using concatenate mode.
	      tcttest misc [-mt] [-tl] [-td|-tb|-tt|-tx] [-nl|-nb] path rnum
		     Perform miscellaneous test of various operations.
	      tcttest wicked [-mt] [-tl] [-td|-tb|-tt|-tx] [-nl|-nb] path rnum
		     Perform updating operations selected at random.

       Options feature the following.

	      -mt : call the function `tctdbsetmutex'.
	      -tl : enable the option `TDBTLARGE'.
	      -td : enable the option `TDBTDEFLATE'.
	      -tb : enable the option `TDBTBZIP'.
	      -tt : enable the option `TDBTTCBS'.
	      -tx : enable the option `TDBTEXCODEC'.
	      -rc num : specify the number of cached records.
	      -lc num : specify the number of cached leaf pages.
	      -nc num : specify the number of cached non-leaf pages.
	      -xm num : specify the size of the extra mapped memory.
	      -df num : specify the unit step number of auto defragmentation.
	      -ip : create the number index for the primary key.
	      -is : create the string index for the column "str".
	      -in : create the number index for the column "num".
	      -it : create the string index for the column "type".
	      -if : create the token inverted index for the column "flag".
	      -ix : create the q-gram inverted index for the column "text".
	      -nl : enable the option `TDBNOLCK'.
	      -nb : enable the option `TDBLCKNB'.
	      -rnd : select keys at random.
	      -pn num : specify the number of patterns.
	      -dai : use the function `tctdbaddint' instead of `tctdbputcat'.
	      -dad : use the function `tctdbadddouble' instead of `tctdbputcat'.
	      -rl : set the length of values at random.
	      -ru : select update operations at random.

       This command returns 0 on success, another on failure.

SEE ALSO

       tctmttest(1), tctmgr(1), tctdb(3), tokyocabinet(3)

Man Page							    2012-08-18								TCTTEST(1)