Filter file to remove duplicate values in first column Post: 302977895

Sponsored Content

Top Forums Shell Programming and Scripting Filter file to remove duplicate values in first column Post 302977895 by LMHmedchem on Saturday 23rd of July 2016 02:15:21 AM

07-23-2016

Registered User

Filter file to remove duplicate values in first column

Hello,

I have a script that is generating a tab delimited output file.

Code:

num     Name            PCA_A1     PCA_A2       PCA_A3
0       compound_00     -3.5054     -1.1207     -2.4372
1       compound_01     -2.2641     0.4287      -1.6120
3       compound_03     -1.3053     1.8495      -1.0224
0       compound_00     -3.5054     -1.1207     -2.4372
4       compound_04     -1.1845     -0.3377     -2.9453
7       compound_07     -0.2988     1.3539      -1.6114
8       compound_08     2.6872     -1.3726      -5.9732
9       compound_09     -1.4546     -0.8284     -3.5016
4       compound_04     -1.1845     -0.3377     -2.9453
7       compound_07     -0.2988     1.3539      -1.6114
8       compound_08     2.6872     -1.3726      -5.9732

I need to trim this down so that there a no duplicates in the first column. Actually, the entire row would be a duplicate, but I don't see any reason to look at anything other than the index value. There is no particular rational to the order and there could be any number of duplicates of a given row.

The final results should look like this,

Code:

num     Name            PCA_A1     PCA_A2       PCA_A3
0       compound_00     -3.5054     -1.1207     -2.4372
1       compound_01     -2.2641     0.4287      -1.6120
3       compound_03     -1.3053     1.8495      -1.0224
4       compound_04     -1.1845     -0.3377     -2.9453
7       compound_07     -0.2988     1.3539      -1.6114
8       compound_08     2.6872     -1.3726      -5.9732
9       compound_09     -1.4546     -0.8284     -3.5016

I need one, and only one, instance of each index value ("num" column value) in the file, not just the lines with num values that appear only once. There always seems to be some confusion about that with discussions of "unique" lines.

The only thing I could think of was to sort the rows on the num column value and then loop through checking if the num value was equal to the previous line. If it is not equal, copy it to a new array, etc.

Any suggestions? There always seems to be some simple one line solution that I don't know about.

LMHmedchem

LMHmedchem

View Public Profile for LMHmedchem

Find all posts by LMHmedchem

9 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Filter/remove duplicate .dat file with certain criteria

I am a beginner in Unix. Though have been asked to write a script to filter(remove duplicates) data from a .dat file. File is very huge containig billions of records. contents of file looks like 30002157,40342424,OTC,mart_rec,100, ,0 30002157,40343369,OTC,mart_rec,95, ,0...

2. UNIX for Dummies Questions & Answers

[SOLVED] remove lines that have duplicate values in column two

Hi, I've got a file that I'd like to uniquely sort based on column 2 (values in column 2 begin with "comp"). I tried sort -t -nuk2,3 file.txtBut got: sort: multi-character tab `-nuk2,3' "man sort" did not help me out Any pointers? Input: Output:

3. Shell Programming and Scripting

Check to identify duplicate values at first column in csv file

Hello experts, I have a requirement where I have to implement two checks on a csv file: 1. Check to see if the value in first column is duplicate, if any value is duplicate script should exit. 2. Check to verify if the value at second column is between "yes" or "no", if it is anything else...

4. Linux

Filter a .CSV file based on the 5th column values

I have a .CSV file with the below format: "column 1","column 2","column 3","column 4","column 5","column 6","column 7","column 8","column 9","column 10 "12310","42324564756","a simple string with a , comma","string with or, without commas","string 1","USD","12","70%","08/01/2013",""...

5. Shell Programming and Scripting

Identify duplicate values at first column in csv file

Input 1,ABCD,no 2,system,yes 3,ABCD,yes 4,XYZ,no 5,XYZ,yes 6,pc,noCode used to find duplicate with regard to 2nd column awk 'NR == 1 {p=$2; next} p == $2 { print "Line" NR "$2 is duplicated"} {p=$2}' FS="," ./input.csv Now is there a wise way to de-duplicate the entire line (remove...

6. Shell Programming and Scripting

Remove duplicate values in a column(not in the file)

Hi Gurus, I have a file(weblog) as below abc|xyz|123|agentcode=sample code abcdeeess,agentcode=sample code abcdeeess,agentcode=sample code abcdeeess|agentadd=abcd stereet 23343,agentadd=abcd stereet 23343 sss|wwq|999|agentcode=sample1 code wqwdeeess,gentcode=sample1 code...

7. Shell Programming and Scripting

Find duplicate values in specific column and delete all the duplicate values

Dear folks I have a map file of around 54K lines and some of the values in the second column have the same value and I want to find them and delete all of the same values. I looked over duplicate commands but my case is not to keep one of the duplicate values. I want to remove all of the same...

8. Shell Programming and Scripting

Filter duplicate records from csv file with condition on one column

I have csv file with 30, 40 columns Pasting just three column for problem description I want to filter record if column 1 matches CN or DN then, check for values in column 2 if column contain 1235, 1235 then in column 3 values must be sequence of 2345, 2345 and if column 2 contains 6789, 6789...

9. Shell Programming and Scripting

CSV File:Filter duplicate records from column1 & another column having unique record

Hi Experts, I have csv file with 30, 40 columns Pasting just 2 column for problem description. Need to print error if below combination is not present in file check for column-1 (DocumentNumber) and filter columns where value in DocumentNumber field is same. For all such rows, the field...

LEARN ABOUT HPUX

tabs

tabs(1) 						      General Commands Manual							   tabs(1)

NAME

       tabs - set tabs on a terminal

SYNOPSIS

       [tabspec] n] type]

DESCRIPTION

       sets  the  tab  stops  on the user's terminal according to the tab specification tabspec, after clearing any previous settings.	The user's
       terminal must have remotely-settable hardware tabs.

       If you are using a non-HP terminal, you should keep in mind that behavior will vary for some tab settings.

       Four types of tab specification are accepted for tabspec: ``canned'', repetitive, arbitrary, and file.  If no is given, the  default  value
       is  i.e.,  UNIX ``standard'' tabs.  The lowest column number is 1.  Note that for tabs, column 1 always refers to the left-most column on a
       terminal, even one whose column markers begin at 0.

       Gives the name of one of a set of ``canned'' tabs.
	       Recognized codes and their meanings are as follows:

		     1,10,16,36,72
			   Assembler, IBM S/370, first format

		     1,10,16,40,72
			   Assembler, IBM S/370, second format

		     1,8,12,16,20,55
			   COBOL, normal format

		     1,6,10,14,49
			   COBOL compact format (columns 1-6 omitted).	Using this code, the first typed character corresponds to card	column	7,
			   one	space  gets you to column 8, and a tab reaches column 12.  Files using this tab setup should have specify a format
			   specification file as defined by below.  The file should have the following format specification:

		     1,6,10,14,18,22,26,30,34,38,42,46,50,54,58,62,67
			   COBOL compact format (columns 1-6 omitted), with more tabs than This is the recommended format for COBOL.   The  appro-
			   priate format specification is:

		     1,7,11,15,19,23
			   FORTRAN

		     1,5,9,13,17,21,25,29,33,37,41,45,49,53,57,61
			   PL/I

		     1,10,55
			   SNOBOL

		     1,12,20,44
			   UNIVAC 1100 Assembler

       In addition to these ``canned'' formats, three other types exist:

       A repetitive specification requests tabs at columns
		   1+n,  1+2xn,  etc.	Of  particular	importance is the value this represents the UNIX ``standard'' tab setting, and is the most
		   likely tab setting to be found at a terminal.  Another special case is the value implying no tabs at all.

       The arbitrary format permits the user to type any
		   chosen set of numbers, separated by commas, in ascending order.  Up to 40 numbers are allowed.  If any number (except the first
		   one) is preceded by a plus sign, it is taken as an increment to be added to the previous value.  Thus, the tab lists 1,10,20,30
		   and 1,10,+10,+10 are considered identical.

       If the name of a file is given,
		   reads the first line of the file, searching for a format specification.  If it finds one there, it sets the tab stops according
		   to  it,  otherwise  it sets them as This type of specification can be used to ensure that a tabbed file is printed with correct
		   tab settings, and is suitable for use with the command (see pr(1)):

       Any of the following can be used also; if a given option occurs more than once, the last value given takes effect:

       usually needs to know the type of terminal in order to set tabs
		   and always needs to know the type to set margins.  type is a name listed in term(5).  If no option is  supplied,  searches  for
		   the	value in the environment (see environ(5)).  If is not defined in the environment, tries a sequence that will work for many
		   terminals.

       The margin argument can be used for some terminals.
		   It causes all tabs to be moved over n columns by making column n+1 the left margin.	If is given without  a	value  of  n,  the
		   value  assumed  is  10.   The normal (left-most) margin on most terminals is obtained by The margin for most terminals is reset
		   only when the option is given explicitly.

       Tab and margin setting is performed via the standard output.

EXTERNAL INFLUENCES

   Environment Variables
       determines the interpretation of text within file as single- and/or multi-byte characters.

       determines the language in which messages are displayed.

       If or is not specified in the environment or is set to the empty string, the value of is used as a default for each  unspecified  or  empty
       variable.  If is not specified or is set to the empty string, a default of "C" (see lang(5)) is used instead of

       If  any	internationalization  variable	contains an invalid setting, behaves as if all internationalization variables are set to "C".  See
       environ(5).

   International Code Set Support
       Single- and multi-byte character code sets are supported.

DIAGNOSTICS

       Arbitrary tabs are ordered incorrectly.

       A zero or missing increment found in an arbitrary specification.

       A ``canned'' code cannot be found.

       option was used and file cannot be opened.

       option was used and the specification in that file
	      points to yet another file.  Indirection of this form is not permitted.

WARNINGS

       There is no consistency among different terminals regarding ways of clearing tabs and setting the left margin.

       It is generally impossible to usefully change the left margin without also setting tabs.

       clears only 20 tabs (on terminals requiring a long sequence), but is willing to set 64.

SEE ALSO

       nroff(1), pr(1), tset(1), environ(5), term(5).

STANDARDS CONFORMANCE

																	   tabs(1)