Remove duplicates from a file

01-21-2010

Registered User

26, 0

Join Date: Apr 2009

Last Activity: 7 May 2010, 3:27 PM EDT

Posts: 26

Thanks Given: 0

Thanked 0 Times in 0 Posts

Remove duplicates from a file

Hi,

I need to remove duplicates from a file. The file will be like this

Code:

0003 10101 20100120 abcdefghi
0003 10101 20100121 abcdefghi
0003 10101 20100122 abcdefghi
0003 10102 20100120 abcdefghi
0003 10103 20100120 abcdefghi
0003 10103 20100121 abcdefghi

Here if the first colum and second column is repeating i need to pick the first record. If not repeating i need to pick the record.

The output shd be.

Code:

0003 10101 20100120 abcdefghi
0003 10102 20100120 abcdefghi
0003 10103 20100120 abcdefghi

Thanks in advance for the help. The script could be in Perl or Unix.

gpaulose

View Public Profile for gpaulose

Find all posts by gpaulose

01-21-2010

Registered User

5,690, 630

Join Date: Jan 2007

Last Activity: 9 January 2017, 4:40 AM EST

Location: Варна, България / Milano, Italia

Posts: 5,690

Thanks Given: 184

Thanked 630 Times in 587 Posts

If the values are ordered, their format is fixed and the uniq implementation on your platform supports the w option:

Code:

uniq -w10 infile

Otherwise use awk:

Code:

awk '!_[$1,$2]++' infile

On Solaris you should use gawk, nawk or /usr/xpg4/bin/awk.

Or Perl:

Code:

perl -ane'print unless $_{$F[0], $F[1]}++' infile

Last edited by radoulov; 01-28-2010 at 05:24 AM.. Reason: corrected

radoulov

View Public Profile for radoulov

Find all posts by radoulov

01-22-2010

Registered User

131, 18

Join Date: Jan 2010

Last Activity: 2 April 2019, 12:28 PM EDT

Posts: 131

Thanks Given: 64

Thanked 18 Times in 18 Posts

Code:

 sort -u -k1,2 infile

ni2

View Public Profile for ni2

Find all posts by ni2

01-22-2010

Registered User

331, 13

Join Date: Aug 2009

Last Activity: 8 December 2017, 5:27 AM EST

Location: izmir

Posts: 331

Thanks Given: 33

Thanked 13 Times in 13 Posts

Hello,
i found this code on a web page which is said to be valid for only gnu linux and delete all lines except duplicate ones, i hope it works (sorry im using solaris10, couldnt try)

# delete all lines except duplicate lines (emulates "uniq -d").

Code:

sed '$!N; s/^\(.*\)\n\1$/\1/; t; D' infile

EAGL�

View Public Profile for EAGL�

Find all posts by EAGL�

01-27-2010

Registered User

26, 0

Join Date: Apr 2009

Last Activity: 7 May 2010, 3:27 PM EDT

Posts: 26

Thanks Given: 0

Thanked 0 Times in 0 Posts

Could you explain the perl code please?

Quote:

Originally Posted by radoulov

If the values are ordered, their format is fixed and the uniq implementation on your platform supports the w option:

Code:

uniq -w10 infile

Otherwise use awk:

Code:

awk '!_[$1,$2]++' infile

On Solaris you should use gawk, nawk or /usr/xpg4/bin/awk.

Or Perl:

Code:

perl -ane'print unless $_{@F[0..1]}++' infile

gpaulose

View Public Profile for gpaulose

Find all posts by gpaulose

01-28-2010

Registered User

5,690, 630

Join Date: Jan 2007

Last Activity: 9 January 2017, 4:40 AM EST

Location: Варна, България / Milano, Italia

Posts: 5,690

Thanks Given: 184

Thanked 630 Times in 587 Posts

Quote:

Originally Posted by gpaulose

Could you explain the perl code please?

Yes,
first of all, the code is wrong

It should be:

Code:

unless $_{$F[0],$F[1]}++

... not:

Code:

unless $_{@F[0..1]}++

So the script becomes:

Code:

perl -ane'print unless $_{$F[0],$F[1]}++' infile

First the command line switches:

Quote:

-a turns on autosplit mode when used with a -n or -p. An implicit
split command to the @F array is done as the first thing inside
the implicit while loop produced by the -n or -p.

Quote:

-e commandline
may be used to enter one line of program. If -e is given, Perl
will not look for a filename in the argument list. Multiple -e
commands may be given to build up a multi-line script. Make sure
to use semicolons where you would in a normal program.

Quote:

-n causes Perl to assume the following loop around your program,
which makes it iterate over filename arguments somewhat like sed
-n or awk:

LINE:
while (<>) {
... # your program goes here
}

Note that the lines are not printed by default. See -p to have
lines printed. If a file named by an argument cannot be opened
for some reason, Perl warns you about it and moves on to the next
file.

So we have the input file read line by line and the @F array automatically populated.

Code:

print unless $_{$F[0],$F[1]}++

Print the current record unless the expression $_{$F[0],$F[1]}++ returns true in boolean context. We build the hash %_ whose keys (the concatenation of the first two fields with the subscript separator) are associated with auto-incremented integers. When we see a given key ($F[0] $; $F[1] - the first and the second fields) for the first time, because of the post-incrementing (k++ and not ++k ) its value is 0, i.e. false, so it prints the record.

This will make the concept clear:

Code:

% perl -lane'print  $_ ," -> ", $_{$F[0],$F[1]}++' infile
0003 10101 20100120 abcdefghi -> 0
0003 10101 20100121 abcdefghi -> 1
0003 10101 20100122 abcdefghi -> 2
0003 10102 20100120 abcdefghi -> 0
0003 10103 20100120 abcdefghi -> 0
0003 10103 20100121 abcdefghi -> 1

We want only the records with value 0.

Hope this helps.

radoulov

View Public Profile for radoulov

Find all posts by radoulov

01-28-2010

Registered User

645, 19

Join Date: May 2008

Last Activity: 7 August 2017, 4:42 AM EDT

Location: Amman, Jordan

Posts: 645

Thanks Given: 2

Thanked 19 Times in 19 Posts

another approach in perl:-

Code:

perl -ane ' ! $_{$F[0],$F[1]}++ and print ' infile.txt

perl -ane ' ! $_{$F[0],$F[1]}++ && print '  infile.txt

perl -ane ' print while ! $_{$F[0],$F[1]}++ '  infile.txt

ahmad.diab

View Public Profile for ahmad.diab

Find all posts by ahmad.diab

Shell Programming and Scripting

Remove duplicates from a file

10 More Discussions You Might Find Interesting

1. UNIX for Advanced & Expert Users

Remove duplicates in flat file

Discussion started by: samjoshuab

2. Shell Programming and Scripting

To remove duplicates from pipe delimited file

Discussion started by: ginrkf

3. UNIX for Dummies Questions & Answers

Remove duplicates and keep them in a separate file

Discussion started by: flacchy

4. UNIX for Dummies Questions & Answers

Remove duplicates from a file

Discussion started by: saga20

5. Shell Programming and Scripting

How to remove duplicates from the .dat file

Discussion started by: Oracle_User

6. Shell Programming and Scripting

Search based on 1,2,4,5 columns and remove duplicates in the same file.

Discussion started by: onesuri

7. Shell Programming and Scripting

Remove duplicates from end of file

Discussion started by: lavnayas

8. Shell Programming and Scripting

Shell script to remove duplicates lines in a file

Discussion started by: RichElks

9. Shell Programming and Scripting

remove duplicates within a block in a file..help required

Discussion started by: nipun_garg

10. Shell Programming and Scripting

Remove duplicates from File from specific location

Discussion started by: gopikgunda