In a huge file, Delete duplicate lines leaving unique lines Post: 302543885

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Delete lines from huge file

I have to delete 1st 7000 lines of a file which is 12GB large. As it is so large, i can't open in vi and delete these lines. Also I found one post here which gave solution using perl, but I don't have perl installed. Also some solutions were redirecting the o/p to a different file and renaming it....

2. Shell Programming and Scripting

delete semi-duplicate lines from file?

Ok here's what I'm trying to do. I need to get a listing of all the mountpoints on a system into a file, which is easy enough, just using something like "mount | awk '{print $1}'" However, on a couple of systems, they have some mount points looking like this: /stage /stand /usr /MFPIS...

3. UNIX for Dummies Questions & Answers

Delete duplicate lines and print to file

OK, I have read several things on how to do this, but can't make it work. I am writing this to a vi file then calling it as an awk script. So I need to search a file for duplicate lines, delete duplicate lines, then write the result to another file, say /home/accountant/files/docs/nodup ...

4. UNIX for Dummies Questions & Answers

How to delete or remove duplicate lines in a file

Hi please help me how to remove duplicate lines in any file. I have a file having huge number of lines. i want to remove selected lines in it. And also if there exists duplicate lines, I want to delete the rest & just keep one of them. Please help me with any unix commands or even fortran...

5. UNIX for Dummies Questions & Answers

Delete lines with duplicate strings based on date

Hey all, a relative bash/script newbie trying solve a problem. I've got a text file with lots of lines that I've been able to clean up and format with awk/sed/cut, but now I'd like to remove the lines with duplicate usernames based on time stamp. Here's what the data looks like 2007-11-03...

6. UNIX for Dummies Questions & Answers

How to delete partial duplicate lines unix

hi :) I need to delete partial duplicate lines I have this in a file sihp8027,/opt/cf20,1980182 sihp8027,/opt/oracle/10gRelIIcd,155200016 sihp8027,/opt/oracle/10gRelIIcd,155200176 sihp8027,/var/opt/ERP,10376312 and need to leave it like this: sihp8027,/opt/cf20,1980182...

7. Shell Programming and Scripting

Delete lines in file containing duplicate strings, keeping longer strings

The question is not as simple as the title... I have a file, it looks like this <string name="string1">RZ-LED</string> <string name="string2">2.0</string> <string name="string2">Version 2.0</string> <string name="string3">BP</string> I would like to check for duplicate entries of...

8. Shell Programming and Scripting

Delete duplicate lines... with a twist!

Hi, I'm sorry I'm no coder so I came here, counting on your free time and good will to beg for spoonfeeding some good code. I'll try to be quick and concise! Got file with 50k lines like this: "Heh, heh. Those darn ninjas. They're _____."*wacky The "canebrake", "timber" & "pygmy" are types...

9. UNIX for Beginners Questions & Answers

How to delete identical lines while leaving one undeleted?

Hi, I have a file as follows. file1 Hello Hi His Hi Hi Hungry hi so I want to delete identical lines while leaving one of them undeleted. So desired output will be Hello Hi

10. UNIX for Beginners Questions & Answers

Delete duplicate like pattern lines

Hi I need to delete duplicate like pattern lines from a text file containing 2 duplicates only (one being subset of the other) using sed or awk preferably. Input: FM:Chicago:Development FM:Chicago:Development:Score SR:Cary:Testing:Testcases PM:Newyork:Scripting PM:Newyork:Scripting:Audit...

LEARN ABOUT MOJAVE

sort

sort(3pm)						 Perl Programmers Reference Guide						 sort(3pm)

NAME

       sort - perl pragma to control sort() behaviour

SYNOPSIS

	   use sort 'stable';	       # guarantee stability
	   use sort '_quicksort';      # use a quicksort algorithm
	   use sort '_mergesort';      # use a mergesort algorithm
	   use sort 'defaults';        # revert to default behavior
	   no  sort 'stable';	       # stability not important

	   use sort '_qsort';	       # alias for quicksort

	   my $current;
	   BEGIN {
	       $current = sort::current();     # identify prevailing algorithm
	   }

DESCRIPTION

       With the "sort" pragma you can control the behaviour of the builtin "sort()" function.

       In Perl versions 5.6 and earlier the quicksort algorithm was used to implement "sort()", but in Perl 5.8 a mergesort algorithm was also
       made available, mainly to guarantee worst case O(N log N) behaviour: the worst case of quicksort is O(N**2).  In Perl 5.8 and later,
       quicksort defends against quadratic behaviour by shuffling large arrays before sorting.

       A stable sort means that for records that compare equal, the original input ordering is preserved.  Mergesort is stable, quicksort is not.
       Stability will matter only if elements that compare equal can be distinguished in some other way.  That means that simple numerical and
       lexical sorts do not profit from stability, since equal elements are indistinguishable.	However, with a comparison such as

	  { substr($a, 0, 3) cmp substr($b, 0, 3) }

       stability might matter because elements that compare equal on the first 3 characters may be distinguished based on subsequent characters.
       In Perl 5.8 and later, quicksort can be stabilized, but doing so will add overhead, so it should only be done if it matters.

       The best algorithm depends on many things.  On average, mergesort does fewer comparisons than quicksort, so it may be better when
       complicated comparison routines are used.  Mergesort also takes advantage of pre-existing order, so it would be favored for using "sort()"
       to merge several sorted arrays.	On the other hand, quicksort is often faster for small arrays, and on arrays of a few distinct values,
       repeated many times.  You can force the choice of algorithm with this pragma, but this feels heavy-handed, so the subpragmas beginning with
       a "_" may not persist beyond Perl 5.8.  The default algorithm is mergesort, which will be stable even if you do not explicitly demand it.
       But the stability of the default sort is a side-effect that could change in later versions.  If stability is important, be sure to say so
       with a

	 use sort 'stable';

       The "no sort" pragma doesn't forbid what follows, it just leaves the choice open.  Thus, after

	 no sort qw(_mergesort stable);

       a mergesort, which happens to be stable, will be employed anyway.  Note that

	 no sort "_quicksort";
	 no sort "_mergesort";

       have exactly the same effect, leaving the choice of sort algorithm open.

CAVEATS

       As of Perl 5.10, this pragma is lexically scoped and takes effect at compile time. In earlier versions its effect was global and took
       effect at run-time; the documentation suggested using "eval()" to change the behaviour:

	 { eval 'use sort qw(defaults _quicksort)'; # force quicksort
	   eval 'no sort "stable"';	 # stability not wanted
	   print sort::current . "
";
	   @a = sort @b;
	   eval 'use sort "defaults"';	 # clean up, for others
	 }
	 { eval 'use sort qw(defaults stable)';     # force stability
	   print sort::current . "
";
	   @c = sort @d;
	   eval 'use sort "defaults"';	 # clean up, for others
	 }

       Such code no longer has the desired effect, for two reasons.  Firstly, the use of "eval()" means that the sorting algorithm is not changed
       until runtime, by which time it's too late to have any effect. Secondly, "sort::current" is also called at run-time, when in fact the
       compile-time value of "sort::current" is the one that matters.

       So now this code would be written:

	 { use sort qw(defaults _quicksort); # force quicksort
	   no sort "stable";	  # stability not wanted
	   my $current;
	   BEGIN { $current = sort::current; }
	   print "$current
";
	   @a = sort @b;
	   # Pragmas go out of scope at the end of the block
	 }
	 { use sort qw(defaults stable);     # force stability
	   my $current;
	   BEGIN { $current = sort::current; }
	   print "$current
";
	   @c = sort @d;
	 }

perl v5.18.2							    2013-11-04								 sort(3pm)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Delete lines from huge file

Discussion started by: rahulrathod

2. Shell Programming and Scripting

delete semi-duplicate lines from file?

Discussion started by: paqman

3. UNIX for Dummies Questions & Answers

Delete duplicate lines and print to file

Discussion started by: bfurlong

4. UNIX for Dummies Questions & Answers

How to delete or remove duplicate lines in a file

Discussion started by: reva