01-04-2013
The sort -u has less scaling difficulty than associative arrays in a shell. If the order of lines is important, number them first, then sort out the duplicates, then sort it back to original order and then remove numbering. The sort will save the first of any duplicates.
JAVA and C++ PPIs like RogueWave H++ have the same associative array functionality, more accurately called a hash map container. The file could be mmap64() mapped to the program so the hash map just has hash keys and offsets. The clean file could be recreated or even updated in place and truncated with fcntl(). The shell has noticably more overhead. Big jobs often need sharper tools.
Last edited by DGPickett; 01-04-2013 at 04:53 PM..
This User Gave Thanks to DGPickett For This Post:
10 More Discussions You Might Find Interesting
1. UNIX for Dummies Questions & Answers
i have a file with some 1000 entries it will contain entries like
1000,ram
2000,pankaj
1001,rahim
1000,ram
2532,govind
2000,pankaj
3000,venkat
2532,govind
what i want is i want to extract only the distinct rows from this file
so my output should contain only
1000,ram... (2 Replies)
Discussion started by: trichyselva
2 Replies
2. UNIX for Dummies Questions & Answers
hey all,
I need some help.
I have a text file with names in it.
My target is that if a particular pattern exists in that file more than once..then i want to rename all the occurences of that pattern by alternate patterns..
for e.g if i have PATTERN occuring 5 times then i want to... (3 Replies)
Discussion started by: ashisharora
3 Replies
3. Shell Programming and Scripting
I have a log file with posts looking like this:
--
Messages can be delivered by different systems at different times. The id number is used to sort out duplicate messages. What I need is to strip the arrival time from each post, sort posts by id number, and reattach arrival time to respective... (2 Replies)
Discussion started by: Ilja
2 Replies
4. Shell Programming and Scripting
Hi Experts,
Please check the following new requirement. I got data like the following in a file.
FILE_HEADER
01cbbfde7898410| 3477945| home| 1
01cbc275d2c122| 3478234| WORK| 1
01cbbe4362743da| 3496386| Rich Spare| 1
01cbc275d2c122| 3478234| WORK| 1
This is pipe separated file with... (3 Replies)
Discussion started by: tinufarid
3 Replies
5. Shell Programming and Scripting
Hi,
I have a file that I want to change the format of. It is a large file in rows but I want it to be comma separated (comma then a space).
The current file looks like this:
HI, Joe, Bob, Jack, Jack
After I would want to remove any duplicates so it would look like this:
HI, Joe,... (2 Replies)
Discussion started by: kylle345
2 Replies
6. HP-UX
I have 2 files; one file (say, details.txt) contains the details of employees and another file (say, emp.txt) has some selected employee names. I am extracting employee details from details.txt by using emp.txt and the corresponding code is:
while read line
do
emp_name=`echo $line`
grep -e... (7 Replies)
Discussion started by: arb_1984
7 Replies
7. UNIX for Dummies Questions & Answers
Hi All,
I am merging files coming from 2 different systems ,while doing that I am getting duplicates entries in the merged file
I,01,000131,764,2,4.00
I,01,000131,765,2,4.00
I,01,000131,772,2,4.00
I,01,000131,773,2,4.00
I,01,000168,762,2,2.00
I,01,000168,763,2,2.00... (5 Replies)
Discussion started by: Sri3001
5 Replies
8. Shell Programming and Scripting
i hav two files like
i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3 (2 Replies)
Discussion started by: sagar_1986
2 Replies
9. Shell Programming and Scripting
i hav two files like
i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3.I have tried previous post also,but in that complete line must be similar.In this case i have to verify first column only regardless what is the content in succeeding columns. (3 Replies)
Discussion started by: sagar_1986
3 Replies
10. Shell Programming and Scripting
I am trying to remove whitespaces from a file containing sample data as:
457 <EOFD> Mar 1 2007 12:00:00:000AM <EOFD> Mar 31 2007 12:00:00:000AM <EOFD> system <EORD> 458 <EOFD> Mar 1 2007 12:00:00:000AM<EOFD>agf <EOFD> Apr 20 2007 9:10:56:036PM <EOFD> prodiws<EORD> . Basically these... (11 Replies)
Discussion started by: amvip
11 Replies
LEARN ABOUT DEBIAN
cz::sort
Cz::Sort(3pm) User Contributed Perl Documentation Cz::Sort(3pm)
NAME
Cz::Sort - Czech sort
SYNOPSIS
use Cz::Sort;
my $result = czcmp("_x j&a", "_&p");
my @sorted = czsort qw(plachta plaoka Planieka planieka plani);
print "@sorted
";
DESCRIPTION
Implements czech sorting conventions, indepentent on current locales in effect, which are often bad. Does the four-pass sort. The idea and
the base of the conversion table comes from Petr Olsak's program csr and the code is as compliant with CSN 97 6030 as possible.
The basic function provided by this module, is czcmp. If compares two scalars and returns the (-1, 0, 1) result. The function can be called
directly, like
my $result = czcmp("_x j&a", "_&p");
But for convenience and also because of compatibility with older versions, there is a function czsort. It works on list of strings and
returns that list, hmm, sorted. The function is defined simply like
sub czsort
{ sort { czcmp($a, $b); } @_; }
standard use of user's function in sort. Hashes would be simply sorted
@sorted = sort { czcmp($hash{$a}, $hash{$b}) }
keys %hash;
Both czcmp and czsort are exported into caller's namespace by default, as well as cscmp and cssort that are just aliases.
This module comes with encoding table prepared for ISO-8859-2 (Latin-2) encoding. If your data come in different one, you might want to
check the module Cstocs which can be used for reencoding of the list's data prior to calling czsort, or reencode this module to fit your
needs.
VERSION
0.68
SEE ALSO
perl(1), Cz::Cstocs(3).
AUTHOR
(c) 1997--2000 Jan Pazdziora <adelton@fi.muni.cz>, http://www.fi.muni.cz/~adelton/
at Faculty of Informatics, Masaryk University, Brno
perl v5.10.1 2000-05-16 Cz::Sort(3pm)