Performance problem with removing duplicates in a huge file (50+ GB) Post: 302752727

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

i have a file with some 1000 entries it will contain entries like 1000,ram 2000,pankaj 1001,rahim 1000,ram 2532,govind 2000,pankaj 3000,venkat 2532,govind what i want is i want to extract only the distinct rows from this file so my output should contain only 1000,ram...

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

hey all, I need some help. I have a text file with names in it. My target is that if a particular pattern exists in that file more than once..then i want to rename all the occurences of that pattern by alternate patterns.. for e.g if i have PATTERN occuring 5 times then i want to...

3. Shell Programming and Scripting

Removing duplicates from log file?

I have a log file with posts looking like this: -- Messages can be delivered by different systems at different times. The id number is used to sort out duplicate messages. What I need is to strip the arrival time from each post, sort posts by id number, and reattach arrival time to respective...

4. Shell Programming and Scripting

Removing Duplicates from file

5. Shell Programming and Scripting

formatting a file and removing duplicates

Hi, I have a file that I want to change the format of. It is a large file in rows but I want it to be comma separated (comma then a space). The current file looks like this: HI, Joe, Bob, Jack, Jack After I would want to remove any duplicates so it would look like this: HI, Joe,...

6. HP-UX

Performance issue with 'grep' command for huge file size

I have 2 files; one file (say, details.txt) contains the details of employees and another file (say, emp.txt) has some selected employee names. I am extracting employee details from details.txt by using emp.txt and the corresponding code is: while read line do emp_name=`echo $line` grep -e...

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Hi All, I am merging files coming from 2 different systems ,while doing that I am getting duplicates entries in the merged file I,01,000131,764,2,4.00 I,01,000131,765,2,4.00 I,01,000131,772,2,4.00 I,01,000131,773,2,4.00 I,01,000168,762,2,2.00 I,01,000168,763,2,2.00...

8. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3

9. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3.I have tried previous post also,but in that complete line must be similar.In this case i have to verify first column only regardless what is the content in succeeding columns.

10. Shell Programming and Scripting

Removing White spaces from a huge file

I am trying to remove whitespaces from a file containing sample data as: 457 <EOFD> Mar 1 2007 12:00:00:000AM <EOFD> Mar 31 2007 12:00:00:000AM <EOFD> system <EORD> 458 <EOFD> Mar 1 2007 12:00:00:000AM<EOFD>agf <EOFD> Apr 20 2007 9:10:56:036PM <EOFD> prodiws<EORD> . Basically these...

LEARN ABOUT DEBIAN

ppmd

ppmd(1) 							       utils								   ppmd(1)

NAME

       ppmd - file-to-file compressor

SYNTAX

       ppmd [e|d] [switches] filename...|wildcard...

DESCRIPTION

       It  is  written	for  embedding in user programs mainly and it is not intended for immediate use. I was interested in speed and performance
       improvements of abstract PPM model [1-6] only, without tuning it to particular data types,  therefore  compressor  works  good  enough  for
       texts, but it is not so good for nonhomogeneous files (executables) and for noisy analog data (sounds, pictures etc.). Program is very mem-
       ory consuming, you can choose balance between execution speed and memory economy, on one hand,  and  compression  performance,  on  another
       hand, with the help of model order selection option (-o).

OPTIONS

       -d     Delete file[s] after processing, default: disabled.

       -s     Silent mode.

       -fName Set output file name to Name.

       -mN    Use  N  MB memory - [1,256], default: 10. The PPMII algorithm might need a lot of memory, especially when used on large files and/or
	      used with large model order. If ppmd needs more memory than you give it, the compression will be worse. The exact effect	is  depen-
	      dent on the -r option.

       -oN    Set  model  order  to N - [2,16], default: 4. Bigger model orders almost surely results in better compression and surely more memory
	      and CPU usage.

       -r{0,1,2}
	      Methods of restoration of model correctness at memory insufficiency:
		  '-r0 - restart model from scratch'. This method is not optimal for any type of data sources, but it works fast and efficient	in
	      average, so it is the recommended method (default).
		  '-r1	-  cut	off model'. This method is optimal for quasistationary sources when the period of stationarity is much larger than
	      period between cutoffs.  As a rule, it gives better results, but it is slower than other methods and it is unstable against fragmen-
	      tation of memory heap at high model orders and low memory.
		  '-r2	- freeze model'. This method is optimal for stationary sources (show me such source when You will find it ;-)). It is fast
	      and efficient for such sources.

EXAMPLES

       To run this program the standard way type:

       ppmd e /tmp/myfile

       Alternatively you can run it as:

       ppmd -e -o 16 /tmp/myfile

AUTHORS

       PPMd was written by Dmitry Shkarin <dmitry.shkarin@mtu-net.ru> and Dmitry Subbotin.

SEE ALSO

       gzip(1), bzip2(1), lzma(1).

10.1								    2011-07-25								   ppmd(1)

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

Discussion started by: trichyselva

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

Discussion started by: ashisharora

3. Shell Programming and Scripting

Removing duplicates from log file?

Discussion started by: Ilja

4. Shell Programming and Scripting

Removing Duplicates from file

Discussion started by: tinufarid

5. Shell Programming and Scripting

formatting a file and removing duplicates

Discussion started by: kylle345

6. HP-UX

Performance issue with 'grep' command for huge file size

Discussion started by: arb_1984

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Discussion started by: Sri3001

8. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

9. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

10. Shell Programming and Scripting

Removing White spaces from a huge file

Discussion started by: amvip

LEARN ABOUT DEBIAN

ppmd