Performance problem with removing duplicates in a huge file (50+ GB) Post: 302751839

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

i have a file with some 1000 entries it will contain entries like 1000,ram 2000,pankaj 1001,rahim 1000,ram 2532,govind 2000,pankaj 3000,venkat 2532,govind what i want is i want to extract only the distinct rows from this file so my output should contain only 1000,ram...

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

hey all, I need some help. I have a text file with names in it. My target is that if a particular pattern exists in that file more than once..then i want to rename all the occurences of that pattern by alternate patterns.. for e.g if i have PATTERN occuring 5 times then i want to...

3. Shell Programming and Scripting

Removing duplicates from log file?

I have a log file with posts looking like this: -- Messages can be delivered by different systems at different times. The id number is used to sort out duplicate messages. What I need is to strip the arrival time from each post, sort posts by id number, and reattach arrival time to respective...

4. Shell Programming and Scripting

Removing Duplicates from file

5. Shell Programming and Scripting

formatting a file and removing duplicates

Hi, I have a file that I want to change the format of. It is a large file in rows but I want it to be comma separated (comma then a space). The current file looks like this: HI, Joe, Bob, Jack, Jack After I would want to remove any duplicates so it would look like this: HI, Joe,...

6. HP-UX

Performance issue with 'grep' command for huge file size

I have 2 files; one file (say, details.txt) contains the details of employees and another file (say, emp.txt) has some selected employee names. I am extracting employee details from details.txt by using emp.txt and the corresponding code is: while read line do emp_name=`echo $line` grep -e...

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Hi All, I am merging files coming from 2 different systems ,while doing that I am getting duplicates entries in the merged file I,01,000131,764,2,4.00 I,01,000131,765,2,4.00 I,01,000131,772,2,4.00 I,01,000131,773,2,4.00 I,01,000168,762,2,2.00 I,01,000168,763,2,2.00...

8. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3

9. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3.I have tried previous post also,but in that complete line must be similar.In this case i have to verify first column only regardless what is the content in succeeding columns.

10. Shell Programming and Scripting

Removing White spaces from a huge file

I am trying to remove whitespaces from a file containing sample data as: 457 <EOFD> Mar 1 2007 12:00:00:000AM <EOFD> Mar 31 2007 12:00:00:000AM <EOFD> system <EORD> 458 <EOFD> Mar 1 2007 12:00:00:000AM<EOFD>agf <EOFD> Apr 20 2007 9:10:56:036PM <EOFD> prodiws<EORD> . Basically these...

LEARN ABOUT PLAN9

dd

DD(1)							      General Commands Manual							     DD(1)

NAME

       dd - convert and copy a file

SYNOPSIS

       dd [ option value ] ...

DESCRIPTION

       Dd  copies  the specified input file to the specified output with possible conversions.	The standard input and output are used by default.
       The input and output block size may be specified to take advantage of raw physical I/O.	The options are

       -if f   Open file f for input.

       -of f   Open file f for output.

       -ibs n  Set input block size to n bytes (default 512).

       -obs n  Set output block size (default 512).

       -bs n   Set both input and output block size, superseding ibs and obs.  If no conversion  is  specified,  preserve  the	input  block  size
	       instead of packing short blocks into the output buffer.	This is particularly efficient since no in-core copy need be done.

       -cbs n  Set conversion buffer size.

       -skip n Skip n input records before copying.

       -iseek n
	       Seek n records forward on input file before copying.

       -files n
	       Catenate n input files (useful only for magnetic tape or similar input device).

       -oseek n
	       Seek n records from beginning of output file before copying.

       -count n
	       Copy only n input records.

       -conv ascii    Convert EBCDIC to ASCII.

	    ebcdic   Convert ASCII to EBCDIC.

	    ibm      Like ebcdic but with a slightly different character map.

	    block    Convert variable length ASCII records to fixed length.

	    unblock  Convert fixed length ASCII records to variable length.

	    lcase    Map alphabetics to lower case.

	    ucase    Map alphabetics to upper case.

	    swab     Swap every pair of bytes.

	    noerror  Do not stop processing on an error.

	    sync     Pad every input record to ibs bytes.

       Where  sizes are specified, a number of bytes is expected.  A number may end with or to specify multiplication by 1024 or 512 respectively;
       a pair of numbers may be separated by to indicate a product.  Multiple conversions may be specified in the style:

       is used only if or conversion is specified.  In the first two cases, n characters are copied into  the  conversion  buffer,  any  specified
       character  mapping  is  done, trailing blanks are trimmed and new-line is added before sending the line to the output.  In the latter three
       cases, characters are read into the conversion buffer and blanks are added to make up an output record of size n.   If  is  unspecified	or
       zero,  the  and	options  convert the character set without changing the block structure of the input file; the and options become a simple
       file copy.

SOURCE

       /sys/src/cmd/dd.c

SEE ALSO

       cp(1)

DIAGNOSTICS

       Dd reports the number of full + partial input and output blocks handled.

																	     DD(1)

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

Discussion started by: trichyselva

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

Discussion started by: ashisharora

3. Shell Programming and Scripting

Removing duplicates from log file?

Discussion started by: Ilja

4. Shell Programming and Scripting

Removing Duplicates from file

Discussion started by: tinufarid

5. Shell Programming and Scripting

formatting a file and removing duplicates

Discussion started by: kylle345

6. HP-UX

Performance issue with 'grep' command for huge file size

Discussion started by: arb_1984

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Discussion started by: Sri3001

8. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

9. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

10. Shell Programming and Scripting

Removing White spaces from a huge file

Discussion started by: amvip

LEARN ABOUT PLAN9

dd