Sponsored Content
Top Forums UNIX for Advanced & Expert Users Performance problem with removing duplicates in a huge file (50+ GB) Post 302752727 by Corona688 on Monday 7th of January 2013 12:31:39 PM
Old 01-07-2013
I don't think there is a super-fast way to handle 50G of data.

Taking advantage of a database index sounds as good a way as any, a proper DB is designed to duplicate-check data larger than memory.
 

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

i have a file with some 1000 entries it will contain entries like 1000,ram 2000,pankaj 1001,rahim 1000,ram 2532,govind 2000,pankaj 3000,venkat 2532,govind what i want is i want to extract only the distinct rows from this file so my output should contain only 1000,ram... (2 Replies)
Discussion started by: trichyselva
2 Replies

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

hey all, I need some help. I have a text file with names in it. My target is that if a particular pattern exists in that file more than once..then i want to rename all the occurences of that pattern by alternate patterns.. for e.g if i have PATTERN occuring 5 times then i want to... (3 Replies)
Discussion started by: ashisharora
3 Replies

3. Shell Programming and Scripting

Removing duplicates from log file?

I have a log file with posts looking like this: -- Messages can be delivered by different systems at different times. The id number is used to sort out duplicate messages. What I need is to strip the arrival time from each post, sort posts by id number, and reattach arrival time to respective... (2 Replies)
Discussion started by: Ilja
2 Replies

4. Shell Programming and Scripting

Removing Duplicates from file

Hi Experts, Please check the following new requirement. I got data like the following in a file. FILE_HEADER 01cbbfde7898410| 3477945| home| 1 01cbc275d2c122| 3478234| WORK| 1 01cbbe4362743da| 3496386| Rich Spare| 1 01cbc275d2c122| 3478234| WORK| 1 This is pipe separated file with... (3 Replies)
Discussion started by: tinufarid
3 Replies

5. Shell Programming and Scripting

formatting a file and removing duplicates

Hi, I have a file that I want to change the format of. It is a large file in rows but I want it to be comma separated (comma then a space). The current file looks like this: HI, Joe, Bob, Jack, Jack After I would want to remove any duplicates so it would look like this: HI, Joe,... (2 Replies)
Discussion started by: kylle345
2 Replies

6. HP-UX

Performance issue with 'grep' command for huge file size

I have 2 files; one file (say, details.txt) contains the details of employees and another file (say, emp.txt) has some selected employee names. I am extracting employee details from details.txt by using emp.txt and the corresponding code is: while read line do emp_name=`echo $line` grep -e... (7 Replies)
Discussion started by: arb_1984
7 Replies

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Hi All, I am merging files coming from 2 different systems ,while doing that I am getting duplicates entries in the merged file I,01,000131,764,2,4.00 I,01,000131,765,2,4.00 I,01,000131,772,2,4.00 I,01,000131,773,2,4.00 I,01,000168,762,2,2.00 I,01,000168,763,2,2.00... (5 Replies)
Discussion started by: Sri3001
5 Replies

8. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3 (2 Replies)
Discussion started by: sagar_1986
2 Replies

9. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3.I have tried previous post also,but in that complete line must be similar.In this case i have to verify first column only regardless what is the content in succeeding columns. (3 Replies)
Discussion started by: sagar_1986
3 Replies

10. Shell Programming and Scripting

Removing White spaces from a huge file

I am trying to remove whitespaces from a file containing sample data as: 457 <EOFD> Mar 1 2007 12:00:00:000AM <EOFD> Mar 31 2007 12:00:00:000AM <EOFD> system <EORD> 458 <EOFD> Mar 1 2007 12:00:00:000AM<EOFD>agf <EOFD> Apr 20 2007 9:10:56:036PM <EOFD> prodiws<EORD> . Basically these... (11 Replies)
Discussion started by: amvip
11 Replies
RPLAY(1)						      General Commands Manual							  RPLAY(1)

NAME
rplay - play, pause, continue, and stop sounds SYNOPSIS
rplay [options] [sound ...] DESCRIPTION
rplay is client that communicates with rplayd to play, pause, continue, and stop sounds using both the RPLAY and RPTP protocols. Sound files can be played by rplayd directly if available on the local system or sounds can be sent over the network using UDP or TCP/IP. rplay will attempt to determine whether or not the server has the sound before using the network. OPTIONS
-b BYTES, --buffer-size=BYTES Use of a buffer size of BYTES when playing sounds using RPTP flows. The default is 8K. -c, --continue Continue sounds. -n N, --count=N Number of times to play the sound, default = 1. -N N, --list-count=N Number of times to play all the sounds, default = 1. --list-name=NAME Name this list NAME. rplayd appends sounds with the same NAME into the same sound list -- it plays them sequentially. --help Display helpful information. -h HOST, --host=HOST, --hosts=HOST Specify the rplay host, default = localhost. -i INFO, --info=INFO Audio information for a sound file. This option is intended to be used when sounds are read from standard input. INFO must be of the form: `format,sample-rate,bits,channels,byte-order,offset' Examples: ulaw,8000,8,1,big-endian,0 gsm,8000 Shorthand info is provided for Sun's audio devices using the following options: --info-amd, --info-dbri, --info-cs4231. There's also: --info-ulaw and --info-gsm. -p, --pause Pause sounds. --port=PORT Use PORT instead of the default RPLAY/UDP or RPTP/TCP port. -P N, --priority=N Play sounds at priority N (0 <= N <= 255), default = 0. -r, --random Randomly choose one of the given sounds. --reset Tell the server to reset itself. --rplay, --RPLAY Force the use of the RPLAY protocol. The default protocol to be used is determined by checking whether or not the server has local access to the specified sounds. RPLAY is used when sounds are accessible, otherwise RPTP and possibly flows are used. RPLAY will also be used when sound accessibility cannot be determined. --rptp, --RPTP Force the use of the RPTP protocol. See `--rplay' for more information about protocols. -R N, --sample-rate=N Play sounds at sample rate N, default = 0. -s, --stop Stop sounds. --version Print the rplay version and exit. -v N, --volume=N Play sounds at volume N (0 <= N <= 255), default = 127. SEE ALSO
rplayd(8), rptp(1) 6/29/98 RPLAY(1)
All times are GMT -4. The time now is 04:21 AM.
Unix & Linux Forums Content Copyright 1993-2022. All Rights Reserved.
Privacy Policy