Performance problem with removing duplicates in a huge file (50+ GB) Post: 302752709

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

i have a file with some 1000 entries it will contain entries like 1000,ram 2000,pankaj 1001,rahim 1000,ram 2532,govind 2000,pankaj 3000,venkat 2532,govind what i want is i want to extract only the distinct rows from this file so my output should contain only 1000,ram...

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

hey all, I need some help. I have a text file with names in it. My target is that if a particular pattern exists in that file more than once..then i want to rename all the occurences of that pattern by alternate patterns.. for e.g if i have PATTERN occuring 5 times then i want to...

3. Shell Programming and Scripting

Removing duplicates from log file?

I have a log file with posts looking like this: -- Messages can be delivered by different systems at different times. The id number is used to sort out duplicate messages. What I need is to strip the arrival time from each post, sort posts by id number, and reattach arrival time to respective...

4. Shell Programming and Scripting

Removing Duplicates from file

5. Shell Programming and Scripting

formatting a file and removing duplicates

Hi, I have a file that I want to change the format of. It is a large file in rows but I want it to be comma separated (comma then a space). The current file looks like this: HI, Joe, Bob, Jack, Jack After I would want to remove any duplicates so it would look like this: HI, Joe,...

6. HP-UX

Performance issue with 'grep' command for huge file size

I have 2 files; one file (say, details.txt) contains the details of employees and another file (say, emp.txt) has some selected employee names. I am extracting employee details from details.txt by using emp.txt and the corresponding code is: while read line do emp_name=`echo $line` grep -e...

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Hi All, I am merging files coming from 2 different systems ,while doing that I am getting duplicates entries in the merged file I,01,000131,764,2,4.00 I,01,000131,765,2,4.00 I,01,000131,772,2,4.00 I,01,000131,773,2,4.00 I,01,000168,762,2,2.00 I,01,000168,763,2,2.00...

8. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3

9. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3.I have tried previous post also,but in that complete line must be similar.In this case i have to verify first column only regardless what is the content in succeeding columns.

10. Shell Programming and Scripting

Removing White spaces from a huge file

I am trying to remove whitespaces from a file containing sample data as: 457 <EOFD> Mar 1 2007 12:00:00:000AM <EOFD> Mar 31 2007 12:00:00:000AM <EOFD> system <EORD> 458 <EOFD> Mar 1 2007 12:00:00:000AM<EOFD>agf <EOFD> Apr 20 2007 9:10:56:036PM <EOFD> prodiws<EORD> . Basically these...

LEARN ABOUT SUSE

funindex

funindex(1)							SAORD Documentation						       funindex(1)

NAME

       funindex - create an index for a column of a FITS binary table

SYNOPSIS

       funindex <switches>  <iname> [oname]

OPTIONS

	 NB: these options are not compatible with Funtools processing. Please
	 use the defaults instead.
	 -c	   # compress output using gzip"
	 -a	   # ASCII output, ignore -c (default: FITS table)"
	 -f	   # FITS table output (default: FITS table)"
	 -l	   # long output, i.e. with key value(s) (default: long)"
	 -s	   # short output, i.e. no key value(s) (default: long)"

DESCRIPTION

       The funindex script creates an index for the specified column (key) by running funtable -s (sort) and then saving the column value and the
       record number for each sorted row. This index will be used automatically
	by funtools filtering of that column, provided the index file's modification date is later than that of the data file.

       The first required argument is the name of the FITS binary table to index. Please note that text files cannot be indexed at this time.  The
       second required argument is the column (key) name to index. While multiple keys can be specified in principle, the funtools index process-
       ing assume a single key and will not recognize files containing multiple keys.

       By default, the output index file name is [root]_[key].idx, where [root] is the root of the input file. Funtools looks for this specific
       file name when deciding whether to use an index for faster filtering. Therefore, the optional third argument (output file name) should not
       be used for funtools processing.

       For example, to create an index on column Y for a given FITS file, use:

	 funindex foo.fits Y

       This will generate an index named foo_y.idx, which will be used by funtools for filters involving the Y column.

SEE ALSO

       See funtools(7) for a list of Funtools help pages

version 1.4.2							  January 2, 2008						       funindex(1)

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

Discussion started by: trichyselva

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

Discussion started by: ashisharora

3. Shell Programming and Scripting

Removing duplicates from log file?

Discussion started by: Ilja

4. Shell Programming and Scripting

Removing Duplicates from file

Discussion started by: tinufarid

5. Shell Programming and Scripting

formatting a file and removing duplicates

Discussion started by: kylle345

6. HP-UX

Performance issue with 'grep' command for huge file size

Discussion started by: arb_1984

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Discussion started by: Sri3001

8. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

9. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

10. Shell Programming and Scripting

Removing White spaces from a huge file

Discussion started by: amvip

LEARN ABOUT SUSE

funindex