Filter file to remove duplicate values in first column
Hello,
I have a script that is generating a tab delimited output file.
I need to trim this down so that there a no duplicates in the first column. Actually, the entire row would be a duplicate, but I don't see any reason to look at anything other than the index value. There is no particular rational to the order and there could be any number of duplicates of a given row.
The final results should look like this,
I need one, and only one, instance of each index value ("num" column value) in the file, not just the lines with num values that appear only once. There always seems to be some confusion about that with discussions of "unique" lines.
The only thing I could think of was to sort the rows on the num column value and then loop through checking if the num value was equal to the previous line. If it is not equal, copy it to a new array, etc.
Any suggestions? There always seems to be some simple one line solution that I don't know about.
I am a beginner in Unix. Though have been asked to write a script to filter(remove duplicates) data from a .dat file. File is very huge containig billions of records.
contents of file looks like
30002157,40342424,OTC,mart_rec,100, ,0
30002157,40343369,OTC,mart_rec,95, ,0... (6 Replies)
Hi, I've got a file that I'd like to uniquely sort based on column 2 (values in column 2 begin with "comp").
I tried sort -t -nuk2,3 file.txtBut got:
sort: multi-character tab `-nuk2,3'
"man sort" did not help me out
Any pointers?
Input:
Output: (5 Replies)
Hello experts,
I have a requirement where I have to implement two checks on a csv file:
1. Check to see if the value in first column is duplicate, if any value is duplicate script should exit.
2. Check to verify if the value at second column is between "yes" or "no", if it is anything else... (4 Replies)
I have a .CSV file with the below format:
"column 1","column 2","column 3","column 4","column 5","column 6","column 7","column 8","column 9","column 10
"12310","42324564756","a simple string with a , comma","string with or, without commas","string 1","USD","12","70%","08/01/2013",""... (2 Replies)
Input
1,ABCD,no
2,system,yes
3,ABCD,yes
4,XYZ,no
5,XYZ,yes
6,pc,noCode used to find duplicate with regard to 2nd column
awk 'NR == 1 {p=$2; next} p == $2 { print "Line" NR "$2 is duplicated"} {p=$2}' FS="," ./input.csv
Now is there a wise way to de-duplicate the entire line (remove... (4 Replies)
Hi Gurus,
I have a file(weblog) as below
abc|xyz|123|agentcode=sample code abcdeeess,agentcode=sample code abcdeeess,agentcode=sample code abcdeeess|agentadd=abcd stereet 23343,agentadd=abcd stereet 23343
sss|wwq|999|agentcode=sample1 code wqwdeeess,gentcode=sample1 code... (4 Replies)
Dear folks
I have a map file of around 54K lines and some of the values in the second column have the same value and I want to find them and delete all of the same values. I looked over duplicate commands but my case is not to keep one of the duplicate values. I want to remove all of the same... (4 Replies)
I have csv file with 30, 40 columns
Pasting just three column for problem description
I want to filter record if column 1 matches CN or DN then,
check for values in column 2 if column contain 1235, 1235 then in column 3 values must be sequence of 2345, 2345
and if column 2 contains 6789, 6789... (5 Replies)
Hi Experts,
I have csv file with 30, 40 columns
Pasting just 2 column for problem description.
Need to print error if below combination is not present in file
check for column-1 (DocumentNumber) and filter columns where value in DocumentNumber field is same.
For all such rows, the field... (7 Replies)
Discussion started by: as7951
7 Replies
LEARN ABOUT CENTOS
tcttest
TCTTEST(1) Tokyo Cabinet TCTTEST(1)NAME
tcttest - test cases of the table database API
DESCRIPTION
The command `tcttest' is a utility for facility test and performance test. This command is used in the following format. `path' specifies
the path of a database file. `rnum' specifies the number of iterations. `bnum' specifies the number of buckets. `apow' specifies the
power of the alignment. `fpow' specifies the power of the free block pool.
tcttest write [-mt] [-tl] [-td|-tb|-tt|-tx] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-ip] [-is] [-in] [-it] [-if] [-ix]
[-nl|-nb] [-rnd] path rnum [bnum [apow [fpow]]]
Store records with columns "str", "num", "type", and "flag".
tcttest read [-mt] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-nl|-nb] [-rnd] path
Retrieve all records of the database above.
tcttest remove [-mt] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-nl|-nb] [-rnd] path
Remove all records of the database above.
tcttest rcat [-mt] [-tl] [-td|-tb|-tt|-tx] [-rc num] [-lc num] [-nc num] [-xm num] [-df num] [-ip] [-is] [-in] [-it] [-if] [-ix]
[-nl|-nb] [-pn num] [-dai|-dad|-rl|-ru] path rnum [bnum [apow [fpow]]]
Store records with partway duplicated keys using concatenate mode.
tcttest misc [-mt] [-tl] [-td|-tb|-tt|-tx] [-nl|-nb] path rnum
Perform miscellaneous test of various operations.
tcttest wicked [-mt] [-tl] [-td|-tb|-tt|-tx] [-nl|-nb] path rnum
Perform updating operations selected at random.
Options feature the following.
-mt : call the function `tctdbsetmutex'.
-tl : enable the option `TDBTLARGE'.
-td : enable the option `TDBTDEFLATE'.
-tb : enable the option `TDBTBZIP'.
-tt : enable the option `TDBTTCBS'.
-tx : enable the option `TDBTEXCODEC'.
-rc num : specify the number of cached records.
-lc num : specify the number of cached leaf pages.
-nc num : specify the number of cached non-leaf pages.
-xm num : specify the size of the extra mapped memory.
-df num : specify the unit step number of auto defragmentation.
-ip : create the number index for the primary key.
-is : create the string index for the column "str".
-in : create the number index for the column "num".
-it : create the string index for the column "type".
-if : create the token inverted index for the column "flag".
-ix : create the q-gram inverted index for the column "text".
-nl : enable the option `TDBNOLCK'.
-nb : enable the option `TDBLCKNB'.
-rnd : select keys at random.
-pn num : specify the number of patterns.
-dai : use the function `tctdbaddint' instead of `tctdbputcat'.
-dad : use the function `tctdbadddouble' instead of `tctdbputcat'.
-rl : set the length of values at random.
-ru : select update operations at random.
This command returns 0 on success, another on failure.
SEE ALSO tctmttest(1), tctmgr(1), tctdb(3), tokyocabinet(3)Man Page 2012-08-18 TCTTEST(1)