Removing duplicates in a sorted file by field. Post: 302190380

Sponsored Content

Top Forums Shell Programming and Scripting Removing duplicates in a sorted file by field. Post 302190380 by kinksville on Tuesday 29th of April 2008 12:22:05 PM

04-29-2008

Registered User

Removing duplicates in a sorted file by field.

I have data like this:
It's sorted by the 2nd field (TID).
envoy,90000000000000634600010001,04/11/2008,23:19:27,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/12/2008,04:23:45,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/12/2008,23:14:25,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/13/2008,04:23:39,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/13/2008,22:41:58,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/13/2008,22:42:44,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/13/2008,22:49:43,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/13/2008,22:50:45,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/13/2008,22:53:23,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/14/2008,12:38:40,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/14/2008,12:52:22,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000693200010001,04/17/2008,09:07:09,RB00060,0009,ENVOY,ERROR,26
envoy,90000000000000693200010001,04/18/2008,10:27:13,RB00083,0009,ENVOY,ERROR,26
envoy,90000000000000693200010001,04/18/2008,11:36:27,RB00084,0009,ENVOY,ERROR,26
envoy,90000000000001034800010001,04/01/2008,23:59:15,RB00294,0030,ENVOY,ERROR,57
envoy,90000000000001034800010001,04/02/2008,23:59:12,RB00295,0030,ENVOY,ERROR,57
envoy,90000000000001034800010001,04/03/2008,23:59:11,RB00296,0030,ENVOY,ERROR,57
envoy,90000000000001034800010001,04/04/2008,23:59:08,RB00297,0030,ENVOY,ERROR,57
envoy,90000000000001034800010001,04/05/2008,23:59:04,RB00297,0030,ENVOY,ERROR,57
envoy,90000000000001034800010001,04/06/2008,22:59:06,RB00297,0030,ENVOY,ERROR,57

I want to do the following:
Check the second field to see if the TID is the same as the previous line. If the TID has been seen before then check the 7th field to see if that is the same as the previous line. If both are the same, I want to remove the line and increment a counter.

My ideal output would look something like this.
11,envoy,90000000000000634600010001,04/11/2008,23:19:27,RB00266,0015,DETAIL,ERROR,
3,envoy,90000000000000693200010001,04/17/2008,09:07:09,RB00060,0009,ENVOY,ERROR,26

etc.

I figure I actually need to do an awk script rather than a 1 liner. The other option is for the last 3 fields to be one field and compare by the TID field and the error field, then split them into 3 on the output.

Any thoughts? I've looked at other stuff removing dups with awk and it's mostly one liners.

I'd love to ask get some explanation of WHY it works so that I can mod it if need be.

kinksville

View Public Profile for kinksville

Find all posts by kinksville

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

i have a file with some 1000 entries it will contain entries like 1000,ram 2000,pankaj 1001,rahim 1000,ram 2532,govind 2000,pankaj 3000,venkat 2532,govind what i want is i want to extract only the distinct rows from this file so my output should contain only 1000,ram...

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

hey all, I need some help. I have a text file with names in it. My target is that if a particular pattern exists in that file more than once..then i want to rename all the occurences of that pattern by alternate patterns.. for e.g if i have PATTERN occuring 5 times then i want to...

3. Shell Programming and Scripting

Removing duplicates from log file?

I have a log file with posts looking like this: -- Messages can be delivered by different systems at different times. The id number is used to sort out duplicate messages. What I need is to strip the arrival time from each post, sort posts by id number, and reattach arrival time to respective...

4. Shell Programming and Scripting

Removing Duplicates from file

5. Shell Programming and Scripting

formatting a file and removing duplicates

Hi, I have a file that I want to change the format of. It is a large file in rows but I want it to be comma separated (comma then a space). The current file looks like this: HI, Joe, Bob, Jack, Jack After I would want to remove any duplicates so it would look like this: HI, Joe,...

6. Shell Programming and Scripting

Removing duplicates depending on file size

Hi all, I am working with a huge amount of files in a Linux environment and I was trying to filter my data. Here's what my data looks like Name............................Size OLUSDN.gf.gif-1.JPEG.......5 kb LKJFDA01.gf.gif-1.JPEG.....3 kb LKJFDA01.gf.gif-2.JPEG.....1 kb...

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Hi All, I am merging files coming from 2 different systems ,while doing that I am getting duplicates entries in the merged file I,01,000131,764,2,4.00 I,01,000131,765,2,4.00 I,01,000131,772,2,4.00 I,01,000131,773,2,4.00 I,01,000168,762,2,2.00 I,01,000168,763,2,2.00...

8. UNIX for Dummies Questions & Answers

Grep from pattern file without removing duplicates?

9. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3

10. Shell Programming and Scripting

Removing duplicates from new file

i hav two files like i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3.I have tried previous post also,but in that complete line must be similar.In this case i have to verify first column only regardless what is the content in succeeding columns.

LEARN ABOUT OSX

db_upgrade

db_upgrade(1)						    BSD General Commands Manual 					     db_upgrade(1)

NAME

     db_upgrade

SYNOPSIS

     db_upgrade [-NsV] [-h home] [-P password] file ...

DESCRIPTION

     The db_upgrade utility upgrades the Berkeley DB version of one or more files and the databases they contain to the current release version.

     The options are as follows:

     -h
	Specify a home directory for the database environment; by default, the current working directory is used.

     -N
	Do not acquire shared region mutexes while running. Other problems, such as potentially fatal errors in Berkeley DB, will be ignored as
	well. This option is intended only for debugging errors, and should not be used under any other circumstances.

     -P
	Specify an environment password. Although Berkeley DB utilities overwrite password strings as soon as possible, be aware there may be a
	window of vulnerability on systems where unprivileged users can see command-line arguments or where utilities are not able to overwrite
	the memory containing the command-line arguments.

     -s
	This flag is only meaningful when upgrading databases from releases before the Berkeley DB 3.1 release.

	As part of the upgrade from the Berkeley DB 3.0 release to the 3.1 release, the on-disk format of duplicate data items changed. To cor-
	rectly upgrade the format requires that applications specify whether duplicate data items in the database are sorted or not. Specifying
	the -s flag means that the duplicates are sorted; otherwise, they are assumed to be unsorted. Incorrectly specifying the value of this
	flag may lead to database corruption.

	Because the db_upgrade utility upgrades a physical file (including all the databases it contains), it is not possible to use db_upgrade to
	upgrade files where some of the databases it includes have sorted duplicate data items, and some of the databases it includes have
	unsorted duplicate data items. If the file does not have more than a single database, if the databases do not support duplicate data
	items, or if all the databases that support duplicate data items support the same style of duplicates (either sorted or unsorted),
	db_upgrade will work correctly as long as the -s flag is correctly specified. Otherwise, the file cannot be upgraded using db_upgrade, and
	must be upgraded manually using the db_dump and db_load utilities.

     -V
	Write the library version number to the standard output, and exit.

     It is important to realize that Berkeley DB database upgrades are done in place, and so are potentially destructive. This means that if the
     system crashes during the upgrade procedure, or if the upgrade procedure runs out of disk space, the databases may be left in an inconsistent
     and unrecoverable state. See Upgrading databases for more information.

     The db_upgrade utility may be used with a Berkeley DB environment (as described for the -h option, the environment variable DB_HOME, or
     because the utility was run in a directory containing a Berkeley DB environment). In order to avoid environment corruption when using a
     Berkeley DB environment, db_upgrade should always be given the chance to detach from the environment and exit gracefully. To cause db_upgrade
     to release all environment resources and exit cleanly, send it an interrupt signal (SIGINT).

     The db_upgrade utility exits 0 on success, and >0 if an error occurs.

ENVIRONMENT

     DB_HOME  If the -h option is not specified and the environment variable DB_HOME is set, it is used as the path of the database home, as
	      described in DB_ENV->open.

SEE ALSO

     db_archive(1), db_checkpoint(1), db_deadlock(1), db_dump(1), db_load(1), db_printlog(1), db_recover(1), db_stat(1), db_verify(1)

Darwin								 December 3, 2003							    Darwin

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

removing duplicates from a file

Discussion started by: trichyselva

2. UNIX for Dummies Questions & Answers

removing duplicates of a pattern from a file

Discussion started by: ashisharora

3. Shell Programming and Scripting

Removing duplicates from log file?

Discussion started by: Ilja

4. Shell Programming and Scripting

Removing Duplicates from file

Discussion started by: tinufarid

5. Shell Programming and Scripting

formatting a file and removing duplicates

Discussion started by: kylle345

6. Shell Programming and Scripting

Removing duplicates depending on file size

Discussion started by: Error404

7. UNIX for Dummies Questions & Answers

Removing duplicates from a file

Discussion started by: Sri3001

8. UNIX for Dummies Questions & Answers

Grep from pattern file without removing duplicates?

Discussion started by: Mauve

9. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

10. Shell Programming and Scripting

Removing duplicates from new file

Discussion started by: sagar_1986

LEARN ABOUT OSX

db_upgrade