Remove duplicate files based on text string? Post: 302331720

Sponsored Content

Top Forums Shell Programming and Scripting Remove duplicate files based on text string? Post 302331720 by spangberg on Tuesday 7th of July 2009 12:02:02 PM

07-07-2009

Registered User

Remove duplicate files based on text string?

Hi

I have been struggling with a script for removing duplicate messages from a shared mailbox.
I would like to search for duplicate messages based on the “Message-ID” string within the messages files.

I have managed to find the duplicate “Message-ID” strings and (if I would like) delete the files in which they where found.
My problem is who to preserve one of each file.

My script so far:

--------------------
#!/bin/tcsh
set dir=/my/maildir

foreach file (`grep -h "Message-ID: <" $dir/* | uniq -d |xargs -i \grep -l "{}" $dir/*`)

rm -f "$file"

end

--------------------

Any ideas?

Thanks // Tomas

---------- Post updated at 06:02 PM ---------- Previous update was at 10:18 AM ----------

Fyi, solved
-------------------
#!/bin/tcsh
set maildir=/my/maildir
foreach dupstring ("`grep -m 1 -h -R "^Message-ID:" $maildir/ | sort | uniq -d`")
grep -l -R "$dupstring" $maildir/ |sed 1d |xargs -i \rm -f "{}"
end
-------------------

// Tomas

spangberg

View Public Profile for spangberg

Find all posts by spangberg

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

How to get remove duplicate of a file based on many conditions

Hii Friends.. I have a huge set of data stored in a file.Which is as shown below a.dat: RAO 1869 12 19 0 0 0.00 17.9000 82.3000 10.0 0 0.00 0 3.70 0.00 0.00 0 0.00 3.70 4 NULL LEE 1870 4 11 1 0 0.00 30.0000 99.0000 0.0 0 0.00 0 0.00 0.00 0.00 0 ...

2. Shell Programming and Scripting

Remove duplicate based on Group

3. Shell Programming and Scripting

Remove duplicate value based on two field $4 and $5

Hi All, i have input file like below... CA009156;20091003;M;AWBKCA72;123;;CANADIAN WESTERN BANK;EDMONTON;;2300, 10303, JASPER AVENUE;;T5J 3X6;; CA009156;20091003;M;AWBKCA72;321;;CANADIAN WESTERN BANK;EDMONTON;;2300, 10303, JASPER AVENUE;;T5J 3X6;; CA009156;20091003;M;AWBKCA72;231;;CANADIAN...

4. Shell Programming and Scripting

How To Remove Duplicate Based on the Value?

Hi , Some time i got duplicated value in my files , bundle_identifier= B Sometext=ABC bundle_identifier= A bundle_unit=500 Sometext123=ABCD bundle_unit=400 i need to check if there is a duplicated values or not if yes , i need to check if the value is A or B when Bundle_Identified ,...

5. Shell Programming and Scripting

Remove duplicate entries based on the range

I have file like this: chr start end chr15 99874874 99875874 chr15 99875173 99876173 aa1 chr15 99874923 99875923 chr15 99875173 99876173 aa1 chr15 99874962 99875962 chr15 99875173 99876173 aa1 chr1 ...

6. Shell Programming and Scripting

Remove not only the duplicate string but also the keyword of the string in Perl

Hi Perl users, I have another problem with text processing in Perl. I have a file below: Linux Unix Linux Windows SUN MACOS SUN SUN HP-AUX I want the result below: Unix Windows SUN MACOS HP-AUX so the duplicate string will be removed and also the keyword of the string on...

7. Shell Programming and Scripting

Remove duplicate rows based on one column

Dear members, I need to filter a file based on the 8th column (that is id), and does not mather the other columns, because I want just one id (1 line of each id) and remove the duplicates lines based on this id (8th column), and does not matter wich duplicate will be removed. example of my file...

8. Windows & DOS: Issues & Discussions

Remove duplicate lines from text files.

So, I have text files, one "fail.txt" And one "color.txt" I now want to use a command line (DOS) to remove ANY line that is PRESENT IN BOTH from each text file. Afterwards there shall be no duplicate lines.

9. Shell Programming and Scripting

Remove duplicate lines from file based on fields

Dear community, I have to remove duplicate lines from a file contains a very big ammount of rows (milions?) based on 1st and 3rd columns The data are like this: Region 23/11/2014 09:11:36 41752 Medio 23/11/2014 03:11:38 4132 Info 23/11/2014 05:11:09 4323...

10. Shell Programming and Scripting

Remove sections based on duplicate first line

Hi, I have a file with many sections in it. Each section is separated by a blank line. The first line of each section would determine if the section is duplicate or not. if the section is duplicate then remove the entire section from the file. below is the example of input and output....

LEARN ABOUT REDHAT

strndupa

STRDUP(3)						     Linux Programmer's Manual							 STRDUP(3)

NAME

       strdup, strndup, strdupa, strndupa - duplicate a string

SYNOPSIS

       #include <string.h>

       char *strdup(const char *s);

       #define _GNU_SOURCE
       #include <string.h>

       char *strndup(const char *s, size_t n);
       char *strdupa(const char *s);
       char *strndupa(const char *s, size_t n);

DESCRIPTION

       The  strdup()  function returns a pointer to a new string which is a duplicate of the string s.	Memory for the new string is obtained with
       malloc(3), and can be freed with free(3).

       The strndup() function is similar, but only copies at most n characters. If s is longer than n, only n characters are copied, and a  termi-
       nating NUL is added.

       strdupa	and strndupa are similar, but use alloca(3) to allocate the buffer. They are only available when using the GNU GCC suite, and suf-
       fer from the same limitations described in alloca(3).

RETURN VALUE

       The strdup() function returns a pointer to the duplicated string, or NULL if insufficient memory was available.

ERRORS

       ENOMEM Insufficient memory available to allocate duplicate string.

CONFORMING TO

       SVID 3, BSD 4.3.  strndup(), strdupa(), and strndupa() are GNU extensions.

SEE ALSO

       alloca(3), calloc(3), free(3), malloc(3), realloc(3)

GNU
								    1993-04-12								 STRDUP(3)

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

How to get remove duplicate of a file based on many conditions

Discussion started by: reva

2. Shell Programming and Scripting

Remove duplicate based on Group

Discussion started by: yale_work

3. Shell Programming and Scripting

Remove duplicate value based on two field $4 and $5

Discussion started by: mohan sharma

4. Shell Programming and Scripting

How To Remove Duplicate Based on the Value?

Discussion started by: OTNA