Remove duplicate occurrences of text pattern

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Remove comments like pattern from text

Hi , We need to remove comment like pattern from a code text. The possible comment expressions are as follows. Input BizComment : Special/*@ Name:bzt_53_3aea640a_51783afa_5d64_0 BizHidden:true @*/ /* lookup Disease Category Therapuetic Class */ a=b;...

2. Shell Programming and Scripting

How to remove duplicate text blocks from a file?

Hi All I have a list of files which will have duplicate list of blocks of text. Following is a sample of the file, I have removed the sensitive information from the file. All the code samples starts from <TR BGCOLOR="white"> and Ends with IP address and two html tags like this. 10.14.22.22...

3. Windows & DOS: Issues & Discussions

Remove duplicate lines from text files.

So, I have text files, one "fail.txt" And one "color.txt" I now want to use a command line (DOS) to remove ANY line that is PRESENT IN BOTH from each text file. Afterwards there shall be no duplicate lines.

4. Shell Programming and Scripting

Remove duplicate line starting with a pattern

HI, I have the below input file /* ----------------- cmdsDlyStartFWJ -----------------*/ UNIX_JOB CMDS065J RUN ANY CMDNAME sleep 5 AGENT CMDSHP USER proddata RUN MON,TUE,WED,THU,FRI DELAYSUB 02:00 /* "Triggers daily file watcher jobs" */ ENVAR...

5. Shell Programming and Scripting

Filter or remove duplicate block of text without distinguishing marks or fields

Hello, Although I have found similar questions, I could not find advice that could help with our problem. The issue: We have several hundreds text files containing repeated blocks of text (I guess back at the time they were prepared like that to optmize printing). The block of texts...

6. Shell Programming and Scripting

Help with remove last text of a file that have specific pattern

Input file matrix-remodelling_associated_8_ aurora_interacting_1_ L20 von_factor_A_domain_1 ATP_containing_3B_ . . Output file matrix-remodelling_associated_8 aurora_interacting_1 L20 von_factor_A_domain_1 ATP_containing_3B . .

7. Shell Programming and Scripting

How to remove all text except pattern

i have nasty html file with 2000+ simbols in 1 row...i need to remove whole the code except title="Some title..." and store those into file with titles (the whole text is in variable text) i've tried something like this: echo $text | sed 's/.*$title=\".+\"$.*/\1/' > titles.html BUT it does...

8. Shell Programming and Scripting

Count the number of occurrences of a pattern between each occurrence of a different pattern

I need to count the number of occurrences of a pattern, say 'key', between each occurrence of a different pattern, say 'lu'. Here's a portion of the text I'm trying to parse: lu S1234L_149_m1_vg.6, part-att 1, vdp-att 1 p-reserver IID 0xdb registrations: key 4156 4353 0000 0000 ...

9. Shell Programming and Scripting

Remove duplicate files based on text string?

Hi I have been struggling with a script for removing duplicate messages from a shared mailbox. I would like to search for duplicate messages based on the “Message-ID” string within the messages files. I have managed to find the duplicate “Message-ID” strings and (if I would like) delete...

10. Shell Programming and Scripting

Remove duplicate text

Hello, I have a log file which is generated by a script which looks like this: userid: 7 starttime: Sat May 24 23:24:13 CEST 2008 endtime: Sat May 24 23:26:57 CEST 2008 total time spent: 2.73072 minutes / 163.843 seconds date: Sat Jun 7 16:09:03 CEST 2008 userid: 8 starttime: Sun May...

LEARN ABOUT BSD

sed

SED(1)							      General Commands Manual							    SED(1)

NAME

       sed - stream editor

SYNOPSIS

       sed [ -n ] [ -e script ] [ -f sfile ] [ file ] ...

DESCRIPTION

       Sed copies the named files (standard input default) to the standard output, edited according to a script of commands.  The -f option causes
       the script to be taken from file sfile; these options accumulate.  If there is just one -e option and no -f's, the flag -e may be  omitted.
       The -n option suppresses the default output.

       A script consists of editing commands, one per line, of the following form:

	      [address [, address] ] function [arguments]

       In  normal  operation  sed  cyclically  copies  a  line of input into a pattern space (unless there is something left after a `D' command),
       applies in sequence all commands whose addresses select that pattern space, and at the end of the script copies the pattern  space  to  the
       standard output (except under -n) and deletes the pattern space.

       An  address is either a decimal number that counts input lines cumulatively across files, a `$' that addresses the last line of input, or a
       context address, `/regular expression/', in the style of ed(1) modified thus:

	      The escape sequence `
' matches a newline embedded in the pattern space.

       A command line with no addresses selects every pattern space.

       A command line with one address selects each pattern space that matches the address.

       A command line with two addresses selects the inclusive range from the first pattern space that matches the first address through the  next
       pattern	space  that matches the second.  (If the second address is a number less than or equal to the line number first selected, only one
       line is selected.)  Thereafter the process is repeated, looking again for the first address.

       Editing commands can be applied only to non-selected pattern spaces by use of the negation function `!' (below).

       In the following list of functions the maximum number of permissible addresses for each function is indicated in parentheses.

       An argument denoted text consists of one or more lines, all but the last of which end with `' to hide the newline.   Backslashes  in  text
       are  treated  like  backslashes in the replacement string of an `s' command, and may be used to protect initial blanks and tabs against the
       stripping that is done on every script line.

       An argument denoted rfile or wfile must terminate the command line and must be preceded by exactly one blank.  Each wfile is created before
       processing begins.  There can be at most 10 distinct wfile arguments.

       (1)a
       text
	      Append.  Place text on the output before reading the next input line.

       (2)b label
	      Branch to the `:' command bearing the label.  If label is empty, branch to the end of the script.

       (2)c
       text
	      Change.	Delete	the  pattern  space.  With 0 or 1 address or at the end of a 2-address range, place text on the output.  Start the
	      next cycle.

       (2)d   Delete the pattern space.  Start the next cycle.

       (2)D   Delete the initial segment of the pattern space through the first newline.  Start the next cycle.

       (2)g   Replace the contents of the pattern space by the contents of the hold space.

       (2)G   Append the contents of the hold space to the pattern space.

       (2)h   Replace the contents of the hold space by the contents of the pattern space.

       (2)H   Append the contents of the pattern space to the hold space.

       (1)i
       text
	      Insert.  Place text on the standard output.

       (2)n   Copy the pattern space to the standard output.  Replace the pattern space with the next line of input.

       (2)N   Append the next line of input to the pattern space with an embedded newline.  (The current line number changes.)

       (2)p   Print.  Copy the pattern space to the standard output.

       (2)P   Copy the initial segment of the pattern space through the first newline to the standard output.

       (1)q   Quit.  Branch to the end of the script.  Do not start a new cycle.

       (2)r rfile
	      Read the contents of rfile.  Place them on the output before reading the next input line.

       (2)s/regular expression/replacement/flags
	      Substitute the replacement string for instances of the regular expression in the pattern space.  Any character may be  used  instead
	      of `/'.  For a fuller description see ed(1).  Flags is zero or more of

	      g      Global.  Substitute for all nonoverlapping instances of the regular expression rather than just the first one.

	      p      Print the pattern space if a replacement was made.

	      w wfile
		     Write.  Append the pattern space to wfile if a replacement was made.

       (2)t label
	      Test.   Branch  to  the  `:' command bearing the label if any substitutions have been made since the most recent reading of an input
	      line or execution of a `t'.  If label is empty, branch to the end of the script.

       (2)w wfile
	      Write.  Append the pattern space to wfile.

       (2)x   Exchange the contents of the pattern and hold spaces.

       (2)y/string1/string2/
	      Transform.  Replace all occurrences of characters in string1 with the corresponding character in string2.  The  lengths  of  string1
	      and string2 must be equal.

       (2)! function
	      Don't.  Apply the function (or group, if function is `{') only to lines not selected by the address(es).

       (0): label
	      This command does nothing; it bears a label for `b' and `t' commands to branch to.

       (1)=   Place the current line number on the standard output as a line.

       (2){   Execute the following commands through a matching `}' only when the pattern space is selected.

       (0)    An empty command is ignored.

SEE ALSO

       ed(1), grep(1), awk(1), lex(1)

7th Edition							  April 29, 1985							    SED(1)

Shell Programming and Scripting