[Solved] Removing duplicates from the file and saving as new file
Dear All
I have 200 data files and each files has many duplicates.
I am looking for the automated awk script such that it checks and removes the duplicates from the each file and saving them as new files for all 200 files in the respective folder.
For example my data looks like this..
I have to remove the line where "4427" is repeated twice and save as new file.
Kindly advice.
Many Thanks
Balaji
Last edited by Corona688; 11-22-2012 at 02:39 PM..
i have a file with some 1000 entries it will contain entries like
1000,ram
2000,pankaj
1001,rahim
1000,ram
2532,govind
2000,pankaj
3000,venkat
2532,govind
what i want is i want to extract only the distinct rows from this file
so my output should contain only
1000,ram... (2 Replies)
I have data like this:
It's sorted by the 2nd field (TID).
envoy,90000000000000634600010001,04/11/2008,23:19:27,RB00266,0015,DETAIL,ERROR,
envoy,90000000000000634600010001,04/12/2008,04:23:45,RB00266,0015,DETAIL,ERROR,... (1 Reply)
hey all,
I need some help.
I have a text file with names in it.
My target is that if a particular pattern exists in that file more than once..then i want to rename all the occurences of that pattern by alternate patterns..
for e.g if i have PATTERN occuring 5 times then i want to... (3 Replies)
I have a log file with posts looking like this:
--
Messages can be delivered by different systems at different times. The id number is used to sort out duplicate messages. What I need is to strip the arrival time from each post, sort posts by id number, and reattach arrival time to respective... (2 Replies)
Hi Experts,
Please check the following new requirement. I got data like the following in a file.
FILE_HEADER
01cbbfde7898410| 3477945| home| 1
01cbc275d2c122| 3478234| WORK| 1
01cbbe4362743da| 3496386| Rich Spare| 1
01cbc275d2c122| 3478234| WORK| 1
This is pipe separated file with... (3 Replies)
Hi,
I have a file that I want to change the format of. It is a large file in rows but I want it to be comma separated (comma then a space).
The current file looks like this:
HI, Joe, Bob, Jack, Jack
After I would want to remove any duplicates so it would look like this:
HI, Joe,... (2 Replies)
Hi All,
I am merging files coming from 2 different systems ,while doing that I am getting duplicates entries in the merged file
I,01,000131,764,2,4.00
I,01,000131,765,2,4.00
I,01,000131,772,2,4.00
I,01,000131,773,2,4.00
I,01,000168,762,2,2.00
I,01,000168,763,2,2.00... (5 Replies)
I have been using grep to output whole lines using a pattern file with identifiers (fileA):
fig|562.2322.peg.1
fig|562.2322.peg.3
fig|562.2322.peg.3
fig|562.2322.peg.3
fig|562.2322.peg.7
From fileB with corresponding identifiers in the second column:
NODE_0 fig|562.2322.peg.1 peg ... (2 Replies)
i hav two files like
i want to remove/delete all the duplicate lines in file2 which are viz unix,unix2,unix3.I have tried previous post also,but in that complete line must be similar.In this case i have to verify first column only regardless what is the content in succeeding columns. (3 Replies)
Discussion started by: sagar_1986
3 Replies
LEARN ABOUT DEBIAN
apertium-tagger
apertium-tagger(1)apertium-tagger(1)NAME
apertium-tagger - This application is part of ( apertium )
This tool is part of the apertium open-source machine translation architecture: http://www.apertium.org.
SYNOPSIS
apertium-tagger --train|-t {n} DIC CRP TSX PROB [--debug|-d]
apertium-tagger --supervised|-s {n} DIC CRP TSX PROB HTAG UNTAG [--debug|-d]
apertium-tagger --retrain|-r {n} CRP PROB [--debug|-d]
apertium-tagger --tagger|-g [--first|-f] PROB [--debug|-d] [INPUT [OUTPUT]]
DESCRIPTION
apertium-tagger is the application responsible for the apertium part-of-speech tagger training or tagging, depending on the calling
options. This command only reads from the standard input if the option --tagger or -g is used.
OPTIONS -t {n}, --train {n}
Initializes parameters through the Kupiec's method (unsupervised), then performs n iterations of the Baum-Welch training algorithm
(unsupervised).
-s {n}, --supervised {n}
Initializes parameters against a hand-tagged text (supervised) through the maximum likelihood estimate method, then performs n iter-
ations of the Baum-Welch training algorithm (unsupervised)
-r {n}, --retrain {n}
Retrains the model with n additional Baum-Welch iterations (unsupervised).
-g, --tagger
Tags input text by means of Viterbi algorithm.
-p, --show-superficial
Prints the superficial form of the word along side the lexical form in the output stream.
-f, --first
Used if conjuntion with -g (--tagger) makes the tagger to give all lexical forms of each word, being the choosen one in the first
place (after the lemma)
-d, --debug
Print error (if any) or debug messages while operating.
-m, --mark
Mark disambiguated words.
-h, --help
Display a help message.
FILES
These are the kinds of files used with each option:
DIC Full expanded dictionary file
CRP Training text corpus file
TSX Tagger specification file, in XML format
PROB Tagger data file, built in the training and used while tagging
HTAG Hand-tagged text corpus
UNTAG Untagged text corpus, morphological analysis of HTAG corpus to use both jointly with -s option
INPUT Input file, stdin by default
OUTPUT Output file, stdout by default
SEE ALSO lt-proc(1), lt-comp(1), lt-expand(1), apertium-translator(1), apertium(1).
BUGS
Lots of...lurking in the dark and waiting for you!
AUTHOR
Copyright (c) 2005, 2006 Universitat d'Alacant / Universidad de Alicante. This is free software. You may redistribute copies of it under
the terms of the GNU General Public License <http://www.gnu.org/licenses/gpl.html>.
2006-08-30 apertium-tagger(1)