Let me say first that there's incredibly refined and sophisticated algorithms out there, used by e.g. the various search engines to analyse all the internet sites around the globe and to hand you the results in a split second, so anything posted here is a clumsy approach cobbled together without any optimisation. Anyhow, try
Code:
awk '
FNR==NR {T[$1]
next
}
{CNT=0
n=split (tolower($0), L)
for (i=1; i<=n; i++) if (L[i] in T) CNT++
# print CNT, CNT/n
if (CNT/n >= 0.8) print $0
}
' list text
You may want/need to get rid of punctuation first in a real world sample.
Hi All,
I have an input below. I tried to use the awk below but it seems that it ;s not working. Can anybody help ?
My concept here is to find the 2nd field of the last occurrence of such pattern " ** XXX ccc ccc cc cc ccc 2007 " . In this case, the 2nd field is " XXX ". With this "XXX" term... (20 Replies)
Hi ,
i have a text file that contain a story
How do i extract the out all the sentences that contain the word Mon. in C++
I only want to show those sentences that contain the word mon
eg.
Monkey on a tree.
Rabbit jumping around the tree.
I am very rich, I have lots of money.
Today... (1 Reply)
This is my first post, please be nice. I have tried to google and read different tutorials.
The task at hand is:
Input file input.txt (example)
abc123defhij-E-1234jslo
456ujs-W-abXjklp
From this file the task is to grep the -E- and -W- strings that are unique and write a new file... (5 Replies)
Hi All,
I am trying to extract data from a large text file , I want to extract lines which contains a five digit number followed by a hyphen , like
12345- , i tried with egrep ,eg : egrep "+" text.txt
but which returns all the lines which contains any number of digits followed by hyhen ,... (19 Replies)
I have an xml file with IP addresses all over the show. I want to print only the IP addresses and cut off any text before or after the IP address.
Example:
Note: The IP addresses (x.x.x.x) do not consistently appear in the xml file as per the pattern below. Sometimes there are text before... (8 Replies)
I sat down yesterday to write this script and have just realised that my methodology is broken........
In essense I have.....
----------------------------------------------------------------- (This line really is in the file)
Service ID: 12345 ... (7 Replies)
Hi
This is my first post and I'm just a beginner. So please be nice to me.
I have a couple of html files where a pattern beginning with "http://www.site.com" and ending with "/resource.dat" is present on every 241st line. How do I extract this to a new text file?
I have tried sed -n 241,241p... (13 Replies)
Hi
I have two text files. The first file is TEXTFILEONE.txt as given below:
<Text Text_ID="10155645315851111_10155645333076543" From="460350337461111" Created="2011-03-16T17:05:37+0000" use_count="123">This is the first text</Text>
<Text Text_ID="10155645315851111_10155645317023456"... (7 Replies)
Hi,
I have to extract the whole set if a pattern matches.i have a file called input.txt
input.txt
------------
CREATE TABLE ABC
(
A,
B,
C
);
CREATE TABLE XYZ
(
X,
Y,
Z,
P,
Q
); (6 Replies)
hi all,
trying this using shell/bash with sed/awk/grep
I have two files, one containing one column, the other containing multiple columns (comma delimited).
file1.txt
abc12345
def12345
ghi54321
...
file2.txt
abc1,text1,texta
abc,text2,textb
def123,text3,textc
gh,text4,textd... (6 Replies)
Discussion started by: shogun1970
6 Replies
LEARN ABOUT DEBIAN
dadadodo
dadadodo(1) General Commands Manual dadadodo(1)NAME
dadadodo - exterminate all rational thought
SYNOPSIS
dadadodo [ options ] [ input-files ]
DESCRIPTION
dadadodo is a program that analyses texts for Markov chains of word probabilities and then generates random sentences based on those proba-
bilities. Sometimes these sentences are nonsense, but sometimes they cut right through to the heart of the matter and reveal hidden mean-
ings.
OPTIONS
dadadodo accepts the following options:
-c, -count n
Generate n sentences.
-h, -help
Show summary of options and exit.
-html Output HTML instead of plain text.
-l, -load file
Load compiled data from file ('-' for standard input).
-o, -output file
Save compiled data in file ('-' for standard output).
-p, -pause s
Delay s seconds between paragraphs.
-w, -columns columns
Format output for a device columns character cells in width. If not specified, the value of the environment variable COLUMNS is
used to determine the width. If that variable is not defined, a width of 72 is assumed.
NOTES
Non-option arguments are input files. These should be text files, but may be mail folders or HTML. MIME messages are handled sensibly.
When no output file is specified, sentences will be generated from the input data directly. However, loading a saved file is far faster
than re-parsing the text files each time.
ENVIRONMENT
COLUMNS
Determines the width (in character cells) of the output if the -w, -columns option is not used. If not set, a width of 72 is
assumed.
SEE ALSO
dadadodo's upstream website is http://www.jwz.org/dadadodo/.
AUTHOR
dadadodo was written by Jamie Zawinski.
This manual page was written by Sudhakar Chandrasekharan <thaths@netscape.com>, based on the program's usage message.
dadadodo(1)