Extract all the sentences from a text file that matches a pattern list Post: 302948030

Sponsored Content

Top Forums Shell Programming and Scripting Extract all the sentences from a text file that matches a pattern list Post 302948030 by my_Perl on Wednesday 24th of June 2015 11:59:00 PM

06-25-2015

Registered User

Extract all the sentences from a text file that matches a pattern list

Hi

I have a big text file. I want to extract all the sentences that matches at least 70% (seventy percent) of the words from each sentence based on a word list called A.

Say the format of the text file is as given below:

Code:

This is the first sentence which consists of fifteen words including AAA, BBB and CCC.
This is the second sentence with twelve consisting of XXX and YYY.
This is the third with nine consisting of KKK.
The last sentence consist of  ZZZ, DDD, FFF, EEE, GGG and HHH.

The output format is based on the availability of all the words from the words list A. If the sentence matches all the words from the word list, then the sentence will be extracted. The condition is that at least seventy percent of the words in the sentence must be matched from the words list A.

Assuming that all the capital letter words such as AAA, BBB, CCC, KKK, XXX, YYY, ZZZ, DDD, FFF, EEE, GGG and HHH are not found in the words list A. After the extraction, the output will look like as given below.

Code:

This is the first sentence which consists of fifteen words including AAA, BBB and CCC.
This is the second sentence with twelve consisting of XXX and YYY.
This is the third with nine consisting of KKK.

I need help to write a script for the above problem. A sample script will be really helpful to me. Thanks in advance. Smilie

Last edited by Don Cragun; 06-25-2015 at 03:08 AM.. Reason: Change HTML tags to CODE tags.

my_Perl

View Public Profile for my_Perl

Find all posts by my_Perl

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Extract if pattern matches

Hi All, I have an input below. I tried to use the awk below but it seems that it ;s not working. Can anybody help ? My concept here is to find the 2nd field of the last occurrence of such pattern " ** XXX ccc ccc cc cc ccc 2007 " . In this case, the 2nd field is " XXX ". With this "XXX" term...

2. Programming

How to extract a sentences of word from a text file.

Hi , i have a text file that contain a story How do i extract the out all the sentences that contain the word Mon. in C++ I only want to show those sentences that contain the word mon eg. Monkey on a tree. Rabbit jumping around the tree. I am very rich, I have lots of money. Today...

3. Shell Programming and Scripting

sed: Find start of pattern and extract text to end of line, including the pattern

This is my first post, please be nice. I have tried to google and read different tutorials. The task at hand is: Input file input.txt (example) abc123defhij-E-1234jslo 456ujs-W-abXjklp From this file the task is to grep the -E- and -W- strings that are unique and write a new file...

4. Shell Programming and Scripting

extract unique pattern from large text file

Hi All, I am trying to extract data from a large text file , I want to extract lines which contains a five digit number followed by a hyphen , like 12345- , i tried with egrep ,eg : egrep "+" text.txt but which returns all the lines which contains any number of digits followed by hyhen ,...

5. Shell Programming and Scripting

Extract list of IP addresses from a text file.

I have an xml file with IP addresses all over the show. I want to print only the IP addresses and cut off any text before or after the IP address. Example: Note: The IP addresses (x.x.x.x) do not consistently appear in the xml file as per the pattern below. Sometimes there are text before...

6. Shell Programming and Scripting

Blocks of text in a file - extract when matches...

I sat down yesterday to write this script and have just realised that my methodology is broken........ In essense I have..... ----------------------------------------------------------------- (This line really is in the file) Service ID: 12345 ...

7. Shell Programming and Scripting

Extract specific line in an html file starting and ending with specific pattern to a text file

Hi This is my first post and I'm just a beginner. So please be nice to me. I have a couple of html files where a pattern beginning with "http://www.site.com" and ending with "/resource.dat" is present on every 241st line. How do I extract this to a new text file? I have tried sed -n 241,241p...

8. Shell Programming and Scripting

Extract sentence and its details from a text file based on another file of sentences

Hi I have two text files. The first file is TEXTFILEONE.txt as given below: <Text Text_ID="10155645315851111_10155645333076543" From="460350337461111" Created="2011-03-16T17:05:37+0000" use_count="123">This is the first text</Text> <Text Text_ID="10155645315851111_10155645317023456"...

9. UNIX for Beginners Questions & Answers

Extract the whole set if a pattern matches

Hi, I have to extract the whole set if a pattern matches.i have a file called input.txt input.txt ------------ CREATE TABLE ABC ( A, B, C ); CREATE TABLE XYZ ( X, Y, Z, P, Q );

10. Shell Programming and Scripting

Match text to lines in a file, iterate backwards until text or text substring matches, print to file

hi all, trying this using shell/bash with sed/awk/grep I have two files, one containing one column, the other containing multiple columns (comma delimited). file1.txt abc12345 def12345 ghi54321 ... file2.txt abc1,text1,texta abc,text2,textb def123,text3,textc gh,text4,textd...

LEARN ABOUT DEBIAN

mbt

mbt(1)							      General Commands Manual							    mbt(1)

NAME

       MBT - Memory Based Tagger

SYNOPSYS

       mbt [options]

DESCRIPTION

       mbt is a memory-based tagger that can tag sequences, based on training files generated by mbtg.

OPTIONS

       -h or --help
	      show help

       -s settingsfile
	      use a settingsfile as generated by mbtg

       Or:

       -l <lexiconfile>

       -r <ambitagfile>

       -k <known words case base>

       -u <unknown words case base>

       -D <loglevel>
	      Possible options levels are LogNormal , LogDebug , LogHeavy and LogExtreme

       -e <sentence delimiter> (default '<utt>')

       -E <enriched tagged testfile>

       -t <testfile>

       -T <tagged testfile> (default is untagged stdin)

       -o <outputfile> (default stdout)

       -Otimbl options
	       (Note: there is NO SPACE between O and the options)
		<options>   classifier options for both known and unknown words instance bases
		K: <options>   classifier options for known words instance base
		U: <options>   classifier options for unknown words instance base
		valid timbl options are: a d k m q v w x -

       -B <beamsize for search> (default = 1)

       -v di
	       add distance to output

       -v db
	       add distribution to output

       -v c
	       add confidence to output

       -V or --version
	      show version info.

       -L <file with list of frequent words>

BUGS

       possibly

AUTHORS

       Ko van der Sloot Timbl@uvt.nl

       Antal van den Bosch Timbl@uvt.nl

SEE ALSO

       timbl(1) mbtg(1) mbtserver(1)

								   2011 march 21							    mbt(1)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Extract if pattern matches

Discussion started by: Raynon

2. Programming

How to extract a sentences of word from a text file.

Discussion started by: xiaojesus

3. Shell Programming and Scripting

sed: Find start of pattern and extract text to end of line, including the pattern

Discussion started by: TestTomas

4. Shell Programming and Scripting

extract unique pattern from large text file

Discussion started by: shijujoe

5. Shell Programming and Scripting

Extract list of IP addresses from a text file.

Discussion started by: lewk

6. Shell Programming and Scripting

Blocks of text in a file - extract when matches...

Discussion started by: Bashingaway

7. Shell Programming and Scripting

Extract specific line in an html file starting and ending with specific pattern to a text file

Discussion started by: dejavo