Extract all the sentences from a text file that matches a pattern list

06-25-2015

Registered User

109, 1

Join Date: Jun 2009

Last Activity: 23 February 2016, 3:52 PM EST

Location: India

Posts: 109

Thanks Given: 40

Thanked 1 Time in 1 Post

Extract all the sentences from a text file that matches a pattern list

Hi

I have a big text file. I want to extract all the sentences that matches at least 70% (seventy percent) of the words from each sentence based on a word list called A.

Say the format of the text file is as given below:

Code:

This is the first sentence which consists of fifteen words including AAA, BBB and CCC.
This is the second sentence with twelve consisting of XXX and YYY.
This is the third with nine consisting of KKK.
The last sentence consist of  ZZZ, DDD, FFF, EEE, GGG and HHH.

The output format is based on the availability of all the words from the words list A. If the sentence matches all the words from the word list, then the sentence will be extracted. The condition is that at least seventy percent of the words in the sentence must be matched from the words list A.

Assuming that all the capital letter words such as AAA, BBB, CCC, KKK, XXX, YYY, ZZZ, DDD, FFF, EEE, GGG and HHH are not found in the words list A. After the extraction, the output will look like as given below.

Code:

This is the first sentence which consists of fifteen words including AAA, BBB and CCC.
This is the second sentence with twelve consisting of XXX and YYY.
This is the third with nine consisting of KKK.

I need help to write a script for the above problem. A sample script will be really helpful to me. Thanks in advance.

Last edited by Don Cragun; 06-25-2015 at 03:08 AM.. Reason: Change HTML tags to CODE tags.

my_Perl

View Public Profile for my_Perl

Find all posts by my_Perl

06-25-2015

Registered User

12,315, 4,560

Join Date: Jul 2012

Last Activity: 22 November 2019, 4:29 PM EST

Location: San Jose, CA, USA

Posts: 12,315

Thanks Given: 952

Thanked 4,560 Times in 3,818 Posts

Is this a homework assignment? (If it is not homework, please explain why you need to do this!)

What have you tried to solve this problem?

How do you define "word"?

How do you define "sentence"?

Instead of having us assume that a class of words are not in your list, provide us with an actual sample of your "words list A"!

Is your real "words list A" case sensitive?

What operating system are you using?

What shell are you using? (Or, more importantly, what shells are you willing to use to solve this problem?)

This User Gave Thanks to Don Cragun For This Post:

Don Cragun

View Public Profile for Don Cragun

Find all posts by Don Cragun

06-25-2015

Registered User

109, 1

Join Date: Jun 2009

Last Activity: 23 February 2016, 3:52 PM EST

Location: India

Posts: 109

Thanks Given: 40

Thanked 1 Time in 1 Post

Sorry, I have to explain it again.

This is not an assignment at all. This is for my personal interest for processing text with different coding.

I have tried this using C code instead of scripts but I would prefer script because they are comparatively faster and convenient to put in the pipeline. Moreover, unix scripts are convenient for processing text. Bash shell is the one I use.

There is no actual sample of word list or token list.

To clarify, I would say tokens instead of words and sentence as a set of tokens. I use Ubuntu OS.

Please let me know if you need more details.
Thanks in advance.

my_Perl

View Public Profile for my_Perl

Find all posts by my_Perl

06-25-2015

Registered User

15,129, 5,008

Join Date: Jul 2012

Last Activity: 4 May 2020, 4:31 PM EDT

Location: Aachen, Germany

Posts: 15,129

Thanks Given: 735

Thanked 5,008 Times in 4,483 Posts

Let me say first that there's incredibly refined and sophisticated algorithms out there, used by e.g. the various search engines to analyse all the internet sites around the globe and to hand you the results in a split second, so anything posted here is a clumsy approach cobbled together without any optimisation. Anyhow, try

Code:

awk '
FNR==NR         {T[$1]
                 next
                }
                {CNT=0
                 n=split (tolower($0), L)
                 for (i=1; i<=n; i++) if (L[i] in T) CNT++
#                print CNT, CNT/n
                 if (CNT/n >= 0.8) print $0
                }
' list text

You may want/need to get rid of punctuation first in a real world sample.

This User Gave Thanks to RudiC For This Post:

RudiC

View Public Profile for RudiC

Find all posts by RudiC

06-25-2015

Registered User

12,315, 4,560

Join Date: Jul 2012

Last Activity: 22 November 2019, 4:29 PM EST

Location: San Jose, CA, USA

Posts: 12,315

Thanks Given: 952

Thanked 4,560 Times in 3,818 Posts

Changing "word" to "token" and "sentence" to "set" doesn't clarify anything. Changing two undefined terms to two other undefined terms still leaves us with no defined terms. If you refuse to explain what you want your code to do, there is no reason for any of us to waste our time trying to guess at your requirements, nor to try to write code when we don't know what the code is supposed to do. Why did you explicitly say that you had a "word list A" if there is no word list? Your original requirement was:

Quote:

which can now be restated as:

Quote:

The output format is based on the availability of all of the <undefined term1>s from the <undefined term1>s in a non-existent list. If the <undefined term2> matches all of the <undefined term1>s from the non-existent list, then the <undefined term2> will be extracted. The condition is that at least 70% of the <undefined term1>s in the <undefined term2> must be matched from the <undefined term1>s in the non-existent list.

With requirements like this, it looks like a homework assignment that you want us to complete for you.

If you already have a way to do this and just want to write it in a different language, show us the C code that you have written that you now want to translate to shell code. They we would be able to deduce your definitions from you C code and know what it is that we're trying to do. (But, don't claim that you are converting from C to shell to make the code faster; for any particular task, well crafted C will almost certainly be faster than a corresponding shell script. And, there is absolutely no reason to claim that C code can't be used in a pipeline. Almost all of the standard utilities on UNIX and Linux systems are written in C and many of them are perfectly capable of being used in a pipeline. Changing C code that can't be used as a filter to a shell script won't magically turn it into a filter.)

RudiC made a valiant effort to help you get a start on your problem, but it ignores the fact that you don't have a list, assumes that tokens (or words) include punctuation, assumes that <sentence> or <set> and <line in a text file> are synonymous, ignores the requirement to ignore uppercase <words> or <tokens> from you nonexistent list, and uses 80% instead of 70% as the threshold.

Last edited by Don Cragun; 06-25-2015 at 04:42 AM.. Reason: Fix typo and note difference in percentages.

This User Gave Thanks to Don Cragun For This Post:

Don Cragun

View Public Profile for Don Cragun

Find all posts by Don Cragun

Shell Programming and Scripting

Extract all the sentences from a text file that matches a pattern list

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Match text to lines in a file, iterate backwards until text or text substring matches, print to file

Discussion started by: shogun1970

2. UNIX for Beginners Questions & Answers

Extract the whole set if a pattern matches

Discussion started by: raju2016

3. Shell Programming and Scripting

Extract sentence and its details from a text file based on another file of sentences

Discussion started by: my_Perl

4. Shell Programming and Scripting

Extract specific line in an html file starting and ending with specific pattern to a text file

Discussion started by: dejavo

5. Shell Programming and Scripting

Blocks of text in a file - extract when matches...

Discussion started by: Bashingaway

6. Shell Programming and Scripting

Extract list of IP addresses from a text file.

Discussion started by: lewk

7. Shell Programming and Scripting

extract unique pattern from large text file

Discussion started by: shijujoe

8. Shell Programming and Scripting

sed: Find start of pattern and extract text to end of line, including the pattern

Discussion started by: TestTomas

9. Programming

How to extract a sentences of word from a text file.

Discussion started by: xiaojesus

10. Shell Programming and Scripting

Extract if pattern matches

Discussion started by: Raynon