Linguistic project: extract co-occurrences from text corpus Post: 302661089

Sponsored Content

Top Forums Shell Programming and Scripting Linguistic project: extract co-occurrences from text corpus Post 302661089 by bobylapointe on Sunday 24th of June 2012 06:25:55 PM

06-24-2012

Registered User

I'm sorry I wasn't really clear in my first post. In more concrete words, I'm trying to see with what word the word of my choice is most commonly associated with - on its left, that is to say: word wordofmychoice - within a corpus.

The textfile looks like this:

Are you one of those people who prefer larger dogs? Do you know someone who has told you that they prefer larger dogs because small dogs are yappy and snappy? Whether you are a large-dog person or a small-dog person, one thing we all would agree on is that a larger percentage of small dogs tend to have a different type of temperament than medium and large dogs. Small dogs have earned the reputation of being yappy, snappy, jealous, protective, wary of strangers and not the greatest child companions.

Let's say I'm interested in the word "dogs". The output would be:
larger dogs
small dogs
large dogs

But I want to count how many times each association appears:
larger dogs 2
small dogs 3
large dogs 1

And, I only want to keep (print in a new file) associations appearing at least 3 times. Therefore, the final result (in a new textfile) I want to obtain would be:
small dogs 3

That's it basically. If possible, now, but this is not a priority, I would like to make sure no association contain any punctuation in the middle, to avoid getting what I would call false results. For instance, let's say I'm looking for "small" and its associations (with one word on the left) in the previous text:

"dogs. Small"

This is what I want to avoid. But once again, that's not a priority.

Thanks for your answers guys, I hope it was a bit clearer

bobylapointe

View Public Profile for bobylapointe

Find all posts by bobylapointe

6 More Discussions You Might Find Interesting

1. Programming

c program to extract text between two delimiters from some text file

needa c program to extract text between two delimiters from some text file. and then storing them in to diffrent variables ? text file like 0: abc.txt ========= aaaaaa|11111111|sssssssssss|333333|ddddddddd|34343454564|asass aaaaaa|11111111|sssssssssss|333333|ddddddddd|34343454564|asass...

2. Shell Programming and Scripting

Text Substitution Project

History: large open source PHP project, school management program. Comprises about 200 scripts. Had another developer for awhile, and he wanted a version in German, so he edited all the scripts and replaced text that would show up in the browser with variables (i.e. instead of "Click Here",...

3. Shell Programming and Scripting

Creating Frequency of words from a file by accessing a corpus

Hello, I have a large file of syllables /strings in Urdu. Each word is on a separate line. Example in English: be at for if being attract I need to identify the frequency of each of these strings from a large corpus (which I cannot attach unfortunately because of size limitations) and...

4. Shell Programming and Scripting

Grepping verbal forms from a large corpus

I want to extract verbal forms from a large corpus of English. I have identified a certain number of patterns. Each pattern has the following structure SPACE word_CATEGORY where word refers to the verbal form and CATEGORY refers to the class of the verb The categories are identified as per the...

5. Shell Programming and Scripting

Remove duplicate occurrences of text pattern

Hi folks! I have a file which contains a 1000 lines. On each line i have multiple occurrences ( 26 to be exact ) of pattern folder#/folder#. # is depicting the line number in the file some text here folder1/folder1 some text here folder1/folder1 some text here folder1/folder1 some text...

6. Shell Programming and Scripting

Alignment tool to join text files in 2 directories to create a parallel corpus

I have two directories called English and Hindi. Each directory contains the same number of files with the only difference being that in the case of the English Directory the tag is .english and in the Hindi one the tag is .Hindi The file may contain either a single text or more than one text...

LEARN ABOUT DEBIAN

apertium-gen-stopwords-lextor

apertium-gen-stopwords-lextor(1)										  apertium-gen-stopwords-lextor(1)

NAME

       apertium-gen-stopwords-lextor - This application is part of ( apertium )

       This tool is part of the apertium machine translation architecture: http://apertium.org.

SYNOPSIS

       apertium-gen-stopwords-lextor n input_file output_file

DESCRIPTION

       apertium-gen-stopwords-lextor  is  the  application  responsible  for generating the list of stopwords used by the lexical selection module
       (apertium-lextor). Stopwords are ignored as they cannot have multiple translations.

OPTIONS

       n the desired number of stopwords.

FILES

       These are the kinds of parameters and files used with this tool:

       input_file contains a large preprocessed corpus (see apertium-preprocess-corpus-lextor).

       output_file The file which gets the generated stopwords.

SEE ALSO

       apertium-gen-lextorbil(1),   apertium-gen-lextormono(1),    apertium-preprocess-corpus-lextor(1),    apertium-gen-wlist-lextor(1),    aper-
       tium-gen-wlist-lextor-translation(1), apertium-lextor-eval(1), apertium-lextor(1).

BUGS

       Lots of...lurking in the dark and waiting for you!

AUTHOR

       (c) 2005,2006 Universitat d'Alacant / Universidad de Alicante. All rights reserved.

								    2006-12-12					  apertium-gen-stopwords-lextor(1)

6 More Discussions You Might Find Interesting

1. Programming

c program to extract text between two delimiters from some text file

Discussion started by: kukretiabhi13

2. Shell Programming and Scripting

Text Substitution Project

Discussion started by: dougp23

3. Shell Programming and Scripting

Creating Frequency of words from a file by accessing a corpus

Discussion started by: gimley

4. Shell Programming and Scripting

Grepping verbal forms from a large corpus

Discussion started by: gimley

5. Shell Programming and Scripting

Remove duplicate occurrences of text pattern

Discussion started by: martinsmith

6. Shell Programming and Scripting

Alignment tool to join text files in 2 directories to create a parallel corpus

Discussion started by: gimley

LEARN ABOUT DEBIAN

apertium-gen-stopwords-lextor