Extremely Fast Text Feature Extraction for Classification and Indexing

Tags

Special Forums News, Links, Events and Announcements UNIX and Linux RSS News Extremely Fast Text Feature Extraction for Classification and Indexing

08-22-2008

Registered User

26,240, 27

Join Date: Sep 2000

Last Activity: 1 August 2008, 3:09 PM EDT

Posts: 26,240

Thanks Given: 0

Thanked 27 Times in 26 Posts

Extremely Fast Text Feature Extraction for Classification and Indexing

HPL-2008-91R1 Extremely Fast Text Feature Extraction for Classification and Indexing - Forman, George; Kirshenbaum, Evan
Keyword(s): text mining, text indexing, bag-of-words, feature engineering, feature extraction, document categorization, text tokenization
Abstract: Most research in speeding up text mining involves algorithmic improvements to induction algorithms, and yet for many large scale applications, such as classifying or indexing large document repositories, the time spent extracting word features from texts can itself greatly exceed the initial trainin ...
Full Report

More...

Linux Bot

View Public Profile for Linux Bot

Find all posts by Linux Bot

Previous Thread | Next Thread

6 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Text extraction

Dear All, I am trying to extract text from a file containing cron entries. cat /var/tmp/cron_backups/debmed_tmp < * * * * * /bell > * * * * * /belly what I am trying to do is create two text files containing all entries that begin with < and another text files containing entries with > ....

2. Shell Programming and Scripting

sed text extraction between 2 patterns using variables

Hi everyone! I'm writting a function in .bashrc to extract some text from a file. The file looks like this: " random text Begin CG step 1 random text Begin CG step 2 ... Begin CG step 100 random text" For a given number, let's say 70, I want all the text between "Begin CG...

3. UNIX for Dummies Questions & Answers

fast sequence extraction

Hi everyone, I have a large text file containing DNA sequences in fasta format as follows: >someseq GAACTTGAGATCCGGGGAGCAGTGGATCTC CACCAGCGGCCAGAACTGGTGCACCTCCAG GCCAGCCTCGTCCTGCGTGTC >another seq GGCATTTTTGTGTAATTTTTGGCTGGATGAGGT GACATTTTCATTACTACCATTTTGGAGTACA >seq3450...

4. Programming

Fast string removal from large text collection

Hi All, I don't want any codes for this problem. Just suggestions: I have a huge collection of text files (around 300,000) which look like this: 1.fil orange apple dskjdsk computer skjks The entire text collection (referenced above) has about 1 billion words. I have created...

5. UNIX for Dummies Questions & Answers

String extraction from a text file

The following script code works great for extracting 'postmaster' from a line of text stored in a variable named string: string="PenaltyError:=554 5.7.1 Error, send your mail to postmaster@LOCALDOMAIN" stuff=$( echo $string | cut -d@ -f1 | awk '{ print $NF }' ) echo $stuff However, I need to be...

6. Shell Programming and Scripting

extraction of perfect text from file.

Hi All, I have a file of the following format. <?xml version='1.0' encoding='utf-8'?> <tomcat-users> <role rolename="tomcat"/> <role rolename="role1"/> <role rolename="manager"/> <role rolename="admin"/> <user username="tomcat" password="tomcat" roles="tomcat"/> <user...

LEARN ABOUT DEBIAN

xml::dom::text

XML::DOM::Text(3pm)					User Contributed Perl Documentation				       XML::DOM::Text(3pm)

NAME

       XML::DOM::Text - A piece of XML text in XML::DOM

DESCRIPTION

       XML::DOM::Text extends XML::DOM::CharacterData, which extends XML::DOM::Node.

       The Text interface represents the textual content (termed character data in XML) of an Element or Attr. If there is no markup inside an
       element's content, the text is contained in a single object implementing the Text interface that is the only child of the element.  If
       there is markup, it is parsed into a list of elements and Text nodes that form the list of children of the element.

       When a document is first made available via the DOM, there is only one Text node for each block of text. Users may create adjacent Text
       nodes that represent the contents of a given element without any intervening markup, but should be aware that there is no way to represent
       the separations between these nodes in XML or HTML, so they will not (in general) persist between DOM editing sessions. The normalize()
       method on Element merges any such adjacent Text objects into a single node for each block of text; this is recommended before employing
       operations that depend on a particular document structure, such as navigation with XPointers.

       METHODS

       splitText (offset)
	   Breaks this Text node into two Text nodes at the specified offset, keeping both in the tree as siblings. This node then only contains
	   all the content up to the offset point. And a new Text node, which is inserted as the next sibling of this node, contains all the con-
	   tent at and after the offset point.

	   Parameters:
	    offset  The offset at which to split, starting from 0.

	   Return Value: The new Text node.

	   DOMExceptions:

	   * INDEX_SIZE_ERR
	       Raised if the specified offset is negative or greater than the number of characters in data.

	   * NO_MODIFICATION_ALLOWED_ERR
	       Raised if this node is readonly.

perl v5.8.8							    2008-02-03						       XML::DOM::Text(3pm)

UNIX and Linux RSS News