Phrase XML with Huge Data Post: 302967738

9 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

How to extract data from a huge file?

Hi, I have a huge file of bibliographic records in some standard format.I need a script to do some repeatable task as follows: 1. Needs to create folders as the strings starts with "item_*" from the input file 2. Create a file "contents" in each folders having "license.txt(tab...

2. Shell Programming and Scripting

Splitting huge XML Files into fixsized wellformed parts

Hi, I need to split xml-files with sizes greater than 2 gb into smaler chunks. As I dont want to end up with billions of files, I want those splitted files to have configurable sizes like 250 MB. Each file should be well formed having an exact copy of the header (and footer as the closing of the...

3. Shell Programming and Scripting

splitting huge xml into multiple files

hi all i have a some huge html files (500MB to 1GB). Each file has multiple <html></html> tags <html> ................. .................... .................... </html> <html> ................. .................... .................... </html> <html> ....................

4. Shell Programming and Scripting

Split a huge data into few different files?!

Input file data contents: >seq_1 MSNQSPPQSQRPGHSHSHSHSHAGLASSTSSHSNPSANASYNLNGPRTGGDQRYRASVDA >seq_2 AGAAGRGWGRDVTAAASPNPRNGGGRPASDLLSVGNAGGQASFASPETIDRWFEDLQHYE >seq_3 ATLEEMAAASLDANFKEELSAIEQWFRVLSEAERTAALYSLLQSSTQVQMRFFVTVLQQM ARADPITALLSPANPGQASMEAQMDAKLAAMGLKSPASPAVRQYARQSLSGDTYLSPHSA...

5. Shell Programming and Scripting

convert huge .xml file in .csv with specific column.

I have huge xml file in server and i want to convert it to .csv with specific column ... i have search in blog but i didn't get any usefully command. Thanks in advance

6. Shell Programming and Scripting

How to find a phrase and pull all lines that follow until the phrase occurs again?

I want to burst a report by using the page number value in the report header. Each section starts with *PAGE NO:* 1 Each section might have several pages, but the next section always starts back at 1. So I want to find the "*PAGE NO:* 1" value and pull all lines that follow until "*PAGE NO:* 1"...

7. Shell Programming and Scripting

Aggregation of huge data

Hi Friends, I have a file with sample amount data as follows: -89990.3456 8788798.990000128 55109787.20 -12455558989.90876 I need to exclude the '-' symbol in order to treat all values as an absolute one and then I need to sum up.The record count is around 1 million. How...

8. Solaris

The Fastest for copy huge data

Dear Experts, I would like to know what's the best method for copy data around 3 mio (spread in a hundred folders, size each file around 1kb) between 2 servers? I already tried using Rsync and tar command. But using these command is too long. Please advice. Thanks Edy

9. UNIX and Linux Applications

How to delete a data starting with a phrase in a table - SQL?

Hello, I am trying to remove some rows in a table, which are including a phrase at a defined column but i could not find the unique result for this. What I wish to do is to remove all lines including http://xx.yy at link column ...

LEARN ABOUT DEBIAN

xml::sax::byrecord

XML::SAX::ByRecord(3pm) 				User Contributed Perl Documentation				   XML::SAX::ByRecord(3pm)

NAME

       XML::SAX::ByRecord - Record oriented processing of (data) documents

SYNOPSIS

	   use XML::SAX::Machines qw( ByRecord ) ;

	   my $m = ByRecord(
	       "My::RecordFilter1",
	       "My::RecordFilter2",
	       ...
	       {
		   Handler => $h, ## optional
	       }
	   );

	   $m->parse_uri( "foo.xml" );

DESCRIPTION

       XML::SAX::ByRecord is a SAX machine that treats a document as a series of records.  Everything before and after the records is emitted as-
       is while the records are excerpted in to little mini-documents and run one at a time through the filter pipeline contained in ByRecord.

       The output is a document that has the same exact things before, after, and between the records that the input document did, but which has
       run each record through a filter.  So if a document has 10 records in it, the per-record filter pipeline will see 10 sets of (
       start_document, body of record, end_document ) events.  An example is below.

       This has several use cases:

       o   Big, record oriented documents

	   Big documents can be treated a record at a time with various DOM oriented processors like XML::Filter::XSLT.

       o   Streaming XML

	   Small sections of an XML stream can be run through a document processor without holding up the stream.

       o   Record oriented style sheets / processors

	   Sometimes it's just plain easier to write a style sheet or SAX filter that applies to a single record at at time, rather than having to
	   run through a series of records.

   Topology
       Here's how the innards look:

	  +-----------------------------------------------------------+
	  |		     An XML:SAX::ByRecord		      |
	  |    Intake						      |
	  |   +----------+    +---------+	  +--------+  Exhaust |
	--+-->| Splitter |--->| Stage_1 |-->...-->| Merger |----------+----->
	  |   +----------+    +---------+	  +--------+	      |
	  |		  			       ^	      |
	  |		   			       |	      |
	  |		    +---------->---------------+	      |
	  |		      Events not in any records 	      |
	  |							      |
	  +-----------------------------------------------------------+

       The "Splitter" is an XML::Filter::DocSplitter by default, and the "Merger" is an XML::Filter::Merger by default.  The line that bypasses
       the "Stage_1 ..." filter pipeline is used for all events that do not occur in a record.	All events that occur in a record pass through the
       filter pipeline.

   Example
       Here's a quick little filter to uppercase text content:

	   package My::Filter::Uc;

	   use vars qw( @ISA );
	   @ISA = qw( XML::SAX::Base );

	   use XML::SAX::Base;

	   sub characters {
	       my $self = shift;
	       my ( $data ) = @_;
	       $data->{Data} = uc $data->{Data};
	       $self->SUPER::characters( @_ );
	   }

       And here's a little machine that uses it:

	   $m = Pipeline(
	       ByRecord( "My::Filter::Uc" ),
	       $out,
	   );

       When fed a document like:

	   <root> a
	       <rec>b</rec> c
	       <rec>d</rec> e
	       <rec>f</rec> g
	   </root>

       the output looks like:

	   <root> a
	       <rec>B</rec> c
	       <rec>C</rec> e
	       <rec>D</rec> g
	   </root>

       and the My::Filter::Uc got three sets of events like:

	   start_document
	   start_element: <rec>
	   characters:	  'b'
	   end_element:   </rec>
	   end_document

	   start_document
	   start_element: <rec>
	   characters:	  'd'
	   end_element:   </rec>
	   end_document

	   start_document
	   start_element: <rec>
	   characters:	 'f'
	   end_element:   </rec>
	   end_document

METHODS

       new
	       my $d = XML::SAX::ByRecord->new( @channels, \%options );

	   Longhand for calling the ByRecord function exported by XML::SAX::Machines.

CREDIT

       Proposed by Matt Sergeant, with advise by Kip Hampton and Robin Berjon.

Writing an aggregator.
       To be written.  Pretty much just that "start_manifold_processing" and "end_manifold_processing" need to be provided.  See
       XML::Filter::Merger and it's source code for a starter.

perl v5.10.0							    2009-06-11						   XML::SAX::ByRecord(3pm)