Extracting String Post: 302301560

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

extracting from a string

How do I extract 5th to 10th characters of string as given below stored in a shell variable. "ab cd ef gh ij kl" How is cut to be used on this? Thanks for any help.

2. Shell Programming and Scripting

Hi all, I have to extract only the second part of a database column (VARCHAR) and the value is seperated by a "~" xyz~ chxyz36r~ abder~000082685 mnops~000083554 fulfil302~00026 Above are some examples of the values and for each record I have to extract the value after "~" , if there is a...

3. Shell Programming and Scripting

Extracting a string from one file and searching the same string in other files

Hi, Need to extract a string from one file and search the same in other files. Ex: I have file1 of hundred lines with no delimiters not even space. I have 3 more files. I should get 1 to 10 characters say substring from each line of file1 and search that string in rest of the files and get...

4. Shell Programming and Scripting

extracting a string

Hi All, I am writing a shell script for which I am stuck with an extraction part. I arrived till extraction of a path of file. Lets take an example. I now have a file which contains following one line: 2348/home/userid/mydir/any_num_dir/myfile.text Now I want to extract only...

5. Shell Programming and Scripting

extracting a number from a string

Hi everyone, I have a string as follow ts1n65ulpa4096x16m16_130a_ss1p08v125c i would like to extract 4096 and 16 from string and save it into two variable. but this string could also have the form ts1n65ulpa32x16m16_130a_ss1p08v125c therefore the number before "x" could be 2 or 3 digits. i use...

6. Shell Programming and Scripting

extracting only numerals from string.

Hi!!! i have two files "tushar20090429200000.txt" and "tushar_err20090429200000.txt" The numeric part here is date and time. So this part of file keeps changing after every hour. I want to extract the numeric part from the both file names and compare them whether they are equal or not. ...

7. Shell Programming and Scripting

Extracting particular string in a file and storing matched string in output file

Hi , I have input file and i want to extract below strings <msisdn xmlns="">0492001956</ msisdn> => numaber inside brackets <resCode>3000</resCode> => 3000 needs to be extracted <resMessage>Request time getBalances_PSM.c(37): d out</resMessage></ns2:getBalancesResponse> => the word...

8. Shell Programming and Scripting

Help on extracting portion of string

Hi Gurus, I've some sample of my log information as shown below. -> Processing ABCD123456 This is tp version 372.04.57 (release 700, unicode enabled) This is R3trans version 6.14 (release 700 - 05.03.09 - 08:28:00). unicode enabled version R3trans finished (0000). Warning: Parameter...

9. Shell Programming and Scripting

Extracting the part of string

I have a string: 2015-04-16 07:30:05,625000 +0900 xxxx.com I just want to extract the time from the above line I am using the below syntax x=~ /(.*) (\d+)\:(\d+)\:(\d+),(.*)\.com/ $time = $2 . ':' . $3 . ':' . $4; print $time But it is not working. Can some1 please help

10. Shell Programming and Scripting

Extracting substring within string between 2 token within the string

Hello. First best wishes for everybody. here is the input file ("$INPUT1") contents : BASH_FUNC_message_begin_script%%=() { local -a L_ARRAY; BASH_FUNC_message_debug%%=() { local -a L_ARRAY; BASH_FUNC_message_end_script%%=() { local -a L_ARRAY; BASH_FUNC_message_error%%=() { local...

LEARN ABOUT DEBIAN

html::linkextractor

LinkExtractor(3pm)					User Contributed Perl Documentation					LinkExtractor(3pm)

NAME

       HTML::LinkExtractor - Extract links from an HTML document

DESCRIPTION

       HTML::LinkExtractor is used for extracting links from HTML.  It is very similar to HTML::LinkExtor, except that besides getting the URL,
       you also get the link-text.

       Example ( please run the examples ):

	   use HTML::LinkExtractor;
	   use Data::Dumper;

	   my $input = q{If <a href="http://perl.com/"> I am a LINK!!! </a>};
	   my $LX = new HTML::LinkExtractor();

	   $LX->parse($input);

	   print Dumper($LX->links);
	   __END__
	   # the above example will yield
	   $VAR1 = [
		     {
		       '_TEXT' => '<a href="http://perl.com/"> I am a LINK!!! </a>',
		       'href' => bless(do{(my $o = 'http://perl.com/')}, 'URI::http'),
		       'tag' => 'a'
		     }
		   ];

       "HTML::LinkExtractor" will also correctly extract nested link-type tags.

SYNOPSIS

	   ## the demo
	   perl LinkExtractor.pm
	   perl LinkExtractor.pm file.html othefile.html

	   ## or if the module is installed, but you don't know where

	   perl -MHTML::LinkExtractor -e" system $^X, $INC{q{HTML/LinkExtractor.pm}} "
	   perl -MHTML::LinkExtractor -e' system $^X, $INC{q{HTML/LinkExtractor.pm}} '

	   ## or

	   use HTML::LinkExtractor;
	   use LWP qw( get ); #     use LWP::Simple qw( get );

	   my $base = 'http://search.cpan.org';
	   my $html = get($base.'/recent');
	   my $LX = new HTML::LinkExtractor();

	   $LX->parse($html);

	   print qq{<base href="$base">
};

	   for my $Link( @{ $LX->links } ) {
	   ## new modules are linked  by /author/NAME/Dist
	       if( $$Link{href}=~ m{^/author/w+} ) {
		   print $$Link{_TEXT}."
";
	       }
	   }

	   undef $LX;
	   __END__

	   ## or

	   use HTML::LinkExtractor;
	   use Data::Dumper;

	   my $input = q{If <a href="http://perl.com/"> I am a LINK!!! </a>};
	   my $LX = new HTML::LinkExtractor(
	       sub {
		   print Data::Dumper::Dumper(@_);
	       },
	       'http://perlFox.org/',
	   );

	   $LX->parse($input);
	   $LX->strip(1);
	   $LX->parse($input);
	   __END__

	   #### Calculate to total size of a web-page
	   #### adds up the sizes of all the images and stylesheets and stuff

	   use strict;
	   use LWP; #	  use LWP::Simple;
	   use HTML::LinkExtractor;
							       #
	   my $url  = shift || 'http://www.google.com';
	   my $html = get($url);
	   my $Total = length $html;
							       #
	   print "initial size $Total
";
							       #
	   my $LX = new HTML::LinkExtractor(
	       sub {
		   my( $X, $tag ) = @_;
							       #
		   unless( grep {$_ eq $tag->{tag} } @HTML::LinkExtractor::TAGS_IN_NEED ) {
							       #
	   print "$$tag{tag}
";
							       #
		       for my $urlAttr ( @{$HTML::LinkExtractor::TAGS{$$tag{tag}}} ) {
			   if( exists $$tag{$urlAttr} ) {
			       my $size = (head( $$tag{$urlAttr} ))[1];
			       $Total += $size if $size;
	   print "adding $size
" if $size;
			   }
		       }
		   }
	       },
	       $url,
	       0
	   );
							       #
	   $LX->parse($html);
							       #
	   print "The total size of 
$url
 is $Total bytes
";
	   __END__

METHODS

   "$LX->new([&callback, [$baseUrl, [1]]])"
       Accepts 3 arguments, all of which are optional.	If for example you want to pass a $baseUrl, but don't want to have a callback invoked,
       just put "undef" in place of a subref.

       This is the only class method.

       1.  a callback ( a sub reference, as in "sub{}", or "&sub") which is to be called each time a new LINK is encountered ( for
	   @HTML::LinkExtractor::TAGS_IN_NEED this means
	    after the closing tag is encountered )

	   The callback receives an object reference($LX) and a link hashref.

       2.  and a base URL ( URI->new, so its up to you to make sure it's valid which is used to convert all relative URI's to absolute ones.

	       $ALinkP{href} = URI->new_abs( $ALink{href}, $base );

       3.  A "boolean" (just stick with 1).  See the example in "DESCRIPTION".	Normally, you'd get back _TEXT that looks like

	       '_TEXT' => '<a href="http://perl.com/"> I am a LINK!!! </a>',

	   If you turn this option on, you'll get the following instead

	       '_TEXT' => ' I am a LINK!!! ',

	   The private utility function "_stripHTML" does this by using HTML::TokeParsers method get_trimmed_text.

	   You can turn this feature on an off by using "$LX->strip(undef || 0 || 1)"

   "$LX->parse( $filename || *FILEHANDLE || $FileContent )"
       Each time you call "parse", you should pass it a $filename a *FILEHANDLE or a "$FileContent"

       Each time you call "parse" a new "HTML::TokeParser" object is created and stored in "$this->{_tp}".

       You shouldn't need to mess with the TokeParser object.

   "$LX->links()"
       Only after you call "parse" will this method return anything.  This method returns a reference to an ArrayOfHashes, which basically looks
       like (Data::Dumper output)

	   $VAR1 = [ { tag => 'img', src => 'image.png' }, ];

       Please note that if yo provide a callback this array will be empty.

   "$LX->strip( [ 0 || 1 ])"
       If you pass in "undef" (or nothing), returns the state of the option.  Passing in a true or false value sets the option.

       If you wanna know what the option does see "$LX->new([&callback, [$baseUrl, [1]]])"

WHAT'S A LINK-type tag
       Take a look at %HTML::LinkExtractor::TAGS to see what I consider to be link-type-tag.

       Take a look at @HTML::LinkExtractor::VALID_URL_ATTRIBUTES to see all the possible tag attributes which can contain URI's (the links!!)

       Take a look at @HTML::LinkExtractor::TAGS_IN_NEED to see the tags for which the '_TEXT' attribute is provided, like "<a href="#"> TEST
       </a>"

   How can that be?!?!
       I took at look at %HTML::Tagset::linkElements and the following URL's

	   http://www.blooberry.com/indexdot/html/tagindex/all.htm

	   http://www.blooberry.com/indexdot/html/tagpages/a/a-hyperlink.htm
	   http://www.blooberry.com/indexdot/html/tagpages/a/applet.htm
	   http://www.blooberry.com/indexdot/html/tagpages/a/area.htm

	   http://www.blooberry.com/indexdot/html/tagpages/b/base.htm
	   http://www.blooberry.com/indexdot/html/tagpages/b/bgsound.htm

	   http://www.blooberry.com/indexdot/html/tagpages/d/del.htm
	   http://www.blooberry.com/indexdot/html/tagpages/d/div.htm

	   http://www.blooberry.com/indexdot/html/tagpages/e/embed.htm
	   http://www.blooberry.com/indexdot/html/tagpages/f/frame.htm

	   http://www.blooberry.com/indexdot/html/tagpages/i/ins.htm
	   http://www.blooberry.com/indexdot/html/tagpages/i/image.htm
	   http://www.blooberry.com/indexdot/html/tagpages/i/iframe.htm
	   http://www.blooberry.com/indexdot/html/tagpages/i/ilayer.htm
	   http://www.blooberry.com/indexdot/html/tagpages/i/inputimage.htm

	   http://www.blooberry.com/indexdot/html/tagpages/l/layer.htm
	   http://www.blooberry.com/indexdot/html/tagpages/l/link.htm

	   http://www.blooberry.com/indexdot/html/tagpages/o/object.htm

	   http://www.blooberry.com/indexdot/html/tagpages/q/q.htm

	   http://www.blooberry.com/indexdot/html/tagpages/s/script.htm
	   http://www.blooberry.com/indexdot/html/tagpages/s/sound.htm

	   And the special cases

	   <!DOCTYPE HTML SYSTEM "http://www.w3.org/DTD/HTML4-strict.dtd">
	   http://www.blooberry.com/indexdot/html/tagpages/d/doctype.htm
	   '!doctype'  is really a process instruction, but is still listed
	   in %TAGS with 'url' as the attribute

	   and

	   <meta HTTP-EQUIV="Refresh" CONTENT="5; URL=http://www.foo.com/foo.html">
	   http://www.blooberry.com/indexdot/html/tagpages/m/meta.htm
	   If there is a valid url, 'url' is set as the attribute.
	   The meta tag has no 'attributes' listed in %TAGS.

SEE ALSO

       HTML::LinkExtor, HTML::TokeParser, HTML::Tagset.

AUTHOR

       D.H (PodMaster)

       Please use http://rt.cpan.org/ to report bugs.

       Just go to http://rt.cpan.org/NoAuth/Bugs.html?Dist=HTML-Scrubber to see a bug list and/or repot new ones.

LICENSE

       Copyright (c) 2003, 2004 by D.H. (PodMaster).  All rights reserved.

       This module is free software; you can redistribute it and/or modify it under the same terms as Perl itself.  The LICENSE file contains the
       full text of the license.

perl v5.10.1							    2005-01-07							LinkExtractor(3pm)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

extracting from a string

Discussion started by: preetikate

2. Shell Programming and Scripting

Extracting part of a string

Discussion started by: sam_78_nyc

3. Shell Programming and Scripting

Extracting a string from one file and searching the same string in other files

Discussion started by: mohancrr

4. Shell Programming and Scripting

extracting a string

Discussion started by: start_shell