extracting Line between HTML tag Post: 302603719

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

How do I extract text only from html file without HTML tag

I have a html file called myfile. If I simply put "cat myfile.html" in UNIX, it shows all the html tags like <a href=r/26><img src="http://www>. But I want to extract only text part. Same problem happens in "type" command in MS-DOS. I know you can do it by opening it in Internet Explorer,...

2. Shell Programming and Scripting

how to use html tag in shell scripting

Hai friends I have a small doubt.. how can we use html tag in shell scripting code : echo "<html>" echo "<body>" echo " welcome to peace world " echo "</body>" echo "</html>" output displayed like this: <html> <body> welcome to peace world </body> </html>

3. Shell Programming and Scripting

How can i delete html attributes from tag ?

Input: <table class="pixelBorderTable faqTable" width="100%" border="1" cellpadding="3" cellspacing="0"> <tbody><tr> <td class="pixelBorderTableHeaderTd" valign="top" width="20%" bgcolor="#666666"><p> </p></td> <td class="pixelBorderTableHeaderTd" valign="top"...

4. Shell Programming and Scripting

Script to delete HTML tag

Guys, I have a little script that I got of the internet and that I use in Squid to block ads. I used that script with linux but now i have moved my servers to freebsd. I have a step learning curve there but it is fun: Back to the script issue. The script used to work i with linux but...

5. Shell Programming and Scripting

How to retrieve the value from XML tag whose end tag is in next line

Hi All, Find the following code: <Universal>D38x82j1JJ </Universal> I want to retrieve the value of <Universal> tag as below: Please help me.

6. Shell Programming and Scripting

Add the html tag first and last line the file

Hi, i have 30 html files and i want to add the html tag first (<html>) and end of the line </html> tag..How to do it in script. Thanks,

7. Shell Programming and Scripting

Extracting a string from html tag

Hi I am new to string extractions in shell script... I am trying to extract a string such as #1753 from html tag looks like below. <a class="model-link tl-tr" href="lastSuccessfulBuild/">Last successful build (#1753), 40 min ago</a> and want the value as 1753 Could someone help me to...

8. Shell Programming and Scripting

Search for a html tag and print the entire tag

I want to print from <fruits> to </fruits> tag which have <fruit> as mango. Also i want both <fruits> and </fruits> in output. Please help eg. <fruits> <fruit id="111">mango<fruit> . another 20 lines . </fruits>

9. Shell Programming and Scripting

Print Value between desired html tag

Hi, I have a html line as below :-...

10. Shell Programming and Scripting

Extracting data between two tag pairs

In a huge log file (43MB, 43k lines) I am trying to extract data between two tag pairs on same line and export it to a file so I can pull it into Excel for a report. One Pair is <Text>data I need</Text> Other pair follows on same line and is <TimeStamp>more data I need</TimeStamp> I would need...

LEARN ABOUT OSX

html::linkextor

HTML::LinkExtor(3)					User Contributed Perl Documentation					HTML::LinkExtor(3)

NAME

       HTML::LinkExtor - Extract links from an HTML document

SYNOPSIS

	require HTML::LinkExtor;
	$p = HTML::LinkExtor->new(&cb, "http://www.perl.org/");
	sub cb {
	    my($tag, %links) = @_;
	    print "$tag @{[%links]}
";
	}
	$p->parse_file("index.html");

DESCRIPTION

       HTML::LinkExtor is an HTML parser that extracts links from an HTML document.  The HTML::LinkExtor is a subclass of HTML::Parser. This means
       that the document should be given to the parser by calling the $p->parse() or $p->parse_file() methods.

       $p = HTML::LinkExtor->new
       $p = HTML::LinkExtor->new( $callback )
       $p = HTML::LinkExtor->new( $callback, $base )
	   The constructor takes two optional arguments. The first is a reference to a callback routine. It will be called as links are found. If
	   a callback is not provided, then links are just accumulated internally and can be retrieved by calling the $p->links() method.

	   The $base argument is an optional base URL used to absolutize all URLs found.  You need to have the URI module installed if you provide
	   $base.

	   The callback is called with the lowercase tag name as first argument, and then all link attributes as separate key/value pairs.  All
	   non-link attributes are removed.

       $p->links
	   Returns a list of all links found in the document.  The returned values will be anonymous arrays with the following elements:

	     [$tag, $attr => $url1, $attr2 => $url2,...]

	   The $p->links method will also truncate the internal link list.  This means that if the method is called twice without any parsing
	   between them the second call will return an empty list.

	   Also note that $p->links will always be empty if a callback routine was provided when the HTML::LinkExtor was created.

EXAMPLE

       This is an example showing how you can extract links from a document received using LWP:

	 use LWP::UserAgent;
	 use HTML::LinkExtor;
	 use URI::URL;

	 $url = "http://www.perl.org/";  # for instance
	 $ua = LWP::UserAgent->new;

	 # Set up a callback that collect image links
	 my @imgs = ();
	 sub callback {
	    my($tag, %attr) = @_;
	    return if $tag ne 'img';  # we only look closer at <img ...>
	    push(@imgs, values %attr);
	 }

	 # Make the parser.  Unfortunately, we don't know the base yet
	 # (it might be different from $url)
	 $p = HTML::LinkExtor->new(&callback);

	 # Request document and parse it as it arrives
	 $res = $ua->request(HTTP::Request->new(GET => $url),
			     sub {$p->parse($_[0])});

	 # Expand all image URLs to absolute ones
	 my $base = $res->base;
	 @imgs = map { $_ = url($_, $base)->abs; } @imgs;

	 # Print them out
	 print join("
", @imgs), "
";

SEE ALSO

       HTML::Parser, HTML::Tagset, LWP, URI::URL

COPYRIGHT

       Copyright 1996-2001 Gisle Aas.

       This library is free software; you can redistribute it and/or modify it under the same terms as Perl itself.

perl v5.16.2							    2011-10-15							HTML::LinkExtor(3)

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

How do I extract text only from html file without HTML tag

Discussion started by: los111

2. Shell Programming and Scripting

how to use html tag in shell scripting

Discussion started by: jrex1983

3. Shell Programming and Scripting

How can i delete html attributes from tag ?

Discussion started by: cola

4. Shell Programming and Scripting

Script to delete HTML tag

Discussion started by: zongo