Retrieve information Text/Word from HTML code using awk/sed Post: 302897219

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

How to use sed to remove html tags including text between them

How to use sed to remove html tags including text between them? Example: User <b> rolvak </b> is stupid. It does not using <b>OOP</b>! and should output: User is stupid. It does not using ! Thank you..

2. UNIX for Dummies Questions & Answers

retrieve lines using sed, grep or awk

Hi, I'm looking for a command to retrieve a block of lines using sed or grep, probably awk if that can do the job. In below example, By searching for words "Third line2" i'm expecting to retrieve the full block starting with 'BEGIN' and ending with 'END' of the search. Example: ...

3. Shell Programming and Scripting

SED to extract HTML text data, not quite right!

I am attempting to extract weather data from the following website, but for the Victoria area only: Text Forecasts - Environment Canada I use this: sed -n "/Greater Victoria./,/Fraser Valley./p" But that phrasing does not sometimes get it all and think perhaps the website has more...

4. Shell Programming and Scripting

How to retrieve digital string using sed or awk

Hi, I have filename in the following format: YUENLONG_20070818.DMP HK_20070818_V0.DMP WANCHAI_20070820.DMP KWUNTONG_20070820_V0.DMP How to retrieve only the digital part with sed or awk and return the following format: 20070818 20070818 20070820 20070820 Thanks! Victor

5. Shell Programming and Scripting

sed/awk to retrieve max year in column

I am trying to retrieve that max 'year' in a text file that is delimited by tilde (~). It is the second column and the values may be in Char format (double quoted) and have duplicate values. Please help.

6. Shell Programming and Scripting

Execute a C program and retrieve information

Hi I have the following script: #!/bin/sh gcc -o program program.c ./program & PID=$! where i execute a C program and i get its pid. I want to retrieve information about this program (e.g memory consumption) using command top. So far i have: top -d 1.0 -p $PID But i dont know how to...

7. Shell Programming and Scripting

cut, sed, awk too slow to retrieve line - other options?

Hi, I have a script that, basically, has two input files of this type: file1 key1=value1_1_1 key2=value1_2_1 key4=value1_4_1 ... file2 key2=value2_2_1 key2=value2_2_2 key3=value2_3_1 key4=value2_4_1 ... My files are 10k lines big each (approx). The keys are strings that don't...

8. Shell Programming and Scripting

Extract word from text (sed,awk, etc...)

Hello, I need some help extracting the number after the RBA e.g 15911688 from the below block of text (e.g: grep RBA |sed .......). The code should be valid for blocks if text generated at different times as well and not for the below text only. ...

9. Shell Programming and Scripting

Perl code to retrieve text from website

perl -MLWP::Simple -le '$s=shift;$c=get("http://www.google.com/intl/en/chrome/devices/chromecast/$s/");$c=~/meta content=(.*?)name=\"Remote free\"/msg; print length($1),"\t$1"' ?gclid=CJDg27OdnL0CFcFlOgodFD8A6Q >output.txt output.txt should be: Chromecast works with devices you already own,...

10. Shell Programming and Scripting

Awk/sed HTML extract

I'm extracting text between table tags in HTML <th><a href="/wiki/Buick_LeSabre" title="Buick LeSabre">Buick LeSabre</a></th> using this: awk -F "</*th>" '/<\/*th>/ {print $2}' auto2 > auto3 then this (text between a href): sed -e 's/$<*>$//g' auto3 > auto4 How to shorten this into one...

LEARN ABOUT DEBIAN

hxnormalize

HXNORMALIZE(1)							  HTML-XML-utils						    HXNORMALIZE(1)

NAME

       hxnormalize - pretty-print an HTML file

SYNOPSIS

       hxnormalize [ -x ] [ -e ] [ -d ] [ -s ] [ -L ] [ -i indent ] [ -l line-length ] [ -c commentmagic ] [ file-or-URL ]

DESCRIPTION

       The  hxnormalize  command  pretty-prints  an HTML file, and also tries to fix small errors. The output is the same HTML, but with a maximum
       line length and with optional indentation to indicate the nesting level of each line.

OPTIONS

       The following options are supported:

       -x	 Use XML conventions: empty elements are written with a slash at the end: <IMG />. Implies -e.

       -e	 Always insert endtags, even if HTML does not require them (for example: </p> and </li>).

       -d	 Omit the DOCTYPE from the output.

       -i indent Set the number of spaces to indent each nesting level. Default is 2.  Not all elements cause an indent. In general, elements that
		 can  occur  in a block environment are started on a new line and cause an indent, but inline elements, such as EM and SPAN do not
		 cause an indent.

       -l line-length
		 Sets the maximum length of lines.  hxnormalize will wrap lines so that all lines are as long as possible, but no longer than this
		 length. Default is 72. Words that are longer than the line length will not be broken, and will extend past this length. A

		 content of the STYLE, SCRIPT and PRE elements will not be line-wrapped.

       -s	 Omit <span> tags that don't have any attributes.

       -L	 Remove redundant "lang" and "xml:lang" attributes. (I.e., those whose value is the same as the language inherited from the parent
		 element.)

       -c commentmagic
		 Comments are normally placed right after the preceding text. That is usually correct for short comments, but  some  comments  are
		 meant	to  be on a separate line.  commentmagic is a string and when that string occurs inside a comment, hxnormalize will output
		 an empty line before that comment. E.g. -c "====" can be used to put all comments that contain "====" on a  separate  line,  pre-
		 ceded by an empty line. By default, no comments are treated that way.

OPERANDS

       The following operand is supported:

       file-or-URL
		 The name or URL of an HTML file. If absent, standard input is read instead.

EXIT STATUS

       The following exit values are returned:

       0	 Successful completion.

       > 0	 An error occurred in the parsing of the HTML file.  hxnormalize will try to correct the error and produce output anyway.

ENVIRONMENT

       To use a proxy to retrieve remote files, set the environment variables http_proxy and ftp_proxy.  E.g., http_proxy="http://localhost:8080/"

BUGS

       The error recovery for incorrect HTML is primitive.

       hxnormalize will not omit an endtag if the white space after it could possibly be significant. E.g., it will not remove the first </p> from
       "<div><p>text</p> <p>text</p></div>".

       hxnormalize can currently only retrieve remote files over HTTP. It doesn't handle password-protected files, nor files whose content depends
       on HTTP "cookies."

SEE ALSO

       asc2xml(1), xml2asc(1), UTF-8 (RFC 2279)

6.x								    10 Jul 2011 						    HXNORMALIZE(1)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

How to use sed to remove html tags including text between them

Discussion started by: alphagon

2. UNIX for Dummies Questions & Answers

retrieve lines using sed, grep or awk

Discussion started by: learning_linux

3. Shell Programming and Scripting

SED to extract HTML text data, not quite right!

Discussion started by: lagagnon

4. Shell Programming and Scripting

How to retrieve digital string using sed or awk

Discussion started by: victorcheung