Retrieve information Text/Word from HTML code using awk/sed Post: 302897215

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

How to use sed to remove html tags including text between them

How to use sed to remove html tags including text between them? Example: User <b> rolvak </b> is stupid. It does not using <b>OOP</b>! and should output: User is stupid. It does not using ! Thank you..

2. UNIX for Dummies Questions & Answers

retrieve lines using sed, grep or awk

Hi, I'm looking for a command to retrieve a block of lines using sed or grep, probably awk if that can do the job. In below example, By searching for words "Third line2" i'm expecting to retrieve the full block starting with 'BEGIN' and ending with 'END' of the search. Example: ...

3. Shell Programming and Scripting

SED to extract HTML text data, not quite right!

I am attempting to extract weather data from the following website, but for the Victoria area only: Text Forecasts - Environment Canada I use this: sed -n "/Greater Victoria./,/Fraser Valley./p" But that phrasing does not sometimes get it all and think perhaps the website has more...

4. Shell Programming and Scripting

How to retrieve digital string using sed or awk

Hi, I have filename in the following format: YUENLONG_20070818.DMP HK_20070818_V0.DMP WANCHAI_20070820.DMP KWUNTONG_20070820_V0.DMP How to retrieve only the digital part with sed or awk and return the following format: 20070818 20070818 20070820 20070820 Thanks! Victor

5. Shell Programming and Scripting

sed/awk to retrieve max year in column

I am trying to retrieve that max 'year' in a text file that is delimited by tilde (~). It is the second column and the values may be in Char format (double quoted) and have duplicate values. Please help.

6. Shell Programming and Scripting

Execute a C program and retrieve information

Hi I have the following script: #!/bin/sh gcc -o program program.c ./program & PID=$! where i execute a C program and i get its pid. I want to retrieve information about this program (e.g memory consumption) using command top. So far i have: top -d 1.0 -p $PID But i dont know how to...

7. Shell Programming and Scripting

cut, sed, awk too slow to retrieve line - other options?

Hi, I have a script that, basically, has two input files of this type: file1 key1=value1_1_1 key2=value1_2_1 key4=value1_4_1 ... file2 key2=value2_2_1 key2=value2_2_2 key3=value2_3_1 key4=value2_4_1 ... My files are 10k lines big each (approx). The keys are strings that don't...

8. Shell Programming and Scripting

Extract word from text (sed,awk, etc...)

Hello, I need some help extracting the number after the RBA e.g 15911688 from the below block of text (e.g: grep RBA |sed .......). The code should be valid for blocks if text generated at different times as well and not for the below text only. ...

9. Shell Programming and Scripting

Perl code to retrieve text from website

perl -MLWP::Simple -le '$s=shift;$c=get("http://www.google.com/intl/en/chrome/devices/chromecast/$s/");$c=~/meta content=(.*?)name=\"Remote free\"/msg; print length($1),"\t$1"' ?gclid=CJDg27OdnL0CFcFlOgodFD8A6Q >output.txt output.txt should be: Chromecast works with devices you already own,...

10. Shell Programming and Scripting

Awk/sed HTML extract

I'm extracting text between table tags in HTML <th><a href="/wiki/Buick_LeSabre" title="Buick LeSabre">Buick LeSabre</a></th> using this: awk -F "</*th>" '/<\/*th>/ {print $2}' auto2 > auto3 then this (text between a href): sed -e 's/$<*>$//g' auto3 > auto4 How to shorten this into one...

LEARN ABOUT OSF1

g3cat

g3cat(1)						       mgetty+sendfax manual							  g3cat(1)

NAME

       g3cat - concatenate multiple g3 documents

SYNOPSIS

       g3cat [-l] [-a] g3-file1 ...

DESCRIPTION

       g3cat  concatenates  g3	files.	These can either be 'raw', that is, bitmaps packed according to the CCITT T.4 standard for one-dimensional
       bitmap encoding, or 'digifax' files, created by GNU's GhostScript package with the digifax drivers. Its output is a  concatenation  of  all
       the input files, in raw G3 format, with two white lines in between.

       If a - is given as input file, stdin is used.

       If the input data is malformed, a warning is printed to stderr, and the output file will have a blank line at this place.

OPTIONS

       -l     separate files with a one-pixel wide black line.

       -h <blank lines>
	      specifies the number of blank lines g3cat should prepend to each page. Default is 0.

       -L <lines>
	      limit lenght of output page to maximum <lines> lines.

SPECIAL-CASE OPTIONS
       -w <width>
	      specifies  the  desired  page width in pixels per line. Default is 1728 PELs, and this is mandatory if you want to send the fax to a
	      standard fax machine.  If one of the input files doesn't match this line width (for example because it was created by  a	broken	G3
	      creator), a warning is printed, and the line width is transparently fixed.

       -a     byte-align the end-of-line codes (EOL) in the file. Every EOL will end at a byte boundary, that is, with a  01 byte.

       -p <pad>
	      specifies a minimum number of bytes that each output line must be padded to.  Padding is done with 0-bits before the EOL code.

       -R     suppress output of end-of-page code (RTC).

Example
       The following example will put a header line on a given g3 page, 'page1' and put the result into 'page2':

       echo '$header' | pbmtext | pbm2g3 | g3cat - page1 >page2

FILES

       --

BUGS

       Hopefully none :-).

SEE ALSO

       g32pbm(1), sendfax(8), faxspool(1)

AUTHORS

       g3cat is Copyright (C) 1993 by Gert Doering, <gert@greenie.muc.de>

greenie 							     27 Oct 93								  g3cat(1)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

How to use sed to remove html tags including text between them

Discussion started by: alphagon

2. UNIX for Dummies Questions & Answers

retrieve lines using sed, grep or awk

Discussion started by: learning_linux

3. Shell Programming and Scripting

SED to extract HTML text data, not quite right!

Discussion started by: lagagnon

4. Shell Programming and Scripting

How to retrieve digital string using sed or awk

Discussion started by: victorcheung