Retrieve information Text/Word from HTML code using awk/sed

04-11-2014

Registered User

9, 1

Join Date: Apr 2014

Last Activity: 14 November 2014, 1:43 PM EST

Location: Chicago

Posts: 9

Thanks Given: 3

Thanked 1 Time in 1 Post

Retrieve information Text/Word from HTML code using awk/sed

awk/sed newbie here. I have a HTML file and from that file and I would like to retrieve a text word.

HTML Code:

<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc_process.txt>abc</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc01_process.txt>abc01</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc045_process.txt>abc045</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/cdf_process.txt>cdf</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/Manhattan_process.txt>Manhattan</a> NDK Version:  4.0 </li>

For eg. From the 1st line I would like to retrieve abc placed between: txt>abc</a>

I have used the following command but as you can see that number of letters in the word keeps changing abc, abc01, abc045, cdf, Manhattan.
awk -F\/ '{print substr($4,0,3)}' list.html

So this command is getting the output for only the 3 letter word. However I want to extract the same information (abc01, abc045, cdf, Manhattan) from all the lines in the HTML code. Please help.

sk2code

View Public Profile for sk2code

Find all posts by sk2code

04-11-2014

Moderator

3,689, 1,352

Join Date: Jan 2012

Last Activity: 22 August 2020, 11:29 PM EDT

Location: Galactic Empire

Posts: 3,689

Thanks Given: 268

Thanked 1,352 Times in 1,258 Posts

Code:

awk -F'[<>]' '{ print $7 }' file.html

Yoda

View Public Profile for Yoda

Visit Yoda's homepage!

Find all posts by Yoda

04-11-2014

Registered User

9, 1

Join Date: Apr 2014

Last Activity: 14 November 2014, 1:43 PM EST

Location: Chicago

Posts: 9

Thanks Given: 3

Thanked 1 Time in 1 Post

I Just ran this but it is giving me no output. Just blank lines. This HTML file is having 5 lines and when I run the command you mentioned I am just getting 5 blank lines.

sk2code

View Public Profile for sk2code

Find all posts by sk2code

04-11-2014

Moderator

3,689, 1,352

Join Date: Jan 2012

Last Activity: 22 August 2020, 11:29 PM EDT

Location: Galactic Empire

Posts: 3,689

Thanks Given: 268

Thanked 1,352 Times in 1,258 Posts

Ok, here is what I got:

Code:

$ cat file.html
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc_process.txt>abc</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc01_process.txt>abc01</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc045_process.txt>abc045</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/cdf_process.txt>cdf</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/Manhattan_process.txt>Manhattan</a> NDK Version:  4.0 </li>

Code:

$ awk -F'[<>]' '{ print $7 }' file.html
abc
abc01
abc045
cdf
Manhattan

You can also try:

Code:

sed 's#.*txt>##;s#<.*##' file.html

This User Gave Thanks to Yoda For This Post:

Yoda

View Public Profile for Yoda

Visit Yoda's homepage!

Find all posts by Yoda

04-11-2014

Registered User

23,310, 4,623

Join Date: Aug 2005

Last Activity: 7 July 2020, 11:47 AM EDT

Location: Saskatchewan

Posts: 23,310

Thanks Given: 1,331

Thanked 4,623 Times in 4,217 Posts

Quote:

Originally Posted by sk2code

I Just ran this but it is giving me no output. Just blank lines. This HTML file is having 5 lines and when I run the command you mentioned I am just getting 5 blank lines.

Does the HTML actually look like the data you pasted, or did you pretty it up? Many times when XML/HTML comes up, 5 "lines" is later found to mean tags not necessarily organized into lines at all.

Corona688

View Public Profile for Corona688

Visit Corona688's homepage!

Find all posts by Corona688

04-12-2014

Registered User

559, 160

Join Date: Jul 2012

Last Activity: 20 September 2019, 7:24 AM EDT

Location: India, Hyderabad

Posts: 559

Thanks Given: 11

Thanked 160 Times in 148 Posts

Code:

awk 'BEGIN{FS = "</a>"} {n=split($1, a, ">"); print a[n]}' file

Code:

srinus@ubuntu:~$ cat sam
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc_process.txt>abc</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc01_process.txt>abc01</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/abc045_process.txt>abc045</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/cdf_process.txt>cdf</a> NDK Version:  4.0 </li>
<font face=arial size=-1><li><a href=/value_for_clients/Tokyo/Manhattan_process.txt>Manhattan</a> NDK Version:  4.0 </li>
srinus@ubuntu:~$ awk 'BEGIN{FS = "</a>"} {n=split($1, a, ">"); print a[n]}' sam
abc
abc01
abc045
cdf
Manhattan
srinus@ubuntu:~$

SriniShoo

View Public Profile for SriniShoo

Find all posts by SriniShoo

04-14-2014

Registered User

9, 1

Join Date: Apr 2014

Last Activity: 14 November 2014, 1:43 PM EST

Location: Chicago

Posts: 9

Thanks Given: 3

Thanked 1 Time in 1 Post

Yoda and everyone thanks for your help.
The awk command still results in blank output but the following command helped me.

Code:

sed 's#.*txt>##;s#<.*##' file.html

Thank you!!

sk2code

View Public Profile for sk2code

Find all posts by sk2code

Shell Programming and Scripting

Retrieve information Text/Word from HTML code using awk/sed

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Awk/sed HTML extract

Discussion started by: p1ne

2. Shell Programming and Scripting

Perl code to retrieve text from website

Discussion started by: cmccabe

3. Shell Programming and Scripting

Extract word from text (sed,awk, etc...)

Discussion started by: drbiloukos

4. Shell Programming and Scripting

cut, sed, awk too slow to retrieve line - other options?

Discussion started by: fzd

5. Shell Programming and Scripting

Execute a C program and retrieve information

Discussion started by: nteath

6. Shell Programming and Scripting

sed/awk to retrieve max year in column

Discussion started by: CKT_newbie88

7. Shell Programming and Scripting

How to retrieve digital string using sed or awk

Discussion started by: victorcheung

8. Shell Programming and Scripting

SED to extract HTML text data, not quite right!

Discussion started by: lagagnon

9. UNIX for Dummies Questions & Answers

retrieve lines using sed, grep or awk

Discussion started by: learning_linux

10. Shell Programming and Scripting

How to use sed to remove html tags including text between them

Discussion started by: alphagon