Awk: print all URL addresses between iframe tags without repeating an already printed URL

Thread Tools Search this Thread
# 15  
Yes, thank you. Appears cleaner now, although I prefer to use sort -u instead of the if.

For now I use the following:

find . -name '*.html' -or -name '*.htm' -or -name '*.php' -type f| xargs awk -F\" -v RS='<' '/^iframe src=/ {print $2}' | sort -u

My next goal in this case would actually be to strip everything else but the hostname.

Meaning whatever is between "http://" and "/".

For example, here is a sample output I have now:


I would like the output to be just:


I will then consolidate this command with another one in a script in order to implement an easy way to list all unique Iframes in the user's web space and selectively remove those of unknown source (hacked).

I am starting to understand the concept of awk and sed, but there is just so much more to learn...

Thank you for your great help. I really appreciate all the effort!
# 16  
Come to think of it.... If you find enough files, as in thousands, you probably would need to use the sort -u anyway.

---------- Post updated at 05:28 PM ---------- Previous update was at 05:25 PM ----------

You can tack this on before the sort -u to remove http:// before and /.* after:
... | sed 's#http://##;s#/.*##' | sort -u

# 17  
Yeah, thanks for that. Here is what I got going:

So I will be using

find . -name '*.html' -or -name '*.htm' -or -name '*.php' -type f| xargs awk -F\" -v RS='<' '/^iframe src=/ {print $2}'|sed 's#http://##;s#/.*##' | sort -u

to get addresses. Then I prompt the user for deletion on each entry (haven't figured that out), and process each (Yes) with the following:

find . -name "*php*" -or -name "*htm*" |xargs grep -rl "HOSTNAME" |xargs sed -i 's/[echo "]*<iframe src=[\\]*.http:\/\/HOSTNAME[^>]*>[\w]*<\/iframe>[";]*//g'

along with something like

    while [ -d /proc/$1 ]
        printf "$SP_COLOUR\e7  %${SP_WIDTH}s  \e8\e[0m" "$SP_STRING"
        sleep ${SP_DELAY:-.2}

## Adjust to taste (or leave empty)
SP_WIDTH=1.1  ## Try: SP_WIDTH=5.5

sleep 3 &
spinner "$!" '.o0Oo'

while each HOSTNAME is getting cleaned out. If the user selects to not remove some of the found domains, the script skips to the next domain until there are no domains left.

# 18  
Oh, I get it.

Why wait for sleep though? Why not wait for the thing you want to wait for?
# 19  
the purpose of sleep 3 & is examplatory

Previous Thread | Next Thread
Thread Tools Search this Thread
Search this Thread:
Advanced Search

Test Your Knowledge in Science: Gadgets
Difficulty: Easy
Microphones can be used not only to pick up sound, but also to project sound similar to a speaker.
True or False?

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Reading URL using Mechanize and dump all the contents of the URL to a file

Hello, Am very new to perl , please help me here !! I need help in reading a URL from command line using PERL:: Mechanize and needs all the contents from the URL to get into a file. below is the script which i have written so far , #!/usr/bin/perl use LWP::UserAgent; use... (2 Replies)
Discussion started by: scott_cog
2 Replies

2. Shell Programming and Scripting

awk and or sed command to sum the value in repeating tags in a XML

I have a XML in which <Amt Ccy="EUR">3.1</Amt> tag repeats. This is under another tag <Main>. I need to sum all the values of <Amt Ccy=""> (Ccy may vary) coming under <Main> using awk and or sed command. can some help? Sample looks like below <root> <Main> ... (6 Replies)
Discussion started by: bk_12345
6 Replies

3. UNIX for Dummies Questions & Answers

URL decoding with awk

The challenge: Decode URL's, i.e. convert %HEX to the corresponding special characters, using only UNIX base utilities, and without having to type out each special character. I have an anonymous C code snippet where the author assigns each hex digit a number from 0 to 16 and then does some... (2 Replies)
Discussion started by: uiop44
2 Replies

4. Web Development

Regex to rewrite URL to another URL based on HTTP_HOST?

I am trying to find a way to test some code, but I need to rewrite a specific URL only from a specific HTTP_HOST The call goes out to http://SUB.DOMAIN.COM/showAssignment/7bde10b45efdd7a97629ef2fe01f7303/jsmodule/Nevow.Athena The ID in the middle is always random due to the cookie. I... (5 Replies)
Discussion started by: EXT3FSCK
5 Replies

5. Shell Programming and Scripting

Extract URL from RSS Feed in AWK

Hi, I have following data file; <outline title="Matt Cutts" type="rss" version="RSS" xmlUrl="" htmlUrl=""/> <outline title="Stone" text="Stone" type="rss" version="RSS" xmlUrl=""... (8 Replies)
Discussion started by: fahdmirza
8 Replies

6. Shell Programming and Scripting

how to judge wether a url is valid or not using awk

rt 3ks:confused: (6 Replies)
Discussion started by: rainboisterous
6 Replies

7. UNIX for Advanced & Expert Users

Need to grab URL and place between <A></A> Tags

my output looks like: <A HREF=""> </A> <A HREF=""> </A> <A HREF=""> </A> <A HREF=""> </A> <A... (3 Replies)
Discussion started by: glev2005
3 Replies

8. UNIX for Dummies Questions & Answers

ReDirecting a URL to another URL - Linux

Hello, I need to redirect an existing URL, how can i do that? There's a current web address to a GUI that I have to redirect to another webaddress. Does anyone know how to do this? This is on Unix boxes Linux. example: make it so the... (3 Replies)
Discussion started by: SkySmart
3 Replies

9. Shell Programming and Scripting

url calling and parameter passing to url in script

Hi all, I need to write a unix script in which need to call a url. Then need to pass parameters to that url. please help. Regards, gander_ss (1 Reply)
Discussion started by: gander_ss
1 Replies

10. UNIX for Advanced & Expert Users

url calling and parameter passing to url in script

Hi all, I need to write a unix script in which need to call a url. Then need to pass parameters to that url. please help. Regards, gander_ss (1 Reply)
Discussion started by: gander_ss
1 Replies

Featured Tech Videos