Parse multiple html files in directory

12-16-2014

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

Parse multiple html files in directory

I have downloaded source code for 97 files using:

Code:

 wget -x -i link.txt

then run a rename loop:

Code:

 for file in *
do
   mv $file $file.txt
done

to keep the html tags but make the file a text that can be parsed.

In each of the 97 txt files the gene # is variable, but the gene is associated or should have a corresponding OMIM #. They are all in,

Code:

 C:\Users\cmccabe\Desktop\list\geneticslab.emory.edu.txt\tests_txt

Is there a way to search the source code for these gene names and OMIM #’s?

For example, in the attached file there are 26 genes:

Output (tab-delimited)

Code:

A            B
Gene     OMIM
AKT1	164730
ALK	105590
APC	611731

The gene names seem to be after

Code:

 target = '_blank'>AKT1</a>

and the OMIM # seem to be

Code:

 style = 'margin-bottom:10px;'><a href =

I think

Code:

sed

can parse html but I am not familiar enough to know how to code it for multiple files in a directory.

Thank you to all for the help

CM080.txt (17.2 KB)

Last edited by Scrutinizer; 12-16-2014 at 03:16 PM.. Reason: Extra code tags

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

12-16-2014

Registered User

15,129, 5,008

Join Date: Jul 2012

Last Activity: 4 May 2020, 4:31 PM EDT

Location: Aachen, Germany

Posts: 15,129

Thanks Given: 735

Thanked 5,008 Times in 4,483 Posts

You don't need to rename the files to .txt to parse them with sed. Try:

Code:

awk     '/^<h2 id="genes"/      {getline
                                 for (i=1; i<=NF; i++)
                                        {n1=gsub ("[, ]*<a href = .http://www.omim.org/entry/", "", $i)
                                         n2=gsub (". target = ._blank.", "", $i)
                                         n3=gsub ("</a", "", $(i+1))
                                         if (n1 + n2) print $(i+1) "\t" $i
                                        }
                                 exit
                                }
        ' FS=">" /tmp/CM080
AKT1    164730
ALK     105590
APC     611731
BRAF    164757
CDH1    192090
CTNNB1  116806
EGFR    131550
ERBB2   164870
FBXW7   606278
FGFR2   176943
FOXL2   605597
GNAQ    600998
GNAS    139320
KIT     164920
KRAS    190070
MAP2K1  176872
MET     164860
MSH6    600678
NRAS    164790
PDGFRA  173490
PIK3CA  171834
PTEN    601728
SMAD4   600993
SRC     190090
STK11   602216
TP53    191170

RudiC

View Public Profile for RudiC

Find all posts by RudiC

12-16-2014

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

Can a loop be used to parse each of the 97 files in

Code:

 C:\Users\cmccabe\Desktop\list\geneticslab.emory.edu.txt\tests

keeping the original file and the newly created parsed.txt? Thank you

.

So, CM080 (original) - CM080parsed.txt

Also, if I run the command to write an output file it takes a very long time

Code:

 awk     '/^<h2 id="genes"/      {getline
                                 for (i=1; i<=NF; i++)
                                        {n1=gsub ("[, ]*<a href = .http://www.omim.org/entry/", "", $i)
                                         n2=gsub (". target = ._blank.", "", $i)
                                         n3=gsub ("</a", "", $(i+1))
                                         if (n1 + n2) print $(i+1) "\t" $i
                                        }
                                 exit
                                }
        ' FS=">" CM080 > output.txt

, but if there is no output it runs quickly.

Code:

 awk     '/^<h2 id="genes"/      {getline
                                 for (i=1; i<=NF; i++)
                                        {n1=gsub ("[, ]*<a href = .http://www.omim.org/entry/", "", $i)
                                         n2=gsub (". target = ._blank.", "", $i)
                                         n3=gsub ("</a", "", $(i+1))
                                         if (n1 + n2) print $(i+1) "\t" $i
                                        }
                                 exit
                                }
        ' FS=">" CM080

Thank you

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

12-17-2014

Registered User

15,129, 5,008

Join Date: Jul 2012

Last Activity: 4 May 2020, 4:31 PM EDT

Location: Aachen, Germany

Posts: 15,129

Thanks Given: 735

Thanked 5,008 Times in 4,483 Posts

There's no reason it should run slower when printing to a file, and it doesn't if I try. For your multiple files, try

Code:

awk     '/^<h2 id="genes"/      {getline
                                 for (i=1; i<=NF; i++)
                                        {n1=gsub (".*omim.org/entry/", "", $i)
                                         n2=gsub (". target = ._blank.>", "\t", $i)
                                         n3=gsub ("</a>.*", "", $i)
                                         print $i  >  FILENAME".parsed"
                                        }
                                }
        ' FS="," /tmp/CM*

after having adapted the input files' path. It requires that ALL input files have the same structure as CM080 does.

RudiC

View Public Profile for RudiC

Find all posts by RudiC

12-17-2014

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

Code:

 awk     '/^<h2 id="genes"/      {getline
>                                  for (i=1; i<=NF; i++)
>                                         {n1=gsub (".*omim.org/entry/", "", $i)
>                                          n2=gsub (". target = ._blank.>", "\t", $i)
>                                          n3=gsub ("</a>.*", "", $i)
>                                          print $i  >  FILENAME".parsed"
>                                         }
>                                 }
>         ' FS="," C:\Users\cmccabe\Desktop\list\geneticslab.emory.edu.txt\parse
awk: fatal: cannot open file `C:UserscmccabeDesktoplistgeneticslab.emory.edu.txtparse' for reading (No such file or directory)

I am using cygwin on a windows machine. Should I try Ubuntu? Thank you

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

12-17-2014

Registered User

15,129, 5,008

Join Date: Jul 2012

Last Activity: 4 May 2020, 4:31 PM EDT

Location: Aachen, Germany

Posts: 15,129

Thanks Given: 735

Thanked 5,008 Times in 4,483 Posts

Looks like it suppressed all "\" in the input file path. Try to escape or quote.
You did not specify any wildcard, so it would work on file "parse" only. Is this what you want? Or is "parse" a directory?

RudiC

View Public Profile for RudiC

Find all posts by RudiC

12-17-2014

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

I put all 97 of the files to be parsed in a directory:

Code:

 C:\Users\cmccabe\Desktop\list\geneticslab.emory.edu.txt\parse

So maybe that should be

Code:

 "C:\Users\cmccabe\Desktop\list\geneticslab.emory.edu.txt\parse
"*

Not sure how to escape.

Thank you

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

Shell Programming and Scripting

Parse multiple html files in directory

10 More Discussions You Might Find Interesting

1. UNIX for Beginners Questions & Answers

Merge Multiple html files into one

Discussion started by: Snehasish

2. Shell Programming and Scripting

awk Parse And Create Multiple Files Based on Field Value

Discussion started by: ec012

3. Shell Programming and Scripting

Parse html

Discussion started by: cmccabe

4. UNIX for Advanced & Expert Users

Mutt for html body and multiple html & pdf attachments

Discussion started by: raggmopp

5. Shell Programming and Scripting

Read multiple files, parse data and append to a file

Discussion started by: empyrean

6. Shell Programming and Scripting

Parse files in directory and compare with another file

Discussion started by: empyrean

7. Shell Programming and Scripting

Need script to remove millions of tmp files in /html/cache/ directory

Discussion started by: andymc1

8. Shell Programming and Scripting

sed to parse html

Discussion started by: prasanna1157

9. Shell Programming and Scripting

to parse a directory and its subdirectories and find owner name of files

Discussion started by: vyasa

10. Shell Programming and Scripting

Multiple edits to a bunch of html files

Discussion started by: dheian