awk to print fields that match using conditions and a default value for non-matching in two files

03-18-2017

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

awk to print fields that match using conditions and a default value for non-matching in two files

Trying to use awk to match the contents of each line in file1 with $5 in file2. Both files are tab-delimited and there may be a space or special character in the name being matched in file2, for example in file1 the name is BRCA1 but in file2 the name is BRCA 1 or in file1 name is BCR but in file2 the name is BCR/ABL.

If there is a match and $5 of file2 and $7 has full gene sequence in it, then $5 and $4 are printed separated by a tab. If there is no match found then the name that was not matched and 279 are printed separated by a tab. The awk below does execute, but the output is not correct. Also I am not sure how to add in the condition to ensure $7 is full gene sequence.

The names in file2 may be partial matchto file1, but in file1 they will always be complete. Like in the BRCA1 in file1 that matches the BRCA 1, BRCA2 in file2. The full gene sequence in $7 may also be partial in file2 as is the case for BCRA1. The file2 is not a controlled document so the case may be different as in fbn1 from file1 matching FBN1 in file2. The awk seems close but not all conditions are accounted for. Thank you

Code:

awk 'BEGIN{FS=OFS="\t"}
  FNR==NR{
      if(NR>1){
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"/")
          d[v[1]] = $4          #from split, first element as key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1

BRCA1	279
BCR	806
SCN1A	279
fbn1	85

Code:

BRCA1	81
BCR	806
SCN1A	279
fbn1	85

file1

Code:

BRCA1
BCR
SCN1A
fbn1

file2

Code:

Tier	explanation	.	List code	gene	gene name	methodology	disease
Tier 1	.	.	811	DMD	dystrophin	deletion analysis and duplication analysis, if performed Publication Date: January 1, 2014	Duchenne/Becker muscular dystrophy
Tier 1	.	Jan-16	81	BRCA 1, BRCA2	breast cancer 1 and 2	full gene sequence and full deletion/duplication analysis	hereditary breast and ovarian cancer
Tier 1	.	Jan-16	70	ABL1	ABL1	gene analysis variants in the kinse domane	acquired imatinib tyrosine kinase inhibitor
Tier 1	.	.	806	BCR/ABL 1 	t(9;22)	major breakpoint, qualitative or quantitative	chronic myelogenous leukemia CML
Tier 1	.	Jan-16	85	FBN1	Fibrillin	full gene sequencing	heart disease
Tier 1	.	Jan-16	95	FBN1	fibrillin	del/dup	heart disease

Last edited by cmccabe; 03-18-2017 at 11:31 AM.. Reason: fixed format

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

03-19-2017

Moderator

3,791, 1,452

Join Date: Oct 2010

Last Activity: 1 August 2020, 1:38 AM EDT

Posts: 3,791

Thanks Given: 183

Thanked 1,452 Times in 1,302 Posts

Try these changes:

Code:

awk 'BEGIN{FS=OFS="\t"}
{$0=toupper($0)}
FNR==NR{
   if(NR>1 && ($7 ~ "FULL GENE SEQUENC")) {
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"/")
          d[v[1]] = $4          #from split, first element as key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1

This User Gave Thanks to Chubler_XL For This Post:

Chubler_XL

View Public Profile for Chubler_XL

Find all posts by Chubler_XL

03-19-2017

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

I made a typo in on of the file1 lines, BCRA1 should be BCRA2.

file1

Code:

BRCA2
BCR
SCN1A
fbn1

current output:

Code:

awk 'BEGIN{FS=OFS="\t"}
{$0=toupper($0)}
FNR==NR{
   if(NR>1 && ($7 ~ "FULL GENE SEQUENC")) {
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"/")
          d[v[1]] = $4          #from split, first element as key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1
BRCA2    279
BCR    279
SCN1A    279
FBN1    85

FULL GENE SEQUENC could also be case in sensitive so I added a check in for that... why isn't FULL GENE SEQUENCE used, when I try that I get all the names with a value of 279.

Code:

awk 'BEGIN{FS=OFS="\t"}
{$0=toupper($0)} {$7=toupper($7)}
FNR==NR{
   if(NR>1 && ($7 ~ "FULL GENE SEQUENC")) {
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"/")
          d[v[1]] = $4          #from split, first element as key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1
BRCA2    279 
BCR    279
SCN1A    279
FBN1    85

desired output

Code:

BRCA2    81   - match in line 2 of $5 in file 2, BRCA 1, BRCA2
BCR    279     - match in line 2 of $5 in file but $7 is not full gene sequence
SCN1A    279
fbn1    85

Thank you

Last edited by cmccabe; 03-19-2017 at 06:09 PM.. Reason: fixed format

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

03-19-2017

Moderator

3,791, 1,452

Join Date: Oct 2010

Last Activity: 1 August 2020, 1:38 AM EDT

Posts: 3,791

Thanks Given: 183

Thanked 1,452 Times in 1,302 Posts

Trying to match full gene sequencing and full gene sequence as this is in the data.

This User Gave Thanks to Chubler_XL For This Post:

Chubler_XL

View Public Profile for Chubler_XL

Find all posts by Chubler_XL

03-19-2017

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

Makes sense now, thank you.

In order to capture BRCA2 from file1 with BRCA 1, BRCA2 from file2, would:

Code:

 gsub(" ","",$5)       #removing white space

have to be to capture the matching name anywhere in that $5, I believe it is currently only looking in the first position, but BRCA2 is in position 2. Thank you

Code:

 gsub(" ","",""$5)       #removing white space

or is the problem that $7 is full gene sequence and full deletion/duplication analysis, that is full gene sequence is a partial match to the full line in $7?

Last edited by cmccabe; 03-19-2017 at 08:55 PM.. Reason: added details

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

03-20-2017

Moderator

3,791, 1,452

Join Date: Oct 2010

Last Activity: 1 August 2020, 1:38 AM EDT

Posts: 3,791

Thanks Given: 183

Thanked 1,452 Times in 1,302 Posts

You could go thru all the values in $5 using comma , and slash / as separators.

Code:

awk 'BEGIN{FS=OFS="\t"}
{$0=toupper($0)}
FNR==NR{
   if(NR>1 && ($7 ~ "FULL GENE SEQUENC")) {
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"[,/]")
          for(i=1; i<=n; i++)
             d[v[i]] = $4          # from split, use each element as a key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1

This User Gave Thanks to Chubler_XL For This Post:

Chubler_XL

View Public Profile for Chubler_XL

Find all posts by Chubler_XL

03-23-2017

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

Thank you very much for all your help

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

Shell Programming and Scripting

awk to print fields that match using conditions and a default value for non-matching in two files

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Matching two fields in two csv files, create new file and append match

Discussion started by: dis0wned

2. UNIX for Beginners Questions & Answers

awk match two fields in two files

Discussion started by: geomarine

3. UNIX for Beginners Questions & Answers

Match Fields between two files, print portions of each file together when matched in ([g]awk)'

Discussion started by: jvoot

4. Shell Programming and Scripting

awk to print match or non-match and select fields/patterns for non-matches

Discussion started by: cmccabe

5. UNIX for Beginners Questions & Answers

Awk: matching multiple fields between 2 files

Discussion started by: beca123456

6. Shell Programming and Scripting

Print matching fields (if they exist) from two text files

Discussion started by: gacanepa

7. Shell Programming and Scripting

awk to combine all matching fields in input but only print line with largest value in specific field

Discussion started by: cmccabe

8. Shell Programming and Scripting

awk to match field between two files and use conditions on match

Discussion started by: cmccabe

9. UNIX for Advanced & Expert Users

awk print all fields except matching regex

Discussion started by: glev2005

10. Shell Programming and Scripting

AWK Matching Fields and Combining Files

Discussion started by: Michelangelo