awk to print fields that match using conditions and a default value for non-matching in two files

03-18-2017

Registered User

1,393, 20

Join Date: Nov 2013

Last Activity: 1 May 2020, 2:35 PM EDT

Location: Chicago

Posts: 1,393

Thanks Given: 901

Thanked 20 Times in 19 Posts

awk to print fields that match using conditions and a default value for non-matching in two files

Trying to use awk to match the contents of each line in file1 with $5 in file2. Both files are tab-delimited and there may be a space or special character in the name being matched in file2, for example in file1 the name is BRCA1 but in file2 the name is BRCA 1 or in file1 name is BCR but in file2 the name is BCR/ABL.

If there is a match and $5 of file2 and $7 has full gene sequence in it, then $5 and $4 are printed separated by a tab. If there is no match found then the name that was not matched and 279 are printed separated by a tab. The awk below does execute, but the output is not correct. Also I am not sure how to add in the condition to ensure $7 is full gene sequence.

The names in file2 may be partial matchto file1, but in file1 they will always be complete. Like in the BRCA1 in file1 that matches the BRCA 1, BRCA2 in file2. The full gene sequence in $7 may also be partial in file2 as is the case for BCRA1. The file2 is not a controlled document so the case may be different as in fbn1 from file1 matching FBN1 in file2. The awk seems close but not all conditions are accounted for. Thank you

Code:

awk 'BEGIN{FS=OFS="\t"}
  FNR==NR{
      if(NR>1){
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"/")
          d[v[1]] = $4          #from split, first element as key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1

BRCA1	279
BCR	806
SCN1A	279
fbn1	85

Code:

BRCA1	81
BCR	806
SCN1A	279
fbn1	85

file1

Code:

BRCA1
BCR
SCN1A
fbn1

file2

Code:

Tier	explanation	.	List code	gene	gene name	methodology	disease
Tier 1	.	.	811	DMD	dystrophin	deletion analysis and duplication analysis, if performed Publication Date: January 1, 2014	Duchenne/Becker muscular dystrophy
Tier 1	.	Jan-16	81	BRCA 1, BRCA2	breast cancer 1 and 2	full gene sequence and full deletion/duplication analysis	hereditary breast and ovarian cancer
Tier 1	.	Jan-16	70	ABL1	ABL1	gene analysis variants in the kinse domane	acquired imatinib tyrosine kinase inhibitor
Tier 1	.	.	806	BCR/ABL 1 	t(9;22)	major breakpoint, qualitative or quantitative	chronic myelogenous leukemia CML
Tier 1	.	Jan-16	85	FBN1	Fibrillin	full gene sequencing	heart disease
Tier 1	.	Jan-16	95	FBN1	fibrillin	del/dup	heart disease

Last edited by cmccabe; 03-18-2017 at 11:31 AM.. Reason: fixed format

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

Shell Programming and Scripting

awk to print fields that match using conditions and a default value for non-matching in two files

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Matching two fields in two csv files, create new file and append match

Discussion started by: dis0wned

2. UNIX for Beginners Questions & Answers

awk match two fields in two files

Discussion started by: geomarine

3. UNIX for Beginners Questions & Answers

Match Fields between two files, print portions of each file together when matched in ([g]awk)'

Discussion started by: jvoot

4. Shell Programming and Scripting

awk to print match or non-match and select fields/patterns for non-matches

Discussion started by: cmccabe

5. UNIX for Beginners Questions & Answers

Awk: matching multiple fields between 2 files

Discussion started by: beca123456

6. Shell Programming and Scripting

Print matching fields (if they exist) from two text files

Discussion started by: gacanepa

7. Shell Programming and Scripting

awk to combine all matching fields in input but only print line with largest value in specific field

Discussion started by: cmccabe

8. Shell Programming and Scripting

awk to match field between two files and use conditions on match

Discussion started by: cmccabe

9. UNIX for Advanced & Expert Users

awk print all fields except matching regex

Discussion started by: glev2005

10. Shell Programming and Scripting

AWK Matching Fields and Combining Files

Discussion started by: Michelangelo