awk to print fields that match using conditions and a default value for non-matching in two files Post: 302994137

Sponsored Content

Top Forums Shell Programming and Scripting awk to print fields that match using conditions and a default value for non-matching in two files Post 302994137 by cmccabe on Sunday 19th of March 2017 05:08:33 PM

03-19-2017

Registered User

I made a typo in on of the file1 lines, BCRA1 should be BCRA2.

file1

Code:

BRCA2
BCR
SCN1A
fbn1

current output:

Code:

awk 'BEGIN{FS=OFS="\t"}
{$0=toupper($0)}
FNR==NR{
   if(NR>1 && ($7 ~ "FULL GENE SEQUENC")) {
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"/")
          d[v[1]] = $4          #from split, first element as key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1
BRCA2    279
BCR    279
SCN1A    279
FBN1    85

FULL GENE SEQUENC could also be case in sensitive so I added a check in for that... why isn't FULL GENE SEQUENCE used, when I try that I get all the names with a value of 279.

Code:

awk 'BEGIN{FS=OFS="\t"}
{$0=toupper($0)} {$7=toupper($7)}
FNR==NR{
   if(NR>1 && ($7 ~ "FULL GENE SEQUENC")) {
          gsub(" ","",$5)       #removing white space
          n=split($5,v,"/")
          d[v[1]] = $4          #from split, first element as key
      }
      next
}{print $1, ($1 in d?d[$1]:279)}' file2 file1
BRCA2    279 
BCR    279
SCN1A    279
FBN1    85

desired output

Code:

BRCA2    81   - match in line 2 of $5 in file 2, BRCA 1, BRCA2
BCR    279     - match in line 2 of $5 in file but $7 is not full gene sequence
SCN1A    279
fbn1    85

Thank you

Last edited by cmccabe; 03-19-2017 at 06:09 PM.. Reason: fixed format

cmccabe

View Public Profile for cmccabe

Find all posts by cmccabe

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

AWK Matching Fields and Combining Files

Hello! I am writing a program to run through two large lists of data (~300,000 rows), find where rows in one file match another, and combine them based on matching fields. Due to the large file sizes, I'm guessing AWK will be the most efficient way to do this. Overall, the input and output I'm...

2. UNIX for Advanced & Expert Users

awk print all fields except matching regex

grep -v will exclude matching lines, but I want something that will print all lines but exclude a matching field. The pattern that I want excluded is '/mnt/svn' If there is a better solution than awk I am happy to hear about it, but I would like to see this done in awk as well. I know I can...

3. Shell Programming and Scripting

awk to match field between two files and use conditions on match

I am trying to look for $2 of file1 (skipping the header) in $2 of file2 (skipping the header) and if they match and the value in $10 is > 30 and $11 is > 49, then print the line from file1 to a output file. If no match is foung the line is not printed. Both the input and output are tab-delimited....

4. Shell Programming and Scripting

awk to combine all matching fields in input but only print line with largest value in specific field

In the below I am trying to use awk to match all the $13 values in input, which is tab-delimited, that are in $1 of gene which is just a single column of text. However only the line with the greatest $9 value in input needs to be printed. So in the example below all the MECP2 and LTBP1...

5. Shell Programming and Scripting

Print matching fields (if they exist) from two text files

Hi everyone, Given two files (test1 and test2) with the following contents: test1: 80263760,I71 80267369,M44 80274628,L77 80276793,I32 80277390,K05 80277391,I06 80279206,I43 80279859,K37 80279866,K35 80279867,J16 80280346,I14and test2: 80263760,PT18 80279867,PT01I need to do some...

6. UNIX for Beginners Questions & Answers

Awk: matching multiple fields between 2 files

Hi, I have 2 tab-delimited input files as follows. file1.tab: green A apple red B apple file2.tab: apple - A;Z Objective: Return $1 of file1 if, . $1 of file2 matches $3 of file1 and, . any single element (separated by ";") in $3 of file2 is present in $2 of file1 In order to...

7. Shell Programming and Scripting

awk to print match or non-match and select fields/patterns for non-matches

In the awk below I am trying to output those lines that Match between file1 and file2, those Missing in file1, and those missing in file2. Using each $1,$2,$4,$5 value as a key to match on, that is if those 4 fields are found in both files the match, but if those 4 fields are not found then missing...

8. UNIX for Beginners Questions & Answers

Match Fields between two files, print portions of each file together when matched in ([g]awk)'

I've written an awk script to compare two fields in two different files and then print portions of each file on the same line when matched. It works reasonably well, but every now and again, I notice some errors and cannot seem to figure out what the issue may be and am turning to you for help. ...

9. UNIX for Beginners Questions & Answers

awk match two fields in two files

Hi, I have two TEST files t.xyz and a.xyz which have three columns each. a.xyz have more rows than t.xyz. I will like to output rows at which $1 and $2 of t.xyz match $1 and $2 of a.xyz. Total number of output rows should be equal to that of t.xyz. It works fine, but when I apply it to large...

10. Shell Programming and Scripting

Matching two fields in two csv files, create new file and append match

I am trying to parse two csv files and make a match in one column then print the entire file to a new file and append an additional column that gives description from the match to the new file. If a match is not made, I would like to add "NA" to the end of the file Command that Ive been using...

LEARN ABOUT HPUX

merge

merge(1)						      General Commands Manual							  merge(1)

NAME

       merge - three-way file merge

SYNOPSIS

       file1 file2 file3

DESCRIPTION

       combines  two  files  that are revisions of a single original file.  The original file is file2, and the revised files are file1 and file3.
       identifies all changes that lead from file2 to file3 and from file2 to file1, then deposits the merged text into file1.	If the	option	is
       used, the result goes to standard output instead of file1.

       An  overlap  occurs  if both file1 and file3 have changes in the same place.  prints how many overlaps occurred, and includes both alterna-
       tives in the result.  The alternatives are delimited as follows:

	      lines in file1
	      lines in file3

       If there are overlaps, edit the result in file1 and delete one of the alternatives.

       This command is particularly useful for revision control, especially if file1 and file3 are the ends of two branches that have file2  as  a
       common ancestor.

EXAMPLES

       A typical use for is as follows:

	      1.   To  merge  an  RCS branch into the trunk, first check out the three different versions from RCS (see co(1)) and rename them for
		   their revision numbers: 5.2, 5.11, and 5.2.3.3.  File 5.2.3.3 is the end of an RCS branch that split off the trunk at file 5.2.

	      2.   For this example, assume file 5.11 is the latest version on the trunk, and is also a revision  of  the  "original"  file,  5.2.
		   Merge the branch into the trunk with the command:

	      3.   File  5.11  now  contains  all  changes  made on the branch and the trunk, and has markings in the file to show all overlapping
		   changes.

	      4.   Edit file 5.11 to correct the overlaps, then use the command to check the file back in (see ci(1)).

WARNINGS

       uses the ed(1) system editor.  Therefore, the file size limits of ed(1) apply to

AUTHOR

       was developed by Walter F. Tichy.

SEE ALSO

       diff3(1), diff(1), rcsmerge(1), co(1).

																	  merge(1)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

AWK Matching Fields and Combining Files

Discussion started by: Michelangelo

2. UNIX for Advanced & Expert Users

awk print all fields except matching regex

Discussion started by: glev2005

3. Shell Programming and Scripting

awk to match field between two files and use conditions on match

Discussion started by: cmccabe

4. Shell Programming and Scripting

awk to combine all matching fields in input but only print line with largest value in specific field

Discussion started by: cmccabe

5. Shell Programming and Scripting

Print matching fields (if they exist) from two text files

Discussion started by: gacanepa

6. UNIX for Beginners Questions & Answers

Awk: matching multiple fields between 2 files

Discussion started by: beca123456

7. Shell Programming and Scripting

awk to print match or non-match and select fields/patterns for non-matches

Discussion started by: cmccabe

8. UNIX for Beginners Questions & Answers

Match Fields between two files, print portions of each file together when matched in ([g]awk)'

Discussion started by: jvoot

9. UNIX for Beginners Questions & Answers

awk match two fields in two files

Discussion started by: geomarine

10. Shell Programming and Scripting

Matching two fields in two csv files, create new file and append match

Discussion started by: dis0wned

LEARN ABOUT HPUX

merge