Print matching fields (if they exist) from two text files Post: 302988535

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

AWK Matching Fields and Combining Files

Hello! I am writing a program to run through two large lists of data (~300,000 rows), find where rows in one file match another, and combine them based on matching fields. Due to the large file sizes, I'm guessing AWK will be the most efficient way to do this. Overall, the input and output I'm...

2. Shell Programming and Scripting

comparing two files for matching fields

I am newbie to unix and would please like some help to solve the task below I have two files, file_a.text and file_b.text that I want to evaluate. file_a.text 1698.74 1711.88 6576.25 899.41 3205.63 4187.98 697.35 1551.83 ...

3. Shell Programming and Scripting

Matching multiple fields from two files and then some?

Hi, I am working with two tab-delimited files with multiple columns, formatted as follows: File 1: >chrom 1 100 A G 20 …(10 columns) >chrom 1 104 G C 18 …(10 columns) >chrom 2 28 T C ...

4. UNIX for Advanced & Expert Users

awk print all fields except matching regex

grep -v will exclude matching lines, but I want something that will print all lines but exclude a matching field. The pattern that I want excluded is '/mnt/svn' If there is a better solution than awk I am happy to hear about it, but I would like to see this done in awk as well. I know I can...

5. Shell Programming and Scripting

How to merge two or more fields from two different files where there is non matching column?

Hi, Please excuse for often requesting queries and making R&D, I am trying to work out a possibility where i have two files field separated by pipe and another file containing only one field where there is no matching columns, Could you please advise how to merge two files. $more...

6. Shell Programming and Scripting

awk to combine all matching fields in input but only print line with largest value in specific field

In the below I am trying to use awk to match all the $13 values in input, which is tab-delimited, that are in $1 of gene which is just a single column of text. However only the line with the greatest $9 value in input needs to be printed. So in the example below all the MECP2 and LTBP1...

7. UNIX for Beginners Questions & Answers

Awk: matching multiple fields between 2 files

Hi, I have 2 tab-delimited input files as follows. file1.tab: green A apple red B apple file2.tab: apple - A;Z Objective: Return $1 of file1 if, . $1 of file2 matches $3 of file1 and, . any single element (separated by ";") in $3 of file2 is present in $2 of file1 In order to...

8. Shell Programming and Scripting

awk to print fields that match using conditions and a default value for non-matching in two files

Trying to use awk to match the contents of each line in file1 with $5 in file2. Both files are tab-delimited and there may be a space or special character in the name being matched in file2, for example in file1 the name is BRCA1 but in file2 the name is BRCA 1 or in file1 name is BCR but in file2...

9. UNIX for Beginners Questions & Answers

Matching fields between two files, repeated records

In two previous posts (here) and (here), I received help from forum members comparing multiple fields across two files and selectively printing portions of each as output based upon would-be matches using awk. I had been fairly comfortable populating awk arrays with fields and using awk's special...

10. Shell Programming and Scripting

Comparing two files by two matching fields

Long time listener first time poster. Hope someone can advise. I have two files, 1000+ lines in each, two fields in each file. After performing a sort, what is the best way to find exact matches where field $1 and $2 in file1 are also present in file2 on the same line, then output only those...

LEARN ABOUT CENTOS

join

JOIN(1) 							   User Commands							   JOIN(1)

NAME

       join - join lines of two files on a common field

SYNOPSIS

       join [OPTION]... FILE1 FILE2

DESCRIPTION

       For  each  pair of input lines with identical join fields, write a line to standard output.  The default join field is the first, delimited
       by whitespace.  When FILE1 or FILE2 (not both) is -, read standard input.

       -a FILENUM
	      also print unpairable lines from file FILENUM, where FILENUM is 1 or 2, corresponding to FILE1 or FILE2

       -e EMPTY
	      replace missing input fields with EMPTY

       -i, --ignore-case
	      ignore differences in case when comparing fields

       -j FIELD
	      equivalent to '-1 FIELD -2 FIELD'

       -o FORMAT
	      obey FORMAT while constructing output line

       -t CHAR
	      use CHAR as input and output field separator

       -v FILENUM
	      like -a FILENUM, but suppress joined output lines

       -1 FIELD
	      join on this FIELD of file 1

       -2 FIELD
	      join on this FIELD of file 2

       --check-order
	      check that the input is correctly sorted, even if all input lines are pairable

       --nocheck-order
	      do not check that the input is correctly sorted

       --header
	      treat the first line in each file as field headers, print them without trying to pair them

       -z, --zero-terminated
	      end lines with 0 byte, not newline

       --help display this help and exit

       --version
	      output version information and exit

       Unless -t CHAR is given, leading blanks separate fields and are ignored, else fields are separated by CHAR.  Any FIELD is  a  field  number
       counted	from 1.  FORMAT is one or more comma or blank separated specifications, each being 'FILENUM.FIELD' or '0'.  Default FORMAT outputs
       the join field, the remaining fields from FILE1, the remaining fields from FILE2, all separated by CHAR.  If FORMAT is the keyword  'auto',
       then the first line of each file determines the number of fields output for each line.

       Important:  FILE1  and  FILE2 must be sorted on the join fields.  E.g., use "sort -k 1b,1" if 'join' has no options, or use "join -t ''" if
       'sort' has no options.  Note, comparisons honor the rules specified by 'LC_COLLATE'.  If the input is not sorted and some lines	cannot	be
       joined, a warning message will be given.

       GNU coreutils online help: <http://www.gnu.org/software/coreutils/> Report join translation bugs to <http://translationproject.org/team/>

AUTHOR

       Written by Mike Haertel.

COPYRIGHT

       Copyright (C) 2013 Free Software Foundation, Inc.  License GPLv3+: GNU GPL version 3 or later <http://gnu.org/licenses/gpl.html>.
       This is free software: you are free to change and redistribute it.  There is NO WARRANTY, to the extent permitted by law.

SEE ALSO

       comm(1), uniq(1)

       The  full documentation for join is maintained as a Texinfo manual.  If the info and join programs are properly installed at your site, the
       command

	      info coreutils 'join invocation'

       should give you access to the complete manual.

GNU coreutils 8.22						     June 2014								   JOIN(1)