Matching 2 files based on one column Post: 302471171

Sponsored Content

Top Forums Shell Programming and Scripting Matching 2 files based on one column Post 302471171 by Scrutinizer on Friday 12th of November 2010 06:56:57 AM

11-12-2010

Moderator

Like this?

Code:

awk 'NR==FNR{A[$1]=$0;next}{if(A[$1]){sub(/[^,]*/,"",A[$1]);$2=$2 A[$1]}else $2=$2 ",,,"}1' FS=, OFS=, file1 FS="[ \t]*" file2

---------- Post updated at 12:30 ---------- Previous update was at 12:21 ----------

Or shorter:

Code:

awk 'NR==FNR{p=$1;$1=x;A[p]=$0;next}{$2=$2(A[$1]?A[$1]:",,,")}1'  FS=, OFS=, file1 FS="[ \t]*" file2

---------- Post updated at 12:56 ---------- Previous update was at 12:30 ----------

Quote:

Originally Posted by swvanderlaan

[..]I just have some questions about the code though: can you explain the parts? I don't fully understand what each part does, and than if I'd understand I could learn maybe new commands to work my files. Smilie

NR==FNR	If we are reading the first file (The variables NR and FNR are only equal when reading the first file)
A[$1]=$2	store the second field in array a using the index of the first field
next	proceed to read the next record
A[$1]	if A[$1] exists (using the $1 of the second file)
$2=A[$1] FS $2;print	then append FS (a comma) followed by A[$1] (using the $1 of the second file)
Fs=,	set the input file seperator to ","
OFS=,	set the output file seperator to ","

Scrutinizer

View Public Profile for Scrutinizer

Find all posts by Scrutinizer

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Compare files column to column based on keys

Here is my situation. I need to compare two tab separated files (diff is not useful since there could be known difference between files). I have found similar posts , but not fully matching.I was thinking of writing a shell script using cut and grep and while loop but after going thru posts it...

2. Shell Programming and Scripting

Matching words based on column headers

Hi , Pls help on this. Input file: NAME1 BSC1 TEXT ID 1 MAINSFAIL TEXT ID 2 DGON TEXT ID 3 lOADONDG NAME2 BSC2 TEXT ID 1 DGON TEXT ID 3 lOADONG

3. UNIX for Dummies Questions & Answers

Removing Lines based on matching first column

I have a file1 that looks like this: File 1 a b b c c e d e and a file 2 that looks like this: File 2 b c e e Note that file 2 is the right hand column from file1. I want to remove any lines from file1 that begin with the column in file2. In this case the desired output...

4. Shell Programming and Scripting

awk print non matching lines based on column

My item was not answered on previous thread as code given did not work I wanted to print records from file2 where comparing column 1 and 16 for both files find rows where column 16 in file 1 does not match column 16 in file 2 Here was CODE give to issue ~/unix.com$ cat f1...

5. UNIX for Dummies Questions & Answers

How to fetch files right below based on some matching criteria?

I have a requirement where in i need to select records right below the search criteria qwertykeyboard white 10 20 30 30 40 50 60 70 80 qwertykeyboard black 40 50 60 70 90 100 qwertykeyboard and white are headers separated by a tab. when i execute my script..i would be searching...

6. Shell Programming and Scripting

Based on column in file1, find match in file2 and print matching lines

file1: file2: I need to find matches for any lines in file1 that appear in file2. Desired output is '>' plus the file1 term, followed by the line after the match in file2 (so the title is a little misleading): This is honestly beyond what I can do without spending the whole night on it, so I'm...

7. Shell Programming and Scripting

Matching two files per column

Hi, I hope somebody can help me with this problem, since I would like to solve this problem using awk, but im not experienced enough with this. I have two files which i want to match, and output the matching column name and row number. One file contains 4 columns like this: FILE1: a ...

8. Shell Programming and Scripting

Insert value of column based on file name matching

At the top of the XYZ file, I need to insert the ABC data value of column 2 only when ABC column 1 matches the prefix XYZ file name (not the ".txt"). Is there an awk solution for this? ABC Data 0101 0.54 0102 0.48 0103 1.63 XYZ File Name 0101.txt 0102.txt 0103.txt ...

9. Linux

Merge two files based on matching criteria

Hi, I am trying to merge two csv files based on matching criteria: File description is as below : Key_File : 000|��|Key_HF|��|Key_FName 001|��|Key_11|��|Sort_Key22|��|Key_31 002|��|Key_12|��|Sort_Key23|��|Key_32 003|��|Key_13|��|Sort_Key24|��|Key_33 050|��|Key_15|��|Sort_Key25|��|Key_34...

10. UNIX for Beginners Questions & Answers

Matching 2 files based on key

Hi all I have two files I need to match record from first file and second file on column 1,8 and and output only match records on file1 File1: 020059801803180116130926800002090000800231000245204003160000000002000461OUNCE000000350000100152500BM01007W0000 ...

LEARN ABOUT MINIX

join

JOIN(1) 						      General Commands Manual							   JOIN(1)

NAME

       join - relational database operator

SYNOPSIS

       join [-an] [-e s] [-o list] [-tc] file1 file2

DESCRIPTION

       Join  forms,  on the standard output, a join of the two relations specified by the lines of file1 and file2.  If file1 is `-', the standard
       input is used.

       File1 and file2 must be sorted in increasing ASCII collating sequence on the fields on which they are to be joined, normally the  first	in
       each line.

       There  is  one line in the output for each pair of lines in file1 and file2 that have identical join fields.  The output line normally con-
       sists of the common field, then the rest of the line from file1, then the rest of the line from file2.

       Fields are normally separated by blank, tab or newline.	In this case, multiple separators count as one, and leading  separators  are  dis-
       carded.

       These options are recognized:

       -an    In addition to the normal output, produce a line for each unpairable line in file n, where n is 1 or 2.

       -e s   Replace empty output fields by string s.

       -o list
	      Each output line comprises the fields specified in list, each element of which has the form n.m, where n is a file number and m is a
	      field number.

       -tc    Use character c as a separator (tab character).  Every appearance of c in a line is significant.

SEE ALSO

       sort(1), comm(1), awk(1).

BUGS

       With default field separation, the collating sequence is that of sort -b; with -t, the sequence is that of a plain sort.

       The conventions of join, sort, comm, uniq, look and awk(1) are wildly incongruous.

7th Edition							  April 29, 1985							   JOIN(1)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Compare files column to column based on keys

Discussion started by: blackjack101

2. Shell Programming and Scripting

Matching words based on column headers

Discussion started by: bha148

3. UNIX for Dummies Questions & Answers

Removing Lines based on matching first column

Discussion started by: kschiltz55

4. Shell Programming and Scripting

awk print non matching lines based on column

Discussion started by: sigh2010