04-24-2015
Dear Rudic
I did not bring the whole data set of file1 and file2 here. That's what you think these two file do not match together. The question is file2 have unique number of id in the second column but in file1 there could be repeated time of ids in column first. I want to know why after merging these two file, I do not get the same line in file1 which it had before.
10 More Discussions You Might Find Interesting
1. UNIX for Dummies Questions & Answers
Hi,
I have a big file of 50GB size. I need copy it to a second ftp from a ftp. I am not able to do the full 50GB transfer as it timesout after some time. SO i am trying to split the file into 5gb each 10 files with the below command.
split -b 5368709120 pack.tar.gz backup.gz
After I... (2 Replies)
Discussion started by: venu_nbk
2 Replies
2. UNIX for Dummies Questions & Answers
Hello,
My apologies if this has been posted elsewhere, I have had a look at several threads but I am still confused how to use these functions. I have two files, each with 5 columns:
File A: (tab-delimited)
PDB CHAIN Start End Fragment
1avq A 171 176 awyfan
1avq A 172 177 wyfany
1c7k A 2 7... (3 Replies)
Discussion started by: InfoSeeker
3 Replies
3. Shell Programming and Scripting
I am trying to join a few hundred files using join. Is there a way to use while read or something else to automate this. My problem is the following.
Day 1
City Temp
ABC 20
DEF 30
HIJ 15
Day 2
City Temp
ABC 22
DEF 29
KLM 5
Day 3 (3 Replies)
Discussion started by: theFinn
3 Replies
4. Shell Programming and Scripting
Hi all,
I searched through the forum but i can't manage to find a solution. I need to join a set of files placed in a directory (~1600) by column, and obtain an output with first and second column common to each file, but following columns are taken from the file in the list (precisely the fourth... (10 Replies)
Discussion started by: macsx82
10 Replies
5. Shell Programming and Scripting
Is it possible to join all the files with input1 based on 1st column?
input1
a
b
c
d
e
f
input2
a
b
input3
a
e
input4
c (2 Replies)
Discussion started by: quincyjones
2 Replies
6. UNIX for Dummies Questions & Answers
Hi,
I have 20 tab delimited text files that have a common column (column 1). The files are named GSM1.txt through GSM20.txt. Each file has 3 columns (2 other columns in addition to the first common column).
I want to write a script to join the files by the first common column so that in the... (5 Replies)
Discussion started by: evelibertine
5 Replies
7. Shell Programming and Scripting
Hi there,
I am trying to join 24 files (i showed example of 3 files below). They all have 2 columns. The first columns is common to all. The files are tab delimited eg
file 1
rs0001 100e-34
rs0003 2.8e-01
rs008 1.9e-90
file 2
rs0001 1.98e-22
rs0004 3.77e-10... (4 Replies)
Discussion started by: fat
4 Replies
8. Shell Programming and Scripting
Please help, I want to join multiple files based on column 1, and put the missing values as 0. Also the colname in the output should say which file the values came from.
FILE1
1 11
2 12
3 13
FILE2
2 22
3 23
4 24
FILE3
1 31
3 33
4 34
FILE1 FILE2 FILE3
1 11 0 31 (1 Reply)
Discussion started by: newbie83
1 Replies
9. Shell Programming and Scripting
I have 2 files namely branch.txt file & RXD.txt file as below
Ex:Branch.txt
=========================
B1,Branchname1,city,country
B2,Branchname2,city,country
B3,Branchname3,city,country
B4,Branchname4,city,country
B5,Branchname5,city,country
RXD file : will... (11 Replies)
Discussion started by: satece
11 Replies
10. Shell Programming and Scripting
Hello all,
I want to join 2 tabbed files on the first 2 fields, and filling the missing values with 0. The 3rd column in each file is constant for the entire file.
file1
12658699 ST5 XX2720 0 1 0 1
53039541 ST5 XX2720 1 0 1.5 1
file2 ... (6 Replies)
Discussion started by: sheetalk
6 Replies
comm(1) General Commands Manual comm(1)
NAME
comm - Compares two sorted files.
SYNOPSIS
comm [-123] file1 file2
STANDARDS
Interfaces documented on this reference page conform to industry standards as follows:
command: XCU5.0
Refer to the standards(5) reference page for more information about industry standards and associated tags.
OPTIONS
Suppresses output of the first column (lines in file1 only). Suppresses output of the second column (lines in file2 only). Suppresses
output of the third column (lines common to file1 and file2).
The command comm -123 produces no output.
OPERANDS
A pathname of the first file to be compared. If file1 is a hyphen (-), the standard input is used. A pathname of the second file to be
compared. If file2 is a hyphen (-), the standard input is used.
If both file1 and file2 refer to standard input or to the same FIFO special, block special or character special file, the results are unde-
fined.
DESCRIPTION
The comm command reads file1 and file2 and writes three columns to standard output, showing which lines are common to the files and which
are unique to each.
The leftmost column of standard output includes lines that are in file1 only. The middle column includes lines that are in file2 only.
The rightmost column includes lines that are in both file1 and file2.
If you specify a hyphen (-) in place of one of the file names, comm reads standard input.
Generally, file1 and file2 should be sorted according to the collating sequence specified by the LC_COLLATE environment variable. (See
sort(1).) If the input files are not sorted properly, the output of comm might not be useful.
EXIT STATUS
Successful completion. Error occurred.
EXAMPLES
In the following examples, file1 contains the following sorted list of North American cities:
Anaheim Baltimore Boston Chicago Cleveland Dallas Detroit Kansas City Milwaukee Minneapolis New York Oakland Seattle Toronto
The second file, file2, contains this sorted list:
Atlanta Chicago Cincinnati Houston Los Angeles Montreal New York Philadelphia Pittsburgh San Diego San Francisco St. Louis
To display the lines unique to each file and common to the two files, enter: comm file1 file2
This command results in the following output: Anaheim Atlanta Baltimore Boston Chicago Cincinnati Cleveland Dal-
las Detroit Houston Kansas City Los Angeles Milwaukee Minneapolis Montreal New York Oakland Philadel-
phia Pittsburgh San Diego San Francisco Seattle St. Louis Toronto
The leftmost column contains lines in file1 only, the middle column contains lines in file2 only, and the rightmost column contains
lines common to both files. To display any one or two of the three output columns, include the appropriate flags to suppress the
columns you do not want. For example, the following command displays columns 1 and 2 only: comm -3 file1 file2
Anaheim
Atlanta Baltimore Boston
Cincinnati Cleveland Dallas Detroit
Houston Kansas City
Los Angeles Milwaukee Minneapolis
Montreal Oakland
Philadelphia
Pittsburgh
San Diego
San Francisco Seattle
St. Louis Toronto
The following command displays output from only the second column: comm -13 file1 file2
Atlanta Cincinnati Houston Los Angeles Montreal Philadelphia Pittsburgh San Diego San Francisco St. Louis
The following command displays output from only the third column: comm -12 file1 file2
Chicago New York
SEE ALSO
Commands: cmp(1), diff(1), sdiff(1), sort(1), uniq(1)
comm(1)