Find common numbers from two very large files using awk or the like Post: 302799629

9 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

Get un common numbers from two files

Hi, I have two files: abc : 50040 123123 31703 cde: 104 97 50040 123123 31703 36609 50534

2. Shell Programming and Scripting

To find all common lines from 'n' no. of files

Hi, I have one situation. I have some 6-7 no. of files in one directory & I have to extract all the lines which exist in all these files. means I need to extract all common lines from all these files & put them in a separate file. Please help. I know it could be done with the help of...

3. UNIX for Dummies Questions & Answers

Grep alternative to handle large numbers of files

I am looking for a file with 'MCR0000000716214' in it. I tried the following command: grep MCR0000000716214 * The problem is that the folder I am searching in has over 87000 files and I am getting the following: bash: /bin/grep: Arg list too long Is there any command I can use that can...

4. Shell Programming and Scripting

Drop common lines at head/tail of a large set of files

Hi! I have a large set of pairs of text files (each pair in their own subdirectory) and each pair shares head/tail (a couple of first and last lines) but differs in the middle part. I need to delete the heads/tails and keep only the middle portions in which they differ. The lengths of heads/tails...

5. UNIX for Advanced & Expert Users

Find common Strings in two large files

Hi , I have a text file in the format DB2: DB2: WB: WB: WB: WB: and a second text file of the format Time=00:00:00.473 Time=00:00:00.436 Time=00:00:00.016 Time=00:00:00.027 Time=00:00:00.471 Time=00:00:00.436 the last string in both the text files is of the...

6. Shell Programming and Scripting

finding common numbers (contents) across 2 or 3 files

I have 3 files which are tab delimited and have numbers in it. file 1 1 2 3 4 5 6 7 File 2 3 5 7 8 File 3 1

7. Shell Programming and Scripting

Find common numbers and print yes or no

Hi I have 2 files with following data First file, sp|Q676U5|A16L1_HUMAN, Autophagy-related protein 16-1 OS=Homo sapiens GN=ATG16L1 PE=1 SV=2, Maximum coiled-coil residue probability: 0.657 in position 163. Maximum dimeric residue probability: 0.288 in position 163. ...

8. Shell Programming and Scripting

Find Common Values Across Two Files

9. Shell Programming and Scripting

Find common files between two directories

I have two directories Dir 1 /home/sid/release1 Dir 2 /home/sid/release2 I want to find the common files between the two directories Dir 1 files /home/sid/release1>ls -lrt total 16 -rw-r--r-- 1 sid cool 0 Jun 19 12:53 File123 -rw-r--r-- 1 sid cool 0 Jun 19 12:53...

LEARN ABOUT CENTOS

uniq

UNIQ(1) 							   User Commands							   UNIQ(1)

NAME

       uniq - report or omit repeated lines

SYNOPSIS

       uniq [OPTION]... [INPUT [OUTPUT]]

DESCRIPTION

       Filter adjacent matching lines from INPUT (or standard input), writing to OUTPUT (or standard output).

       With no options, matching lines are merged to the first occurrence.

       Mandatory arguments to long options are mandatory for short options too.

       -c, --count
	      prefix lines by the number of occurrences

       -d, --repeated
	      only print duplicate lines, one for each group

       -D, --all-repeated[=METHOD]
	      print all duplicate lines groups can be delimited with an empty line METHOD={none(default),prepend,separate}

       -f, --skip-fields=N
	      avoid comparing the first N fields

       --group[=METHOD]
	      show all items, separating groups with an empty line METHOD={separate(default),prepend,append,both}

       -i, --ignore-case
	      ignore differences in case when comparing

       -s, --skip-chars=N
	      avoid comparing the first N characters

       -u, --unique
	      only print unique lines

       -z, --zero-terminated
	      end lines with 0 byte, not newline

       -w, --check-chars=N
	      compare no more than N characters in lines

       --help display this help and exit

       --version
	      output version information and exit

       A field is a run of blanks (usually spaces and/or TABs), then non-blank characters.  Fields are skipped before chars.

       Note:  'uniq'  does  not  detect  repeated  lines unless they are adjacent.  You may want to sort the input first, or use 'sort -u' without
       'uniq'.	Also, comparisons honor the rules specified by 'LC_COLLATE'.

       GNU coreutils online help: <http://www.gnu.org/software/coreutils/> Report uniq translation bugs to <http://translationproject.org/team/>

AUTHOR

       Written by Richard M. Stallman and David MacKenzie.

COPYRIGHT

       Copyright (C) 2013 Free Software Foundation, Inc.  License GPLv3+: GNU GPL version 3 or later <http://gnu.org/licenses/gpl.html>.
       This is free software: you are free to change and redistribute it.  There is NO WARRANTY, to the extent permitted by law.

SEE ALSO

       comm(1), join(1), sort(1)

       The full documentation for uniq is maintained as a Texinfo manual.  If the info and uniq programs are properly installed at your site,  the
       command

	      info coreutils 'uniq invocation'

       should give you access to the complete manual.

GNU coreutils 8.22						     June 2014								   UNIQ(1)