01-10-2010
compare 2 very large lists of different length
I have two very large datasets (>100MB) in a simple vertical list format. They are of different size and with different order and formatting (e.g. whitespace and some other minor cruft that would thwart easy regex).
Let's call them set1 and set2.
I want to check set2 to see if it contains any of the data entries in set1. I think of this as individual greps of set2 using each line of set1.
(NB- I could, with some work, manipulate the sets to make the order and formatting the same.)
In your opinion, what is the best tool to use for this search of set2 using the data in set1?
- comm?
- a looping shell script, or xargs, that calls grep?
- grep -f?
- diff?
- combine the sets (after making format the same) then sort and print only duplicate lines? uniq -d, sed or awk
10 More Discussions You Might Find Interesting
1. Shell Programming and Scripting
If I had a list of numbers in two different files, what would be the fastest and easiest way to find out which numbers in list B are not in list A without reading each number in list B one at a time and using grep thousands of times against list A?
I have two very long lists of numbers and the... (4 Replies)
Discussion started by: keelba
4 Replies
2. UNIX for Dummies Questions & Answers
Hi ,
I have a peculiar case, where my sed command is working on a file which contains lines of small length.
sed "s/XYZ:1/XYZ:3/g" abc.txt > xyz.txt
when abc.txt contains lines of small length(currently around 80 chars) , this sed command is working fine.
when abc.txt contains lines of... (3 Replies)
Discussion started by: thanuman
3 Replies
3. UNIX for Dummies Questions & Answers
hello all,
I wonder if anybody might be able to help with this.
I have file 1 and file2.
Both files may contain thousands of lines that have variable contents.
file1
234GH
5234BTW
89er
678tfg
234
234YT
tfg456
wert
78gt
gh23444 (7 Replies)
Discussion started by: Garrred
7 Replies
4. Shell Programming and Scripting
The following bash script does not work because the java/groovy code always thinks there are four arguments even if there are only 1 or 2. As you can see from my hideous backslashes, I am using cygwin bash on windows.
export... (1 Reply)
Discussion started by: siegfried
1 Replies
5. Programming
Hi.
I am trying to write a Python programme that compares two different text files which both contain a list of words. Each word has its own line
worda
wordb
wordc
I want to compare textfile 2 with textfile 1, and if there's a word in textfile 2 that is NOT in textfile 1, I want to... (6 Replies)
Discussion started by: Bloomy
6 Replies
6. Shell Programming and Scripting
hi,
I have 2 large lists:
LIST A: containes 6 fields of many entries (VARIABLE number), like:
2011-07-10 | 18:19:47 | 38037300 | 9647808003122 | 2 | success
LIST B: containes 3 fields & 183 entries (FIXED number), like:
9647805651885 9647805651885 SCP_10
What I want is a... (8 Replies)
Discussion started by: amurib
8 Replies
7. Shell Programming and Scripting
Hi,
I do little bash scripting so sorry for my ignorance.
How do I compare if the two variable not match and if they do not match run a command.
I was thinking a for loop but then I need another for loop for the 2nd list and I do not think that would work as in the real world there could... (2 Replies)
Discussion started by: GermanJulian
2 Replies
8. Shell Programming and Scripting
Hi everybody!
I'm trying to delete some elements from a list with two elements on each row agreeing with the elements in another list. Pratically I want a perl script able to take each element of the second list (that is a single column list), compare it with both elements of each row from the... (3 Replies)
Discussion started by: gabrysfe
3 Replies
9. Shell Programming and Scripting
I have two files A and B listing ip addresses
and all the ip addresses in B are in A, and A includes other ip addresses
now I want to get the list of the ip addresses that are in A but not in B
how to achieve this? thanks (1 Reply)
Discussion started by: esolvepolito
1 Replies
10. Homework & Coursework Questions
Hello,
I'm new to the python programming, and I have a question.
I have to write a program that prints a receipt for a restaurant. The input is a list which looks like:
product1
product3
product8
....
In the other input file there is a list which looks like:
product1 coffee 5,00... (1 Reply)
Discussion started by: dagendy
1 Replies
LEARN ABOUT DEBIAN
paraview
PARAVIEW(1) General Commands Manual PARAVIEW(1)
NAME
Paraview - Rendering and displaying program for small and large, three dimensional datasets.
SYNOPSIS
paraview [-cc | --cave-configuration FILE] [--compare-view OPT] [--connect-id ID] [--data DATA] [--data-directory DIR] [-dr | --disable-
registry] [--exit] [/? | --help] [--image-threshold THRESH] [-m | --machines FILE] [--run-test CASE] [--run-test-init CASE] [-s | --server
NAME] [--stereo] [--test-directory DIR] [-V | --version]
DESCRIPTION
Paraview is a program for displaying and rendering of small to large datasets in two or three dimensions. It runs on a single computer or
on a cluster of nodes with distributed or shared memory. The Visualization Toolkit (VTK) is used as processing and rendering machine.
Options
-cc | --cave-configuration FILE
Specify the file FILE that defines the displays for a cave. It is used only with CaveRenderModule.
--compare-view OPT
Compare the viewport to a reference image, and exit.
--connect-id ID
Set the ID of the server and client to make sure they match.
--data DATA
Load the specified DATA.
--data-directory DIR
Set the data DIR directory where test-case data are.
-dr | --disable-registry
Do not use registry when running ParaView (for testing).
--exit Exit application when testing is done. Use for testing.
/? | --help
Displays available command line arguments.
--image-threshold THRESH
Set the threshold THRESH beyond which viewport-image comparisons fail.
-m | --machines FILE
Specify the network configurations file FILE for the render server.
--run-test CASE
Run a recorded test case CASE.
--run-test-init CASE
Run a recorded test initialization case CASE.
-s | --server NAME
Set the name NAME of the server resource to connect with when the client starts.
--stereo
Tell the application to enable stereo rendering (only when running on a single process).
--test-directory DIR
Set the temporary directory DIR where test-case output will be stored.
-V | --version
Give the version number and exit.
SEE ALSO
On-line documentation from the program help menu, wiki pages at http://paraview.org/Wiki/ParaView, FAQ at http://paraview.org/Wiki/Par-
aView:FAQ and mailing list at http://public.kitware.com/mailman/listinfo/paraview
AUTHOR
Gerber van der Graaf
21 May 2008 PARAVIEW(1)