I have 3 column csv files with ~25 million rows and considerable redundancy.
Brief description of data: columns 1 and 2 contain variables, column 3 contains their correlation. For simplicity and discussion, let's use this example:
I would like to remove the redundant variable combinations with a script (i.e. the correlation of "a" with "b" is the same as the correlation of "b" with "a". I typically use awk for manipulating csv files, but am open to all suggestions.
Currently, I create a new csv where columns 1 and 2 are flipped, cat it and the original file, then remove duplicates using
However, the inefficiency of this method is problematic with a 25 million row csv file that is then doubled. Does anyone have any suggestions for doing this more cleverly? Something that would check for the existence of a combination of variables, perhaps?
Thanks in advance.
Last edited by R3353; 07-10-2012 at 05:08 PM..
Reason: added CODE tags
I am writing a shell script.
Now i need to read in a string and send it to an awk file to compare and search for compatible record.
I wrote it like tat:
read serial | awk -f generate.awk data.dat
p/s: the data file got 6 field.
According to an expert, we can write it like tat:
read... (1 Reply)
i'm trying to pass a numerical argument with function xyz to print specfic lines of filename, but my 'awk' syntax is incorrect.
ie
xyx 3 (prints the 3rd line, separated by ':' of filename)
function xyz() {
arg1=$1
cat filename | awk -F: -v x=$arg1 '{print $x}'
}
any ideas? (4 Replies)
Hello,
I have a file like
was123##abcdefg abddef
was123##xuzaghg agdfgg
was133##CGHAKS DKGJG
from the file i need to print the line after ## where the serach value is passed by an env variable called luster (which is currently set to was123):
i tried using the below code but it... (7 Replies)
Dear Conerned,
I am facing a situation where i need to pass an argument which is non-awk variable like
day=090319
awk '/TID:R/ && /TTIN:/' transaction.log
I want to add this day variable like below
awk '/TID:R$day/ && /TTIN:/' transaction.log
But it is not working. :confused: (1 Reply)
Hi all
I have got a file digits.data containing the following data
1 3 4
2 4 9
7 3 1
7 3 10
I am writing a script that will pass an argument from C-shell to nawk command. But it seems the values in the nawk comman does not get set. the program does not print no values out. Here is the... (1 Reply)
I'm trying to figure out what's getting passed as the argument when I try to pass a directory as an argument, and I'm getting incredibly strange behavior. For example, from the command line I'm typing:
nawk -f ./test.awk ~
test.awk contains the following:
{
directory = $NF
print... (13 Replies)
I have the awk script below and things go wrong when I do
awk -v dsrmx=25 -f ./checkSRDry.awk --usage
I basically want to override the usual --usage and --help that awk gives.
How do people usually handle this situation when you also want to supply your own usage and help
concerning the... (2 Replies)
Hi,
I need to check whether a particular file exists ot not using awk.
Can anyone help me please?
For Example:script that i am using:
awk '{filename =$NF;
rc=(system("test -r filename")) print $rc;}' "$1"
is not working.
Here I am passing a text file as input whose last word contains a... (6 Replies)
Hello,
I want to execute remote command with ssh.
For exemple, i have a variable
SERVERS=lpar1,lpar2,lpar3
I want to execute some commands like:
ssh -q lpar1 ls /
ssh -q lpar2 ls /
ssh -q lpar3 ls /
Can you help me with awk command ?
Thank you :) (6 Replies)
Discussion started by: khalidou13
6 Replies
LEARN ABOUT REDHAT
csv
csv(n) CSV processing csv(n)
NAME
csv - Procedures to handle CSV data.
SYNOPSIS
package require Tcl 8.3
package require csv ?0.3?
::csv::join values {sepChar ,}
::csv::joinlist values {sepChar ,}
::csv::read2matrix chan m {sepChar ,} {expand none}
::csv::read2queue chan q {sepChar ,}
::csv::report cmd matrix ?chan?
::csv::split line {sepChar ,}
::csv::split2matrix m line {sepChar ,} {expand none}
::csv::split2queue q line {sepChar ,}
::csv::writematrix m chan {sepChar ,}
::csv::writequeue q chan {sepChar ,}
DESCRIPTION
The csv package provides commands to manipulate information in CSV FORMAT (CSV = Comma Separated Values).
COMMANDS
The following commands are available:
::csv::join values {sepChar ,}
Takes a list of values and returns a string in CSV format containing these values. The separator character can be defined by the
caller, but this is optional. The default is ",".
::csv::joinlist values {sepChar ,}
Takes a list of lists of values and returns a string in CSV format containing these values. The separator character can be defined
by the caller, but this is optional. The default is ",". Each element of the outer list is considered a record, these are separated
by newlines in the result. The elements of each record are formatted as usual (via ::csv::join).
::csv::read2matrix chan m {sepChar ,} {expand none}
A wrapper around ::csv::split2matrix (see below) reading CSV-formatted lines from the specified channel (until EOF) and adding them
to the given matrix. For an explanation of the expand argument see ::csv::split2matrix.
::csv::read2queue chan q {sepChar ,}
A wrapper around ::csv::split2queue (see below) reading CSV-formatted lines from the specified channel (until EOF) and adding them
to the given queue.
::csv::report cmd matrix ?chan?
A report command which can be used by the matrix methods format 2string and format 2chan. For the latter this command delegates the
work to ::csv::writematrix. cmd is expected to be either printmatrix or printmatrix2channel. The channel argument, chan, has to be
present for the latter and must not be present for the first.
::csv::split line {sepChar ,}
converts a line in CSV format into a list of the values contained in the line. The character used to separate the values from each
other can be defined by the caller, via sepChar, but this is optional. The default is ",".
::csv::split2matrix m line {sepChar ,} {expand none}
The same as ::csv::split, but appends the resulting list as a new row to the matrix m, using the method add row. The expansion mode
specified via expand determines how the command handles a matrix with less columns than contained in line. The allowed modes are:
none This is the default mode. In this mode it is the responsibility of the caller to ensure that the matrix has enough columns to
contain the full line. If there are not enough columns the list of values is silently truncated at the end to fit.
empty In this mode the command expands an empty matrix to hold all columns of the specified line, but goes no further. The overall
effect is that the first of a series of lines determines the number of columns in the matrix and all following lines are
truncated to that size, as if mode none was set.
auto In this mode the command expands the matrix as needed to hold all columns contained in line. The overall effect is that after
adding a series of lines the matrix will have enough columns to hold all columns of the longest line encountered so far.
::csv::split2queue q line {sepChar ,}
The same as ::csv::split, but appending the resulting list as a single item to the queue q, using the method put.
::csv::writematrix m chan {sepChar ,}
A wrapper around ::csv::join taking all rows in the matrix m and writing them CSV formatted into the channel chan.
::csv::writequeue q chan {sepChar ,}
A wrapper around ::csv::join taking all items in the queue q (assumes that they are lists) and writing them CSV formatted into the
channel chan.
FORMAT
Each record of a csv file (comma-separated values, as exported e.g. by Excel) is a set of ASCII values separated by ",". For other lan-
guages it may be ";" however, although this is not important for this case (The functions provided here allow any separator character).
If a value contains itself the separator ",", then it (the value) is put between "".
If a value contains ", it is replaced by "".
EXAMPLE
The record
123,"123,521.2","Mary says ""Hello, I am Mary"""
is parsed as follows:
a) 123
b) 123,521.2
c) Mary says "Hello, I am Mary"
SEE ALSO
matrix, queue
KEYWORDS
csv, matrix, queue, package, tcllib
csv 0.3 csv(n)