Frequency Count of chunked data Post: 302928332

Sponsored Content

Top Forums Shell Programming and Scripting Frequency Count of chunked data Post 302928332 by gimley on Thursday 11th of December 2014 04:37:13 AM

12-11-2014

Registered User

Frequency Count of chunked data

Dear all,
I have an AWK script which provides frequency of words. However I am interested in getting the frequency of chunked data. This means that I have already produced valid chunks of running text, with each chunk on a line. What I need is a script to count the frequencies of each string. A pseudo sample is provided below

Code:

this interesting event
has been going on
since years
in this country
the two actors
met
one another
in this country
Mary
met
her husband
in this country

The output would be

Code:

Mary	1
has been going on	1
her husband	1
in this country	3
met	2
one another	1
since years	1
the two actors	1
this interesting event	1

I have been able to sort the data so that all similar strings are clubbed together

Code:

Mary	
has been going on
her husband
in this country
in this country
in this country
met
met
one another
since years
the two actors
this interesting event

My question is how do I manipulate a script so that a whole line is treated as an entity and lines that match (I have come till there) can be treated as one unit and a frequency counter set up.
My awk script handles space as delimiter but I do not know how to make it recognise start of line and end of line CRLF as delimiters.
I am sure this tool will be useful to people who work with chunked big data.
Many thanks

gimley

View Public Profile for gimley

Find all posts by gimley

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Splitting Chunked-FullNames Nightmare

I've got a problem i'm hoping other more experienced programmers have had to deal with sometime in their careers and can help me: how to get fullnames that were chunked together into one field in an old database into separate more meaningful fields. I'd like to get the records that nicely fit...

2. Shell Programming and Scripting

Count field frequency in a '|' delimited file

I have a large file with fields delimited by '|', and I want to run some analysis on it. What I want to do is count how many times each field is populated, or list the frequency of population for each field. I am in a Sun OS environment. Thanks, - CB

3. Shell Programming and Scripting

Help with checking reference data frequency count

reference data GHTAS QER CC N input data NNWQERPROEGHTASTTTGHTASNCC Desired output GHTAS 2 QER 1 CC 1 N 3

4. Shell Programming and Scripting

Extracting high frequency data-lines

Hi, I have a very large log file in the following format: 198.28.0.0 - - 200 348 244.48.0.0 - - 200 211 198.28.0.0 - - 200 191 4.48.0.0 - - 200 1131 244.48.0.0 - - 200 1131 244.48.0.0 - - 200 1131 4.48.0.0 - - 200 1131 244.48.0.0 - - 200 211 4.48.0.0 - - 200 1131 ...

5. Shell Programming and Scripting

count horizontal data

dear all.. i need help i have data ID,A,B,C,D,E,F,G,H --> header 917188,4,1,2,1,4,6,3,5 --> data i want output : ID,OUT1,OUT2,OUT3 --> header 917188,3,3,2 where OUT1 is count of 1 and 2 from $2-$9 OUT2 is count of 3 and 4 from $2-$9...

6. Shell Programming and Scripting

count frequency of words in a file

I need to write a shell script "cmn" that, given an integer k, print the k most common words in descending order of frequency. Example Usage: user@ubuntu:/$ cmn 4 < example.txt :b:

7. Shell Programming and Scripting

Count column data

Hi Guys, B07 U51C A1 44 B1 44 Yes B07 L64U A2 44 B1 44 Yes B07 L62U A2 44 B1 44 Yes B07 L11C A4 32 B1 44 NO B05 L12Z A1 12 B1 44 NO B01 651Z A2 44 B1 44 NO B04 A51Z A2 12 B1 44 NO L07 B08D A4 12 B1 44 NO B07 RU8D A4 44 B1 44 Yes B07 L58D A4 15 B1 44 No B07 LA8D A4 44 B1 44 Yes B07...

8. Shell Programming and Scripting

frequency count using shell

Hello everyone, please consider the following lines of a matrix 59 32 59 32 59 32 59 32 59 32 59 32 59 32 60 32 60 33 60 33 60 33 60 33 60 33 60 33 60 33 60 33 60 33

9. Shell Programming and Scripting

Code for count the frequency of interacting pairs

Hi all, I am trying to analyze my data, and I will need your experience. I have some files with the below format: res1 = TYR res2 = ASN res1 = ASP res2 = SER res1 = TYR res2 = ASN res1 = THR res2 = LYS res1 = THR res2 = TYR etc (many lines) I am...

10. Shell Programming and Scripting

Count frequency of unique values in specific column

Hi, I have tab-deliminated data similar to the following: dot is-big 2 dot is-round 3 dot is-gray 4 cat is-big 3 hot in-summer 5 I want to count the frequency of each individual "unique" value in the 1st column. Thus, the desired output would be as follows: dot 3 cat 1 hot 1 is...

LEARN ABOUT REDHAT

tzselect

TZSELECT(8)						      System Manager's Manual						       TZSELECT(8)

NAME

       tzselect - select a time zone

SYNOPSIS

       tzselect

DESCRIPTION

       The  tzselect program asks the user for information about the current location, and outputs the resulting time zone description to standard
       output.	The output is suitable as a value for the TZ environment variable.

       All interaction with the user is done via standard input and standard error.

ENVIRONMENT VARIABLES

       AWK    Name of a Posix-compliant awk program (default: awk).

       TZDIR  Name of the directory containing time zone data files (default: /usr/local/etc/zoneinfo).

FILES

       TZDIR/iso3166.tab
	      Table of ISO 3166 2-letter country codes and country names.

       TZDIR/zone.tab
	      Table of country codes, latitude and longitude, TZ values, and descriptive comments.

       TZDIR/TZ
	      Time zone data file for time zone TZ.

EXIT STATUS

       The exit status is zero if a time zone was successfully obtained from the user, nonzero otherwise.

SEE ALSO

       newctime(3), tzfile(5), zdump(8), zic(8)

																       TZSELECT(8)

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Splitting Chunked-FullNames Nightmare

Discussion started by: RacerX

2. Shell Programming and Scripting

Count field frequency in a '|' delimited file

Discussion started by: ChicagoBlues

3. Shell Programming and Scripting

Help with checking reference data frequency count

Discussion started by: perl_beginner

4. Shell Programming and Scripting

Extracting high frequency data-lines

Discussion started by: sajal.bhatia

5. Shell Programming and Scripting

count horizontal data

Discussion started by: buncit8

6. Shell Programming and Scripting

count frequency of words in a file

Discussion started by: mohit_iitk

7. Shell Programming and Scripting

Count column data

Discussion started by: asavaliya

8. Shell Programming and Scripting

frequency count using shell

Discussion started by: xshang

9. Shell Programming and Scripting

Code for count the frequency of interacting pairs

Discussion started by: Tzole

10. Shell Programming and Scripting

Count frequency of unique values in specific column

Discussion started by: owwow14

LEARN ABOUT REDHAT

tzselect