Running uniq -c and sort -nr in parallel.

10-08-2011

Registered User

190, 1

Join Date: Jan 2011

Last Activity: 11 October 2017, 1:35 PM EDT

Location: Nowhere

Posts: 190

Thanks Given: 227

Thanked 1 Time in 1 Post

Running uniq -c and sort -nr in parallel.

Hi All,

I have a huge collection of files in a directory about 200000. I have the command below but it only uses one core of the computer. I want it to do task in parallel.

This is the command that I want to run in parallel:

Code:

sort testfile | uniq -c | sort -nr

I know how to run sort command in parallel:

Code:

cat testfile | parallel --pipe --files sort | parallel -Xj1 sort -m {} ';' rm {} >testfile.sort

Can anyone please help me so that I can run

Code:

sort testfile | uniq -c | sort -nr

command in parallel? I am using Linux with GNU parallel installed.

shoaibjameel123

View Public Profile for shoaibjameel123

Find all posts by shoaibjameel123

10-08-2011

Registered User

3,231, 978

Join Date: Dec 2009

Last Activity: 11 June 2014, 8:40 PM EDT

Posts: 3,231

Thanks Given: 179

Thanked 978 Times in 791 Posts

Quote:

Originally Posted by shoaibjameel123

Can anyone please help me so that I can run

Code:

sort testfile | uniq -c | sort -nr

command in parallel? I am using Linux with GNU parallel installed.

You cannot parallelize that pipeline. Think about it.

sort testfile has to read and process the entire contents of 'testfile' before it can even output a single line. Otherwise, it's possible that a line not yet read from 'testfile' would need to precede an already printed line.

The same goes for sort -nr.

In theory, the only aspect of this pipeline that can be parallelized is uniq -c. However, for that to work, The first sort's output would need to be carefully chopped at the boundaries between sequences of identical lines. Also, the output from the tasks handling the count would have to be recombined before being fed to the numeric sort. Furthermore, you cannot count on this working with pipes if line lengths may exceed PIPE_BUF (this can lead to the interweaving of non-atomic write()s.). Temporary files would be required and some mechanism to know when a temp file is complete (renaming, moving, locking, etc).

In short, you cannot do better than the original pipeline. If you have multiple cores, your system is probably running the processes on different cores. The reason you're probably only seeing one of them utilized is because a sort is running and hasn't yet completed. While that's happening, everything downstream must sleep.

Regards,
Alister

Last edited by alister; 10-08-2011 at 10:08 AM..

alister

View Public Profile for alister

Find all posts by alister

10-08-2011

Registered User

190, 1

Join Date: Jan 2011

Last Activity: 11 October 2017, 1:35 PM EDT

Location: Nowhere

Posts: 190

Thanks Given: 227

Thanked 1 Time in 1 Post

got it...

this is how I can do it:

Code:

cat filenmame | parallel --pipe --files sort | parallel -Xj1 sort -m {} ';' rm {} | parallel --pipe uniq -c

---------- Post updated at 08:54 PM ---------- Previous update was at 08:52 PM ----------

yes, you are right. I got what you mean. Certainly that pipe cannot be paralleled but I have tried to do the best I can with the new command.

shoaibjameel123

View Public Profile for shoaibjameel123

Find all posts by shoaibjameel123

10-08-2011

Registered User

3,231, 978

Join Date: Dec 2009

Last Activity: 11 June 2014, 8:40 PM EDT

Posts: 3,231

Thanks Given: 179

Thanked 978 Times in 791 Posts

Hmmm. I may have leapt before I looked, since I don't know anything about GNU Parallel.

I will have to study that tool's documentation and your solution. Thank you for giving me something interesting to do while a storm runs it course outside.

Regards,
Alister

This User Gave Thanks to alister For This Post:

alister

View Public Profile for alister

Find all posts by alister

10-08-2011

Registered User

190, 1

Join Date: Jan 2011

Last Activity: 11 October 2017, 1:35 PM EDT

Location: Nowhere

Posts: 190

Thanks Given: 227

Thanked 1 Time in 1 Post

My idea is simple. You were right, it cannot be parallelized but individual sorts and uniq's can be made to do tasks in parallel using multiple cores.

shoaibjameel123

View Public Profile for shoaibjameel123

Find all posts by shoaibjameel123

10-08-2011

Registered User

3,231, 978

Join Date: Dec 2009

Last Activity: 11 June 2014, 8:40 PM EDT

Posts: 3,231

Thanks Given: 179

Thanked 978 Times in 791 Posts

Quote:

Originally Posted by shoaibjameel123

got it...

this is how I can do it:

Code:

cat filenmame | parallel --pipe --files sort | parallel -Xj1 sort -m {} ';' rm {} | parallel --pipe uniq -c

I skimmed GNU Parallel's documentation, focusing on the options you've used.

Code:

parallel --pipe --files sort

As I understand it, this will spawn one instance of sort per cpu. parallel will read 1 megabyte of data (give or take the length of a line) from its standard input and feed those chunks to alternating sorts. parallel writes the sorted data to a file and writes that filename to its standard output.

Nifty. Saves us the trouble of manually splitting the original file, spawning multiple sorts, and managing temp files.

Code:

parallel -Xj1 sort -m {} ';' rm {}

Not as nifty. Runs one instance of sort to merge the sorted files whose names were generated by the previous parallel. Then deletes those files.

We can accomplished that with plain xargs: xargs sh -c 'sort -m "$@" && rm "$@"' sh

DANGER AHEAD!!!

Code:

parallel --pipe uniq -c

This component is broken (even if it usually gives the correct result). parallel's --pipe will by default decompose the input into line-oriented chunks of approximately 1 megabyte in size. It has no knowledge of the contents of that input. If a sequence of identical lines spans more than one chunk, the output will show multiple, consecutive counts for the same line (whose values should sum to the correct value), because multiple instances of uniq see those identical lines.

Without a smarter way to distribute the data, you'll have to use a single instance of uniq to achieve reliably correct results.

I did not test as the documentation seemed sufficiently clear on the workings of --pipe and I don't have parallel installed on any machine.

If my analysis is incorrect, I look forward to learning some more.

Regards,
Alister

Last edited by alister; 10-08-2011 at 01:22 PM..

This User Gave Thanks to alister For This Post:

alister

View Public Profile for alister

Find all posts by alister

10-08-2011

Registered User

190, 1

Join Date: Jan 2011

Last Activity: 11 October 2017, 1:35 PM EDT

Location: Nowhere

Posts: 190

Thanks Given: 227

Thanked 1 Time in 1 Post

I've tested uniq -c with parallel. As of now it does give correct results (with some test files I've generated on my own). But if its not reliable I am removing that.

shoaibjameel123

View Public Profile for shoaibjameel123

Find all posts by shoaibjameel123

Shell Programming and Scripting

Running uniq -c and sort -nr in parallel.

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Sort & Uniq -u

Discussion started by: Antony Ankrose

2. UNIX for Dummies Questions & Answers

Uniq and sort -u

Discussion started by: senhia83

3. Shell Programming and Scripting

Uniq or sort -u or similar only between { }

Discussion started by: fugitivus

4. Shell Programming and Scripting

Sort field and uniq

Discussion started by: sabercats

5. Shell Programming and Scripting

Sort and uniq after comparision

Discussion started by: nua7

6. Shell Programming and Scripting

sort | uniq question

Discussion started by: palex

7. Shell Programming and Scripting

Help with Uniq and sort

Discussion started by: pinnacle

8. Shell Programming and Scripting

sort and uniq in perl

Discussion started by: reggiej

9. UNIX for Dummies Questions & Answers

Help with Last,uniq, sort and cut

Discussion started by: jay1228

10. UNIX for Dummies Questions & Answers

sort/uniq

Discussion started by: jimmyflip