Help to improve speed of text processing script

07-27-2009

Registered User

3,216, 33

Join Date: Mar 2005

Last Activity: 4 September 2020, 7:11 AM EDT

Location: classification algos

Posts: 3,216

Thanks Given: 19

Thanked 33 Times in 30 Posts

Quote:

You know the movie matrix?

Since you have remainded me of the movie 'matrix' - am all charged up to answer your question to an extent atleast

Code:

${#lines[@]}

If this is not modified, try assigning it to a variable and reuse that, instead of computing it each time.

Code:

echo "${lines[${i}]}" >> test_$count.txt
done
echo "" >> test_$count.txt

Writing to a file whilst in the loop block will greatly reduce the performance of the script, what happens for every write call is ...

Code:

open file
write data
close file

Ideally what should be done is

Code:

open file
write data
write data
.
.
.
close file

instead write it to the outer block something like

Code:

while [ condition ]
do
# check and do some processing
done > $output_file

In this method, there will be n + 2 (approx) calls to file libs for 'n' units of write

instead of

Code:

while [ condition ]
do
# check and do some processing
# write to a file  > $output_file
done

In this method, there will be 3n calls for 'n' units of write, which will scale badly as 'n' progresses ...

matrixmadhan

View Public Profile for matrixmadhan

Find all posts by matrixmadhan

07-27-2009

Registered User

110, 2

Join Date: Jul 2007

Last Activity: 28 December 2015, 1:11 PM EST

Posts: 110

Thanks Given: 0

Thanked 2 Times in 2 Posts

lorus,

why are we trying to write a split script..!? let split command do the job...

Code:

split -db 1m InFile OutFile

o/p: it creates files like
OutFile00
OutFile01
...

ilan

View Public Profile for ilan

Find all posts by ilan

07-27-2009

Registered User

8, 0

Join Date: Jul 2009

Last Activity: 29 July 2009, 9:43 AM EDT

Posts: 8

Thanks Given: 0

Thanked 0 Times in 0 Posts

Quote:

Originally Posted by matrixmadhan

Since you have remainded me of the movie 'matrix' - am all charged up to answer your question to an extent atleast Smilie

Hehe thats good to know. I have just to bring something about matrix in every post to get your help :-D

Your suggestions makes absolutly sence, so I rewrote it to this

Code:

#!/bin/bash

declare -a lines
OIFS="$IFS"
IFS=$'\n'
set -f   # cf. help set
lines=($(< "input.txt"))
set +f
IFS="$OIFS"

splitsize=102400
i=0
count=0
lines_tot=${#lines[@]}

while [ $i -le $lines_tot ]; do
    count=$[count+1]
    touch output/output_$count.txt
    
    while [ `ls -al output/output_$count.txt | awk '{print $5}'` -le $splitsize -a $i -le ${#lines[@]} ]; do
        i=$[i+1]
        if [ `expr "${lines[${i}]}" : '#Game No.*'` != 0 ]; then
            while [ `expr "${lines[${i}]}" : '.*wins.*'` = 0 ]; do
                i=$[i+1]
                echo "${lines[${i}]}"
            done
            echo ""
        fi
    done >> output/output_$count.txt
    
done

But that file open/close process doesn't seems to be the time-thief

A simple

Code:

while [ $i -le $lines_tot ]; do

    echo "${lines[${i}]}" >> output_test.txt
    
done

processes the whole file in just a few seconds.

So the time thief must be the "expr" command inside the inner loop. Is there maybe any equivalent command that is faster?

Quote:

lorus,

why are we trying to write a split script..!? let split command do the job...

Because "split" cuts at static points and that would destroy the structure of my file, doesn't it?

lorus

View Public Profile for lorus

Find all posts by lorus

07-27-2009

Registered User

3,216, 33

Join Date: Mar 2005

Last Activity: 4 September 2020, 7:11 AM EDT

Location: classification algos

Posts: 3,216

Thanks Given: 19

Thanked 33 Times in 30 Posts

Quote:

Hehe thats good to know. I have just to bring something about matrix in every post to get your help :-D

Good one !

Quote:

But that file open/close process doesn't seems to be the time-thief

This will definitely have an impact and will scale accordingly to larger files.

Quote:

So the time thief must be the "expr" command inside the inner loop. Is there maybe any equivalent command that is faster?

What exactly is the operation performed? Can you please give an example?

matrixmadhan

View Public Profile for matrixmadhan

Find all posts by matrixmadhan

07-27-2009

Registered User

8, 0

Join Date: Jul 2009

Last Activity: 29 July 2009, 9:43 AM EDT

Posts: 8

Thanks Given: 0

Thanked 0 Times in 0 Posts

My input file contains of blocks like the following

Code:

#Game No : 8273167998 
***** Hand History for Game 8273167998 *****
$100 USD NL Texas Hold'em - Saturday, July 25, 11:34:58 EDT 2009
Table Deep Stack #1459548 (No DP) (Real Money)
Seat 6 is the button
Total number of players : 6 
Seat 5: Ducilator ( $128.60 USD )
Seat 4: EvilAdj ( $145.66 USD )
Seat 3: Ice81111 ( $78.60 USD )
Seat 6: RicsterM ( $292.48 USD )
Seat 1: Techno1990 ( $141.06 USD )
Seat 2: pdiloop ( $100 USD )
Techno1990 posts small blind [$0.50 USD].
pdiloop posts big blind [$1 USD].
** Dealing down cards **
Ice81111 folds
EvilAdj folds
Ducilator raises [$4 USD]
RicsterM folds
Techno1990 folds
pdiloop folds
Ducilator does not show cards.
Ducilator wins $5.50 USD

first I search for the start of the block with this expression: ' #Game No.*"

Code:

if [ `expr "${lines[${i}]}" : '#Game No.*'` != 0 ]; then

then I put out all following lines while I find this expression: '.*wins.*'

Code:

 while [ `expr "${lines[${i}]}" : '.*wins.*'` = 0 ]; do
                i=$[i+1]
                echo "${lines[${i}]}"
  done

the loop around this is to check if the current output file size reaches the split limit

Code:

while [ `ls -al output/output_$count.txt | awk '{print $5}'` -le $splitsize -a $i -le ${#lines[@]} ]; do
       ...
done >> output/output_$count.txt

so the `expr "${lines[${i}]}" : '.*wins.*'` command is executed at every single line of the input file. That are ~800.000 times.
0.03sec per iteration means ~7 hours for the whole process.

lorus

View Public Profile for lorus

Find all posts by lorus

07-27-2009

Registered User

11,728, 1,345

Join Date: Feb 2004

Last Activity: 8 May 2020, 9:07 AM EDT

Location: NM

Posts: 11,728

Thanks Given: 903

Thanked 1,345 Times in 1,201 Posts

Have you considered csplit. Assume the average size of your block is 300 bytes 1024000/300 = 3413 block for 1MB

Code:

csplit -k myinputfilename  '/^#Game/-1{3413}'

jim mcnamara

View Public Profile for jim mcnamara

Find all posts by jim mcnamara

07-27-2009

Registered User

8, 0

Join Date: Jul 2009

Last Activity: 29 July 2009, 9:43 AM EDT

Posts: 8

Thanks Given: 0

Thanked 0 Times in 0 Posts

Quote:

Originally Posted by jim mcnamara

Have you considered csplit. Assume the average size of your block is 300 bytes 1024000/300 = 3413 block for 1MB

Code:

csplit -k myinputfilename  '/^#Game/-1{3413}'

I forgot to say, that the length of each block is different

The posted one is just an example.

lorus

View Public Profile for lorus

Find all posts by lorus

Shell Programming and Scripting

Help to improve speed of text processing script

10 More Discussions You Might Find Interesting

1. Solaris

Rsync quite slow (using very little cpu): how to improve its speed?

Discussion started by: priyadarshan

2. Shell Programming and Scripting

Improve script

Discussion started by: jiam912

3. Shell Programming and Scripting

How to improve an script?

Discussion started by: jiam912

4. UNIX for Dummies Questions & Answers

How to improve the performance of this script?

Discussion started by: vikatakavi

5. Programming

awk processing / Shell Script Processing to remove columns text file

Discussion started by: ajayram

6. Shell Programming and Scripting

Help need to improve performance :Parallel processing ideas

Discussion started by: justchill

7. Shell Programming and Scripting

awk, perl Script for processing a single line text file

Discussion started by: hmsadiq

8. Shell Programming and Scripting

Any way to improve performance of this script

Discussion started by: sirababu

9. Shell Programming and Scripting

KSH script -text file processing NULL issues

Discussion started by: geauxsaints

10. Shell Programming and Scripting

Can I improve this script ???

Discussion started by: Cameron