Help to improve speed of text processing script

07-26-2009

Registered User

8, 0

Join Date: Jul 2009

Last Activity: 29 July 2009, 9:43 AM EDT

Posts: 8

Thanks Given: 0

Thanked 0 Times in 0 Posts

Help to improve speed of text processing script

Hey together,

You should know, that I'am relatively new to shell scripting, so my solution is probably a little awkward.

Here is the script:

Code:

#!/bin/bash

live_dir=/var/lib/pokerhands/live

for limit in `find $live_dir/ -type d  | sed -e s#$live_dir/##`; do
    cat $live_dir/$limit/* > $limit
    
    declare -a lines
    OIFS="$IFS"
    IFS=$'\n'
    set -f   # cf. help set
    lines=($(< "$limit"))
    set +f
    IFS="$OIFS"
    
    i=0
    count=0

    while [ $i -le ${#lines[@]} ]; do

        count=$[count+1]
        touch test_$count.txt
        
        while [ `ls -al test_$count.txt | awk '{print $5}'` -le 1048576 -a $i -le ${#lines[@]} ]; do
            i=$[i+1]
            if [ `expr "${lines[${i}]}" : '#Game No.*'` != 0 ]; then
                while [ `expr "${lines[${i}]}" : '.*wins.*'` = 0 ]; do
                    i=$[i+1]
                    echo "${lines[${i}]}" >> test_$count.txt
                done
                echo "" >> test_$count.txt
            fi
        done
        
    done
done

This Script splits a input file into ~1MB Parts, without destroying the data blocks.

The data blocks of the input file look something like this:

Code:

#Game No : 8273167998 
***** Hand History for Game 8273167998 *****
$100 USD NL Texas Hold'em - Saturday, July 25, 11:34:58 EDT 2009
Table Deep Stack #1459548 (No DP) (Real Money)
Seat 6 is the button
Total number of players : 6 
Seat 5: Ducilator ( $128.60 USD )
Seat 4: EvilAdj ( $145.66 USD )
Seat 3: Ice81111 ( $78.60 USD )
Seat 6: RicsterM ( $292.48 USD )
Seat 1: Techno1990 ( $141.06 USD )
Seat 2: pdiloop ( $100 USD )
Techno1990 posts small blind [$0.50 USD].
pdiloop posts big blind [$1 USD].
** Dealing down cards **
Ice81111 folds
EvilAdj folds
Ducilator raises [$4 USD]
RicsterM folds
Techno1990 folds
pdiloop folds
Ducilator does not show cards.
Ducilator wins $5.50 USD

It is working so far, but the problem is the speed ... for a ~20mb input file it runs for some hours ...

What makes it so slow?

Can anyone help me to improve the speed?

lorus

View Public Profile for lorus

Find all posts by lorus

07-26-2009

Registered User

187, 4

Join Date: Jul 2009

Last Activity: 20 February 2013, 8:48 AM EST

Posts: 187

Thanks Given: 0

Thanked 4 Times in 4 Posts

From the code it looks like you ARE an experienced scripter.
But did you try the basic?
Try your script with debug.
ksh -x
It could be any of them.
See which line takes more time. Simple.

Also, why do you need the "sed" in the first line ?
From the first look, it looks like you are removing the directory details but you are adding it up again at the "cat".

edidataguy

View Public Profile for edidataguy

Find all posts by edidataguy

07-26-2009

Registered User

7,747, 559

Join Date: Feb 2007

Last Activity: 20 April 2020, 11:28 AM EDT

Location: The Netherlands

Posts: 7,747

Thanks Given: 139

Thanked 559 Times in 520 Posts

What are the conditions to split the file? There are probably other approaches to speed up the process.

Regards

Franklin52

View Public Profile for Franklin52

Find all posts by Franklin52

07-26-2009

Registered User

2,163, 123

Join Date: Nov 2007

Last Activity: 31 July 2016, 9:42 AM EDT

Location: H3X

Posts: 2,163

Thanks Given: 11

Thanked 123 Times in 116 Posts

Quote:

Originally Posted by lorus

Can anyone help me to improve the speed?

This is the price to pay when you don't follow the rules and use temp file and useless commands .... time

Try to post a sample data file and required output and suggest a different approach.

danmero

View Public Profile for danmero

Find all posts by danmero

07-26-2009

Registered User

8, 0

Join Date: Jul 2009

Last Activity: 29 July 2009, 9:43 AM EDT

Posts: 8

Thanks Given: 0

Thanked 0 Times in 0 Posts

Quote:

Originally Posted by danmero

This is the price to pay when you don't follow the rules and use temp file and useless commands .... time Smilie

Yeah I realized that and so I just asked for some hints to get a better solution to run it faster

Please notice that I'am relatively new to shell scripting. In fact this is my 2nd script try.

Quote:

Originally Posted by Franklin52

What are the conditions to split the file? There are probably other approaches to speed up the process.

Regards

I have a large input text file (aprox. 20mb) and want to split it into single 1MB output files. The input file include contains of data blocks that are'nt allowed to destroy during the split. (posted short sample of this data block in my 1st post)

Quote:

Originally Posted by danmero

Try to post a sample data file and required output and suggest a different approach.

Yeah thats a good idea. I attached my test enviroment to this post.
It contains of the following structure.

Code:

testenv/
|-- output        <-output folder
|   |-- output_1.txt    
|   |-- output_2.txt    <-output files
|   |-- output_3.txt
|   `-- output_4.txt
|-- input.txt        <-input file
`-- split.sh        <-script file

I reduced the script code to the essential things and set the output file size to 100kb to demonstrate you the principle and let you better understand what I want to do.

In this example I let the script run for ~5min and in this time it processes ~400kb of 25mb and put out these 4 files.

Thanks in advance for your help

test_environment.rar (1.44 MB)

lorus

View Public Profile for lorus

Find all posts by lorus

07-27-2009

Registered User

187, 4

Join Date: Jul 2009

Last Activity: 20 February 2013, 8:48 AM EST

Posts: 187

Thanks Given: 0

Thanked 4 Times in 4 Posts

Just run your script as requested earlier.

Code:

ksh -x my your.sh

You can find out by yourself which line is taking more time.

edidataguy

View Public Profile for edidataguy

Find all posts by edidataguy

07-27-2009

Registered User

8, 0

Join Date: Jul 2009

Last Activity: 29 July 2009, 9:43 AM EDT

Posts: 8

Thanks Given: 0

Thanked 0 Times in 0 Posts

Quote:

Originally Posted by edidataguy

Just run your script as requested earlier.

Code:

ksh -x my your.sh

You can find out by yourself which line is taking more time.

You know the movie matrix? In that speed the chars fly through my screen, when I use "ksh -x". That really doesn't helps me much

I think the problem is, that I use the "expr" command on any single line of the input file.
That means every iteration takes ~0.03sec. On an input file with 800.000 lines the whole process take ~7 hours.
Each iteration have to be 0.0003sec to get an acceptable result.

Is there maybe any faster command than "expr" which can do the same (regexp)?

Last edited by lorus; 07-27-2009 at 05:58 AM..

lorus

View Public Profile for lorus

Find all posts by lorus

Shell Programming and Scripting

Help to improve speed of text processing script

10 More Discussions You Might Find Interesting

1. Solaris

Rsync quite slow (using very little cpu): how to improve its speed?

Discussion started by: priyadarshan

2. Shell Programming and Scripting

Improve script

Discussion started by: jiam912

3. Shell Programming and Scripting

How to improve an script?

Discussion started by: jiam912

4. UNIX for Dummies Questions & Answers

How to improve the performance of this script?

Discussion started by: vikatakavi

5. Programming

awk processing / Shell Script Processing to remove columns text file

Discussion started by: ajayram

6. Shell Programming and Scripting

Help need to improve performance :Parallel processing ideas

Discussion started by: justchill

7. Shell Programming and Scripting

awk, perl Script for processing a single line text file

Discussion started by: hmsadiq

8. Shell Programming and Scripting

Any way to improve performance of this script

Discussion started by: sirababu

9. Shell Programming and Scripting

KSH script -text file processing NULL issues

Discussion started by: geauxsaints

10. Shell Programming and Scripting

Can I improve this script ???

Discussion started by: Cameron