Improve the performance of my C++ code

01-15-2015

Registered User

1,015, 157

Join Date: Jun 2009

Last Activity: 25 June 2018, 8:15 AM EDT

Posts: 1,015

Thanks Given: 3

Thanked 157 Times in 149 Posts

If you want to run faster,

1. Insert data in a way that it's always sorted, as Corona688 has already noted.
2. Don't use new/delete in a loop.
3. Don't use C++ I/O routines - use C open/read/close or other low-level routines.

---------- Post updated at 04:34 PM ---------- Previous update was at 04:34 PM ----------

Quote:

Originally Posted by Corona688

It is UNIX text, not Windows text.

I don't always have access to a Unix box.

achenle

View Public Profile for achenle

Find all posts by achenle

01-15-2015

Moderator

3,791, 1,452

Join Date: Oct 2010

Last Activity: 1 August 2020, 1:38 AM EDT

Posts: 3,791

Thanks Given: 183

Thanked 1,452 Times in 1,302 Posts

You say you have a lot of memory so I think you would be best using a hash table. Store both each entry and it's reverse-complementary.

Chubler_XL

View Public Profile for Chubler_XL

Find all posts by Chubler_XL

01-15-2015

Registered User

564, 13

Join Date: Sep 2009

Last Activity: 26 May 2021, 8:59 AM EDT

Location: Saskatchewan, Canada

Posts: 564

Thanks Given: 376

Thanked 13 Times in 12 Posts

My problem is the implementation. I want to try programming using some available "libraries". I appreciate any code example on top of mine. This sounds lazy, but I'm self-learning by practice. Of course googled a while, but did not get any similar example. Thanks a lot!

Last edited by yifangt; 01-16-2015 at 11:30 AM..

yifangt

View Public Profile for yifangt

Find all posts by yifangt

01-16-2015

Registered User

3, 1

Join Date: Jan 2015

Last Activity: 29 January 2015, 3:45 AM EST

Posts: 3

Thanks Given: 0

Thanked 1 Time in 1 Post

I would agree to achenles 1(Corona688) and 3 points.

On top of that I would just use simple dynamic array with pointer to next element, something like :

Code:

typedef struct _sList _sList;
struct _sList{
    _sList    *next;
    int       SEQ_size;
    SEQ       *element;
};

when reading elements from file on complete element just start comparing SEQ_size from start of list, until point where element_size from file equals or smaller to element in list, then in case of smaller - add element to list (unique one). If equal memcmp(list_elemen, file_element, SEQ_size) until list element is greater, equal or SEQ_size of file element is greater, if equal delete - its equal, if greater add to list before greater element.
P.S. this is case when sorting from small to big.

This way you will fast forward to elements of same length, and then will compare only until you will find equal element or storing spot.

In the end you will have sorted list of unique elements.

Need even faster then make more difficult structures from which you could build graph(data tree). In this area only your imagination and content of element will stop from optimizing even more. Usually more complex structures will pay off when having bigger amount of data.

This User Gave Thanks to Lauris_k For This Post:

Lauris_k

View Public Profile for Lauris_k

Find all posts by Lauris_k

01-16-2015

Registered User

23,310, 4,623

Join Date: Aug 2005

Last Activity: 7 July 2020, 11:47 AM EDT

Location: Saskatchewan

Posts: 23,310

Thanks Given: 1,331

Thanked 4,623 Times in 4,217 Posts

Quote:

Originally Posted by yifangt

My problem is the implementation. I want to step to "my second stage" of programming by using those available libraries.

"hash table" isn't exactly a library, it's different enough from other data structures it's often hand-rolled. Generalizing it too much would run the risk of poor performance, you need to pick the right algorithms for your application. It has a lot of restrictions as well (hard to iterate, deletion can cause something like fragmentation, and it can't be sorted). I've seen a few attempts at building a library for it, but nothing I ever liked very much.

In the end it's not that complicated. It's a big array with strict rules about what data gets put in what element. I'd suggest "open chaining" for your table -- basically an array full of lists -- with an index that's not really hashed at all, just converted from ACGT into boolean. Four letters would be 8 bits, for an array 256 long for example. Then you could just look up the first four letters of your sequence, find that list, and speedily check every possible thing which might contain your sequence without having to brute-force it.

Last edited by Corona688; 01-16-2015 at 11:24 AM..

This User Gave Thanks to Corona688 For This Post:

Corona688

View Public Profile for Corona688

Visit Corona688's homepage!

Find all posts by Corona688

Programming

Improve the performance of my C++ code

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

How to improve the performance of this script?

Discussion started by: vikatakavi

2. Shell Programming and Scripting

Improve performance of echo |awk

Discussion started by: chetan.c

3. Programming

Help with improve the performance of grep

Discussion started by: cpp_beginner

4. Shell Programming and Scripting

How to improve the performance of parsers in Perl?

Discussion started by: vanitham

5. Shell Programming and Scripting

Want to improve the performance of script

Discussion started by: poweroflinux

6. Shell Programming and Scripting

Improve the performance of a shell script

Discussion started by: apsprabhu

7. Shell Programming and Scripting

Any way to improve performance of this script

Discussion started by: sirababu

8. UNIX for Dummies Questions & Answers

Improve Performance

Discussion started by: mazhar99

9. Shell Programming and Scripting

How to improve grep performance...

Discussion started by: pooga17

10. UNIX for Advanced & Expert Users

improve performance by using ls better than find

Discussion started by: Nicol