Select distinct sequences from fasta file and list Post: 302918725

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

select distinct row from a file

Hi, buddies out there. I have a text file ( only one column ) which I created using vi editor. The file contains duplicate rows and I would like to select distinct rows, how to go on it using unix command: file content = apple apple orange watermelon apple orange Can it be done...

2. Shell Programming and Scripting

Select distinct values from a flat file

Hi , I have a similar problem. Please can anyone help me with a shell script or a perl. I have a flat file like this fruit country apple germany apple india banana pakistan banana saudi mango india I want to get a output like fruit country apple ...

3. Shell Programming and Scripting

Select distinct rows in a file by last column

Hi, I have the following file: LOG:015608::ERR:2310:map_spsrec:Invalid parameter LOG:015608::ERR:2471:map_dgdrec:Invalid parameter LOG:015608::ERR:2487:map_nnmrec:Invalid number LOG:015608::ERR:2310:map_nmrec:Invalid number LOG:015608::ERR:2438:map_nmrec:Invalid number As a delimiter I...

4. Shell Programming and Scripting

Shell script for changing the accession number of DNA sequences in a FASTA file

Hi, I am having a file of dna sequences in fasta format which look like this: >admin_1_45 atatagcaga >admin_1_46 atatagcagaatatatat with many such thousands of sequences in a single file. I want to the replace the accession Id "admin_1_45" similarly in following sequences to...

5. Shell Programming and Scripting

Extract sequences from a FASTA file based on another file

6. Shell Programming and Scripting

Shorten header of protein sequences in fasta file

I have a fasta file as follows >sp|O15090|FABP4_HUMAN Fatty acid-binding protein, adipocyte OS=Homo sapiens GN=FABP4 PE=1 SV=3 MCDAFVGTWKLVSSENFDDYMKEVGVGFATRKVAGMAKPNMIISVNGDVITIKSESTFKN TEISFILGQEFDEVTADDRKVKSTITLDGGVLVHVQKWDGKSTTIKRKREDDKLVVECVM KGVTSTRVYERA >sp|L18484|AP2A2_RAT AP-2...

7. Shell Programming and Scripting

Getting unique sequences from multiple fasta file

Hi, I have a fasta file with multiple sequences. How can i get only unique sequences from the file. For example my_file.fasta >seq1 TCTCAAAGAAAGCTGTGCTGCATACTGTACAAAACTTTGTCTGGAGAGATGGAGAATCTCATTGACTTTACAGGTGTGGACGGTCTTCAGAGATGGCTCAAGCTAACATTCCCTGACACACCTATAGGGAAAGAGCTAAC >seq2...

8. UNIX for Beginners Questions & Answers

How to count the length of fasta sequences?

I could calculate the length of entire fasta sequences by following command, awk '/^>/{if (l!="") print l; print; l=0; next}{l+=length($0)}END{print l}' unique.fasta But, I need to calculate the length of a particular fasta sequence specified/listed in another txt file. The results to to be...

9. Shell Programming and Scripting

Shorten header of protein sequences in fasta file to only organism name

I have a fasta file as follows >sp|Q8WWQ8|STAB2_HUMAN Stabilin-2 OS=Homo sapiens OX=9606 GN=STAB2 PE=1 SV=3 MMLQHLVIFCLGLVVQNFCSPAETTGQARRCDRKSLLTIRTECRSCALNLGVKCPDGYTM ITSGSVGVRDCRYTFEVRTYSLSLPGCRHICRKDYLQPRCCPGRWGPDCIECPGGAGSPC NGRGSCAEGMEGNGTCSCQEGFGGTACETCADDNLFGPSCSSVCNCVHGVCNSGLDGDGT...

10. UNIX for Beginners Questions & Answers

How to add specific bases at the beginning and ending of all the fasta sequences?

Hi, I have to add 7 bases of specific nucleotide at the beginning and ending of all the fasta sequences of a file. For example, I have a multi fasta file namely test.fasta as given below test.fasta >TalAA18_Xoo_CIAT_NZ_CP033194.1:_2936369-2939570:+1...

LEARN ABOUT DEBIAN

bp_mask_by_search

BP_MASK_BY_SEARCH(1p)					User Contributed Perl Documentation				     BP_MASK_BY_SEARCH(1p)

NAME

       mask_by_search - mask sequence(s) based on its alignment results

SYNOPSIS

	 mask_by_search.pl -f blast genomefile blastfile.bls > maskedgenome.fa

DESCRIPTION

       Mask sequence based on significant alignments of another sequence.  You need to provide the report file and the entire sequence data which
       you want to mask.  By default this will assume you have done a TBLASTN (or TFASTY) and try and mask the hit sequence assuming you've
       provided the sequence file for the hit database.  If you would like to do the reverse and mask the query sequence specify the -t/--type
       query flag.

       This is going to read in the whole sequence file into memory so for large genomes this may fall over.  I'm using DB_File to prevent keeping
       everything in memory, one solution is to split the genome into pieces (BEFORE you run the DB search though, you want to use the exact file
       you BLASTed with as input to this program).

       Below the double dash (--) options are of the form --format=fasta or --format fasta or you can just say -f fasta

       By -f/--format I mean either are acceptable options.  The =s or =n or =c specify these arguments expect a 'string'

       Options:
	   -f/--format=s    Search report format (fasta,blast,axt,hmmer,etc)
	   -sf/--sformat=s  Sequence format (fasta,genbank,embl,swissprot)
	   --hardmask	    (booelean) Hard mask the sequence
			    with the maskchar [default is lowercase mask]
	   --maskchar=c     Character to mask with [default is N], change
			    to 'X' for protein sequences
	   -e/--evalue=n    Evalue cutoff for HSPs and Hits, only
			    mask sequence if alignment has specified evalue
			    or better
	   -o/--out/
	   --outfile=file   Output file to save the masked sequence to.
	   -t/--type=s	    Alignment seq type you want to mask, the
			    'hit' or the 'query' sequence. [default is 'hit']
	   --minlen=n	    Minimum length of an HSP for it to be used
			    in masking [default 0]
	   -h/--help	    See this help information

AUTHOR - Jason Stajich
       Jason Stajich, jason-at-bioperl-dot-org.

perl v5.14.2							    2012-03-02						     BP_MASK_BY_SEARCH(1p)