Select distinct sequences from fasta file and list Post: 302918641

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

select distinct row from a file

Hi, buddies out there. I have a text file ( only one column ) which I created using vi editor. The file contains duplicate rows and I would like to select distinct rows, how to go on it using unix command: file content = apple apple orange watermelon apple orange Can it be done...

2. Shell Programming and Scripting

Select distinct values from a flat file

Hi , I have a similar problem. Please can anyone help me with a shell script or a perl. I have a flat file like this fruit country apple germany apple india banana pakistan banana saudi mango india I want to get a output like fruit country apple ...

3. Shell Programming and Scripting

Select distinct rows in a file by last column

Hi, I have the following file: LOG:015608::ERR:2310:map_spsrec:Invalid parameter LOG:015608::ERR:2471:map_dgdrec:Invalid parameter LOG:015608::ERR:2487:map_nnmrec:Invalid number LOG:015608::ERR:2310:map_nmrec:Invalid number LOG:015608::ERR:2438:map_nmrec:Invalid number As a delimiter I...

4. Shell Programming and Scripting

Shell script for changing the accession number of DNA sequences in a FASTA file

Hi, I am having a file of dna sequences in fasta format which look like this: >admin_1_45 atatagcaga >admin_1_46 atatagcagaatatatat with many such thousands of sequences in a single file. I want to the replace the accession Id "admin_1_45" similarly in following sequences to...

5. Shell Programming and Scripting

Extract sequences from a FASTA file based on another file

6. Shell Programming and Scripting

Shorten header of protein sequences in fasta file

I have a fasta file as follows >sp|O15090|FABP4_HUMAN Fatty acid-binding protein, adipocyte OS=Homo sapiens GN=FABP4 PE=1 SV=3 MCDAFVGTWKLVSSENFDDYMKEVGVGFATRKVAGMAKPNMIISVNGDVITIKSESTFKN TEISFILGQEFDEVTADDRKVKSTITLDGGVLVHVQKWDGKSTTIKRKREDDKLVVECVM KGVTSTRVYERA >sp|L18484|AP2A2_RAT AP-2...

7. Shell Programming and Scripting

Getting unique sequences from multiple fasta file

Hi, I have a fasta file with multiple sequences. How can i get only unique sequences from the file. For example my_file.fasta >seq1 TCTCAAAGAAAGCTGTGCTGCATACTGTACAAAACTTTGTCTGGAGAGATGGAGAATCTCATTGACTTTACAGGTGTGGACGGTCTTCAGAGATGGCTCAAGCTAACATTCCCTGACACACCTATAGGGAAAGAGCTAAC >seq2...

8. UNIX for Beginners Questions & Answers

How to count the length of fasta sequences?

I could calculate the length of entire fasta sequences by following command, awk '/^>/{if (l!="") print l; print; l=0; next}{l+=length($0)}END{print l}' unique.fasta But, I need to calculate the length of a particular fasta sequence specified/listed in another txt file. The results to to be...

9. Shell Programming and Scripting

Shorten header of protein sequences in fasta file to only organism name

I have a fasta file as follows >sp|Q8WWQ8|STAB2_HUMAN Stabilin-2 OS=Homo sapiens OX=9606 GN=STAB2 PE=1 SV=3 MMLQHLVIFCLGLVVQNFCSPAETTGQARRCDRKSLLTIRTECRSCALNLGVKCPDGYTM ITSGSVGVRDCRYTFEVRTYSLSLPGCRHICRKDYLQPRCCPGRWGPDCIECPGGAGSPC NGRGSCAEGMEGNGTCSCQEGFGGTACETCADDNLFGPSCSSVCNCVHGVCNSGLDGDGT...

10. UNIX for Beginners Questions & Answers

How to add specific bases at the beginning and ending of all the fasta sequences?

Hi, I have to add 7 bases of specific nucleotide at the beginning and ending of all the fasta sequences of a file. For example, I have a multi fasta file namely test.fasta as given below test.fasta >TalAA18_Xoo_CIAT_NZ_CP033194.1:_2936369-2939570:+1...

LEARN ABOUT DEBIAN

pynast

VERSION:(1)							   User Commands						       VERSION:(1)

NAME

       PyNAST - alignment of short DNA sequences

SYNOPSIS

       pynast [options] {-i input_fp -t template_fp}

DESCRIPTION

       [] indicates optional input (order unimportant) {} indicates required input (order unimportant)

   Example usage:
	      pynast -i my_input.fasta -t my_template.fasta

OPTIONS

       --version
	      show program's version number and exit

       -h, --help
	      show this help message and exit

       -t TEMPLATE_FP, --template_fp=TEMPLATE_FP
	      path to template alignment file [REQUIRED]

       -i INPUT_FP, --input_fp=INPUT_FP
	      path to input fasta file [REQUIRED]

       -v, --verbose
	      Print status and other information during execution [default: False]

       -p MIN_PCT_ID, --min_pct_id=MIN_PCT_ID
	      minimum percent sequence	identity to consider a sequence a match [default: 75.0]

       -l MIN_LEN, --min_len=MIN_LEN
	      minimum sequence length to include in NAST alignment [default: 1000]

       -m PAIRWISE_ALIGNMENT_METHOD, --pairwise_alignment_method=PAIRWISE_ALIGNMENT_METHOD
	      method for performing pairwise alignment [default: uclust]

       -a FASTA_OUT_FP, --fasta_out_fp=FASTA_OUT_FP
	      path to store resulting alignment file [default: derived from input filepath]

       -g LOG_FP, --log_fp=LOG_FP
	      path to store log file [default: derived from input filepath]

       -f FAILURE_FP, --failure_fp=FAILURE_FP
	      path to store file of seqs which fail to align [default: derived from input filepath]

       -e MAX_E_VALUE, --max_e_value=MAX_E_VALUE
	      Depreciated. Will be removed in PyNAST 1.2

       -d BLAST_DB, --blast_db=BLAST_DB
	      Depreciated. Will be removed in PyNAST 1.2

SEE ALSO

       http://pynast.sourceforge.net

Version: pynast 1.1						    August 2011 						       VERSION:(1)