Hello Unix gurus,
I have a gzipped file where each line contains 2 street addresses in the US. What I want to do is get a count for each state that does not match.
What I have so far is:
$ gzcat matched_10_09.txt.gz |cut -c 106-107,184-185 | head -5
CTCT
CTNY
CTCT
CTFL
CTMA
This cuts... (5 Replies)
here is what i want to achieve... consider a file contains below contents. the file size is large about 60mb
cat dump.sql
INSERT INTO `table1` (`id`, `action`, `date`, `descrip`, `lastModified`) VALUES (1,'Change','2011-05-05 00:00:00','Account Updated','2012-02-10... (10 Replies)
I am a novice writing perl scripts so I'd appreciate any help you guys can offer.
I have a list of 100 words in a file (words.txt) and I need to find them in a second file (data.txt). Whenever one of these words is found I need to write that line to a third file (out.txt) and then continue... (1 Reply)
Hi,
I would like to have the length of a segment based on coordinates of its parts.
Example input file:
chr11 genes_good3.gtf aggregate_gene 1 100 gene1
chr11 genes_good3.gtf exonic_part 1 60
chr11 genes_good3.gtf exonic_part 70 100
chr11 genes_good3.gtf aggregate_gene 200 1000 gene2... (2 Replies)
Use and complete the template provided. The entire template must be completed. If you don't, your post may be deleted!
1. The problem statement, all variables and given/known data:
My goal to find how many requests in 14 days from weblog server. I know to cat a weblog file to wc -l to find the... (8 Replies)
I am trying to add a condition to the below perl that will capture the GTtag and place a specific string in the last field of each line. The problem is that the GT value used is not right after the tag rather it is a few fields away. The values should always be 0/1 or 1/2 and are in bold in the... (12 Replies)
Trying to output a result that uses the data from file to combine and subtract specific lines. If $4 matches in each line then the last $6 value is added to $2 and that becomes the new$3. Each matching line in combined into one with $1 then the original $2 then the new$3 then $5. For the cases... (4 Replies)
I am trying to output a tab-delimited result that uses the data from a tab-delimited file to combine and subtract specific lines.
If $4 matches in each line then the first matching sequential $6 value is added to $2, unless the value is 1, then the original $2 is used (like in the case of line... (3 Replies)
The below awk executes as is and produces the current output. It isvery close but what Ican not seem to do is add the -exon..., the ... portion comes from $1 and the _exon is static and will never change. If there is + sign in $4 then the ... is in acending order or sequential. If there is a - in... (2 Replies)
Discussion started by: cmccabe
2 Replies
LEARN ABOUT DEBIAN
ampliconnoise
AMPLICONNOISE(1) AmpliconNoise Documentation AMPLICONNOISE(1)NAME
AmpliconNoise - remove noise from high throughput nucleotide sequence data
VERSION
This documentation refers to version 1.22
SYNOPSIS
See /usr/share/doc/ampliconnoise/Doc.pdf.gz for details of how to run.
DESCRIPTION
The following tools are included. Most of them have an MPI equivalent, for example SeqNoise has an equivalent SeqNoiseM which can be used
with mpirun.
FastaUnique - dereplicates fasta file
-in string input file name
Options:
FCluster
-in string distance input file name
-out string output file stub
Options:
-r resolution
-a average linkage
-w use weights
-i read identifiers
-s scale dist.
NDist - pairwise Needleman-Wunsch sequence distance matrix from a fasta file
-in string fata file name
Options:
-i output identifiers
Perseus - slays monsters
-sin string seq file name
Options:
-tin string reference sequence file
-a output alignments
-d use imbalance
-rin string lookup file name
PyroDist - pairwise distance matrix from flowgrams
-in string flow file name
-out stub out file stub
Options:
-ni no index in dat file
-rin string lookup file name
PyroNoise - clusters flowgrams without alignments
-din string flow file name
-out string cluster input file name
-lin string list file
Options:
-v verbose
-c double initial cut-off
-ni no index in dat file
-s double precision
-rin file lookup file name
SeqDist - pairwise distance matrix from a fasta file
-in string fasta file name
Options:
-i output identifiers
-rin string lookup file name
SeqNoise - clusters sequences
-in string sequence file name
-din string distance matrix file name
-out string cluster input file name
-lin string list file
Options:
-min mapping file
-v verbose
-c double initial cut-off
-s double precision
-rin string lookup file name
SplitClusterEven
-din string dat filename
-min string map filename
-tin string tree filename
-s split size
-m min size
AUTHOR
All software by Chris Quince (quince@civil.gla.ac.uk) This manpage by Tim Booth (tbooth@ceh.ac.uk)
LICENCE AND COPYRIGHT
Copyright (c) 2009 (quince@civil.gla.ac.uk). All rights reserved.
Released under the Lesser GPL.
Permission is granted for anyone to copy, use, or modify these programs and documents for purposes of research or education, provided this
copyright notice is retained, and note is made of any changes that have been made.
perl v5.12.4 2011-04-28 AMPLICONNOISE(1)