Split a file with no pattern -- Split, Csplit, Awk Post: 302151173

10 More Discussions You Might Find Interesting

1. UNIX for Dummies Questions & Answers

Split files using Csplit

I have an excel file with more than 65K records... Since excel does not take more than 65K records i wan to split the file and send it as two excel files... Could some help me how to use the csplit by specifiying the no of records

2. Shell Programming and Scripting

Split a file based on pattern in awk, grep, sed or perl

Hi All, Can someone please help me write a script for the following requirement in awk, grep, sed or perl. Buuuu xxx bbb Kmmmm rrr ssss uuuu Kwwww zzzz ccc Roooowwww eeee Bxxxx jjjj dddd Kuuuu eeeee nnnn Rpppp cccc vvvv cccc Rhhhhhhyyyy tttt Lhhhh rrrrrssssss Bffff mmmm iiiii Ktttt...

3. Shell Programming and Scripting

Split a file based on a pattern

Dear all, I have a large file which is composed of 8000 frames, what i would like to do is split the file into 8000 single files names file.pdb.1, file.pdb.2 etc etc each frame in the large file is seperated by a "ENDMDL" flag so my thinking is to use this flag a a point to split the files...

4. Shell Programming and Scripting

Split binary file with pattern

Hello! Have some problem with extract files from saved session. File contains any kind of special/printable characters. DATA NumberA DATA DATA Begin DATA1.1 DATA1.2 NumberB1 DATA1.3 DATA1.4 End DATA DATA DATA Begin DATA2.1 DATA2.2 NumberB2 DATA2.3 DATA2.4 End DATA DATA ...

5. Shell Programming and Scripting

Split File by Pattern with File Names in Source File... Awk?

Hi all, I'm pretty new to Shell scripting and I need some help to split a source text file into multiple files. The source has a row with pattern where the file needs to be split, and the pattern row also contains the file name of the destination for that specific piece. Here is an example: ...

6. Shell Programming and Scripting

awk to split one field and print the last two fields within the split part.

Hello; I have a file consists of 4 columns separated by tab. The problem is the third fields. Some of the them are very long but can be split by the vertical bar "|". Also some of them do not contain the string "UniProt", but I could ignore it at this moment, and sort the file afterwards. Here is...

7. Shell Programming and Scripting

Split a file based on pattern and size

Hello, I have a large file (2GB) that I would like to split based on pattern and size. I've used the following command to split the file (token is "HELLO") awk '/HELLO/{i++}{print > "file"i}' input.txt and the output is similar to the following (i included filesize in KB): 10 ...

8. Shell Programming and Scripting

split file by delimiter with csplit

Hello, I want to split a big file into smaller ones with certain "counts". I am aware this type of job has been asked quite often, but I posted again when I came to csplit, which may be simpler to solve the problem. Input file (fasta format): >seq1 agtcagtc agtcagtc ag >seq2 agtcagtcagtc...

9. Shell Programming and Scripting

Split the file based on pattern

Hi , I have huge files around 400 mb, which has clob data and have diffeent scenarios: I am trying to pass scenario number as parameter and and get required modified file based on the scenario number and criteria. Scenario 1: file name : scenario_1.txt ...

10. UNIX for Advanced & Expert Users

Split one file to many based on pattern

Hello All, I have records in a file in a pattern A,B,B,B,B,K,A,B,B,K Is there any command or simple logic I can pull out records into multiple files based on A record? I want output as File1: A,B,B,B,B,K File2: A,B,B,K

LEARN ABOUT OSF1

split

split(1)						      General Commands Manual							  split(1)

NAME

       split - Splits a file into pieces

SYNOPSIS

   Current syntax
       split [-l line_count] [-a suffix_length] [file | -] [prefix]

       split -b  n [k|m] [-a suffix_length] [file | -] [prefix]

   Obsolescent syntax
       split [-number] [-a suffix_length] [file | -] [prefix]

STANDARDS

       Interfaces documented on this reference page conform to industry standards as follows:

       split:  XCU5.0

       Refer to the standards(5) reference page for more information about industry standards and associated tags.

OPTIONS

       Uses  suffix_length  letters  to  form  the suffix portion of the file names of the split file.	If -a is not specified, the default suffix
       length is two letters.  If the sum of the prefix and the suffix arguments would create a file  name  exceeding  NAME_MAX  bytes,  an  error
       occurs.	 In this case, split exits with a diagnostic message and no files are created.	Split a file into pieces n bytes in size.  Split a
       file into pieces n kilobytes (1024 bytes) in size.  Split a file into pieces n megabytes (1048576 bytes) in size.  Specifies the number	of
       lines  in each output file.  The line_count argument is an unsigned decimal integer.  The default value is 1000.  If the input does not end
       with a newline character, the partial line is included in the last output file.	Specifies the number of lines in each  output  file.   The
       default is 1000 lines per output file.  If the input does not end with a newline character, the partial line is included in the last output
       file.  (Obsolescent)

OPERANDS

       The pathname of the file to be split.

	      If you do not specify an input file, or if you specify -, the standard input is used.

DESCRIPTION

       The split command reads file and writes it in number-line pieces (default 1000 lines) to a set of output files.

       The size of the output files can be modified by using the -b or -l options.  Each output file is created with a unique suffix consisting of
       exactly	suffix	lowercase  letters from the POSIX locale.  The letters of the suffix are used as if they were a base-26 digit system, with
       the first suffix to be created consisting of all a characters, the second with b replacing the last a etc., until a name of all zs is  cre-
       ated.   By  default,  the  names  of  the output files are x, followed by a two-character suffix from the character set as described above,
       starting with aa, ab, ac, etc., and continuing until the suffix zz, for a maximum of 676 files.

       The value of prefix cannot be longer than the value of NAME_MAX from <limits.h> minus two.

       If the number of files required is greater than the maximum allowed by the effective suffix length (such that the last allowable file would
       be  larger  than  the  requested  size), split fails after creating the last possible file with a valid suffix.	The split command will not
       delete the files it created with valid suffixes.  If the file limit is not exceeded, the last file created contains the	remainder  of  the
       input file and thus might be smaller than the requested size.

EXIT STATUS

       The following exit values are returned: Successful completion.  An error occurred.

EXAMPLES

       To split a file into 1000-line segments, enter: split book

	      This  splits  book into 1000-line segments named xaa, xab, xac, and so forth.  To split a file into 50-line segments and specify the
	      file name prefix, enter: split -l50 book sect

	      This splits book into 50-line segments named sectaa, sectab, sectac, and so forth.

ENVIRONMENT VARIABLES

       The following environment variables affect the execution of split: Provides a default value for the internationalization variables that are
       unset or null. If LANG is unset or null, the corresponding value from the default locale is used.  If any of the internationalization vari-
       ables contain an invalid setting, the utility behaves as if none of the variables had been defined.  If set to a  non-empty  string  value,
       overrides  the  values of all the other internationalization variables.	Determines the locale for the interpretation of sequences of bytes
       of text data as characters (for example, single-byte as opposed to multibyte characters in arguments  and  input  files).   Determines  the
       locale for the format and contents of diagnostic messages written to standard error.  Determines the location of message catalogues for the
       processing of LC_MESSAGES.

SEE ALSO

       Commands:  bfs(1), csplit(1)

       Standards:  standards(5)

																	  split(1)