Linux shell script to insert new lines based on delimiter count


 
Thread Tools Search this Thread
Top Forums Shell Programming and Scripting Linux shell script to insert new lines based on delimiter count
# 1  
Old 11-18-2016
Linux shell script to insert new lines based on delimiter count

The input file is a .dat file which is delimited by null (^@ in Linux). On a windows PC it looks something like this (numbers are masked with 1).

Image

The entire file is in one row but it has multiple records - each record contains 80 fields i.e. there are 81 counts of the delimiter (null or ^@). After all of the records there is a trailer of 20 fields delimited by ^@

Please suggest a sh script on how to split the file into multiple rows, after every 81 count of the delimiter ^@

The output should be like
Code:
record 1 comprising of 81 count of ^@
/n record 2 comprising of 81 count of ^@
/n ...
/n trailer record (we need not count the trailer as after the last record's 81 count it will remain as is)

# 2  
Old 11-18-2016
split -C might not be the perfect fit but something to look into. Or, try
Code:
sed 's/./&\n/480;s/./&\n/400;s/./&\n/320;s/./&\n/240;s/./&\n/160;s/./&\n/80;' file

This User Gave Thanks to RudiC For This Post:
# 3  
Old 11-18-2016
Does that code start splitting from 480th position i.e. 6th record and work backwards? Apologies for the stupid question but where is it matching the delimiter?

The file can be of any length, ie have any number of records. The records in the file are also of variable length so the only way to identify the end of a record is that it is after each 81st occurrence of the "^@" delimiter.

Ie from the start of the file till 81st delimiter is one record, 82nd till 163rd is the second record, and so on.

The records are currently all in one line and they need to be in multiple lines ie separated by \n

Is there a way to do the 81st delimiter check till end of file
# 4  
Old 11-18-2016
Looks like I've misread/misinterpreted your specification. Sorry for that.
So, in your picture, we're seeing many empty fields (e.g. 15 after the first), 81 fields make up a record of unpredictable length, and there's no <NL> (\n, ^J, 0x0A) char in it.
Are other non-printable, control characters possible, like <TAB>s? Or are all field contents printable alphanumeric characters?

Please note that it is far better to post or attach a sample input file, be it abbreviated, to work (and test) upon, than to show a picture.
This User Gave Thanks to RudiC For This Post:
# 5  
Old 11-18-2016
Apologies, my bad. I should've uploaded the file. Attached is a masked .dat file renamed as .txt for uploading.

On opening it with notepad++ in Windows, the null characters show up as boxes. In a Linux vi the nulls are ^@.

All the records in this file are in one row. This particular file has 2 records followed by the trailer record.
  • First record = starts at the beginning of the file 00000230 (this field gives the length of the record in bytes)
  • Second record = starts at the next 00000230 (it is a coincidence, here both records have same length)
  • Trailer record = starts at 0000096 (the trailer length is of 96 bytes and it also has 80 delimiters of ^@ or null characters. Ignore my earlier post saying trailer has 20 delimiters. It has 80 actually)

As the field lengths are variable so we cannot define a record in terms of total length of its fields or total bytes. This is why we are defining a record as effectively having length of 80 ^@ delimiters.

I require the 1st record in one row, 2nd record in next row and so on till the end of the file, with the trailer in the last row. If there is a way of adding a newline after every 80th ^@ from the beginning till the eof, then perhaps it will work?

The only unprintable character is null ^@, no TABS or other spaces, all other characters are alphanumeric.

Please let me know if any questions. Thanks for the help
# 6  
Old 11-18-2016
The *nix text utilities are NOT the best to deal with binary files like the one that you post. Would this come close to what you need?
Code:
sed 's/\o000/&\n/162;s/\o000/&\n/81' /tmp/MED_BIL_accmasked.DAT.txt
0000023000353123456789272050123456789100UNKNOWN00353123456789101710511-05-2016 01:01:03ABC MEDIATION50100353830000002===EABCTwE2E2k+GFLBBE35383000000200000230
00353123456789272050123456789100UNKNOWN00353123456789101710711-05-2016 01:01:10ABC MEDIATION20100353830000002===AAAAAAAIeQmVGFLBBA35383000000200000096
30000502

Use e.g. od -bc to verify the result.
This User Gave Thanks to RudiC For This Post:
# 7  
Old 11-22-2016
Apologies for the delayed response. While the code here would work that is because we know it has 2 records at the 80th and 160th positions, in actual files there would be thousands or tens of thousands of records. So is there a way to cut or grep each record (from the beginning of the file counting 80 ^@ delimiters), then move it to a temp file, add a \n, then add the second record, and so on, till the end of the file. So that the end result is a file with each record in a row.

Appreciate all your assistance. Thanks.
Login or Register to Ask a Question

Previous Thread | Next Thread

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Split files based on row delimiter count

I have a huge file (around 4-5 GB containing 20 million rows) which has text like: <EOFD>11<EOFD>22<EORD>2<EOFD>2222<EOFD>3333<EORD>3<EOFD>44<EOFD>55<EORD>66<EOFD>888<EOFD>9999<EORD> Actually above is an extracted file from a Sql Server with each field delimited by <EOFD> and each row ends... (8 Replies)
Discussion started by: amvip
8 Replies

2. Shell Programming and Scripting

Shell script count lines and sum numbers from multiple files

I want to count the number of lines, I need this result be a number, and sum the last numeric column, I had done to make this one at time, but I need to make this for a crontab, so, it has to be an script, here is my lines: It counts the number of lines: egrep -i String file_name_201611* |... (5 Replies)
Discussion started by: Elly
5 Replies

3. Shell Programming and Scripting

awk joining multiple lines based on field count

Hi Folks, I have a file with fields as follows which has last field in multiple lines. I would like to combine a line which has three fields with single field line for as shown in expected output. Please help. INPUT hname01 windows appnamec1eda_p1, ... (5 Replies)
Discussion started by: shunya
5 Replies

4. Shell Programming and Scripting

Removing duplicate lines on first column based with pipe delimiter

Hi, I have tried to remove dublicate lines based on first column with pipe delimiter . but i ma not able to get some uniqu lines Command : sort -t'|' -nuk1 file.txt Input : 38376KZ|09/25/15|1.057 38376KZ|09/25/15|1.057 02006YB|09/25/15|0.859 12593PS|09/25/15|2.803... (2 Replies)
Discussion started by: parithi06
2 Replies

5. Shell Programming and Scripting

File Count Based on FileDate using Shell Script

I have file listed in my directory in following format -rwxrwxr-x+ 1 test test 4.9M Oct 3 16:06 test20141002150108.txt -rwxrwxr-x+ 1 test test 4.9M Oct 4 16:06 test20141003150108.txt -rwxrwxr-x+ 1 test test 4.9M Oct 5 16:06 test20141005150108.txt -rwxrwxr-x+ 1 test ... (2 Replies)
Discussion started by: krish2014
2 Replies

6. Shell Programming and Scripting

Insert Columns before the last Column based on the Count of Delimiters

Hi, I have a requirement where in I need to insert delimiters before the last column of the total delimiters is less than a specified number. Say if the delimiters is less than 139, I need to insert 2 columns ( with blanks) before the last field awk -F 'Ç' '{ if (NF-1 < 139)} END { "Insert 2... (5 Replies)
Discussion started by: arunkesi
5 Replies

7. Shell Programming and Scripting

Bash script to count and insert

Hi not sure if this is possible but I need some help with a bash script, I have a text file and on the first line that starts with 7150230 I need it to put a 1 at position 79 and a 2 at position 88, this is where it gets complicated, on the next line it finds that starts with 7150230 I then need it... (8 Replies)
Discussion started by: firefox2k2
8 Replies

8. Shell Programming and Scripting

Shell script to put delimiter for a no delimiter variable length text file

Hi, I have a No Delimiter variable length text file with following schema - Column Name Data length Firstname 5 Lastname 5 age 3 phoneno1 10 phoneno2 10 phoneno3 10 sample data - ... (16 Replies)
Discussion started by: Gaurav Martha
16 Replies

9. Shell Programming and Scripting

insert leading zeroes based on the character count

Hi, I need add leading zeroes to a field in a file based on the character count. The field can be of 1 character to 6 character length. I need to make the field 14bytes. eg: 8351,20,1 8351,234,6 8351,2,0 8351,1234,2 8351,123456,1 8351,12345,2 This should become. ... (3 Replies)
Discussion started by: gpaulose
3 Replies

10. UNIX for Dummies Questions & Answers

Perl/shell script count the lines

Hi Guys, I want to write a perl/shell script do parse the following file input file content NPA-NXX SC 2084549 45 2084552 45 2084563 2007 2084572 45 2084580 45 3278411 45 3278430 45 3278493 530 3278507 530... (3 Replies)
Discussion started by: pistachio
3 Replies
Login or Register to Ask a Question