Convert files to UTF-8 on AIX 7.1

02-26-2018

Registered User

1, 0

Join Date: Feb 2018

Last Activity: 15 March 2018, 7:38 AM EDT

Posts: 1

Thanks Given: 0

Thanked 0 Times in 0 Posts

Convert files to UTF-8 on AIX 7.1

Dears,

I have a shell script - working perfectly on Oracle Linux - that detects the encoding (the charset to be exact) of the files in a specified directory using the "file" command (The file command outputs the charset in Linux, but doesn't do that in AIX), then if the file isn't a UTF-8 text file, it converts it to UTF-8 using "iconv" command.

I searched lots of forums and threads but it seems this is extremely hard to do in AIX, since the "file" command doesn't output the charset.

I also read this useful thread on this forum: Converting Unicode file to UTF8 format.

My problem is that if I want to use the "iconv" command to convert my files to UTF-8, how can I determine the charset of the original file ?

Code:

iconv -f FromCode -t ToCode

(The ToCode can be replaced by UTF-8, but I need to guess the FromCode).

Is there any way to do that ?

Does anyone have a working script on AIX that does what I want to do ?

Thank you,
Regards.

Moderator's Comments:

Please use CODE tags as required by forum rules!

Last edited by RudiC; 02-26-2018 at 11:29 AM.. Reason: Added CODE tags.

JeanM-1

View Public Profile for JeanM-1

Find all posts by JeanM-1

02-27-2018

Registered User

12,315, 4,560

Join Date: Jul 2012

Last Activity: 22 November 2019, 4:29 PM EST

Location: San Jose, CA, USA

Posts: 12,315

Thanks Given: 952

Thanked 4,560 Times in 3,818 Posts

If you don't know what codeset was used to encode a file, there isn't much that can be done to guess at what it might be.
It is easy to guess that it is just ASCII if there aren't any bytes with the high order bit set and there aren't any NUL bytes. It is easy to guess that it might be UTF-16 if every other byte is a NUL byte. Guessing that some text might be encoded in one of the EBCDIC codesets might not be too hard, but correctly guessing which variant is another matter. And, other than that, good luck. The differences between the various 8859-* character sets is only obvious to most people if you know what the text in the file is supposed to be beforehand.

Don Cragun

View Public Profile for Don Cragun

Find all posts by Don Cragun

02-27-2018

Registered User

472, 104

Join Date: Aug 2006

Last Activity: 19 October 2018, 6:30 AM EDT

Posts: 472

Thanks Given: 4

Thanked 104 Times in 95 Posts

Hi,

the file command uses a file called magic to identify the file type. According to the POSIX man page the -m flag can be used to specify an own magic file and I think the AIX file command supports this flag too.
Maybe you can obtain or create a magic file that fits your needs.

cero

View Public Profile for cero

Find all posts by cero

02-27-2018

Registered User

12,315, 4,560

Join Date: Jul 2012

Last Activity: 22 November 2019, 4:29 PM EST

Location: San Jose, CA, USA

Posts: 12,315

Thanks Given: 952

Thanked 4,560 Times in 3,818 Posts

Hi cero,
The file magic file is used to identify things like executable file formats (that have certain fixed binary values at fixed locations in a file). It is great for various a.out files, music files, photographic files, PDF files, word processing files, spreadsheet files, and similar things with fixed headers.

The file utility uses other built-in knowledge when trying to identify the language or codeset used in a text file. It sounds like the Linux file utility has some code built-in that does a better job of guessing at codesets underlying a text file than the AIX file utility for the files that JeanM-1 is processing. Whether or not the GNU file utility source would build correctly on AIX is something JeanM-1 may want to investigate. But, copying a Linux magic file to AIX and having the AIX file utility use it instead of AIX's default magic file isn't likely to make any difference for this issue.

Don Cragun

View Public Profile for Don Cragun

Find all posts by Don Cragun

02-27-2018

Registered User

472, 104

Join Date: Aug 2006

Last Activity: 19 October 2018, 6:30 AM EDT

Posts: 472

Thanks Given: 4

Thanked 104 Times in 95 Posts

Hi Don,
thanks for pointing out that my post was not that clear - I did not suggest to copy the Linux magic file to AIX (and am nearly sure that it would not work), but my post can be interpreted like that.
The magic file would be my starting point if I'd have to solve JeanM-1's problem (and the Byte Order Mark the first thing I'd start playing around with to see if I get anywhere).

cero

View Public Profile for cero

Find all posts by cero

UNIX for Beginners Questions & Answers

Convert files to UTF-8 on AIX 7.1

10 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

Convert UTF-8 file to ASCII/ISO8859-1 OR replace characters

Discussion started by: hemkiran.s

2. Shell Programming and Scripting

Trying to convert utf-8 to WINDOWS-1251

Discussion started by: umen

3. AIX

Install EN_GB UTF-8 on AIX 5.3

Discussion started by: ningy

4. Linux

Help to Convert file from UNIX UTF-8 to Windows UTF-16

Discussion started by: phanidhar6039

5. OS X (Apple)

Changing txt files to pure UTF-8

Discussion started by: sovdia

6. AIX

How to print UTF-8 from AIX (lp)

Discussion started by: burnAF

7. Red Hat

Can't convert 7bit ASCII to UTF-8

Discussion started by: rockf1bull

8. UNIX for Advanced & Expert Users

Convert UTF-8 encoded hex value to a character

Discussion started by: sumirmehta

9. UNIX for Dummies Questions & Answers

Finding files with UTF-8 BOM

Discussion started by: kotoponus

10. Programming

Howto convert Ascii -> UTF-8 & back C++

Discussion started by: macron