# String encodings help

**URL:** <https://discourse.julialang.org/t/string-encodings-help/8188>\
**Category:** General Usage\
**Created:** [January 5, 2018, 9:10pm UTC](https://discourse.julialang.org/t/string-encodings-help/8188 "2018-01-05T21:10:11Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [January 5, 2018, 9:10pm UTC](https://discourse.julialang.org/t/string-encodings-help/8188/1 "2018-01-05T21:10:11Z")

</div>

I’m looking for some help in determining what character sets are compatible with UTF-8 and therefore safe to be transcoded from Vector{UInt8} using the String constructor. I’m not an expert with character coding and not sure how to accurately do that. Any help or pointer to references is much appreciated!

These are the encodings identifiable from a SAS data file using my SASLib.jl package. I believe that US-ASCII, ISO-8859-1 and WINDOWS-1252 are OK.  
What about those variations of ISO-8859-x and WINDOWS-125x?

“BIG-5”  
“CP932”  
“EUC-JP”  
“EUC-KR”  
“EUC-TW”  
“GB18030”  
“ISO-8859-1”  
“ISO-8859-11”  
“ISO-8859-2”  
“ISO-8859-3”  
“ISO-8859-6”  
“ISO-8859-7”  
“ISO-8859-8”  
“ISO-8859-9”  
“US-ASCII”  
“WINDOWS-1250”  
“WINDOWS-1251”  
“WINDOWS-1252”  
“WINDOWS-1253”  
“WINDOWS-1254”  
“WINDOWS-1255”  
“WINDOWS-1256”  
“WINDOWS-1257”

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [January 5, 2018, 11:15pm UTC](https://discourse.julialang.org/t/string-encodings-help/8188/2 "2018-01-05T23:15:02Z")

</div>

I’m mistaken… anything other than US-ASCII would have to be converted to UTF-8 somehow. Just found a great post from before.

> [@Readstring encoding](https://discourse.julialang.org/t/readstring-encoding/4245):
>
> Just got a project with quite a few iso8859-1 encoded files. Is there a concise way to readstring a given file in Julia that is not utf-8? UPDATED: forgot about StringEncodings.jl

@ScottPJones, is your neat conversion utility already in a package or something that I can reuse?

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [January 5, 2018, 11:36pm UTC](https://discourse.julialang.org/t/string-encodings-help/8188/3 "2018-01-05T23:36:18Z")

</div>

I must be blinded to not catch this.  
[https://github.com/JuliaString/Strs.jl](https://github.com/JuliaString/Strs.jl)

Great work, Scott!

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [January 6, 2018, 3:10am UTC](https://discourse.julialang.org/t/string-encodings-help/8188/4 "2018-01-06T03:10:46Z")

</div>

I think [StringEncodings.jl](https://github.com/nalimilan/StringEncodings.jl) is currently the only Julia package for converting non-Unicode encodings to UTF-8. Scott’s Strs.jl package mentions that he is working on a StrEncodings.jl package for conversion also, but as far as I can tell nothing has been posted yet. @nalimilan seems to have kept StringEncodings.jl up-to-date, however, and it is based on the robust and well-tested `iconv` library; I see no reason not to use it.

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [January 6, 2018, 3:33am UTC](https://discourse.julialang.org/t/string-encodings-help/8188/5 "2018-01-06T03:33:43Z")

</div>

If they are really ISO-8859-1, it’s actually very easy - read in the bytes, and then make a vector of `Char`s out of them, and then make a `String` out of that. That works because 8859-1 is a pure 8-bit subset of Unicode (as is ASCII a pure 7-bit subset). With the `Strs.jl` package, can simply read them in as `LatinStr` directly.

The problem comes in if it’s not really 8859-1, but rather Microsoft’s CP-1252, which is almost the same as 8859-1, but adds a bunch of printable characters (such as the Euro sign € at 0x80). Then you’d need a (simple) mapping table for the code points between 0x80 and 0x9f.

It may be a few weeks before I get conversions added, I want to do that in pure Julia, in a way that tables can be loaded on the fly (and also supply some utility functions to look at what something like `iconv` or `ICU` has for mappings, and build a compressed table that can be used later by my `StrEncodings`).  
If you want to get a feel for how that would work, please take a look at what I did to read and process things like the Unicode data file, as well as other sources for HTML entities, LaTeX entities, and Emojis, in some of the other packages I did in the [JuliaString](https://github.com/JuliaString) organization.

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [January 6, 2018, 7:40am UTC](https://discourse.julialang.org/t/string-encodings-help/8188/6 "2018-01-06T07:40:13Z")

</div>

It’s still very WIP now, needs a lot more testing, benchmarking, and optimization done.  
Hopefully shortly it will be something people can make use of.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [January 6, 2018, 8:09am UTC](https://discourse.julialang.org/t/string-encodings-help/8188/7 "2018-01-06T08:09:06Z")

</div>

> [@stevengj](#):
>
> `@`nalimilan seems to have kept StringEncodings.jl up-to-date, however, and it is based on the robust and well-tested iconv library; I see no reason not to use it.

Depending on the problem, it may also be viable to use `iconv` or `recode` directly, _once_, on the file(s), obtain UTF8 and forget about the whole issue afterwards. This is what I usually do, last non-UTF8 files I have seen were from 10+ years ago.

---

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [January 6, 2018, 8:21am UTC](https://discourse.julialang.org/t/string-encodings-help/8188/8 "2018-01-06T08:21:36Z")

</div>

I already depend on StringEncodings.jl but I am trying to optimize performance by avoiding it when there’s a faster path. As Scott pointed out, certain conversions are quite simple & easy to implement so it would be a nice win.
