# Convert string encoding to UTF-8

**URL:** <https://discourse.julialang.org/t/convert-string-encoding-to-utf-8/57782>\
**Category:** New to Julia\
**Tags:** strings\
**Created:** [March 23, 2021, 12:54pm UTC](https://discourse.julialang.org/t/convert-string-encoding-to-utf-8/57782 "2021-03-23T12:54:06Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Thomas\_Lei](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thomas_lei/32/20372_2.png) [@Thomas\_Lei](https://discourse.julialang.org/u/Thomas_Lei)\
**Post date:** [March 23, 2021, 12:54pm UTC](https://discourse.julialang.org/t/convert-string-encoding-to-utf-8/57782/1 "2021-03-23T12:54:06Z")

</div>

Hi all,

Somewhere in my code I read strings which contains unicode characters.  
For example, the reading output is:  
`myString = "caf\xe9"`

I’d like this string to be parsed as  
`myString = "café"`

I played along with [StringEncodings.jl](https://github.com/JuliaStrings/StringEncodings.jl) but with no sucess. Any ideas how I can do that ?

Thanks

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [March 23, 2021, 1:15pm UTC](https://discourse.julialang.org/t/convert-string-encoding-to-utf-8/57782/2 "2021-03-23T13:15:00Z")

</div>

> [@Thomas\_Lei](#):
>
> Somewhere in my code I read strings which contains unicode characters.  
> For example, the reading output is:  
> `myString = "caf\xe9"`

That’s the [Latin-1 encoding](https://en.wikipedia.org/wiki/ISO/IEC_8859-1) of `"café"`, not Unicode. You need to convert this to the UTF-8 encoding of Unicode as used by Julia. I’m guessing that you are on Windows and that this is actually [Windows-1252](https://en.wikipedia.org/wiki/Windows-1252), since Latin-1 is not common anymore elsewhere.

Two options:

1. Read it into Julia as bytes (`Vector{UInt8}` via `read(io)`) and convert the encoding with [StringEncodings.jl](https://github.com/JuliaStrings/StringEncodings.jl) or some similar package. e.g. `decode(Vector{UInt8}("caf\xe9"), "Windows-1252")` gives `"café"`.

2. Change your files to use UTF-8. Windows-1252 is an archaic encoding that can only encode 256 characters, nowhere near all of Unicode. People should really stop using it. See [here](https://stackoverflow.com/questions/2014069/windows-1252-to-utf-8-encoding) for various tools. (e.g. For a single file, you can just open it in Notepad or some other editor and choose “UTF-8” when you save, but there are also batch tools to re-encode many files at once.)

(In any case, this not not technically about “parsing”, which is a distinct concept from [“encoding”](https://en.wikipedia.org/wiki/Character_encoding).)

---

<div class="post-metadata">

**Author:** ![Thomas\_Lei](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thomas_lei/32/20372_2.png) [@Thomas\_Lei](https://discourse.julialang.org/u/Thomas_Lei)\
**Post date:** [March 23, 2021, 4:54pm UTC](https://discourse.julialang.org/t/convert-string-encoding-to-utf-8/57782/3 "2021-03-23T16:54:12Z")

</div>

Indeed, the encoding is “windows-1252”. I can’t change the files encoding to UTF-8 as they are produced by an external system which I don’t have access to…

If anyone has the same question, this works like a charm:

```julia
using StringEncodings
my_string = "caf\xe9"
my_string = decode(Vector{UInt8}(my_string), "windows-1252")

```

Thanks !  
(I also edited the post title)
