# How to get file encoding?

**URL:** <https://discourse.julialang.org/t/how-to-get-file-encoding/72325>\
**Category:** General Usage\
**Tags:** strings, io\
**Created:** [December 1, 2021, 12:00am UTC](https://discourse.julialang.org/t/how-to-get-file-encoding/72325 "2021-12-01T00:00:46Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![mathdeep](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mathdeep/32/29510_2.png) [@mathdeep](https://discourse.julialang.org/u/mathdeep)\
**Post date:** [December 1, 2021, 12:00am UTC](https://discourse.julialang.org/t/how-to-get-file-encoding/72325/1 "2021-12-01T00:00:46Z")

</div>

Hi there,

Is there a way to find out how a file is encoded with functions?  
I’m getting a text file and want to do sume stuff. In order to do right stuff, I have to know how the file is encoded, but I can’t find how to do this. I found that I could open the file, go to “save as” and what ever the standard encoding there is, is what I’m looking for, but I can’t find to way to let a program do this.  
Best case scenario would be, if I could convert everything to UTF8 (in every file are only UTF8 managable characters).  
I tried a little with the StringEncodings package, but it isn’t really working.  
Any ideas? Help would be much appreciated.

Thanks

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [December 1, 2021, 12:12am UTC](https://discourse.julialang.org/t/how-to-get-file-encoding/72325/2 "2021-12-01T00:12:36Z")

</div>

> [@mathdeep](#):
>
> Is there a way to find out how a file is encoded with functions?

not really, on Linux, you can use `file -i`, but in general there’s no way to detect for certain.

* * *

To use `StringEncodings`, you need to know the name of the encoding of course:

```julia
akako@desktop ~/tmp> file -i utf.txt
utf.txt: text/plain; charset=utf-8
akako@desktop ~/tmp> iconv -t GBK utf.txt -o gbk.txt

julia> String(read("./utf.txt"))
"hello甲\n"

julia> String(read("./gbk.txt"))
"hello\xbc\xd7\n"

julia> decode(read("./gbk.txt"), "GBK")
"hello甲\n"

```

---

<div class="post-metadata">

**Author:** ![mathdeep](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mathdeep/32/29510_2.png) [@mathdeep](https://discourse.julialang.org/u/mathdeep)\
**Post date:** [December 1, 2021, 12:33am UTC](https://discourse.julialang.org/t/how-to-get-file-encoding/72325/3 "2021-12-01T00:33:37Z")

</div>

Thanks!

For the first I’d say it might be enough, if this works on Windows. Would this help?

To the second: I thought maybe something like checking whether everything is readable and if not changing the encoding until it is would be possible. Well I don’t know. I just found isvalid() and I know that for some not readable Chars it get’s something like “Malformed UTF-8 Character” displayed. So I thought with one of these it might be possible to check.

Any ideas if there are more restrictions? I really don’t know how to work with these files otherwise.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [December 1, 2021, 7:00am UTC](https://discourse.julialang.org/t/how-to-get-file-encoding/72325/4 "2021-12-01T07:00:52Z")

</div>

Python’s chardet module might help - see [suggestion here](https://superuser.com/a/301562). Maybe you could try calling it with PyCall.

Or with expert help, it should be faster calling the uchardet [C library here](https://github.com/freedesktop/uchardet).

_ **NB** : maybe there is already such Julia package?_

---

<div class="post-metadata">

**Author:** ![mathdeep](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mathdeep/32/29510_2.png) [@mathdeep](https://discourse.julialang.org/u/mathdeep)\
**Post date:** [December 1, 2021, 9:02pm UTC](https://discourse.julialang.org/t/how-to-get-file-encoding/72325/5 "2021-12-01T21:02:35Z")

</div>

Thanks for the tip!  
It works kind of ok, I guess. I was reminded why I don’t like C, but most of it works now. I still sometimes get problems when a char can’t at least be converted to UTF-16. Does something like print(myString,enc"ENCODING") like in StringEncodings exist? Am I correct though that I won’t be able the make an .exe file when calling C?  
Another thing: Imagine you have a text file, which was wrongly encoded, so now you just see something like “?äüß$§%”. If I understand the StringEncodings package correctly, I should be able to try and convert this to every known encoding and look which was is fine, right?? How???

---

<div class="post-metadata">

**Author:** ![ellocco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ellocco/32/31331_2.png) [@ellocco](https://discourse.julialang.org/u/ellocco)\
**Post date:** [March 24, 2026, 9:54am UTC](https://discourse.julialang.org/t/how-to-get-file-encoding/72325/6 "2026-03-24T09:54:46Z")

</div>

Package `chardet` is quite weak, I prefere `cchardet` instead.  
Here my code:

```julia-auto
cchardet = pyimport("cchardet")

function detect_encoding_cchardet(filepath::String)
    data = read(filepath)
    data_bytes = pybytes(data) # Explizite Konvertierung zu Python bytes
    result = cchardet.detect(data_bytes)
    return result
end

fn = raw"C:\temp\your_file.asc"
encoding_result = detect_encoding_cchardet(fn)
println("Encoding: $(encoding_result["encoding"])")
println("Confidence: $(encoding_result["confidence"])")

```

I wonder, if also a suitable Julia function exist.
