# Unicode diacritical marks in filenames

**URL:** <https://discourse.julialang.org/t/unicode-diacritical-marks-in-filenames/124392>\
**Category:** Offtopic\
**Tags:** unicode\
**Created:** [December 23, 2024, 2:55pm UTC](https://discourse.julialang.org/t/unicode-diacritical-marks-in-filenames/124392 "2024-12-23T14:55:59Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![kellertuer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kellertuer/32/220707_2.png) [@kellertuer](https://discourse.julialang.org/u/kellertuer)\
**Post date:** [December 23, 2024, 2:55pm UTC](https://discourse.julialang.org/t/unicode-diacritical-marks-in-filenames/124392/1 "2024-12-23T14:55:59Z")

</div>

Anything with diacritics (dots hats, wiggles above/below) are a bit complicated

- you can write the German ö as exactly [that letter](https://www.compart.com/en/unicode/U+00F6)
- or use [“◌߳” U+07F3 Nko Combining Double Dot Above Unicode Character](https://www.compart.com/en/unicode/U+07F3) before/with an o.

That is probably the main challenge here, and yeah, one should be careful with these. I once had a samba connection from Mac OS to Linux that exactly changed between the two encodings above, which messed up quite some stuff. Probably Browsers struggle the same.

But sure anything single letter, greek, fraktur, caligraphic,… should be fine.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [January 2, 2025, 11:34pm UTC](https://discourse.julialang.org/t/unicode-diacritical-marks-in-filenames/124392/2 "2025-01-02T23:34:35Z")

</div>

> [@kellertuer](#):
>
> or use [“◌߳” U+07F3 Nko Combining Double Dot Above Unicode Character](https://www.compart.com/en/unicode/U+07F3) before/with an o.

According to the Unicode standard, this has a different semantic meaning than the diacritic in `ö` — this mark is used for the [N’Ko script of West-African languages](https://en.wikipedia.org/wiki/N%27Ko_script). If you want ö, you should be using [U+0308 “Combining Diaeresis”](https://www.compart.com/en/unicode/U+0308)

There are still two ways to write `ö`, but `"\u00f6"` and `"o\u0308"` are [canonically equivalent](https://en.wikipedia.org/wiki/Unicode_equivalence) according to Unicode (they have the same semantic meaning), and are converted to one another if you normalize the string. You definitely need to have some understanding of normalization when comparing strings in Unicode or when dealing with Unicode filenames across filesystems.

(I don’t think this has anything to do with support for ö in README.md, so I’ve moved it to a new topic.)

> [@kellertuer](#):
>
> I once had a samba connection from Mac OS to Linux that exactly changed between the two encodings above

This sounds implausible to me. There is no Unicode normalization that should convert the N’Ko combining character into an umlaut or vice versa.

Probably you are mis-remembering, and actually encountered U+0308, which indeed could have gotten normalized away (or normalized into existence) in a filename transferred across systems.

In particular, MacOS HFS+ filenames are NFD-normalized, so `ö` will get converted into `o` followed by a combining character U+0308 when it is saved to a MacOS HFS+ filesystem. Whereas most Linux filesystems perform no normalization by default IIRC.

(It would make a lot more sense for me for a filesystem to be normalization-insensitive _and_ normalization-preserving, but I don’t know if there is any filesystem that does this? _Update_: it appears that [ZFS can do this](https://docs.oracle.com/cd/E19120-01/open.solaris/817-2271/6mhupg6ma/index.html), but only if you set the `normalization=formD` property when the filesystem is created … but this is not the default, probably because it forces filenames to be UTF-8 and [can conflict with legacy filenames](https://utcc.utoronto.ca/~cks/space/blog/linux/ForcedUTF8Filenames).)

> [@kellertuer](#):
>
> Probably Browsers struggle the same

I seriously doubt that any modern browser will be confused by Unicode normalization nowadays when it comes to _displaying_ text. They’ve had to deal with Unicode display for decades now; the main limitation is that they can’t render glyphs that don’t exist in the installed font(s).

---

<div class="post-metadata">

**Author:** ![kellertuer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kellertuer/32/220707_2.png) [@kellertuer](https://discourse.julialang.org/u/kellertuer)\
**Post date:** [January 3, 2025, 6:56am UTC](https://discourse.julialang.org/t/unicode-diacritical-marks-in-filenames/124392/3 "2025-01-03T06:56:00Z")

</div>

Thanks for the detailed answer. I am not sure _which_ unification it was, but yours surely sounds more plausible. It was while moving files from MacOS to Linux (via smb) and afterwards they were no longer found by my scripts. Sure, in the scripts, the unicode was not changed, in the file names it was normalised.

Whether browsers struggle with that, I do not know, that is correct.
