# Warning against Unicode confusables

**URL:** <https://discourse.julialang.org/t/warning-against-unicode-confusables/108734>\
**Category:** Internals & Design\
**Tags:** unicode\
**Created:** [January 9, 2024, 2:19am UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734 "2024-01-09T02:19:49Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [January 12, 2024, 5:13pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/21 "2024-01-12T17:13:28Z")

</div>

Again, I repeat: With a good programming editor, I, l, 1 are distinguishable at first sight. No copying and pasting into the repl to figure it out is necessary. No malicious (or stupid) substitutions are possible.  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/3/3/3335b724bd71601f218a93a6f973be5937c2b08a.png)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [January 12, 2024, 5:14pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/22 "2024-01-12T17:14:49Z")

</div>

> [@Mason](#):
>
> Julia normalizes `ℓ`, and `ℓ` to the same letter

`ℓ` and `ℓ` _are_ the same letter U+2113. Perhaps Petr meant `\ell = ℓ` = U+2113 and `\scrl = 𝓁` = U+1D4C1, which are not normalized to equivalent identifiers (`:ℓ != :𝓁`).

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [January 12, 2024, 5:15pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/23 "2024-01-12T17:15:37Z")

</div>

> [@PetrKryslUCSD](#):
>
> At the danger of repeating myself once too often: Unicode used to write source code (not as data or as commentary) is a BAD IDEA.

I think we need to be careful with statements like this. I believe you want to encourage programmers to rarely use obscure symbols.

However, you are also effectively saying that all programmers should only program in English as supported by ASCII or perhaps Latin-1 (ISO-8859-1). That’s a very western “European-origin” perspective that we need to be be conscious about. While American English has become a defacto standard for international correspondence and thus programming, there is validity in people also writing programs in their native human languages. I know you likely did not intend the statement to discount this.

To the larger question, I think we should aim on making code as readable as possible as a general goal. There’s some question of audience as well.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [January 12, 2024, 5:16pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/24 "2024-01-12T17:16:49Z")

</div>

> [@mkitti](#):
>
> While American English has become a defacto standard for international correspondence and thus programming, there is validity in people also writing programs in their native human languages.

Which might make it impossible for speakers of other languages to understand their programs. 😉

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [January 12, 2024, 5:19pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/25 "2024-01-12T17:19:10Z")

</div>

[Some browsers](https://stackoverflow.com/questions/63791942/prevent-safari-from-normalizing-unicode-when-sending-post-data) perform NFC unicode normalization — much like Julia does — when sending discourse replies. It’s likely there was a third distinct codepoint that was attempted to be entered here… but that’d be a moot point as Julia would normalize it similarly, avoiding the confusion for whatever that particular original codepoint was.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [January 12, 2024, 5:20pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/26 "2024-01-12T17:20:25Z")

</div>

Yes, indeed. A copy-paste error. Sorry.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [January 12, 2024, 5:37pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/27 "2024-01-12T17:37:15Z")

</div>

> [@PetrKryslUCSD](#):
>
> At the danger of repeating myself once too often: Unicode used to write source code (not as data or as commentary) is a BAD IDEA.

And to repeat myself: Unicode improves legibility and clarity of code, so it’s a _good idea_ to use it responsibly.

---

<div class="post-metadata">

**Author:** ![goerz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerz/32/3269_2.png) [@goerz](https://discourse.julialang.org/u/goerz)\
**Post date:** [January 12, 2024, 5:38pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/28 "2024-01-12T17:38:43Z")

</div>

> [@PetrKryslUCSD](#):
>
> However, reading someone else’s code is different: It is not obvious when reading the source code which of the unicode characters they chose.

I use vim, and if I hit `ga` it will show me detailed information about the character the cursor is on. I imagine other editors have similar features.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [January 12, 2024, 5:51pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/29 "2024-01-12T17:51:49Z")

</div>

This seems likely to go in circles forever to no productive end. Unicode exists. It’s not going anywhere. Style guides can and do address the matter. You can choose to endorse or enforce a particular style guide for code you control. And you can kindly suggest others follow some style guides, but with the understanding that it’s an opinion-based choice. You can start #Tooling topics about better linting/editing/refactoring tools.

Arguing on the internet isn’t going to change opinions on the matter… and there are plenty of other places to argue ad nauseam. Let’s set a timer to prevent this from going on indefinitely and potentially escalating unnecessarily here.

---

<div class="post-metadata">

**Author:** ![MilesCranmer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/milescranmer/32/21070_2.png) [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Post date:** [January 12, 2024, 6:19pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/30 "2024-01-12T18:19:32Z")

</div>

I apologize in advance for piling onto this but I thought I’d add a point I didn’t see earlier. I guess the usual point of these types of threads is to push the needle on the “stylistic culture” of Julia in one particular direction, so technically I see this all as productive (in some sense of the word :)).

In general I worry the popular use of unicode in Julia may be ever-so-slightly hurting Julia’s rate of adoption. While it may be supported in other languages, it is much more common in Julia including in the standard library.

I remember when I was first starting out with Julia, I looked at some ODE solver examples, and saw the use of unicode everywhere (for variable names, and some API use like \epsilon). I remember this leaving me with a poor first impression of Julia based on my naive assumptions at that time – “How can I even write Julia code in my editor if it needs unicode? I can’t just remember all the commands…” You have to keep in mind that a person encountering Julia for the first time is just going to skim some snippets and not really read deep into the complexities of the documentation and all the best tooling practices until they’ve already made the leap (if you’ve ever looked at Google analytics for a blog, it can be a bit depressing!) In my mind I only knew that ⌥+S is ß (which in retrospect is the German beta, rather than the math one…) so I got initially discouraged.

Of course it’s clear we should be setting up our tooling correctly to work with this, but I do want to point out that seeing the use of unicode for the first time in code can be a bit of a distraction to the best parts of Julia, and it might leave users with a negative impression. I think the vast majority of beginners haven’t coded in unicode before so it’s quite weird to see it the first time.

The only other language I can think of with such prevalant unicode is Lean: [100 theorems in Lean](https://leanprover-community.github.io/100.html), but this makes a bit more sense to me as it is exclusively about Maths and the visual presentation of theorems, whereas Julia is a more general programming language. Thus Julia should hope to attract a more general audience as well (who might similarly recoil at the sight of unicode math).

* * *

Slightly tangential but I was trying to read through the Julia source code today in the method `abstract_eval_statement_expr`: [julia/base/compiler/abstractinterpretation.jl at 5b6a94da5af35a4aa91759cac2f8db7669a6ec2a · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/blob/5b6a94da5af35a4aa91759cac2f8db7669a6ec2a/base/compiler/abstractinterpretation.jl#L2417)

This has unicode which is apparently failed to render on my very modern MacBook Pro in Firefox:

 ![Screenshot 2024-01-12 at 17.20.21](https://global.discourse-cdn.com/julialang/original/3X/a/c/acc3f9be68ddb41e16cb2f28378e1e7b0dcc1262.png)

I have no idea why GitHub can’t render it but stuff like this is always a bit of a negative for unicode imo.

(Not to mention I still have no idea how to write unicode math on my phone – which I will occasionally use for reviewing PRs)

---

<div class="post-metadata">

**Author:** ![goerz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerz/32/3269_2.png) [@goerz](https://discourse.julialang.org/u/goerz)\
**Post date:** [January 12, 2024, 6:23pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/31 "2024-01-12T18:23:18Z")

</div>

> [@MilesCranmer](#):
>
> In general I worry the popular use of unicode in Julia may every so slightly hurting Julia’s rate of adoption. While it may be supported in other languages, it is much more common in Julia including in the standard library.

As a counter-point, unicode (as well as the other “controversial” feature of 1-based indexing) were among the things that very much attracted me to Julia. As a computational scientist, it was a strong signal that this language is _designed_ for scientific computing and that I should invest in it.

---

<div class="post-metadata">

**Author:** ![MilesCranmer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/milescranmer/32/21070_2.png) [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Post date:** [January 12, 2024, 6:34pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/32 "2024-01-12T18:34:05Z")

</div>

Thanks, that is a useful perspective. I guess we could coarsely simplify this to the following inference problem:

- P(starts using julia | likes unicode)
- P(starts using julia | hates unicode)
- P(likes unicode)
- P(hates unicode)

if we only care about attracting more users (inb4 this is obviously not a complete objective, I am aware), then say we optimize P(starts using julia).

My intuition is that P(hates unicode) \> P(likes unicode) among general programmers, but perhaps P(hates unicode) ~ P(likes unicode) among computational scientists like yourself?

The remaining question is P(starts using julia | likes unicode) and P(starts using julia | hates unicode). I guess, how much does the unicode really matter if you **do** like it? Does it really impact your decision to like Julia?

And the other hand, how much does the unicode matter if you **don’t** like it?

My very rough guesstimation (which could be biased to my own taste!) is that P(starts using julia | hates unicode) is much smaller than P(starts using julia | likes unicode) is large… And therefore we should be more wary about using unicode in APIs and examples based on that group. But I really am not sure in any of this.

I guess the other question is how much we want to get more general programmers into Julia which is historically very heavy on the computational scientist side? (And would that change this analysis…?)

---

<div class="post-metadata">

**Author:** ![goerz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerz/32/3269_2.png) [@goerz](https://discourse.julialang.org/u/goerz)\
**Post date:** [January 12, 2024, 6:41pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/33 "2024-01-12T18:41:51Z")

</div>

> [@MilesCranmer](#):
>
> I guess the other question is how much we want to get more general programmers into Julia which is historically very heavy on the computational scientist side? (And would that change this analysis…?)

I’ve always taken the “Julia is a general purpose language” with a grain of salt. It feels very much like a domain-specific language to me (for the domain of scientific computing, understood broadly). And that’s fine! I believe “embrace the niche” is a good motto.

I’m biased, of course.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [January 12, 2024, 6:43pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/34 "2024-01-12T18:43:26Z")

</div>

I’m sure someone has pointed it out already, but the use of the term “Unicode” here seems to have adopted the meaning “the subset of Unicode that isn’t on my ASCII keyboard.” Even if we dismiss mathematics, physics, and non-Western languages (which we absolutely should not), keyboard layouts will differ even within Western languages. People have figured out how to avoid mixing visually similar Unicode characters for their symbols, and if someone doesn’t immediately recognize which in an unfamiliar context, they need only learn from other people. ASCII-ifying everything not only doesn’t make sense in many contexts, it’s not even _feasible_ to do.

> [@PetrKryslUCSD](#):
>
> > [@mkitti](#):
> >
> > While American English has become a defacto standard for international correspondence and thus programming, there is validity in people also writing programs in their native human languages.
> 
> Which might make it impossible for speakers of other languages to understand their programs. 😉

You can say the same about any language. Julia is undoubtedly based in English and requires some English knowledge to interact with, but even people who learned English may find it more feasible to program and communicate to their colleagues in their native language. With enough aliasing, even most variables can be renamed in other languages, which helps to align with docstrings and comments.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [January 12, 2024, 6:45pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/35 "2024-01-12T18:45:11Z")

</div>

At its core, there’s really nothing special about unicode here. It’s all about picking good names — names that **your audience** will understand. See [Naming is hard, let’s do better](https://discourse.julialang.org/t/kate-gregory-naming-is-hard-lets-do-better-and-what-lesson-we-can-learn-from-it/108216).

I think everyone should be able to appreciate that audiences differ wildly throughout the Julia community, with many natural-language and domain-specific subgroups.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [January 12, 2024, 6:48pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/36 "2024-01-12T18:48:05Z")

</div>

> [@MilesCranmer](#):
>
> I remember this leaving me with a poor first impression of Julia based on my naive assumptions at that time – “How can I even write Julia code in my editor if it needs unicode? I can’t just remember all the commands…”

I agree that this is a real concern — as I wrote in [another post](https://discourse.julialang.org/t/syntax-escape-hatch-for-unicode-haters/108363/80), it’s important for newcomers to realize that they _can_ use Unicode symbols in their code, but they are not _required_ to do so in order to program in Julia.

The question is, what are **actionable ways** to help address this concern? Unicode isn’t going away — many people in computational science are not going to go back to writing `alpha` and `beta` as variable names once they realize that they can write `α` and `β`, nor should they have to. Some possible things to work on might be:

- In tutorial materials aimed at beginners, encourage people to be cautious about Unicode symbols. (You _should_ introduce them at some point, because people will see them in other Julia code, but make sure to emphasize that they are optional, and definitely don’t include them without explanation.) I recently submitted a PR to the _Think Julia_ book to correct just such an issue ([replace "🐢" with "turtle" in examples by stevengj · Pull Request #61 · BenLauwens/ThinkJulia.jl · GitHub](https://github.com/BenLauwens/ThinkJulia.jl/pull/61)), though I’m not sure when/whether it will be merged. Clarifying patches submitted to other tutorials might be helpful too.
- Something in the Julia manual to emphasize this? A new FAQ, or edits somewhere else, a blog post? If you have a good idea for this, please feel free to submit a PR aimed at clarifying this issue to newcomers.
- Maybe an addition to the [style guide](https://docs.julialang.org/en/v1/manual/style-guide/), suggesting that public APIs should typically be ASCII or have ASCII synonyms for accessibility. (Of course, people in specialized fields may choose to disregard this, just as people can disregard anything in the style guide if they wish. The point is not to shame people, but to make sure they appreciate the tradeoffs here.)
- If there are isolated Unicode-only APIs in packages (i.e. some random little thing, not a package designed top-to-bottom to use Unicode symbols), a PR to add an ASCII symbol might well be welcome. (No, [`xor`](https://docs.julialang.org/en/v1/base/math/#Base.xor) doesn’t count here: lack of an infix operator ≠ lack of an API.)
- Implement warnings about confusable symbols in linting tools. (I opened an [issue for Aqua.jl](https://github.com/JuliaTesting/Aqua.jl/issues/259)).
- Others?

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [January 12, 2024, 6:49pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/37 "2024-01-12T18:49:24Z")

</div>

Since it hasn’t been mentioned before, I believe @PetrKryslUCSD s point is what is known as a homograph attack in security circles. See also:

> **[IDN homograph attack](https://en.m.wikipedia.org/wiki/IDN_homograph_attack)**
>
> The internationalized domain name (IDN) homograph attack is a way a malicious party may deceive computer users about what remote system they are communicating with, by exploiting the fact that many different characters look alike (i.e., they are homographs, hence the term for the attack, although technically homoglyph is the more accurate term for different characters that look alike. For example, the Cyrillic, Greek and Latin alphabets each have a letter ⟨o⟩ that has the same shape but different...

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [January 12, 2024, 6:55pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/38 "2024-01-12T18:55:08Z")

</div>

Note that the first examples of a homograph attack all involve mimicking ASCII characters. Confusability is not a reason to differentiate ASCII and the rest of Unicode. Barring malice, people have no issue avoiding similar symbols e.g. typing Cyrillic doesn’t accidentally involve Latin.

---

<div class="post-metadata">

**Author:** ![Ininterrompue](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ininterrompue/32/5594_2.png) [@Ininterrompue](https://discourse.julialang.org/u/Ininterrompue)\
**Post date:** [January 12, 2024, 7:02pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/39 "2024-01-12T19:02:17Z")

</div>

> [@MilesCranmer](#):
>
> My very rough guesstimation (which could be biased to my own taste!) is that P(starts using julia | hates unicode) is much smaller than P(starts using julia | likes unicode) is large… And therefore we should be more wary about using unicode in APIs and examples based on that group. But I really am not sure in any of this.

My speculation is that, because Julia allows for such liberal usage of Unicode characters, it tends to attracts an audience that use them. Along with this, some can (and do) go overboard with their usage. It’s clear to me that not all usage improves readability. Some may dispute this, but it is absolutely true to me and many others. Already in that screenshot of yours I see problems to readability besides the failure to render: the blackboard L easily confused with capital L, and the \subset symbol, which, I know what the symbol is in isolation, but have no idea what it’s doing in this context. Plenty of other examples I could list off the top of my head that are prone to confusion. Anyway, I think your suspicions that it’s hurting first impressions and adoption is right on point.

---

<div class="post-metadata">

**Author:** ![Jake](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jake/32/46007_2.png) [@Jake](https://discourse.julialang.org/u/Jake)\
**Post date:** [January 12, 2024, 7:15pm UTC](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734/40 "2024-01-12T19:15:30Z")

</div>

At the risk of injecting a bit of humor (philosophy?), my dad (born in 1923) would say “Never died when man walked on the moon!” I understand the statement came from people that said something to effect of “that won’t happen until man walks on the moon.”. And we all know what happened…

This Unicode discussion reminded me of his comment.

[Previous page](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734.md?page=1)

[Next page](https://discourse.julialang.org/t/warning-against-unicode-confusables/108734.md?page=3)
