# What happened to parsing Unicode characters?

**URL:** <https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047>\
**Category:** VS Code\
**Tags:** question, unicode\
**Created:** [June 20, 2025, 1:03am UTC](https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047 "2025-06-20T01:03:27Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![jmanthony3](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmanthony3/32/52953_2.png) [@jmanthony3](https://discourse.julialang.org/u/jmanthony3)\
**Post date:** [June 20, 2025, 1:03am UTC](https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047/1 "2025-06-20T01:03:27Z")

</div>

I don’t know what happened between the last time I opened this particular code project (I tabled it for a time to work on something else): but now the Unicode characters look funny. Since I haven’t updated my Julia version in a long while, my best guess is that it is a VS Code extension or some other Julia package that has monkeyed with something in the backend. Unfortunately, I cannot find a specific commit or issue that references updating the parsing of Unicode characters whether in this repository or others. But it used to be that the Unicode for `\dot` appeared directly above the character preceding it: e. g. `\epsilon\dot` would put the dot over epsilon rather than after. And the same goes for the underbars appearing underneath the character before. And the `\scrf` Unicode looks very different.

 ![image](https://global.discourse-cdn.com/julialang/original/3X/b/c/bcfcb9c1105fa399c56aeb6097bc6e53d67ccc19.png)

I do have an older commit in the JuliaSyntax.jl repo for another parsing issue, but I’m not sure if this is related.

> <https://github.com/JuliaLang/JuliaSyntax.jl/issues/546>
>
> \*\*TL;DR:\*\* It seems that not all Unicode characters supported for tab-completion… are considered as valid symbols by the parser in the Julia REPL.
> 
> \## Goal
> In writing a function to model some inelastic deformation of metals, I have found it \*\*very\*\* helpful to include Unicode characters in my variable names to both keep consistent with equations published in literature and keep variables legible throughout the code. For example, writing two \`\\underbar\`s after a variable connotes a second-rank tensor. Periodically, I have to take the magnitude of a second-rank tensor and use it in other calculations so I write that to a variable.
> 
> \## Error
> I would love to use the \`\\Vert\` ($\\Vert$) symbol before and after the variable for a tensor, but the linter in VS Code complains that this is an "unknown unicode character" despite being filled from a tab-completion (see images below). The same messages appear if I attempt to use the \`\\mid\` ($\\mid$) symbol as well.
> 
> \*\*Tab-Completion:\*\*
> !\[Image\](https://github.com/user-attachments/assets/9e0cea05-ccc9-4066-8ffa-a75f6d914b17)
> 
> \*\*Linting Error Message:\*\*
> !\[Image\](https://github.com/user-attachments/assets/71a78447-a6b1-4802-a15f-5ed3648472f9)
> 
> !\[Image\](https://github.com/user-attachments/assets/f5a8979c-f38a-4142-a003-7ac9b9e6740f)
> 
> \## What I Have Found So Far
> It seems from the Julia manual for \[Unicode Input\](https://docs.julialang.org/en/v1/manual/unicode-input/) this character (U+02016 =\> \`\\Vert\` = $\\Vert$) is supported which I am guessing explains the tab-completion. I also tried creating a variable with the \`\\Vert\` symbol in a Julia REPL, but I get the same error message for an "unknown unicode character": 
> 
> !\[Image\](https://github.com/user-attachments/assets/d8fa4c50-b1f5-4225-9dc0-9ee9ba0dbbd8)
> 
> It also does not matter that this character is the first symbol in the variable assignment. It persists even if I put a valid character as the first symbol:
> 
> !\[Image\](https://github.com/user-attachments/assets/77fd6236-392a-42f0-9c9f-00130788813a)
> 
> \## Configuration
> \### Julia Version
> v1.11.4 (but I also saw this issue on v1.11.3)
> 
> \### VS Code
> Version: 1.99.1 (user setup)
> Commit: 7c6fdfb0b8f2f675eb0b47f3d95eeca78962565b
> Date: 2025-04-04T15:58:59.624Z
> Electron: 34.3.2
> ElectronBuildId: 11161073
> Chromium: 132.0.6834.210
> Node.js: 20.18.3
> V8: 13.2.152.41-electron.0
> OS: Windows\_NT x64 10.0.22631
> 
> \### Julia VS Code Extension
> v1.127.2
> 
> \## Conclusion
> That U+02016 maps to \`\\Vert\` ($\\Vert$) in the docs and by tab-completion indicates that there exists some support for this character. But the error messages from the linter for JuliaSyntax.jl and from a \`ParseError\` in a Julia REPL indicates that perhaps the parser doesn't fully map all the Unicode characters as valid symbols?

EDIT: I already tried rolling the VS Code extension all the way back to v1.127.1, but this did not change the formatting of the Unicode characters, as seen in the screenshot.

---

<div class="post-metadata">

**Author:** ![goerz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerz/32/3269_2.png) [@goerz](https://discourse.julialang.org/u/goerz)\
**Post date:** [June 20, 2025, 1:26am UTC](https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047/2 "2025-06-20T01:26:33Z")

</div>

Sounds like a font issue! Make sure you have [JuliaMono](https://juliamono.netlify.app/) installed and set up

---

<div class="post-metadata">

**Author:** ![jmanthony3](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmanthony3/32/52953_2.png) [@jmanthony3](https://discourse.julialang.org/u/jmanthony3)\
**Post date:** [June 20, 2025, 1:31am UTC](https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047/3 "2025-06-20T01:31:46Z")

</div>

Ah! Thank you! That was the thing that had changed. My machine got LCM’d by IT and I had forgotten to switch the font: I did have JuliaMono set before.

 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/1/714396d60de6d7dc00a6b05986ba81d8c088780a.png)

Any idea why I still get this pop-up for the `\Vert` Unicode character in VS Code?

 ![image](https://global.discourse-cdn.com/julialang/original/3X/d/2/d2c924ee4541b466371889efb139f35027b31bf3.png)

---

<div class="post-metadata">

**Author:** ![jmanthony3](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmanthony3/32/52953_2.png) [@jmanthony3](https://discourse.julialang.org/u/jmanthony3)\
**Post date:** [June 20, 2025, 2:37am UTC](https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047/4 "2025-06-20T02:37:37Z")

</div>

From poking around the Discourse a little, it seems as though some Unicode characters are supported and some are not: e. g. [Rationale behind excluding some unicode characters from identifiers - #11 by t-bltg](https://discourse.julialang.org/t/rationale-behind-excluding-some-unicode-characters-from-identifiers/95509/11)

But that doesn’t seem to explain the issue I found with `\Vert` which does TAB-complete: c. f. [Unknown Unicode Characters · Issue #546 · JuliaLang/JuliaSyntax.jl · GitHub](https://github.com/JuliaLang/JuliaSyntax.jl/issues/546)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 20, 2025, 12:26pm UTC](https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047/5 "2025-06-20T12:26:44Z")

</div>

> [@jmanthony3](#):
>
> But that doesn’t seem to explain the issue I found with `\Vert` which does TAB-complete

Just because it tab-completes doesn’t mean that it is a valid character in Julia identifiers.

`‖` is a punctuation character (category Po, “punctuation, other”), and typically punctuation characters are not allowed in identifiers. It’s possible that some future version of Julia might allow it for some other syntactic purpose. And of course, you can use it inside strings.

---

<div class="post-metadata">

**Author:** ![jmanthony3](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmanthony3/32/52953_2.png) [@jmanthony3](https://discourse.julialang.org/u/jmanthony3)\
**Post date:** [June 20, 2025, 12:40pm UTC](https://discourse.julialang.org/t/what-happened-to-parsing-unicode-characters/130047/6 "2025-06-20T12:40:05Z")

</div>

That does make sense that not all possible characters can be used for all purposes. Some are reserved or ambiguous. I was initially confused because it did `TAB`-complete. Thank y’all for helping and clarifying!
