# Tab completion of \\uXXXX in the REPL?

**URL:** <https://discourse.julialang.org/t/tab-completion-of-uxxxx-in-the-repl/108442>\
**Category:** Internals & Design\
**Tags:** unicode\
**Created:** [January 6, 2024, 3:23pm UTC](https://discourse.julialang.org/t/tab-completion-of-uxxxx-in-the-repl/108442 "2024-01-06T15:23:08Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [January 6, 2024, 6:35pm UTC](https://discourse.julialang.org/t/tab-completion-of-uxxxx-in-the-repl/108442/2 "2024-01-06T18:35:56Z")

</div>

> [@devel-chm](#):
>
> While  
> the pictures/glyphs of the characters are good to read, for a computer  
> to say two characters are the same they need to have equal encodings.

It seems like you are using the word “encoding” in a non-standard way — it sounds like you mean the Unicode codepoint value. e.g. U+0041 is the codepoint for the ASCII character `'A'`, not an “encoding”.

In contrast, UTF-8 is an “encoding” of Unicode characters into byte sequences. For example, [U+1F385 `'🎅'`](https://www.compart.com/en/unicode/U+1F385) is encoded as a sequence of 4 bytes in UTF-8:

```julia
julia> codeunits("🎅")
4-element Base.CodeUnits{UInt8, String}:
 0xf0
 0x9f
 0x8e
 0x85

```

whereas [U+03B1 `'α'`](https://www.compart.com/en/unicode/U+03B1) is encoded as two bytes in UTF-8:

```julia
julia> codeunits("α")
2-element Base.CodeUnits{UInt8, String}:
 0xce
 0xb1

```

> [@devel-chm](#):
>
> Similarly, I would like a way to extract the encoding from a visible  
> glyph or character so that it could be “directly” entered into julia via  
> the REPL mode.

I showed how to get the UTF-8 encoding bytes with `codeunits` above, but it sounds like you really want the codepoint value, which you can get easily by e.g.:

```julia
julia> '🎅' # display information about the character in the REPL
'🎅': Unicode U+1F385 (category So: Symbol, other)

julia> UInt32('🎅') # codepoint as an integer value
0x0001f385

```

Note that a “visible glyph” might be more than one character, e.g. `α̂` (a single “grapheme”) is two characters:

```julia
julia> collect("α̂")
2-element Vector{Char}:
 'α': Unicode U+03B1 (category Ll: Letter, lowercase)
 '̂': Unicode U+0302 (category Mn: Mark, nonspacing)

```

and you can get information about how to type it easily by pasting it at the `help?>` prompt:

```julia
help?> α̂
"α̂" can be typed by \alpha<tab>\hat<tab>

```

> [@devel-chm](#):
>
> If we were to add \u+0041 completions to the julia REPL mode  
> that would allow users to enter all possible characters into julia.

You can also type codepoint values as `\uXXXX` escape sequences into a string and then copy-paste it:

```julia
julia> "\u03B1\u0302"
"α̂"

```

You can also add custom tab completions to the REPL, e.g.

```julia
using REPL: REPLCompletions
REPLCompletions.latex_symbols["\\alphahat"] = "α̂" # or "\u03B1\u0302"

```

will let you tab-complete `\alphahat` to `α̂`. (And, of course, all modern operating systems provide a variety of input methods for Unicode characters.)

See [this post](https://discourse.julialang.org/t/syntax-escape-hatch-for-unicode-haters/108363/2) for how to _directly_ use codepoint values as variable names like `uvar"\u03B1\u0302"` in Julia (which in practice will probably be about as popular as [trigraphs](https://en.wikipedia.org/wiki/Digraphs_and_trigraphs)).

> [@devel-chm](#):
>
> for a computer  
> to say two characters are the same they need to have equal encodings

Note that this is not quite true, especially for strings (or “glyphs” or graphemes) that consist of multiple characters. [Unicode equivalence](https://en.wikipedia.org/wiki/Unicode_equivalence) generally involves some form of normalization to do comparisons. (And Julia provides facilities for this. For source-code identifiers, [Julia does NFC normalization](https://docs.julialang.org/en/v1/manual/variables/#man-allowed-variable-names) + some custom normalizations.)

> [@devel-chm](#):
>
> would help for cases where  
> your font is missing characters

It’s 2024 — if your font won’t display characters that you want to use, get a [better font](https://juliamono.netlify.app/). (And if your editor doesn’t support Unicode, stop using [`ed`](https://www.gnu.org/fun/jokes/ed-msg.en.html) and get a better editor.)

---

_[View the full topic](https://discourse.julialang.org/t/tab-completion-of-uxxxx-in-the-repl/108442)._
