# Symbol to String

**URL:** <https://discourse.julialang.org/t/symbol-to-string/92217>\
**Category:** General Usage\
**Tags:** question\
**Created:** [December 28, 2022, 9:31am UTC](https://discourse.julialang.org/t/symbol-to-string/92217 "2022-12-28T09:31:55Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mirsad\_Cosovic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mirsad_cosovic/32/9887_2.png) [@Mirsad\_Cosovic](https://discourse.julialang.org/u/Mirsad_Cosovic)\
**Post date:** [December 28, 2022, 9:31am UTC](https://discourse.julialang.org/t/symbol-to-string/92217/1 "2022-12-28T09:31:55Z")

</div>

I’m trying to convert a special Symbol to a String, but I’m having some trouble. Is the following behavior usual and why?

```julia
julia> String(:Ω) == "Ω"
false
```

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [December 28, 2022, 9:48am UTC](https://discourse.julialang.org/t/symbol-to-string/92217/2 "2022-12-28T09:48:21Z")

</div>

Greek omega and Ohm sign have different unicode characters so they’re not equal under `==` even though they look the same.

---

<div class="post-metadata">

**Author:** ![Mirsad\_Cosovic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mirsad_cosovic/32/9887_2.png) [@Mirsad\_Cosovic](https://discourse.julialang.org/u/Mirsad_Cosovic)\
**Post date:** [December 28, 2022, 9:54am UTC](https://discourse.julialang.org/t/symbol-to-string/92217/3 "2022-12-28T09:54:30Z")

</div>

I understand that, but why unicode is different in those two cases?

---

<div class="post-metadata">

**Author:** ![cormullion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cormullion/32/49131_2.png) [@cormullion](https://discourse.julialang.org/u/cormullion)\
**Post date:** [December 28, 2022, 10:29am UTC](https://discourse.julialang.org/t/symbol-to-string/92217/4 "2022-12-28T10:29:47Z")

</div>

The Unicode Consortium [say](https://www.unicode.org/versions/Unicode4.0.0):

> For compatibility purposes, a few Greek letters are separately encoded as symbols in other character blocks. Examples include U+00B5 μ   in the Latin-1 Supplement character block and U+2126 Ω   in the Letterlike Symbols character block. The ohm sign is canonically equivalent to the capital omega, and normalization would remove any distinction. Its use is therefore discouraged in favor of capital omega. The same equivalence does not exist between micro sign and mu, and use of either character as micro sign is com- mon; for Greek text, only the mu should be used.

which I think is saying “Don’t use U+2126 Ω (or `\ohm` in Julia) - use U+0309 (`\Omega` in Julia)”.

---

<div class="post-metadata">

**Author:** ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)\
**Post date:** [December 28, 2022, 10:51am UTC](https://discourse.julialang.org/t/symbol-to-string/92217/5 "2022-12-28T10:51:00Z")

</div>

Do they also say anything about why we can’t have subscript `\_b`, `\_c`, or `\_d`, but instead we can have `\:turtle:` 🐢 and `\:person_in_steamy_room:` 🧖?

---

<div class="post-metadata">

**Author:** ![Per](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/per/32/10387_2.png) [@Per](https://discourse.julialang.org/u/Per)\
**Post date:** [December 28, 2022, 11:36am UTC](https://discourse.julialang.org/t/symbol-to-string/92217/6 "2022-12-28T11:36:51Z")

</div>

The reason why `:\ohm<tab>` gives the Omega symbol instead of ohm is that symbols are treated like variable names in Julia, and thus it makes sense to normalize them.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [December 28, 2022, 12:27pm UTC](https://discourse.julialang.org/t/symbol-to-string/92217/7 "2022-12-28T12:27:48Z")

</div>

The thing to realize is that the Unicode strings stored internally for Julia symbols (e.g. variable names or quoted symbols) are automatically [normalized to canonical form](https://en.wikipedia.org/wiki/Unicode_equivalence).

So, even if you type `:Ω` using the Ohm symbol U+2126 (e.g. via `:\ohm<tab>`), it will get normalized to Omega U+03A09:

```julia
julia> collect(String(:Ω)) # Ω is Ohm U+2126 (\ohm<tab>)
1-element Vector{Char}:
 'Ω': Unicode U+03A9 (category Lu: Letter, uppercase)

```

Another example would be accented Latin characters like `ë`, which often have two canonically equivalent representations (either a single special character or an unaccented character followed by a “combining” accent character), but you don’t want that to correspond to different variable names depending on how you type it (e.g. different input systems). Canonicalization (technically, NFC normalization) removes that distinction.

This is explained [in the Julia manual](https://docs.julialang.org/en/v1/manual/variables/#man-allowed-variable-names):

> Some Unicode characters are considered to be equivalent in identifiers. Different ways of entering Unicode combining characters (e.g., accents) are treated as equivalent (specifically, Julia identifiers are [NFC](http://www.macchiato.com/unicode/nfc-faq)-normalized). Julia also includes a few non-standard equivalences for characters that are visually similar and are easily entered by some input methods. The Unicode characters `ɛ` (U+025B: Latin small letter open e) and `µ` (U+00B5: micro sign) are treated as equivalent to the corresponding Greek letters. The middle dot `·` (U+00B7) and the Greek [interpunct](https://en.wikipedia.org/wiki/Interpunct) `·` (U+0387) are both treated as the mathematical dot operator `⋅` (U+22C5). The minus sign `−` (U+2212) is treated as equivalent to the hyphen-minus sign `-` (U+002D).

---

<div class="post-metadata">

**Author:** ![cormullion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cormullion/32/49131_2.png) [@cormullion](https://discourse.julialang.org/u/cormullion)\
**Post date:** [December 28, 2022, 1:34pm UTC](https://discourse.julialang.org/t/symbol-to-string/92217/8 "2022-12-28T13:34:53Z")

</div>

> [@uniment](#):
>
> Do they also say anything about why

@stevengj Any progress on your [proposal](https://github.com/stevengj/subsuper-proposal)?

---

<div class="post-metadata">

**Author:** ![sostock](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sostock/32/5546_2.png) [@sostock](https://discourse.julialang.org/u/sostock)\
**Post date:** [December 28, 2022, 1:58pm UTC](https://discourse.julialang.org/t/symbol-to-string/92217/9 "2022-12-28T13:58:37Z")

</div>

> [@stevengj](#):
>
> The thing to realize is that the Unicode strings stored internally for Julia symbols (e.g. variable names or quoted symbols) are automatically [normalized to canonical form](https://en.wikipedia.org/wiki/Unicode_equivalence).
> 
> So, even if you type `:Ω` using the Ohm symbol U+2126 (e.g. via `:\ohm<tab>`), it will get normalized to Omega U+03A09:

However, this does not happen if the the `Symbol` constructor is used (instead of typing `:Ω`):

```julia
julia> collect(String(Symbol(Char(0x2126))))
1-element Vector{Char}:
 'Ω': Unicode U+2126 (category Lu: Letter, uppercase)

```

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [December 28, 2022, 2:02pm UTC](https://discourse.julialang.org/t/symbol-to-string/92217/10 "2022-12-28T14:02:44Z")

</div>

> [@sostock](#):
>
> However, this does not happen if the the `Symbol` constructor is used

That’s right — if you use the `Symbol` constructor then you can make a `Symbol` from any Julia string even if it is not a valid identifier, such as `Symbol(" ")` (as long as the string doesn’t contain `\0`). Because of this it takes the strings literally as-is, with no normalization.
