# Unexpected index of Unicode subscript \`char\` in \`string\`?

**URL:** <https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515>\
**Category:** General Usage\
**Created:** [June 24, 2021, 8:41pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515 "2021-06-24T20:41:55Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![frankwswang](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/frankwswang/32/18561_2.png) [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Post date:** [June 24, 2021, 8:41pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/1 "2021-06-24T20:41:55Z")

</div>

It seems that in some cases the indexing of subscript `char` is wrong:

```julia
julia> "a₁"[1]
'a': ASCII/Unicode U+0061 (category Ll: Letter, lowercase)

julia> "a₁"[2]
'₁': Unicode U+2081 (category No: Number, other)

julia> "α₁"[1]
'α': Unicode U+03B1 (category Ll: Letter, lowercase)

julia> "α₁"[2]
ERROR: StringIndexError: invalid index [2], valid nearby indices [1]=>'α', [3]=>'₁'

julia> length("α₁")
2

```

I’m using Julia 1.6.1 by the way.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [June 24, 2021, 8:48pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/2 "2021-06-24T20:48:04Z")

</div>

Indexes are to code units (bytes) not characters. This is addressed extensively in the string section in the manual:

[https://docs.julialang.org/en/v1/manual/strings/](https://docs.julialang.org/en/v1/manual/strings/)

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [June 24, 2021, 9:08pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/3 "2021-06-24T21:08:07Z")

</div>

There are many things you might want to index.

- **codeunit** : basically the individual bytes
- **codepoint** : a single character, may be multiple bytes in unicode
- **grapheme** : a single display unit. For example 👩‍👩‍👧‍👦 is 1 grapheme but it is made up of 4 codepoints that when placed next to each over display as one 👩‍ 👩‍ 👧‍ 👦

Of all of these: only `codeunit` can be done in O(1) time.  
The rest are O(n).

I don’t actually know how to index by codepoint, or grapheme.  
For codepoint you can call `collect` first, which is OK, but not great.  
For grapheme, I am not sure if it can be known without knowing things about the device displaying it.

---

<div class="post-metadata">

**Author:** ![frankwswang](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/frankwswang/32/18561_2.png) [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Post date:** [June 24, 2021, 9:12pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/4 "2021-06-24T21:12:26Z")

</div>

Now I understand the number of code unit of “α” is 2 instead of 1 like ‘a’:

```julia
julia> codeunits("α")
2-element Base.CodeUnits{UInt8, String}:
 0xce
 0xb1

julia> codeunits("a")
1-element Base.CodeUnits{UInt8, String}:
 0x61

```

which causes this behavior Thank you!

---

<div class="post-metadata">

**Author:** ![frankwswang](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/frankwswang/32/18561_2.png) [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Post date:** [June 24, 2021, 9:36pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/5 "2021-06-24T21:36:03Z")

</div>

A little relevant to this topic but I’d like to know is there a particular reason for `findfirst` to not return all the indices a `char` occupies when it has multiple code units?

```julia
julia> findfirst("α", "α1")
1:1

```

The reason I ask this is that I want to efficiently locate a char inside a string no matter how many code units it contains. I know I can definitely do something like this though:

```julia
findfirst("α", "α1")[1] : ncodeunits("α")
1:2

```

Thank you!

---

<div class="post-metadata">

**Author:** ![tomerarnon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomerarnon/32/3170_2.png) [@tomerarnon](https://discourse.julialang.org/u/tomerarnon)\
**Post date:** [June 24, 2021, 10:04pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/6 "2021-06-24T22:04:55Z")

</div>

Because then that range can’t be used to get the character:

```julia
julia> "α1"[1:2]
ERROR: StringIndexError: invalid index [2], valid nearby indices [1]=>'α', [3]=>'1'
...

julia> "α1"[1:1]
"α"

```

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [June 24, 2021, 11:08pm UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/7 "2021-06-24T23:08:42Z")

</div>

> [@oxinabox](#):
>
> For grapheme, I am not sure if it can be known without knowing things about the device displaying it.

There’s [Unicode.graphemes](https://docs.julialang.org/en/v1/stdlib/Unicode/#Unicode.graphemes) to iterate over them, and I guess one could `collect` that to index them.

---

<div class="post-metadata">

**Author:** ![frankwswang](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/frankwswang/32/18561_2.png) [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Post date:** [June 25, 2021, 12:29am UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/8 "2021-06-25T00:29:28Z")

</div>

If you want to get the location of the `char`, why not just search for it as a `char`, which gives you the starting index? Isn’t this the more appropriate usage?

```julia
julia> str = "α1"
"α1"

julia> i = findfirst('α', str)
1

julia> str[i]
'α': Unicode U+03B1 (category Ll: Letter, lowercase)

```

---

<div class="post-metadata">

**Author:** ![tomerarnon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomerarnon/32/3170_2.png) [@tomerarnon](https://discourse.julialang.org/u/tomerarnon)\
**Post date:** [June 25, 2021, 9:29am UTC](https://discourse.julialang.org/t/unexpected-index-of-unicode-subscript-char-in-string/63515/9 "2021-06-25T09:29:42Z")

</div>

The same problem would occur if searching for a longer string.

```julia
julia> findfirst("α,β", "α,β,γ,δ")
1:4

julia> "α,β,γ,δ"[1:4]
"α,β"

julia> sizeof("α,β")
5

```

It’s simply the convention that characters are indexed by their starting index, rather than by the entire index-range that they occupy. Overall, this leads to fewer inconveniences than the alternative (either the behavior you showed would have to change, or this one does, for consistency)
