# Indexing strings by Unicode code point instead of code unit?

**URL:** <https://discourse.julialang.org/t/indexing-strings-by-unicode-code-point-instead-of-code-unit/55248>\
**Category:** General Usage\
**Tags:** strings\
**Created:** [February 14, 2021, 9:01am UTC](https://discourse.julialang.org/t/indexing-strings-by-unicode-code-point-instead-of-code-unit/55248 "2021-02-14T09:01:08Z")\
**Posts on this page:** 1\
**Showing post:** 12

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 14, 2021, 4:01pm UTC](https://discourse.julialang.org/t/indexing-strings-by-unicode-code-point-instead-of-code-unit/55248/12 "2021-02-14T16:01:24Z")

</div>

> [@johann-petrak](#):
>
> the Julia design decision apparently was to use indexing to retrieve the elements of representation, not the actual elements a string has: characters.

This is **not correct**. Julia string indices retrieve characters, not code units, it’s just that the indices are not consecutive:

```julia
julia> s = "αβγ"
"αβγ"

julia> collect(eachindex(s))
3-element Array{Int64,1}:
 1
 3
 5

julia> s[1]
'α': Unicode U+03B1 (category Ll: Letter, lowercase)

julia> s[3]
'β': Unicode U+03B2 (category Ll: Letter, lowercase)

julia> s[5]
'γ': Unicode U+03B3 (category Ll: Letter, lowercase)

```

(Note that the return value of `s[i]` is a character represented by `Char`, a 4-byte object corresponding to a Unicode codepoint.) String iteration is also over characters:

```julia
julia> for c in s
           display(c) # pretty-print c
       end
'α': Unicode U+03B1 (category Ll: Letter, lowercase)
'β': Unicode U+03B2 (category Ll: Letter, lowercase)
'γ': Unicode U+03B3 (category Ll: Letter, lowercase)

julia> for (i,c) in pairs(s)
           @show i, c, s[i]
       end
(i, c, s[i]) = (1, 'α', 'α')
(i, c, s[i]) = (3, 'β', 'β')
(i, c, s[i]) = (5, 'γ', 'γ')

```

In contrast, the code units (bytes for UTF-8) are retrieved by the `codeunit` function:

```julia
julia> codeunit(s, 3)
0xce

julia> codeunits(s)
6-element Base.CodeUnits{UInt8,String}:
 0xce
 0xb1
 0xce
 0xb2
 0xce
 0xb3

```

Substrings work similarly: you give them code-unit indices, but they still give you the whole string of Unicode characters:

```julia
julia> s[3:5] # a copy
"βγ"

julia> SubString(s, 3:5) # a view
"βγ"

```

So, the point is, once you convert your character offsets to Julia `String` indices, you are still working with characters.

---

_[View the full topic](https://discourse.julialang.org/t/indexing-strings-by-unicode-code-point-instead-of-code-unit/55248)._
