# String getindex problem?

**URL:** <https://discourse.julialang.org/t/string-getindex-problem/41743>\
**Category:** General Usage\
**Created:** [June 19, 2020, 7:18pm UTC](https://discourse.julialang.org/t/string-getindex-problem/41743 "2020-06-19T19:18:55Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ronneesley](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronneesley/32/11072_2.png) [@Ronneesley](https://discourse.julialang.org/u/Ronneesley)\
**Post date:** [June 19, 2020, 7:18pm UTC](https://discourse.julialang.org/t/string-getindex-problem/41743/1 "2020-06-19T19:18:55Z")

</div>

Why Julia have this accentuation problem?  
How to deal with this?

```julia
julia> length("cão")
3

julia> "cão"[3]
ERROR: StringIndexError("cão", 3)
Stacktrace:
 [1] string_index_err(::String, ::Int64) at ./strings/string.jl:12
 [2] getindex_continued(::String, ::Int64, ::UInt32) at ./strings/string.jl:220
 [3] getindex(::String, ::Int64) at ./strings/string.jl:213
 [4] top-level scope at REPL[2]:1

```

Even more weird:

```julia
julia> "cão"[4]
'o': ASCII/Unicode U+006F (category Ll: Letter, lowercase)

```

Thanks.

---

<div class="post-metadata">

**Author:** ![fredrikekre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fredrikekre/32/1688_2.png) [@fredrikekre](https://discourse.julialang.org/u/fredrikekre)\
**Post date:** [June 19, 2020, 7:24pm UTC](https://discourse.julialang.org/t/string-getindex-problem/41743/2 "2020-06-19T19:24:40Z")

</div>

See [Strings · The Julia Language](https://docs.julialang.org/en/v1/manual/strings/#Unicode-and-UTF-8-1)

---

<div class="post-metadata">

**Author:** ![Ronneesley](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronneesley/32/11072_2.png) [@Ronneesley](https://discourse.julialang.org/u/Ronneesley)\
**Post date:** [June 19, 2020, 8:38pm UTC](https://discourse.julialang.org/t/string-getindex-problem/41743/3 "2020-06-19T20:38:07Z")

</div>

Hi @fredrikekre ,

So, if I really want the ith position of a UTF-8 string, like this:

```julia
julia> s = "cão"
"cão"

julia> val(s, 2)
'ã': Unicode U+00E3 (category Ll: Letter, lowercase)

julia> sub(s, 1, 2)
"cã"

julia> sub(s, 1:2)
"cã"

julia> sub(s, 2, 3)
"ão"

julia> sub(s, 2:3)
"ão"

```

I need something like this:

```julia
function ind(s::String, i::Int, k::Int = 1)
    for _ = 1:i - 1
        k = nextind(s, k)
    end    
    return k
end

function interval(s::String, i::Int, f::Int)
    start = ind(s, i)
    stop = ind(s, f - i + 1, start)    
    return start:stop
end

sub(s::String, i::Int, f::Int) =
    s[interval(s, i, f)]

sub(s::String, i::UnitRange) =
    s[interval(s, i.start, i.stop)]

val(s::String, i::Int) =
    s[ind(s, i)]

```

Or there is another way to do it?

Thanks.

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [June 19, 2020, 8:44pm UTC](https://discourse.julialang.org/t/string-getindex-problem/41743/4 "2020-06-19T20:44:02Z")

</div>

Assuming your mental mode is that strings are made of graphemes, you can do:

```julia
julia> import Unicode: graphemes

julia> s = "cão"
"cão"

julia> collect(graphemes(s))[1]
"c"

julia> collect(graphemes(s))[2]
"ã"

julia> collect(graphemes(s))[3]
"o"

```

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 20, 2020, 1:37am UTC](https://discourse.julialang.org/t/string-getindex-problem/41743/5 "2020-06-20T01:37:40Z")

</div>

> [@Ronneesley](#):
>
> So, if I really want the ith position of a UTF-8 string,

The `i`-th codepoint is given by `s[nextind(s, 0, i)]`:

```julia
julia> s = "cão"
"cão"

julia> [s[nextind(s, 0, i)] for i = 1:3]
3-element Array{Char,1}:
 'c'
 'ã'
 'o'

```

However, realize that finding the `i`-th codepoint is O(i) (linear) complexity for the UTF-8 encoding or any variable-width encoding.

The real question is _why_ you want the `i`-th codepoint. Usually, random positions in strings arise from other processing, e.g. searches, in which the index is already computed as a byproduct.

As @johnmyleswhite alluded to, the notion of a “character” in Unicode might not be what you expect. The strings `s = "cão"` and `s2 = "cão"` may look the same, and are [canonically equivalent](https://en.wikipedia.org/wiki/Unicode_equivalence), but `s2` actually has 4 Unicode codepoints (“characters”) even though it has 3 graphemes (what most users would consider “characters”), because in `s2` the `ã` is made from an ASCII `a` followed by a [U+0303](https://www.fileformat.info/info/unicode/char/0303/index.htm) “combining tilde”. So, thinking in terms of the `i`-th “position” in a string may indicate a conceptual misunderstanding of Unicode.
