# String indexing

**URL:** <https://discourse.julialang.org/t/string-indexing/37403>\
**Category:** New to Julia\
**Tags:** indexing\
**Created:** [April 11, 2020, 4:43pm UTC](https://discourse.julialang.org/t/string-indexing/37403 "2020-04-11T16:43:46Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![georgegi86](https://avatars.discourse-cdn.com/v4/letter/g/d2c977/32.png) [@georgegi86](https://discourse.julialang.org/u/georgegi86)\
**Post date:** [April 11, 2020, 4:43pm UTC](https://discourse.julialang.org/t/string-indexing/37403/1 "2020-04-11T16:43:46Z")

</div>

Hi all, playing with Julia today with some limited Python and R experience. How does Julia perform indexing from the end, up to the end minus some value but conditional on the length.

Example: strings = [“name”, “verylargenameforname”]

Python example:  
string[0][-11:] will give me “name”  
string[1][-11:] will give me “nameforname”

R example:  
library(tidyverse)  
strings = c(“name”, “verylargenameforname”)  
str\_sub(strings[1], start = -11)  
str\_sub(strings[2], start = -11)

Julia:  
strings[1][end-10: end] → throws a BoundsError since the string length is less than 10  
strings[2][end-10: end]

---

<div class="post-metadata">

**Author:** ![mbaz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbaz/32/17295_2.png) [@mbaz](https://discourse.julialang.org/u/mbaz)\
**Post date:** [April 11, 2020, 5:49pm UTC](https://discourse.julialang.org/t/string-indexing/37403/2 "2020-04-11T17:49:46Z")

</div>

You may have to define a function to do this:

```julia
julia> function f(s,n)
           try
               return s[end-n:end]
           catch
               return s
           end
       end
f (generic function with 1 method)

julia> f("name",11)
"name"

julia> f("ultralargename",11)
"tralargename"

```

I guess you could avoid the `try` block by doing calculations on the indices and string length, but personally I’m afraid of dealing with codepoints and other string black magic.

---

<div class="post-metadata">

**Author:** ![NiclasMattsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/niclasmattsson/32/21988_2.png) [@NiclasMattsson](https://discourse.julialang.org/u/NiclasMattsson)\
**Post date:** [April 11, 2020, 5:58pm UTC](https://discourse.julialang.org/t/string-indexing/37403/3 "2020-04-11T17:58:14Z")

</div>

Just to show a simple oneliner alternative (not that much black magic involved):

```julia
julia> strings[1][max(end-10,1):end]
"name"

julia> strings[2][max(end-10,1):end]
"nameforname"

julia> endstr(s,n) = s[max(end-n,1):end]
endstr (generic function with 1 method)

julia> endstr(strings[1],10)
"name"

julia> endstr(strings[2],10)
"nameforname"

```

---

<div class="post-metadata">

**Author:** ![mbaz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbaz/32/17295_2.png) [@mbaz](https://discourse.julialang.org/u/mbaz)\
**Post date:** [April 11, 2020, 6:00pm UTC](https://discourse.julialang.org/t/string-indexing/37403/4 "2020-04-11T18:00:15Z")

</div>

Nice, I wish I had thought of using `max` 🤣

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [April 11, 2020, 6:14pm UTC](https://discourse.julialang.org/t/string-indexing/37403/5 "2020-04-11T18:14:16Z")

</div>

Be careful when indexing into strings:

```julia
julia> s = "αβγδϵhelloω"
"αβγδϵhelloω"

julia> s[end-10:end]
ERROR: StringIndexError("αβγδϵhelloω", 6)
Stacktrace:
 [1] string_index_err(::String, ::Int64) at ./strings/string.jl:12
 [2] getindex(::String, ::UnitRange{Int64}) at ./strings/string.jl:249
 [3] top-level scope at REPL[110]:1

```

This won’t work because indices work differently for strings:

```julia
julia> s[1]
'α': Unicode U+03B1 (category Ll: Letter, lowercase)

julia> s[2]
ERROR: StringIndexError("αβγδϵhelloω", 2)
Stacktrace:
 [1] string_index_err(::String, ::Int64) at ./strings/string.jl:12
 [2] getindex_continued(::String, ::Int64, ::UInt32) at ./strings/string.jl:220
 [3] getindex(::String, ::Int64) at ./strings/string.jl:213
 [4] top-level scope at REPL[114]:1

```

I’m not completely confident about this code, but you could try

```julia
julia> foo(str, n) = chop(str; head=max(0, length(str)-n), tail=0)

julia> foo("name", 7)
"name"

julia> foo("nameforname", 7)
"forname"

julia> foo("αβγδϵhelloω", 7)
"ϵhelloω"

```

If you want to use indexing in particular, you should look into the `prevind` and `nextind` functions:

```julia
julia> bar(str, n) = str[prevind(str, lastindex(str), min(length(str), n)-1):end]

julia> bar("name", 7)
"name"

julia> bar("nameforname", 7)
"forname"

julia> bar("αβγδϵhelloω", 7)
"ϵhelloω"

```

---

<div class="post-metadata">

**Author:** ![georgegi86](https://avatars.discourse-cdn.com/v4/letter/g/d2c977/32.png) [@georgegi86](https://discourse.julialang.org/u/georgegi86)\
**Post date:** [April 11, 2020, 8:04pm UTC](https://discourse.julialang.org/t/string-indexing/37403/6 "2020-04-11T20:04:16Z")

</div>

Thanks fellows! I was looking for the simplest solution and Niclas solution fits the bill! Although, I can see that writing a function to accomplish a common operation like that could be an issue to many new comers. I understand that negative indexing is a big no in Julia because of safety issues, which makes sense. But there is a reason why Python and R have those build in - allows for ease of use, data interactivity, and quick prototyping. There should be simple way to grab elements starting from the end (and not just for strings). A middle ground solution would be great like a module similar to R’s stringr, if it doesn’t exist already (did not see this addressed in Strs package).

Looking forward playing with the language!

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [April 11, 2020, 8:41pm UTC](https://discourse.julialang.org/t/string-indexing/37403/7 "2020-04-11T20:41:19Z")

</div>

As I tried to demonstrate, that solution will not work for strings in general.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [April 11, 2020, 8:45pm UTC](https://discourse.julialang.org/t/string-indexing/37403/8 "2020-04-11T20:45:49Z")

</div>

Note that there is a `chop` function which can remove a number of characters from the beginning or end of a string while handling unicode characters correctly:

> ```julia
> chop(s::AbstractString; head::Integer = 0, tail::Integer = 1)
> 
> ```
> 
> Remove the first head and the last tail characters from `s`. The call `chop(s)` removes the last character from `s`. If it is requested to remove more characters than `length(s)` then an empty string is returned.

---

<div class="post-metadata">

**Author:** ![georgegi86](https://avatars.discourse-cdn.com/v4/letter/g/d2c977/32.png) [@georgegi86](https://discourse.julialang.org/u/georgegi86)\
**Post date:** [April 11, 2020, 9:59pm UTC](https://discourse.julialang.org/t/string-indexing/37403/9 "2020-04-11T21:59:12Z")

</div>

Chop would take care of the unicode issue DHF mentioned but it still returns an empty string while the desired return is the string itself (if string is less “chop” → return string). Also, frequently you don’t know where you need to chop at (head = unknown), you just want the a number of characters starting from the end, but keeping the strings that are less than the desired chop amount (out of bounds).

maybe something like this:  
string\_chop(string, from = end - 10 , to = end, max\_length = True)

where max\_length is a Boolean to return the string if end-10 is out of bound. But currently chop doesn’t use the “end” keyword.

Niclas solution works for strings but as DNF pointed out, it doesn’t work for unicode.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [April 11, 2020, 10:03pm UTC](https://discourse.julialang.org/t/string-indexing/37403/10 "2020-04-11T22:03:54Z")

</div>

> [@georgegi86](#):
>
> Chop would take care of the unicode issue DHF mentioned but it still returns an empty string while the desired return is the string itself

I’m missing something. Can you give an example? I thought my suggestions covered your requirements.

---

<div class="post-metadata">

**Author:** ![georgegi86](https://avatars.discourse-cdn.com/v4/letter/g/d2c977/32.png) [@georgegi86](https://discourse.julialang.org/u/georgegi86)\
**Post date:** [April 11, 2020, 10:20pm UTC](https://discourse.julialang.org/t/string-indexing/37403/11 "2020-04-11T22:20:13Z")

</div>

Yes, I am all set as far as strings.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [April 11, 2020, 10:47pm UTC](https://discourse.julialang.org/t/string-indexing/37403/12 "2020-04-11T22:47:15Z")

</div>

The problem statement is unclear (to me at least) so I’m not sure what the desired behavior is. Just mentioning chop which is generally useful for chopping the head or tail off a string.
