# Substring function?

**URL:** <https://discourse.julialang.org/t/substring-function/76675>\
**Category:** New to Julia\
**Tags:** strings, unicode\
**Created:** [February 18, 2022, 7:41am UTC](https://discourse.julialang.org/t/substring-function/76675 "2022-02-18T07:41:49Z")\
**Posts on this page:** 1\
**Showing post:** 10

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 18, 2022, 2:42pm UTC](https://discourse.julialang.org/t/substring-function/76675/10 "2022-02-18T14:42:46Z")

</div>

> Just thought a `substring` function would be useful out of the box, with documentation about how it differs from the `string[start:stop]` form in terms of unicode and performance.

Slicing `a[m:n]` _always_ makes a copy in Julia (at least, with the built-in types), whether for arrays or strings. If you want to use a view (i.e. create a `SubString` object), the easiest way is to [use `@views` on a block of code](https://docs.julialang.org/en/v1/manual/performance-tips/#man-performance-views), e.g.

```julia
julia> s = "αβγł€đŧŧŋ"
"αβγł€đŧŧŋ"

julia> @views s[1:5]
"αβγ"

julia> typeof(ans)
SubString{String}

```

> [@rogerkeays](#):
>
> getting a substring where multibyte unicode characters are involved

Slicing with `@views` works just fine for this.

The real question is, where are you getting these character indices that you want to pass to your `substring` function? Usually you get indices to a substring from some previous iteration over the string, either from your own loop or from something like a `findnext` call, and these give you codeunit indices that you can pass to `s[m:n]` directly.

If you are counting codepoints as “characters”, e.g. you want the “first 3 characters” in a string, then the odds are high that you are making a mistake. For example, `"ü"` is _two_ codepoints (`length("ü") == 2`) because it is `u` followed by a [combining character](https://en.wikipedia.org/wiki/Combining_character) [U+0308](https://www.fileformat.info/info/unicode/char/0308/index.htm). See also this explanation: [Myth: Counting coded characters or code points is important.](http://utf8everywhere.org/#myth.strlen)

Because of Unicode’s complexity, wanting a substring from the `m`-th _codepoint_ (“character”) to the `n`-th codepoint, as opposed to between two string _indices_ (= code units), is actually an extremely uncommon operation (in non-buggy code). This is why it’s not built-in.

---

_[View the full topic](https://discourse.julialang.org/t/substring-function/76675)._
