# Substring function?

**URL:** <https://discourse.julialang.org/t/substring-function/76675>\
**Category:** New to Julia\
**Tags:** strings, unicode\
**Created:** [February 18, 2022, 7:41am UTC](https://discourse.julialang.org/t/substring-function/76675 "2022-02-18T07:41:49Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![rogerkeays](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rogerkeays/32/33927_2.png) [@rogerkeays](https://discourse.julialang.org/u/rogerkeays)\
**Post date:** [February 19, 2022, 5:30am UTC](https://discourse.julialang.org/t/substring-function/76675/21 "2022-02-19T05:30:59Z")

</div>

> [@stevengj](#):
>
> `first` and `last` are defined for _any_ iterator. Since string iteration is over codepoints, they have to be consistent, but I agree that they need to be used with care.

I think that’s getting to bottom of the issue. **Iterators use codepoints, whereas ranges use byte indexes**. I see the need for byte indexing, but being new to the language, this idiosyncracy tripped me up. I expected to be working with codepoints all the way down unless explicitly converting into bytes. For example, in Java:

```java
jshell> "αβγł€đŧŧŋ".substring(0,5)
$1 ==> "αβγł€"
jshell> "αβγł€đŧŧŋ".getBytes()
$2 ==> byte[19] { -50, -79, -50, -78, -50, -77, -59, -126, -30, -126, -84, -60, -111, -59, -89, -59, -89, -59, -117 }

```

It might have been possible to implement strings as a vector of codepoints and still achieve O(1) performance for indexing using a lookup table for indicies, or (to save memory) an `isascii` property. Too late for that though.

Note, I’m not saying that Julia should be “like Java”, but I do agree that the API should be consistent. Looks to me like the [LegacyStrings](https://github.com/JuliaStrings/LegacyStrings.jl) implementation might have been better.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [February 19, 2022, 7:52am UTC](https://discourse.julialang.org/t/substring-function/76675/22 "2022-02-19T07:52:55Z")

</div>

> [@rogerkeays](#):
>
> **Iterators use codepoints, whereas ranges use byte indexes**. I see the need for byte indexing, but being new to the language, this idiosyncracy tripped me up. I expected to be working with codepoints all the way down unless explicitly converting into bytes.

As far as I understand, the reason for this is to be able to support invalid UTF-8 in `String` as well. It’s quite common to have some corrupted data that’s treated as a string. It’s usually seen as a strength to be able to do that - I wouldn’t call it a “idiosyncracy”.

You may also be interested in some of these previous discussions about various parts of the `String` type in julia:

- [How does Char get stored? - #3 by StefanKarpinski](https://discourse.julialang.org/t/how-does-char-get-stored/49366/3)
- [A Python rant about types - #3 by StefanKarpinski](https://discourse.julialang.org/t/a-python-rant-about-types/43294/3)

> [@rogerkeays](#):
>
> For example, in Java

Java is using UTF-16, right? The same problems mentioned in the two links above should apply as well, as it’s a variable length encoding like UTF-8. I don’t know how java would treat those bad encodings though. I think java works around this problem by just not having strings decompose into an iterator of `char` easily, which can get quite hairy to implement in a performant way (I can’t find the links to previous discussions about that though, sorry).

---

<div class="post-metadata">

**Author:** ![rogerkeays](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rogerkeays/32/33927_2.png) [@rogerkeays](https://discourse.julialang.org/u/rogerkeays)\
**Post date:** [February 19, 2022, 12:56pm UTC](https://discourse.julialang.org/t/substring-function/76675/23 "2022-02-19T12:56:22Z")

</div>

> [@Sukera](#):
>
> Java is using UTF-16, right?

This thread is starting to diverge somewhat, but I’m game…

[Java `chars` are 16-bit](https://docs.oracle.com/en/java/javase/17/docs/api/java.base/java/lang/Character.html), and [Strings are encoded internally with UTF-16](https://github.com/openjdk/jdk/blob/d7f31d0d53bfec627edc83ceb75fc6202891e186/src/java.base/share/classes/java/lang/String.java#L106), but the `String` class is (unsuprisingly) implemented as an [array of bytes](https://github.com/openjdk/jdk/blob/d7f31d0d53bfec627edc83ceb75fc6202891e186/src/java.base/share/classes/java/lang/String.java#L145). Before 32-bit characters become a thing, that meant you could do a lot of stuff O(1). Now, I think they have the same problem as Julia. From the javadoc:

```julia
Index values refer to `char` code units, so a supplementary
character uses two positions in a `String`.

```

Which means, as you said, they have the same problem I posted here, except the indexing is every 16 bits instead of 8:

```java
jshell> "🩢🩣🩤".substring(0,1)
$119 ==> "?"

jshell> "🩢🩣🩤".substring(0,2)
$120 ==> "🩢"

jshell> "🩢🩣🩤".length()
$122 ==> 6

```

This is worse than I thought… Julia’s functions which work on codepoints get it right:

```julia
julia> length("🩢🩣🩤")
3

```

And some of the newer Java functions work with variable-length encoding:

```java
jshell> "🩢🩣🩤".codePointCount(0,6)
$137 ==> 3

jshell> "🩢🩣🩤".codePoints().forEach(System.out::println)
129634
129635
129636

jshell> "🩢🩣🩤".codePoints().forEach(c -> System.out.println(Character.toString(c)))
🩢
🩣
🩤

```

Basically it is a mix of new code which works on codepoints and old code which works on 16-bit indexes. A total mess, in other words. I think Julia can do better if it is consistent about using codepoints or byte indexes. I can see that, like Java, you won’t give up O(1) indexing or break backwards compatibility to do this though.

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [February 19, 2022, 1:22pm UTC](https://discourse.julialang.org/t/substring-function/76675/24 "2022-02-19T13:22:49Z")

</div>

> [@Sukera](#):
>
> I wouldn’t call it a “idiosyncracy”.

Why not? Idiosyncrasy just means “: a peculiarity of constitution or temperament : an individualizing characteristic or quality” ([Idiosyncrasy Definition & Meaning - Merriam-Webster](https://www.merriam-webster.com/dictionary/idiosyncrasy)), if Julia is different in this aspect than most languages then the word fits.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [February 19, 2022, 1:28pm UTC](https://discourse.julialang.org/t/substring-function/76675/25 "2022-02-19T13:28:57Z")

</div>

Because it hardly is a individualizing characteristic - lots of languages have indexing into their strings that’s not based on codepoints. I also remember some discussion some time back about only indexing based on graphemes, which would come with a whole different set of problems.

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [February 19, 2022, 1:36pm UTC](https://discourse.julialang.org/t/substring-function/76675/26 "2022-02-19T13:36:08Z")

</div>

I think the idiosyncrasy mentioned is the mix of both?

> [@rogerkeays](#):
>
> **Iterators use codepoints, whereas ranges use byte indexes**.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 19, 2022, 1:55pm UTC](https://discourse.julialang.org/t/substring-function/76675/27 "2022-02-19T13:55:20Z")

</div>

> [@rogerkeays](#):
>
> It might have been possible to implement strings as a vector of codepoints

Not efficiently.

The basic argument is that a variable-width encoding with non-consecutive codepoint indices is a good tradeoff to make (memory efficiency + speed, at the cost of less-intuitive indexing) because “give me the m-th codepoint” or “or give me the substring from codepoints m to n” is extremely uncommon in (correct) string-handling code, as opposed to “give me the substring at opaque indices I found in a previous search/loop”.

That’s why, [in my previous post](https://discourse.julialang.org/t/substring-function/76675/10), I asked you where your `m` and `n` indices come from. You still haven’t given any usecase for your `substring(s, m, n)` function. In what realistic application would someone say "give me the 12th to the 17th characters of this string, where the numbers 12 and 17 just fell out of the sky (not from a previous search/iteration on the string)?

(One option we’ve discussed is literally making string indices an opaque type, so that it no longer resembles a consecutive array index, which is the main source of confusion.)

With UTF-16 as in Java, you have exactly the same issue, except that bugs are harder to catch because surrogate pairs are less common. With UTF-8, you catch the indexing bugs in your code the first time anyone passes you a Unicode string.

---

<div class="post-metadata">

**Author:** ![rogerkeays](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rogerkeays/32/33927_2.png) [@rogerkeays](https://discourse.julialang.org/u/rogerkeays)\
**Post date:** [February 19, 2022, 2:03pm UTC](https://discourse.julialang.org/t/substring-function/76675/28 "2022-02-19T14:03:11Z")

</div>

> [@stevengj](#):
>
> The real question is, where are you getting these character indices that you want to pass to your `substring` function?

I am parsing unicode text in columns. Like unix `cut`.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 19, 2022, 2:07pm UTC](https://discourse.julialang.org/t/substring-function/76675/29 "2022-02-19T14:07:21Z")

</div>

> [@rogerkeays](#):
>
> I am parsing unicode text in columns. Like unix `cut` .

When you parse, you are looping over the string or using `findnext` or similar. Hence you can use the actual string index from that loop/search to subsequently extract substrings — you neither need nor want the codepoint count.

(If you are parsing with ASCII delimiters, you often don’t even need to look at Unicode characters…you literally don’t care what comes between the delimiters when parsing. In this case you can alternatively loop over the raw codeunits/byte array, provided by the `codeunits(str)` function. This still gives you indices that you can use to slice the original string. That’s what e.g. JSON parsing does IIRC.)

---

<div class="post-metadata">

**Author:** ![rogerkeays](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rogerkeays/32/33927_2.png) [@rogerkeays](https://discourse.julialang.org/u/rogerkeays)\
**Post date:** [February 19, 2022, 2:18pm UTC](https://discourse.julialang.org/t/substring-function/76675/30 "2022-02-19T14:18:46Z")

</div>

> [@stevengj](#):
>
> When you parse, you are looping over the string or using `findnext` or similar.

No, I didn’t get that far, as I am just experimenting with the language. I imagine the final implementation would use regexp matching for parsing as the text also contains tags. But I wanted to try out a few things out first… like, hey lets just cut out this field and sum it up. My `substring` function works fine, I was just curious about why you have `chop` instead.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 19, 2022, 2:28pm UTC](https://discourse.julialang.org/t/substring-function/76675/31 "2022-02-19T14:28:53Z")

</div>

> [@rogerkeays](#):
>
> No, I didn’t get that far,

Yes, that’s the problem.

You were probably trying to test things out by coming up with your indices “visually” on some test strings, and noticed that the indices did not coincide with your visual expectation for Unicode strings. But this would be a problem **even if Julia used vectors of codepoints** , as my `"ü"` example showed.

For example, if I gave you the string `"ü,é,â,ỳ"` and asked you to cut out the 3rd and 4th comma-separated fields, analogous to `cut -d -f3-4`, what “characters” (codepoints) do you think that corresponds to? Probably you would guess “the 5th to 7th characters”. But no, look what your `substring` function returns:

```julia
julia> substring("ü,é,â,ỳ", 5,7)
"́,a"

```

Whoops, is your `substring` function buggy? No, it’s just that “codepoints” in Unicode don’t necessarily correspond to what a human reader thinks of as a “character”.

Whereas if you _actually_ implemented a `cut` function, you would do a sequence of searches for the delimiter (e.g. with `findnext`), which would yield a sequence of string indices (≠ codepoint counts), slicing would work just fine, and it would be **absolutely irrelevant** how many codepoints occurred between one delimiter and the next.

But because you hadn’t gotten that far, you jumped to the conclusion that Julia’s string handling is broken and we are missing extremely basic functionality like extracting substrings.

(The UTF-8 encoding that Julia employs is not unusual! It’s taking over most of the internet, it’s used in other modern languages like Go and Swift, and it’s been the subject of many, many discussions and revisions in Julia itself. This is not something we picked out of a hat because we hadn’t thought through basic functionality.)

---

<div class="post-metadata">

**Author:** ![rogerkeays](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rogerkeays/32/33927_2.png) [@rogerkeays](https://discourse.julialang.org/u/rogerkeays)\
**Post date:** [February 19, 2022, 2:49pm UTC](https://discourse.julialang.org/t/substring-function/76675/32 "2022-02-19T14:49:19Z")

</div>

Hey, thanks for the reply. I didn’t mean to criticise Julia, or suggest that Julia’s unicode implementation is broken.

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [February 19, 2022, 3:06pm UTC](https://discourse.julialang.org/t/substring-function/76675/33 "2022-02-19T15:06:59Z")

</div>

```julia
julia> SubString("äöü",2,3)
ERROR: StringIndexError: invalid index [2], valid nearby indices [1]=>'ä', [3]=>'ö'

```

Reading this discussion and I see the reasons. Also knowing about Unicode/UTF-8/… difficulties.  
But…

I would prefer having just

```julia
julia> SubString("äöü",2,3)
"öü"

```

even if it would be not efficient.  
(not discussing function substring or view-like SubString here, doesn’t matter for me)

I don’t think expecting a working `substring` functionality isn’t that uncommon or even wrong like in “you are using substring, you are doing it probably wrong”.

In my opinion: the best reasoning doesn’t help if people just expect it validly. Expecting some kind of `substring` is valid in my opinion for any higher order programming language (of course not for assembler).

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 19, 2022, 3:18pm UTC](https://discourse.julialang.org/t/substring-function/76675/34 "2022-02-19T15:18:19Z")

</div>

> [@oheil](#):
>
> I don’t think expecting a working `substring` functionality isn’t that uncommon or even wrong like in “you are using substring, you are doing it probably wrong”.

Julia _has_ a working substring functionality, which works on string indices.

> [@oheil](#):
>
> I would prefer having just
> 
> ```julia-auto
> julia> SubString("äöü",2,3)
> "öü"
> 
> ```

The proposed character-indexing method does _not_ eliminate the complexities of Unicode. Consider:

```julia-auto
julia> substring(str, start, stop) = str[nextind(str, 0, start):nextind(str, 0, stop)]

julia> substring("äöü", 2,3) # NFC normalized string
"öü"

julia> substring("äöü", 2,3) # NFD normalized string
"̈o"

```

The supposed simplicity of “character indexing” is an illusion.

You pay a big price in performance to reduce apparent confusion on people’s first few days of using strings in Julia, but only postpone your Unicode bugs (because “characters” don’t mean what you think), and you get **zero benefits in the long run** (because indexing codepoint counts is not actually necessary for realistic string processing).

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [February 19, 2022, 3:26pm UTC](https://discourse.julialang.org/t/substring-function/76675/35 "2022-02-19T15:26:59Z")

</div>

> [@Sukera](#):
>
> Java is using UTF-16, right? The same problems mentioned in the two links above should apply as well, as it’s a variable length encoding like UTF-8. I don’t know how java would treat those bad encodings though.

A significant difference from UTF-8 is that there is no such thing as malformed UTF-16, only unpaired surrogate code units, which still have code points, just not ones that correspond to valid Unicode characters. So every UTF-16 string can be iterated as code points, you might just get the code point for an unpaired surrogate if the string is invalid. Compare that with UTF-8 where some byte sequences just don’t follow the right structure at all.

(Another way to put this is that every UTF-16 sequence, valid or invalid, can be represented as a sequence of code points using WTF-8, which is an extension of UTF-8 allowing surrogate pair code points.)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 19, 2022, 3:29pm UTC](https://discourse.julialang.org/t/substring-function/76675/36 "2022-02-19T15:29:06Z")

</div>

We could certainly add a `graphemes(str, m, n)` method that returns a substring of the m-th to n-th graphemes, which is close to a user-perceived character slice.

But is this actual useful functionality in _any_ real application? (When was the last time you used the `graphemes(str)` iterator, for that matter? At least grapheme iteration has some practical uses, e.g. cursor movement, but random grapheme access seems basically useful only for string demos.)

As far as I can tell, its sole practical utility would be to end discussions like this one. (Which might be worth a ~~\< 10~~ ≈25 line function in the Unicode stdlib, I suppose.)

[https://github.com/JuliaLang/julia/pull/44266](https://github.com/JuliaLang/julia/pull/44266)

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [February 19, 2022, 8:51pm UTC](https://discourse.julialang.org/t/substring-function/76675/37 "2022-02-19T20:51:15Z")

</div>

To Steve:

```julia
using Unicode
graphemes(('\U1F44D'*'\U0352')^100)

```

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [February 19, 2022, 9:40pm UTC](https://discourse.julialang.org/t/substring-function/76675/38 "2022-02-19T21:40:25Z")

</div>

> [@stevengj](#):
>
> But is this actual useful functionality in _any_ real application?

Isn’t grapheme slicing the way to generate an icon like this?

![image](https://global.discourse-cdn.com/julialang/original/3X/f/8/f8b5c7150273eecb996e15ce09773504049105f1.png)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 19, 2022, 9:48pm UTC](https://discourse.julialang.org/t/substring-function/76675/39 "2022-02-19T21:48:15Z")

</div>

No, you just need to call `first` on the `graphemes` iterator to get the “first letter”.

In contrast, general slicing means (simulated) _random_ access. No one has so far given any practical use-case for this in which you don’t already have (or can easily get) string indices.

---

<div class="post-metadata">

**Author:** ![davidavdav](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidavdav/32/1065_2.png) [@davidavdav](https://discourse.julialang.org/u/davidavdav)\
**Post date:** [July 18, 2022, 2:26pm UTC](https://discourse.julialang.org/t/substring-function/76675/40 "2022-07-18T14:26:50Z")

</div>

But then, what is the _real_ application of byte-indexing a substring? It looks pretty low-level to me—I understand you might need such indexing at the lowest level, manipulating bytes in doing the encoding/decoding, but I suppose this could also be done by a `convert` to `Vector{UInt8}` or something similar.

And sure, there is the NFC/NFD normalization choice, that users of substring/graphemes should be aware of. In my real applications I sometimes use NFD, when I am interested in diacritics, and otherwise NFC. And sure, I have burnt myself working with un-normalized Unicode strings. Copying from your example `"äöü" != "äöü"`, which is also somewhat counter intuitive.

I can also imaging that modern programming paradigms discourage people from doing any kind of indexing of strings, e.g., by providing the most amazing set of string manipulation and iteration functions.

[Previous page](https://discourse.julialang.org/t/substring-function/76675.md?page=1)

[Next page](https://discourse.julialang.org/t/substring-function/76675.md?page=3)
