# Performance of length(::String)

**URL:** <https://discourse.julialang.org/t/performance-of-length-string/12672>\
**Category:** Performance\
**Created:** [July 26, 2018, 7:28pm UTC](https://discourse.julialang.org/t/performance-of-length-string/12672 "2018-07-26T19:28:37Z")\
**Posts on this page:** 5\
**Page:** 2

<div class="post-metadata">

**Author:** ![tk3369](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tk3369/32/2824_2.png) [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Post date:** [July 28, 2018, 5:14pm UTC](https://discourse.julialang.org/t/performance-of-length-string/12672/21 "2018-07-28T17:14:15Z")

</div>

Argh… makes perfect sense! Thanks for the clarification.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [July 28, 2018, 6:22pm UTC](https://discourse.julialang.org/t/performance-of-length-string/12672/22 "2018-07-28T18:22:05Z")

</div>

The `sizeof` function is a bit of a misrecommendation here. It gives the size in bytes of the string representation. However, what you generally want is the number of code units not the number of bytes, which is given by `ncodeunits(s)`. In the case of `String` the code unit is a byte so these coincide, but indexing in abstract strings in general is in terms of code units, not bytes. Since `ncodeunits` is undefined for characters, the above confusion would not occur: only `ncodeunits(s[1:1])` makes sense and `ncodeunits(s[1])` is an error. I need to rewrite the string docs to explain the general string model but there has been more pressing work until now.

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [July 28, 2018, 8:46pm UTC](https://discourse.julialang.org/t/performance-of-length-string/12672/23 "2018-07-28T20:46:34Z")

</div>

> [@stevengj](#):
>
> If you want the index of the last character in the string, use `lastindex(s)` ; `length` (in any language where it returns the number of codepoints) is wrong in any variable-width encoding. Nor is `length` the size of the string in memory — that’s `sizeof` .

`lastindex()` wouldn’t be able to work on most variable-width encodings, because it depends on being able to move in the reverse direction (the only encodings I know of that can are UTF-8 and UTF-16, and I’m not even sure that it will get the same set of characters if there are invalid sequences as in the forward direction, but `String` no longer makes any guarantees that the string is valid UTF-8).

A number of string handling functions in Julia depend on `lastindex`, `prevind` and `thisind` working, such as creating `SubString`s, which means that they are not useful for handling arbitrary variable-width encodings.

There are also performance issues with having to read 1-4 (or 5 or 6) bytes going backwards to find the first byte (of a _valid_ sequence - it’s a lot more complicated if you have to also deal with possibly invalid sequences)  
Unfortunately, `SubString` doesn’t even have a constructor that allows you to bypass that.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [July 28, 2018, 9:28pm UTC](https://discourse.julialang.org/t/performance-of-length-string/12672/24 "2018-07-28T21:28:24Z")

</div>

> [@ScottPJones](#):
>
> I’m not even sure that it will get the same set of characters if there are invalid sequences as in the forward direction

As far as I remember it was thoroughly checked that you will. To make sure I have just written a randomized test that confirms that all works correctly.

```julia
function rs()
    x = rand(UInt8, 10^8)
    s = String(x)
    i = 2
    curidx = 1
    nextidx = nextind(s, 1)
    while i <= ncodeunits(s)
        if i < nextidx
            thisind(s, i) == curidx || @error i
            if i > curidx
                prevind(s, i) == curidx || @error i
            end
        else
            isvalid(s, i) || @error i
            thisind(s, i) == i || @error i
            curidx = i
            nextidx = nextind(s, i)
        end
    i += 1
    end
end

```

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [July 28, 2018, 9:49pm UTC](https://discourse.julialang.org/t/performance-of-length-string/12672/25 "2018-07-28T21:49:09Z")

</div>

That’s good to know.  
I think there are still some inconsistencies, because of different ways of handling UTF-8 sequences that encode numbers from 0x110000 - 0x7fffffff (those were allowed for the first 10 years of the UTF-8 spec, before they decided to limit things to be consistent with the range of UTF-16, i.e. 0 - 0x10ffff, and many programs, including both internals of PCRE and Julia itself [in the utf8.c that Jeff wrote in 2005], treat those 5-6 byte sequences as a single character, unlike what `String` currently does.

[Previous page](https://discourse.julialang.org/t/performance-of-length-string/12672.md?page=1)
