# Counting special characters ü, å, ø, etc

**URL:** https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858
**Category:** General Usage
**Tags:** strings, unicode
**Created:** [April 1, 2022, 9:49am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858 "2022-04-01T09:49:13Z")
**Posts on this page:** 12
**Page:** 1

<div class="post-metadata">

### Author: ![etas](https://avatars.discourse-cdn.com/v4/letter/e/ecae2f/32.png) [@etas](https://discourse.julialang.org/u/etas)
#### Post date: [April 1, 2022, 9:49am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/1 "2022-04-01T09:49:13Z")

</div>

Hi, I have some strings with a name and a distance such as `"Haugastøl, Km 0"`. I want to remove the km point with the chop() function but what’s interesting is that chop() and length() count ø as one character but not findlast().

```julia
julia> name = "Haugastøl"
"Haugastøl"

julia> length(name)
9

julia> findlast('l',name)
10

```

I don’t know if it’s normal to have these results, maybe findlast() should count special characters as one ?

---

<div class="post-metadata">

### Author: ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)
#### Post date: [April 1, 2022, 9:59am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/2 "2022-04-01T09:59:50Z")

</div>

The reason is that in Julia strings are UTF-8 encoded and you can index into them using byte count or character count. Some functions use the first approach and some the second. This is explained [here in the Julia manual](https://docs.julialang.org/en/v1/manual/strings/#String-Basics) and additionally [here in my blog](https://bkamins.github.io/julialang/2020/08/13/strings.html). If some of the explanations are not clear please comment and I can expand on them.

---

<div class="post-metadata">

### Author: ![josuagrw](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josuagrw/32/1015_2.png) [@josuagrw](https://discourse.julialang.org/u/josuagrw)
#### Post date: [April 1, 2022, 10:00am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/3 "2022-04-01T10:00:37Z")

</div>

Note what happens when you try to access the 9th element with the `Array` interface:

```julia
julia> name[8]
'ø': Unicode U+00F8 (category Ll: Letter, lowercase)

julia> name[9]
ERROR: StringIndexError: invalid index [9], valid nearby indices [8]=>'ø', [10]=>'l'
Stacktrace:
 [1] getindex(s::String, i::Int64)
   @ Base ./strings/string.jl:226
 [2] top-level scope
   @ REPL[14]:1

julia> name[10]
'l': ASCII/Unicode U+006C (category Ll: Letter, lowercase)

```

---

<div class="post-metadata">

### Author: ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)
#### Post date: [April 1, 2022, 10:01am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/4 "2022-04-01T10:01:06Z")

</div>

(actually my blog just reminded me that there is even a third option which is “number of characters displayed” which can be different from the two basic ones I have listed)

---

<div class="post-metadata">

### Author: ![josuagrw](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josuagrw/32/1015_2.png) [@josuagrw](https://discourse.julialang.org/u/josuagrw)
#### Post date: [April 1, 2022, 10:13am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/5 "2022-04-01T10:13:02Z")

</div>

I’m certain this could be done more efficiently but here is a method which returns the number you expected

```julia
julia> findfirst(==(findfirst('l', name)), collect(eachindex(name)))
9

```

---

<div class="post-metadata">

### Author: ![etas](https://avatars.discourse-cdn.com/v4/letter/e/ecae2f/32.png) [@etas](https://discourse.julialang.org/u/etas)
#### Post date: [April 1, 2022, 10:17am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/6 "2022-04-01T10:17:02Z")

</div>

Thank you for your quick responses !

@josuagrw I was exactly wondering what would be the result for the 9th index 🙂  
Thanks for your solution, I will try it.

---

<div class="post-metadata">

### Author: ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)
#### Post date: [April 1, 2022, 10:24am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/7 "2022-04-01T10:24:24Z")

</div>

@etas - going back to your original question. Do you know how to do what you wanted or you need a solution (if it is the latter could you please precisely define what you need then an efficient solution can be proposed). Thank you!

---

<div class="post-metadata">

### Author: ![josuagrw](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josuagrw/32/1015_2.png) [@josuagrw](https://discourse.julialang.org/u/josuagrw)
#### Post date: [April 1, 2022, 10:33am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/8 "2022-04-01T10:33:30Z")

</div>

If you want to do this sort of thing

> [@josuagrw](#):
>
> `julia> findfirst(==(findfirst('l', name)), collect(eachindex(name)))`

many times for the same string, look up the `indexin()` function.

---

<div class="post-metadata">

### Author: ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)
#### Post date: [April 1, 2022, 10:36am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/9 "2022-04-01T10:36:03Z")

</div>

My question above was that most likely what @etas needs is `Regex` matching, but I need to understand exactly what is the pattern that should be identified (most likely not finding the `l` character as it is specific to only a given string and will not work in general).

---

<div class="post-metadata">

### Author: ![josuagrw](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josuagrw/32/1015_2.png) [@josuagrw](https://discourse.julialang.org/u/josuagrw)
#### Post date: [April 1, 2022, 10:36am UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/10 "2022-04-01T10:36:47Z")

</div>

Could also be done with type matching:

```julia
julia> foreach(display, name)
'H': ASCII/Unicode U+0048 (category Lu: Letter, uppercase)
'a': ASCII/Unicode U+0061 (category Ll: Letter, lowercase)
'u': ASCII/Unicode U+0075 (category Ll: Letter, lowercase)
'g': ASCII/Unicode U+0067 (category Ll: Letter, lowercase)
'a': ASCII/Unicode U+0061 (category Ll: Letter, lowercase)
's': ASCII/Unicode U+0073 (category Ll: Letter, lowercase)
't': ASCII/Unicode U+0074 (category Ll: Letter, lowercase)
'ø': Unicode U+00F8 (category Ll: Letter, lowercase)
'l': ASCII/Unicode U+006C (category Ll: Letter, lowercase)

```

---

<div class="post-metadata">

### Author: ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)
#### Post date: [April 1, 2022, 12:39pm UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/11 "2022-04-01T12:39:49Z")

</div>

> [@etas](#):
>
> Hi, I have some strings with a name and a distance such as `"Haugastøl, Km 0"` . I want to remove the km point with the chop() function but what’s interesting is that chop() and length() count ø as one character but not findlast().

Separate from the issue of UTF-8 _indexing_ (as in `findlast`) vs “character counting” (as in `length`), you should also be aware that Unicode is more complicated than you think. For example:

```julia
julia> length("fübâr")
7

julia> collect("fübâr")
7-element Vector{Char}:
 'f': ASCII/Unicode U+0066 (category Ll: Letter, lowercase)
 'u': ASCII/Unicode U+0075 (category Ll: Letter, lowercase)
 '̈': Unicode U+0308 (category Mn: Mark, nonspacing)
 'b': ASCII/Unicode U+0062 (category Ll: Letter, lowercase)
 'a': ASCII/Unicode U+0061 (category Ll: Letter, lowercase)
 '̂': Unicode U+0302 (category Mn: Mark, nonspacing)
 'r': ASCII/Unicode U+0072 (category Ll: Letter, lowercase)

```

This has nothing to do with Julia or how it encodes strings. It’s because a “character” like `ü` might actually be represented by multiple Unicode codepoints (a `u` followed by a “combining accent” in this case).

See also my answer in a previous thread: [Substring function? - #31 by stevengj](https://discourse.julialang.org/t/substring-function/76675/31)

In practice, you mostly find indices in strings by searching, e.g. by doing a regex search for `r", *Km *[0-9]+$"` in this case, in which case these complications are mostly hidden.

But it can be confusing when slicing strings “visually” for things you enter by hand. For working “visually” with a string, the closest thing to a human-perceived “character” is actually something called a “grapheme” in Unicode, and Julia 1.9 should have a [function to slice strings based on grapheme counts](https://github.com/JuliaLang/julia/pull/44266).

---

<div class="post-metadata">

### Author: ![etas](https://avatars.discourse-cdn.com/v4/letter/e/ecae2f/32.png) [@etas](https://discourse.julialang.org/u/etas)
#### Post date: [April 1, 2022, 1:25pm UTC](https://discourse.julialang.org/t/counting-special-characters-u-a-o-etc/78858/12 "2022-04-01T13:25:17Z")

</div>

I completely forgot to use Regex ! I will try that as we don’t need to know the counting methods used by the functions.  
My strings come from a file and I want to remove the last part, beginning by `", Km"`. `Regex` can definitively do the job without complications.

Thank you @bkamins and @stevengj !
