# Efficient way to split string at specific index

**URL:** <https://discourse.julialang.org/t/efficient-way-to-split-string-at-specific-index/83115>\
**Category:** Performance\
**Tags:** question, strings\
**Created:** [June 21, 2022, 11:04am UTC](https://discourse.julialang.org/t/efficient-way-to-split-string-at-specific-index/83115 "2022-06-21T11:04:12Z")\
**Posts on this page:** 1\
**Showing post:** 17

<div class="post-metadata">

**Author:** ![BambOoxX](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bambooxx/32/22179_2.png) [@BambOoxX](https://discourse.julialang.org/u/BambOoxX)\
**Post date:** [June 21, 2022, 5:16pm UTC](https://discourse.julialang.org/t/efficient-way-to-split-string-at-specific-index/83115/17 "2022-06-21T17:16:24Z")

</div>

Just for the sake of the exercise (I do not know if this is really relevant) I wanted to assess the relative performance of the current `split` implementation in two cases :

- A space is added between initial blocks of characters (resulting in a longer string)
- The same but with one character block missing (resulting in an identical initial string, but one less item in the output vector)

```julia
str = join([repeat(letter, 8) for letter in ('a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j')])
str_spaced = join([repeat(letter, 8) for letter in ('a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j')], " ")
str_spaced_short = join([repeat(letter, 8) for letter in ('a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i')], " ")
@benchmark split($str_spaced)
@benchmark split($str_spaced_short)

```

Which gives

```julia
aaaaaaaabbbbbbbbccccccccddddddddeeeeeeeeffffffffgggggggghhhhhhhhiiiiiiiijjjjjjjj
aaaaaaaa bbbbbbbb cccccccc dddddddd eeeeeeee ffffffff gggggggg hhhhhhhh iiiiiiii jjjjjjjj
aaaaaaaa bbbbbbbb cccccccc dddddddd eeeeeeee ffffffff gggggggg hhhhhhhh iiiiiiii

julia> @benchmark split($str_spaced)
BenchmarkTools.Trial: 10000 samples with 10 evaluations.
 Range (min … max): 1.270 μs … 1.429 ms ┊ GC (min … max): 0.00% … 99.78%
 Time (median): 1.910 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 2.171 μs ± 14.289 μs ┊ GC (mean ± σ): 6.57% ± 1.00%

         ▂▅█▆▃▂
  █▆▄▄▃▅███████▆▅▅▄▄▄▄▄▄▃▃▃▃▃▃▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▁▂▂▂▂▂▂▂▂▂▂▂ ▃
  1.27 μs Histogram: frequency by time 4.94 μs <

 Memory estimate: 1.25 KiB, allocs estimate: 3.

julia> @benchmark split($str_spaced_short)
BenchmarkTools.Trial: 10000 samples with 10 evaluations.
 Range (min … max): 1.180 μs … 1.137 ms ┊ GC (min … max): 0.00% … 99.80%
 Time (median): 1.640 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 1.799 μs ± 11.361 μs ┊ GC (mean ± σ): 6.31% ± 1.00%

    ▄▂ ▂ ▁▆█▃▆▄▅▂ 
  ▃▅██▄▄▃▃███████████▇▆▅▅▃▃▃▃▂▂▂▂▂▂▂▂▂▂▂▂▂▂▁▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ ▃
  1.18 μs Histogram: frequency by time 3.05 μs <

 Memory estimate: 1.25 KiB, allocs estimate: 3.

```

I expected that this would be more difficult than splitting on indices given that the delimiter(s) could be anywhere in the string, but I think it also shows that there is room for improvement for index splitting (except the comprehension)

---

_[View the full topic](https://discourse.julialang.org/t/efficient-way-to-split-string-at-specific-index/83115)._
