# Progress towards faster \`sortperm\` for Strings

**URL:** <https://discourse.julialang.org/t/progress-towards-faster-sortperm-for-strings/8505>\
**Category:** Data\
**Tags:** performance, sortperm\
**Created:** [January 21, 2018, 2:04pm UTC](https://discourse.julialang.org/t/progress-towards-faster-sortperm-for-strings/8505 "2018-01-21T14:04:42Z")\
**Posts on this page:** 1\
**Showing post:** 14

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [January 28, 2018, 9:48am UTC](https://discourse.julialang.org/t/progress-towards-faster-sortperm-for-strings/8505/14 "2018-01-28T09:48:32Z")

</div>

> [@ChrisRackauckas](#):
>
> You suggested that DataFrames.jl or other data readers could use interned strings so that way the performance is specialized to both domains. I think that’s a really great idea and would be a great use of Julia’s type system since most/all algorithms should work no matter what string type is internally used, and that context switch is a great way to get performance. I would be interested in seeing that investigated in more depth.

We kind of do this already since CSV.jl creates `CategoricalArray` vectors for columns with a small proportion of unique values. I think it makes sense to use `CategoricalArray`/`PooledArray` to represent values with lots of duplicates, and arrays of strings for “real” strings which are almost all different from one another.

---

_[View the full topic](https://discourse.julialang.org/t/progress-towards-faster-sortperm-for-strings/8505)._
