# What functions/packages should I use to sort and "group by" as fast as possible...?

**URL:** <https://discourse.julialang.org/t/what-functions-packages-should-i-use-to-sort-and-group-by-as-fast-as-possible/18687>\
**Category:** Performance\
**Tags:** sort, dataframes\
**Created:** [December 15, 2018, 1:02pm UTC](https://discourse.julialang.org/t/what-functions-packages-should-i-use-to-sort-and-group-by-as-fast-as-possible/18687 "2018-12-15T13:02:55Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [December 15, 2018, 1:02pm UTC](https://discourse.julialang.org/t/what-functions-packages-should-i-use-to-sort-and-group-by-as-fast-as-possible/18687/1 "2018-12-15T13:02:55Z")

</div>

I’m looking for fast functions to group by or sort large dataframes with several columns of strings (with few different values) or numbers.

There are many packages and options: Base, ShortStrings, SortingLab, FastGroupBy, SortingAlgorithms…

I have only been able to install and use ShortStrings and SortingAlgorithms, the others produce errors.

Some packages say their functionality will be merged onto base Julia o dataframe.

What do yo suggest to use?  
Any special function from those packages? Or just keep with the defaults on dataframe or alternatives such as IndexedTables or CategoricalArrays?

There are other threads about this:

> [@Julia performs poorly on group-by benchmarks](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476):
>
> h20 has published [a set of benchmarks](https://h2oai.github.io/db-benchmark/) that shows Julia’s DataFrames.jl has, in general, the worst group-by performance out of many data packages. JuliaDB.jl was not benchmarked so that may be a good addition. I have done [some work](https://discourse.julialang.org/t/group-by-performance-benchmarks-and-recommendations/9313) before on optimising some benchmarks and I’ve been putting it off until the release of v1.0. Now that v1.0.1 is out, it’s time for me to pick up the work again using [FastGroupBy.jl](https://github.com/xiaodaigh/FastGroupBy.jl).

> [@Progress towards faster \`sortperm\` for Strings](https://discourse.julialang.org/t/progress-towards-faster-sortperm-for-strings/8505/14):
>
> We kind of do this already since CSV.jl creates CategoricalArray vectors for columns with a small proportion of unique values. I think it makes sense to use CategoricalArray/PooledArray to represent values with lots of duplicates, and arrays of strings for “real” strings which are almost all different from one another.

> [@Group-by performance benchmarks and recommendations](https://discourse.julialang.org/t/group-by-performance-benchmarks-and-recommendations/9313):
>
> I have been trying to improve Julia’s DataFrame group-by for a while now and I think I am able to synthesized my thinking into APIs. Here are my synthesized recommendation (as of 25th Feb 2018) Recommendation Why? Benchmark Status If you need to perform group-by A LOT use JuliaDB.jl/IndexedTables.jl You can create indexes on the group-by columns and it is generally faster than other available methods No benchmarks; as the benchmarks was intended to test non-indexing performance If th…

but I don’t want to be told again I’m writing on old threads, even if they speak about the same.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 15, 2018, 10:33pm UTC](https://discourse.julialang.org/t/what-functions-packages-should-i-use-to-sort-and-group-by-as-fast-as-possible/18687/3 "2018-12-15T22:33:17Z")

</div>

Most of them are mt posts. I am learning how to make the package work in v1, but my package can only handle one column sorts well and multi-column sorts would need some work. Basically to make it run fast for stringa you need to radixsort the strings and in the sort one should return the `sortperm` as well which ca n be used to sort the other columns. This is the fastest way I found

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [December 15, 2018, 11:32pm UTC](https://discourse.julialang.org/t/what-functions-packages-should-i-use-to-sort-and-group-by-as-fast-as-possible/18687/4 "2018-12-15T23:32:39Z")

</div>

But do you suggest keeping the data on frameworks and use some package to perform the radix sort? What package (ShortStrings, SortingLab, FastGroupBy, SortingAlgorithms,…)?  
Or should we use other framework such as Indexedtables, Categoricalarrays, Staticarrays… instead?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 15, 2018, 11:47pm UTC](https://discourse.julialang.org/t/what-functions-packages-should-i-use-to-sort-and-group-by-as-fast-as-possible/18687/5 "2018-12-15T23:47:19Z")

</div>

The third post in your original post remains the most up to date advice. I think just use `DataFramesMeta.jl` or just the functions in `DataFrames.jl` directly. They are not fast yet, but they remain the best hope because they might get updated with faster algorithms. Ok. I will devote the next week to work on getting SortingLab.jl and FastGroupBy.jl ready for Julia v1.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [December 16, 2018, 1:05am UTC](https://discourse.julialang.org/t/what-functions-packages-should-i-use-to-sort-and-group-by-as-fast-as-possible/18687/6 "2018-12-16T01:05:07Z")

</div>

While looking at differet benchmarks I’ve seen other promising solutions such as Dask, SciDB and Mapd.
