# Julia performs poorly on group-by benchmarks

**URL:** <https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476>\
**Category:** Data\
**Tags:** performance\
**Created:** [October 18, 2018, 1:29am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476 "2018-10-18T01:29:08Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 18, 2018, 1:29am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/1 "2018-10-18T01:29:08Z")

</div>

h20 has published [a set of benchmarks](https://h2oai.github.io/db-benchmark/) that shows Julia’s DataFrames.jl has, in general, the worst group-by performance out of many data packages. JuliaDB.jl was not benchmarked so that may be a good addition. I have done [some work](https://discourse.julialang.org/t/group-by-performance-benchmarks-and-recommendations/9313) before on optimising some benchmarks and I’ve been putting it off until the release of v1.0. Now that v1.0.1 is out, it’s time for me to pick up the work again using [FastGroupBy.jl](https://github.com/xiaodaigh/FastGroupBy.jl).

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [October 18, 2018, 3:09am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/2 "2018-10-18T03:09:08Z")

</div>

Yeah the benchmarks leave a lot to be desired. It’s being discussed on discourse [here](https://discourse.julialang.org/t/dataframes-operation-scales-badly/16223/20). There are a few soon-to-be-merged PRs that should improve things.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [October 18, 2018, 8:56am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/3 "2018-10-18T08:56:33Z")

</div>

And yes I will get back to this at some point and report on progress with trying out some of the suggestions in that thread…

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [October 18, 2018, 2:22pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/4 "2018-10-18T14:22:41Z")

</div>

I created a nextjournal group for this:

> **[Nextjournal](https://nextjournal.com/julia-data)**
>
> A better way to write, share, remix and collaborate on your research.

Everyone in that group can edit the articles & publish new ones - we can even edit the article together at the same time! I plan to add a minimal julia version for the `sum v1 by id1` benchmark, so collaborative editing will be pretty useful! @xiaodai, seems like you’re in the pole position to add a minimal version? I think it would really help to have more people tune the performance!

The [sum v1 by id1](https://nextjournal.com/julia-data/sum-v1-by-id1) contains all the languages & multiple branches of the same packages in the same article in isolated runtimes which run on the same hardware. That makes it pretty easy to see where we are right now:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/f/2/f2a276b61768eeb278bc1b2fbc46b382664ebd08.png)

I hope we can fill in the performance gaps from here on! I put some instructions into [sum v1 by id1](https://nextjournal.com/julia-data/sum-v1-by-id1) on how to add new benchmarks. Let me know if you have any problems with that!

If you want to be part of the group to add a new version, [signup](https://nextjournal.com/signup?code=juliacon) and let me know so I can add you to julia-data.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [December 4, 2018, 12:21pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/5 "2018-12-04T12:21:15Z")

</div>

The benchmarks could be redone with the new dataframes v0.15.

> [@Release announcements for DataFrames.jl](https://discourse.julialang.org/t/dataframes-v0-15-0-released/18258/12):
>
> I think categorical arrays are not optimized for only 100 observations per group. The typical ratio would be much higher than that. Do you get the same result with 1000 or 10000?

> **[Release notes for the DataFrames.jl package v0.15.0](https://juliasnippets.blogspot.com/2018/12/release-notes-for-dataframesjl-package.html)**
>
> The DataFrames.jl package is getting closer to 1.0 release. In order to reach this level maturity a number of significant changes is introdu...

Another good options is

> **[GitHub - JuliaString/InternedStrings.jl: Fully transparent string interning...](https://github.com/JuliaString/InternedStrings.jl)**
>
> Fully transparent string interning functionality, with proper garbage collection - GitHub - JuliaString/InternedStrings.jl: Fully transparent string interning functionality, with proper garbage col...

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 4, 2018, 12:35pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/6 "2018-12-04T12:35:44Z")

</div>

Need a pr as their code are old

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [December 4, 2018, 2:41pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/7 "2018-12-04T14:41:28Z")

</div>

I actually tried to update the article, but I run out of 15 gb RAM - seems weird, so not sure what’s happening!

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 4, 2018, 5:14pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/8 "2018-12-04T17:14:07Z")

</div>

> [@sdanisch](#):
>
> I actually tried to update the article, but I run out of 15 gb RAM - seems weird, so not sure what’s happening!

With what code exactly? I can’t reproduce locally.

---

<div class="post-metadata">

**Author:** ![ValdarT](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/valdart/32/24146_2.png) [@ValdarT](https://discourse.julialang.org/u/ValdarT)\
**Post date:** [December 4, 2018, 6:57pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/9 "2018-12-04T18:57:44Z")

</div>

The original H2O benchmark should also automatically update to the new `DataFrames` whenever it updates itself but it could use a PR as it uses the `do` syntax that `DataFrames` docs specifically recommend to avoid if performance matters.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 4, 2018, 7:11pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/10 "2018-12-04T19:11:48Z")

</div>

The [PR](https://github.com/h2oai/db-benchmark/pull/51) is already merged so we now have to wait and see how it goes (there are some issues with CSV.jl that might need fixing).

---

<div class="post-metadata">

**Author:** ![ValdarT](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/valdart/32/24146_2.png) [@ValdarT](https://discourse.julialang.org/u/ValdarT)\
**Post date:** [December 8, 2018, 9:44am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/11 "2018-12-08T09:44:50Z")

</div>

And the results are in. Not too good on the performance side but the new syntax is lovely and it’s nice to see that it could also handle the 50GB file (unlike, for example, Pandas)

---

<div class="post-metadata">

**Author:** ![matthieu](https://avatars.discourse-cdn.com/v4/letter/m/da6949/32.png) [@matthieu](https://discourse.julialang.org/u/matthieu)\
**Post date:** [December 8, 2018, 3:21pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/12 "2018-12-08T15:21:24Z")

</div>

These benchmarks are informative but datatable/panda/dplyr heavily optimize sum/mean. For any other function Julia may have similar or better performance (because of how slow it is to repeatedly call a function in R compared to Julia).

On the other hand, doing the benchmarkrs on data with missing value may make Julia slower compared to these packages.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 8, 2018, 3:38pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/13 "2018-12-08T15:38:21Z")

</div>

There are still performance optimizations that we can add to speed it up. But at least in all the benchmarks we got closer to the fastest options.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 9, 2018, 3:45am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/14 "2018-12-09T03:45:17Z")

</div>

> [@bkamins](#):
>
> There are still performance optimizations that we can add to speed it up.

Agree. Because InternedStrings.jl is so much better than before I tried an InternedString approach on 100million rows which yielded this approach

```julia
using SortingAlgorithms, InternedStrings, DataFrames

function createSynDataFrame(N::Int,K::Int)
    pool = "id".*string.(1:K, pad=3)
    pool1 = "id".*string.(1:N÷K,pad=10)
    nums = round.(rand(100).*100, digits = 4)

    df = DataFrame(
        id1 = intern.(rand(pool,N)),
        id2 = intern.(rand(pool,N)),
        id3 = intern.(rand(pool1,N)),
        id4 = rand(1:K,N),
        id5 = rand(1:K,N),
        id6 = rand(1:(N÷K),N),
        v1 = rand(1:5,N),
        v2 = rand(1:5,N),
        v3 = rand(nums,N))
    return df
end

@time df1 = createSynDataFrame(100_000_000, 100)

function sortandperm(x::Vector{String})
    ap = UInt.(pointer.(x))
    ai = sortperm(ap, alg = RadixSort)
    ap, ai
end

using BenchmarkTools
@benchmark sortedid1, ai = sortandperm(df1.id1)

```

> ## BenchmarkTools.Trial: memory estimate: 1.49 GiB allocs estimate: 7
> 
> ## minimum time: 757.153 ms (5.08% GC) median time: 913.535 ms (21.46% GC) mean time: 1.079 s (32.91% GC) maximum time: 2.023 s (64.62% GC)
> 
> samples: 5  
> evals/sample: 1

As you can on my laptop the timing for the most expensive part which is grouping only takes 2 seconds vs 15 seconds it took for the first group by for Julia (5G), and making a sum after should be a piece of cake and take not much time at all. I think Julia should finish within 4 seconds if we use `InternStrings.jl`

---

<div class="post-metadata">

**Author:** ![datnamer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datnamer/32/3471_2.png) [@datnamer](https://discourse.julialang.org/u/datnamer)\
**Post date:** [December 9, 2018, 4:26am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/15 "2018-12-09T04:26:25Z")

</div>

Where are the results?

---

<div class="post-metadata">

**Author:** ![ValdarT](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/valdart/32/24146_2.png) [@ValdarT](https://discourse.julialang.org/u/ValdarT)\
**Post date:** [December 9, 2018, 11:10am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/16 "2018-12-09T11:10:31Z")

</div>

The [link](https://h2oai.github.io/db-benchmark/) is in the very first post of this thread.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [December 9, 2018, 11:22am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/17 "2018-12-09T11:22:57Z")

</div>

Will InternedStrings.jl be included by default on DataFrames.jl for string columns?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [December 9, 2018, 11:29am UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/18 "2018-12-09T11:29:43Z")

</div>

I don’t think it will because it takes time to intern the strings.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 9, 2018, 2:31pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/19 "2018-12-09T14:31:01Z")

</div>

CSV.jl already interns strings, the problem is that we have no way to know whether a `Vector{String}` column only contains interned strings or not. The best solution we have for now is to use CategoricalArrays (using `categorical=true` or `categorical=0.1`, etc.), or PooledArrays (not yet supported). That’s actually even more efficient than interned strings, since we have group indices from 1 to N rather than pointers to strings (which still need to be hashed or sorted).

The H2O benchmark shows poor results for DataFrames.jl because it doesn’t use CategoricalArrays right now, but hopefully it will very soon.

---

<div class="post-metadata">

**Author:** ![ImreSamu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/imresamu/32/20677_2.png) [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Post date:** [December 9, 2018, 6:24pm UTC](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476/20 "2018-12-09T18:24:16Z")

</div>

As I see ( [/h2oai/db-benchmark/juliadf/setup-juliadf.sh](https://github.com/h2oai/db-benchmark/blob/master/juliadf/setup-juliadf.sh))

- it is a **Julia v1.0.0** benchmark ( now `v1.0.2` and `v1.1.0` expected )
- No [precompile](https://docs.julialang.org/en/v1/stdlib/Pkg/#Precompiling-a-project-1).

imho: Maybe it is a little effect … but important for the fresh benchmark.

[Next page](https://discourse.julialang.org/t/julia-performs-poorly-on-group-by-benchmarks/16476.md?page=2)
