# Julia's DataFrames.jl performance on join benchmark

**URL:** <https://discourse.julialang.org/t/julias-dataframes-jl-performance-on-join-benchmark/30788>\
**Category:** Community\
**Tags:** dataframes\
**Created:** [November 6, 2019, 11:26am UTC](https://discourse.julialang.org/t/julias-dataframes-jl-performance-on-join-benchmark/30788 "2019-11-06T11:26:51Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [November 6, 2019, 11:26am UTC](https://discourse.julialang.org/t/julias-dataframes-jl-performance-on-join-benchmark/30788/1 "2019-11-06T11:26:51Z")

</div>

Please see [Database-like ops benchmark](https://h2oai.github.io/db-benchmark/)

There is a new Join benchmark where Julia’s DataFrames.jl is in the last 2 spots most of the time. And takes is 30x slower than R’s {data.table} in some of the benchmarks.

This reflects the fact that Julia’s data ecosystem is not as mature and Julia’s grouping benchmark was in a similar spot not long ago but has since improved greatly.

Hope to see more development in Julia’s data ecosystem. I want to see if Julia can match data.table in some cases. I have done some work on group-bys and managed to match data.table’s performance in some cases.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 6, 2019, 4:09pm UTC](https://discourse.julialang.org/t/julias-dataframes-jl-performance-on-join-benchmark/30788/2 "2019-11-06T16:09:57Z")

</div>

That’s the reason I’m using data.table for most data manipulation I do. It’s really fast and has useful functions.  
I’ll definitely switch to Julia when it gets faster than data.table.
