# Who does "better" than DataFrames?

**URL:** <https://discourse.julialang.org/t/who-does-better-than-dataframes/96946>\
**Category:** Performance\
**Tags:** dataframes\
**Created:** [April 1, 2023, 1:12pm UTC](https://discourse.julialang.org/t/who-does-better-than-dataframes/96946 "2023-04-01T13:12:18Z")\
**Posts on this page:** 1\
**Showing post:** 25

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [April 2, 2023, 5:57pm UTC](https://discourse.julialang.org/t/who-does-better-than-dataframes/96946/25 "2023-04-02T17:57:02Z")

</div>

In the grouping algo an open-addressing hashing table with linear probing is used to find groups. Look at `row_group_slots!` in `src/groupeddataframe/utils.jl` file of DataFrames.

On my computer, even `unique(s)` is slower than the whole `combine` statement (no threading). It is surprising. This is worth digging into, in order to make rest of eco-system more optimized.

UPDATE:

Actually, after a bit more digging, it turns out **DataFrames uses a few more assumptions which allow it to make things faster**. Essentially, it assumes the groups defined by the `:s` column are the integers between its extrema: `(1, 1_000_000)`. Which makes the need for any sorting and hashing spurious. And it allows to construct the result immediately. It’s a useful optimization and isn’t cheating, but this wouldn’t have the generality of any of the other methods attempted here.

The relevant code in DataFrames, is `DataFrames.refpool_and_array(df.s)`.

Here is a benchmark, showing such an optimization allows achieving double the speed of DataFrames in custom-tailored way:

```julia
function testy(s,r)
    res = fill(0.0, 1_000_000)
    @inbounds for i in eachindex(s)
        res[s[i]] = max(res[s[i]], r[i])
    end
    pairs(res)
end

```

and

```julia
julia> @btime combine(groupby(df, :s),:r=>maximum; threads=false);
  162.666 ms (334 allocations: 55.33 MiB)

julia> @btime testy($s,$r);
  80.950 ms (2 allocations: 7.63 MiB)

```

---

_[View the full topic](https://discourse.julialang.org/t/who-does-better-than-dataframes/96946)._
