# Allocations and slow perf for Transform! on GroupedDataFrames

**URL:** <https://discourse.julialang.org/t/allocations-and-slow-perf-for-transform-on-groupeddataframes/60594>\
**Category:** Data\
**Created:** [May 5, 2021, 4:21pm UTC](https://discourse.julialang.org/t/allocations-and-slow-perf-for-transform-on-groupeddataframes/60594 "2021-05-05T16:21:04Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [May 5, 2021, 4:21pm UTC](https://discourse.julialang.org/t/allocations-and-slow-perf-for-transform-on-groupeddataframes/60594/1 "2021-05-05T16:21:04Z")

</div>

The `transform!` operator appears to result in large number of allocation as large slowdown compared to the non-grouped counterpart.

```julia
using DataFrames
using StatsBase: sample
using BenchmarkTools
df1 = DataFrame(rand(1_000_000, 100), :auto)
df1[:, :grp] .= sample(Int.(1:100), 1_000_000)
dfg1 = groupby(df1, ["grp"])

function test1(df)
    transform!(df, "x9" => ((x) -> x .^ 2) => "x9B")
end

```

For regular DataFrame:

```julia
julia> @btime test1($df1);
  1.469 ms (556 allocations: 7.66 MiB)

```

For GroupedDataFrame:

```julia
julia> @btime test1($dfg1);
  388.991 ms (9003310 allocations: 311.65 MiB)

```

As it can be seen, the performance is actually quite bad on the GroupedDataFrame (250X times slower, 40X the allocation size), although my expectation would have been for a relatively modest overhead from operating on the 100 groups. Did I wrongly used `transform!` or is there a real performance issue?

The above was run on `DataFrames v1.1.0`.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [May 5, 2021, 5:15pm UTC](https://discourse.julialang.org/t/allocations-and-slow-perf-for-transform-on-groupeddataframes/60594/2 "2021-05-05T17:15:07Z")

</div>

Thank you very much for reporting. Mostly fixed in [fix performance issue in multirow split-apply-combine by bkamins · Pull Request #2749 · JuliaData/DataFrames.jl · GitHub](https://github.com/JuliaData/DataFrames.jl/pull/2749) (maybe it still can be improved but the major sources of problems are solved):

```julia
julia> @btime test1($df1);
  890.798 μs (556 allocations: 7.66 MiB)

julia> @btime test1($dfg1);
  25.061 ms (65066 allocations: 54.79 MiB)

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [May 5, 2021, 6:41pm UTC](https://discourse.julialang.org/t/allocations-and-slow-perf-for-transform-on-groupeddataframes/60594/3 "2021-05-05T18:41:13Z")

</div>

I am down to:

```julia
julia> @btime test1($df1);
  901.361 μs (554 allocations: 7.66 MiB)

julia> @btime test1($dfg1);
  22.008 ms (4418 allocations: 44.94 MiB)

```

---

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [May 5, 2021, 9:42pm UTC](https://discourse.julialang.org/t/allocations-and-slow-perf-for-transform-on-groupeddataframes/60594/4 "2021-05-05T21:42:53Z")

</div>

Thanks for lot for the quick fix!  
Regarding the remaining ~25X difference in execution time, does such gap falls within expectations? My intuition was that by having the DF sorted by the group, it would have reduced the difference vs the non-grouped DF to a fairly small amount sice each element would be accessed once in a straight sequence.  
However, the performance doesn’t seem to change much compared to where the groups are randomly scattered around:

```julia
julia> @btime test1($dfg1);
  352.164 ms (9003307 allocations: 302.20 MiB)

dfs = sort(df1, [:grp])
dfgs = groupby(dfs, ["grp"])
julia> @btime test1($dfgs);
  296.777 ms (9003406 allocations: 302.20 MiB)

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [May 6, 2021, 8:58am UTC](https://discourse.julialang.org/t/allocations-and-slow-perf-for-transform-on-groupeddataframes/60594/5 "2021-05-06T08:58:17Z")

</div>

> Regarding the remaining ~25X difference in execution time, does such gap falls within expectations?

I would prefer it to be smaller, but I do not see a quick fix for this now (I will have to think).

> the performance doesn’t seem to change much compared to where the groups are randomly scattered around:

This is a good point. In joins we already take advantage of the data being sorted. In grouping we currently do not handle this yet but this is on a to-do list.
