# Fast way to calculate statistics of an array based on values (labels) of another array

**URL:** <https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054>\
**Category:** Performance\
**Tags:** statistics, dataframes, splitapplycombine\
**Created:** [October 19, 2021, 5:19pm UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054 "2021-10-19T17:19:24Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bjorn\_De](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bjorn_de/32/13174_2.png) [@Bjorn\_De](https://discourse.julialang.org/u/Bjorn_De)\
**Post date:** [October 19, 2021, 5:19pm UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/1 "2021-10-19T17:19:24Z")

</div>

Hey,  
I have two arrays of the same size (10000x10000). One array include labels as result of an segmentation. What I want is to compute statistics for each of these segments based on the 2nd array. My approach is very slow.

```julia
function seg_stat(array, labels)
    d = Dict()
    label_ids = unique(labels)
    for label in label_ids
        indices = findall(labels .== label)
        d[label] = median(@view array[indices])
    end
    return d
end

```

Any ideas how to make this more faster?  
Thank you!

---

<div class="post-metadata">

**Author:** ![sijo](https://avatars.discourse-cdn.com/v4/letter/s/da6949/32.png) [@sijo](https://discourse.julialang.org/u/sijo)\
**Post date:** [October 19, 2021, 6:34pm UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/2 "2021-10-19T18:34:28Z")

</div>

You can get a nice speed-up by declaring a precise type for the dictionary with

```julia
d = Dict{eltype(labels), eltype(array)}()

```

In my small test this divides the time by two.

I get another 30% savings by using DataFrames instead:

```julia
using DataFrames

a = rand(1000, 1000)
b = rand(["a", "b", "c", "d"], size(a))
df = DataFrame(array=vec(a), labels=vec(b), copycols=false)
combine(groupby(df, :labels), :array => median)

```

Not sure if you will get the same savings with your real dataset though…

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 19, 2021, 9:00pm UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/3 "2021-10-19T21:00:35Z")

</div>

@sijo, this is a nice example of the power of dataframes but, how to easily use the results?

For example, what is a lighter way to get the scalar results for each label, rather than doing:

```julia
dg = combine(groupby(df, :labels), :array => median)

dg[dg.labels .== "b", :].array_median[1]

```

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 19, 2021, 9:25pm UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/4 "2021-10-19T21:25:48Z")

</div>

How many different labels there are?

---

<div class="post-metadata">

**Author:** ![goerch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/goerch/32/29122_2.png) [@goerch](https://discourse.julialang.org/u/goerch)\
**Post date:** [October 19, 2021, 10:33pm UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/5 "2021-10-19T22:33:22Z")

</div>

Now we have to reconstruct test cases from the description;) BTW: found the question also interesting, but there is not enough input for me.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [October 20, 2021, 5:32am UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/6 "2021-10-20T05:32:41Z")

</div>

A straightforward and performant solution:

```julia
using Statistics
using SplitApplyCombine
using StructArrays

a = rand(1000, 1000)
b = rand(["a", "b", "c", "d"], size(a))

ab = StructArray(; a, b)

# dictionary of per-label medians:
meds = map(groupview(x -> x.b, ab)) do gr
	median(gr.a)
end
# retrieve median for a single label:
meds["c"]

# same code with a nicer piping syntax:
using DataPipes

meds = @p begin
	StructArray(; a, b)
	groupview(_.b)
	map(median(_.a))
end

```

---

<div class="post-metadata">

**Author:** ![Bjorn\_De](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bjorn_de/32/13174_2.png) [@Bjorn\_De](https://discourse.julialang.org/u/Bjorn_De)\
**Post date:** [October 20, 2021, 6:46am UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/7 "2021-10-20T06:46:15Z")

</div>

Thanks a lot @sijo and @aplavin for your suggestions. In this case the segment array consist of ~500000 unique labels. The DataFrame solution performs a little bit faster but both solutions are much faster than my loop. Thank you very much for your help!

```julia
julia> @benchmark combine(groupby(df, :labels), :array => StatsBase.mode)
BenchmarkTools.Trial: 2 samples with 1 evaluation.
 Range (min … max): 3.667 s … 3.797 s ┊ GC (min … max): 0.04% … 8.58%
 Time (median): 3.732 s ┊ GC (median): 4.39%
 Time (mean ± σ): 3.732 s ± 92.293 ms ┊ GC (mean ± σ): 4.39% ± 6.04%

  █ █  
  █▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁█ ▁
  3.67 s Histogram: frequency by time 3.8 s <

 Memory estimate: 2.09 GiB, allocs estimate: 2281604.

julia> @benchmark meds = map(groupview(x -> x.b, ab)) do gr
               median(gr.a)
       end
BenchmarkTools.Trial: 1 sample with 1 evaluation.
 Single result which took 5.612 s (18.35% GC) to evaluate,
 with a memory estimate of 4.99 GiB, over 3688493 allocations.

```

---

<div class="post-metadata">

**Author:** ![sijo](https://avatars.discourse-cdn.com/v4/letter/s/da6949/32.png) [@sijo](https://discourse.julialang.org/u/sijo)\
**Post date:** [October 20, 2021, 7:17am UTC](https://discourse.julialang.org/t/fast-way-to-calculate-statistics-of-an-array-based-on-values-labels-of-another-array/70054/8 "2021-10-20T07:17:42Z")

</div>

> [@rafael.guerra](#):
>
> what is a lighter way to get the scalar results for each label, rather than doing:
> 
> ```julia
> dg = combine(groupby(df, :labels), :array => median)
> 
> dg[dg.labels .== "b", :].array_median[1]
> 
> ```

For a small number of labels you could do something like

```julia
medians = only(permutedims(dg, :labels))

julia> medians.b # or medians[:b]
0.5011299410448031

```

or creating a named tuple manually:

```julia
nt = NamedTuple(Symbol.(dg.labels) .=> dg.array_median)

julia> nt.b
0.5011299410448031

```

In case of many labels it’s probably better to use a Dict:

```julia
d = Dict(dg.labels .=> dg.array_median)

julia> d["b"]
0.5011299410448031

# Or maybe an iterator is better in case of very many labels:
d = Dict(l => m for (l, m) in eachrow(dg))

```

Another option is to keep the groups, so we can use the group keys:

```julia
g = combine(groupby(df, :labels), :array => median, ungroup=false)

julia> g[("b",)][1, :array_median]
0.5011299410448031

```
