# Very best way to concatenate array of arrays, while applying a function

**URL:** <https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [October 3, 2022, 9:59am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151 "2022-10-03T09:59:44Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![BambOoxX](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bambooxx/32/22179_2.png) [@BambOoxX](https://discourse.julialang.org/u/BambOoxX)\
**Post date:** [October 3, 2022, 9:59am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/1 "2022-10-03T09:59:45Z")

</div>

Continuing on this nice [dicussion](https://discourse.julialang.org/t/very-best-way-to-concatenate-an-array-of-arrays/8672) about `reduce` would there be a way to use the same syntax while applying another function, if at al relevant.

Basically, I have some calls that look like `reduce(vcat, fun.(array_of_array))` where I want to apply `fun` to all items in `array_of_array`. Given that I am not interested in the result of a single call of `fun`, would it make a difference to have something like a `reduce(vcat, a -> fun(a), array_of_array)` ?

I am not familiar enough with `reduce` to say if something already exists or not to do that…

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [October 3, 2022, 10:10am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/2 "2022-10-03T10:10:51Z")

</div>

Didn’t look at the other discussion but sounds just like [`mapreduce`](https://docs.julialang.org/en/v1/base/collections/#Base.mapreduce-Tuple%7BAny,%20Any,%20Any%7D) to me.

---

<div class="post-metadata">

**Author:** ![mariok90](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mariok90/32/7076_2.png) [@mariok90](https://discourse.julialang.org/u/mariok90)\
**Post date:** [October 3, 2022, 10:11am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/3 "2022-10-03T10:11:02Z")

</div>

You could use `mapreduce`:

```julia
arr_of_arr = [[rand() for _ in 1:3] for _ in 1:5]

fun(x) = x / 2
mapreduce(fun, vcat, arr_of_arr)

```

edit: I was a little bit too slow

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [October 3, 2022, 10:18am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/4 "2022-10-03T10:18:32Z")

</div>

> [@mariok90](#):
>
> `[[rand() for _ in 1:3] for _ in 1:5]`

```julia
[rand(3) for _ in 1:5]

```

Another issue is if you want the map function to work element-wise on each subarray.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [October 3, 2022, 10:37am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/5 "2022-10-03T10:37:15Z")

</div>

Note that `reduce(vcat, fun.(array))` may well hit a special method for vectors of vectors or matrices, which allocates one array to write into. There is no such method for `mapreduce(fun, vcat, array)`, and thus it will call `vcat` many times, allocating `length(array)-1` arrays in the process. This will probably be much more expensive than the one temporary array from `fun.(array)`.

If `fun(a)` returns a vector, then you can try `vec(stack(fun, array))`, with `using Compat` first.

---

<div class="post-metadata">

**Author:** ![BambOoxX](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bambooxx/32/22179_2.png) [@BambOoxX](https://discourse.julialang.org/u/BambOoxX)\
**Post date:** [October 3, 2022, 10:55am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/6 "2022-10-03T10:55:56Z")

</div>

Damn, I forgot `mapreduce`. I do not have that `map` reflex yet.

> [@mcabbott](#):
>
> Note that `reduce(vcat, fun.(array))` may well hit a special method for vectors of vectors or matrices, which allocates one array to write into. There is no such method for `mapreduce(fun, vcat, array)`, and thus it will call `vcat` many times, allocating `length(array)-1` arrays in the process. This will probably be much more expensive than the one temporary array from `fun.(array)`.

I just benchmarked it quickly and it turns out your are very much right about this. For instance:

```julia
fun = x -> 3*x
arr_of_arr = [[rand() for _ in 1:1000] for _ in 1:100]
 @benchmark reduce(vcat,$fun.($arr_of_arr))
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 136.800 μs … 5.901 ms ┊ GC (min … max): 0.00% … 96.36%
 Time (median): 170.600 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 224.968 μs ± 309.902 μs ┊ GC (mean ± σ): 19.03% ± 13.03%

  ██▄▂▁▂▁ ▁ ▁ ▂
  ███████▇▇█▇▆▅▄▄▃▄▃▄▄▁▃▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▄▅▇▆▆█ █
  137 μs Histogram: log(frequency) by time 2.09 ms <

 Memory estimate: 1.54 MiB, allocs estimate: 104.

julia> @benchmark mapreduce($fun,vcat,$arr_of_arr)
BenchmarkTools.Trial: 1462 samples with 1 evaluation.
 Range (min … max): 2.166 ms … 15.543 ms ┊ GC (min … max): 0.00% … 46.46%
 Time (median): 3.730 ms ┊ GC (median): 35.78%
 Time (mean ± σ): 3.402 ms ± 1.167 ms ┊ GC (mean ± σ): 23.45% ± 17.79%

  ▂▅▆█▅▂▂▂ ▁ ▁▇▇▇▆▄▃▂ ▁ ▁
  ████████████▇▆▇▄▆▆█████████▇█▅▇▇▇▆▆▆▆▆▄▁▅▁▆▁▄▅▁▅▁▁▄▁▄▁▄▁▁▄ █
  2.17 ms Histogram: log(frequency) by time 6.8 ms <

 Memory estimate: 39.30 MiB, allocs estimate: 297.

```

Do you know why `mapreduce` cannot hit the same methods than reduce ?

---

<div class="post-metadata">

**Author:** ![Per](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/per/32/10387_2.png) [@Per](https://discourse.julialang.org/u/Per)\
**Post date:** [October 3, 2022, 11:22am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/7 "2022-10-03T11:22:04Z")

</div>

> [@BambOoxX](#):
>
> Do you know why `mapreduce` cannot hit the same methods than reduce ?

That’s becaues `reduce` can check the sizes beforehand to know how much memory to allocate, but `mapreduce` cannot know the sizes of arrays output by the user function before it is called.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [October 3, 2022, 12:19pm UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/8 "2022-10-03T12:19:11Z")

</div>

This is a strong statement and I only half mean it, but… I kind of wish I had never learned about `mapreduce`. I reach for it quite often now but I usually find out that it is doing things in a pretty inefficient way (because of the shortcomings listed here).

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [October 3, 2022, 1:36pm UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/9 "2022-10-03T13:36:10Z")

</div>

using compat

```julia
julia> @benchmark reduce(vcat,$fun.($arr_of_arr))
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 154.500 μs … 4.178 ms ┊ GC (min … max): 0.00% … 83.20%        
 Time (median): 215.400 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 284.858 μs ± 259.627 μs ┊ GC (mean ± σ): 14.05% ± 14.42%        

  ▄▇█▇▆▅▄▄▃▃▂▁ ▁ ▂
  ███████████████▇▇▆▆▆▆▆▆▅▄▄▄▄▄▃▄▄▄▄▃▅▃▂▅▄▅▄▃▄▄▅▃▄▆██▇▇▇▆▆▆▇▆▆▆ █
  154 μs Histogram: log(frequency) by time 1.43 ms <

 Memory estimate: 1.54 MiB, allocs estimate: 104.

julia> @benchmark mapreduce($fun,vcat,$arr_of_arr)
BenchmarkTools.Trial: 922 samples with 1 evaluation.
 Range (min … max): 3.434 ms … 15.972 ms ┊ GC (min … max): 0.00% … 44.82%
 Time (median): 4.903 ms ┊ GC (median): 15.35%
 Time (mean ± σ): 5.394 ms ± 1.416 ms ┊ GC (mean ± σ): 14.93% ± 6.12%

          █▅▄▁▂
  ▂▃▄▂▃▃▃▅█████▇▆▄▅▄▃▃▃▃▃▃▂▃▃▂▃▂▃▃▂▃▃▂▃▃▄▃▃▂▂▃▁▂▂▂▂▂▁▁▁▁▁▁▁▂ ▃
  3.43 ms Histogram: frequency by time 10.8 ms <

 Memory estimate: 39.30 MiB, allocs estimate: 297.

julia> @benchmark vec(stack($fun, $arr_of_arr))
BenchmarkTools.Trial: 10000 samples with 1 evaluation.
 Range (min … max): 116.700 μs … 3.675 ms ┊ GC (min … max): 0.00% … 85.37%        
 Time (median): 194.900 μs ┊ GC (median): 0.00%
 Time (mean ± σ): 252.690 μs ± 194.317 μs ┊ GC (mean ± σ): 11.77% ± 14.80%        

  ▁ ▂██▆▅▅▄▄▄▃▂▁▁ ▂
  █████████████████▇▅▇▆▆▆▆▅▅▅▅▄▅▄▅▅▆▆▆▆▆▇▇▇▇▆▆▇▆▇▆▆▅▆▆▅▅▅▅▅▅▃▃▅ █
  117 μs Histogram: log(frequency) by time 1.15 ms <

 Memory estimate: 1.54 MiB, allocs estimate: 104.

```

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 3, 2022, 3:25pm UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/10 "2022-10-03T15:25:16Z")

</div>

> [@mcabbott](#):
>
> If `fun(a)` returns a vector, then you can try `vec(stack(fun, array))`, with `using Compat` first.

By the way, this will only work if `fun` always returns a vector of the same length.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [October 4, 2022, 12:20am UTC](https://discourse.julialang.org/t/very-best-way-to-concatenate-array-of-arrays-while-applying-a-function/88151/11 "2022-10-04T00:20:58Z")

</div>

> [@Per](#):
>
> `reduce` can check the sizes beforehand to know how much memory to allocate, but `mapreduce` cannot know the sizes of arrays output by the user function before it is called.

This is true, although it could still do better than pairwise `vcat`, for instance by guessing that all are the same size as the first. Or by repeated doubling like `append!`.

> [@Mason](#):
>
> this will only work if `fun` always returns a vector of the same length

Indeed, this is one of the things which makes `stack` easier.

> [@BambOoxX](#):
>
> Do you know why `mapreduce` cannot hit the same methods than reduce ?

Mostly nobody has got around to it. [PR 31636](https://github.com/JuliaLang/julia/issues/31636) had a go. Another option would simply be to send this to `reduce(vcat, map(f, xs))`, as this will usually be better. (This is also what happens for `mapreduce(f, op, xs, ys)` at present.)

The other option would be to write a `flatten` companion for `stack`, since these clever `reduce` overloads are hard to discover & quite fragile – it’s easy to step off the fast path, as with [init keyword](https://discourse.julialang.org/t/how-can-i-do-list-comprehension-for-vcat/84178/11) for example.
