# Mapreduce, pass extra arguments to reduce/vcat of DataFrames

**URL:** <https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347>\
**Category:** General Usage\
**Tags:** dataframes\
**Created:** [March 3, 2022, 11:02am UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347 "2022-03-03T11:02:01Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![baptnz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baptnz/32/24519_2.png) [@baptnz](https://discourse.julialang.org/u/baptnz)\
**Post date:** [March 3, 2022, 11:02am UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/1 "2022-03-03T11:02:01Z")

</div>

I’m using a map+reduce combination to apply a function returning a DataFrame and combine all the results with vcat,

```julia
f(p) = DataFrame(x = collect(1:10), y = p[1]*rand(10), z = p[2]*rand(10))  
tmp = map(f, ([1,2], [3,4], [5,6]))
all = reduce(vcat, tmp, source="id")

```

the allocation of temporary results in `tmp` isn’t ideal, as `?mapreduce` suggests, but I’ve been unable to find how to pass the equivalent `source = "id"` argument in `mapreduce`, to keep track of each block’s origin. Any idea?

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [March 3, 2022, 2:13pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/2 "2022-03-03T14:13:18Z")

</div>

uhh, the `Base.reduce` has no parameter `source` except some package you import change this?

---

<div class="post-metadata">

**Author:** ![baptnz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baptnz/32/24519_2.png) [@baptnz](https://discourse.julialang.org/u/baptnz)\
**Post date:** [March 3, 2022, 2:27pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/3 "2022-03-03T14:27:53Z")

</div>

It’s the method defined for DataFrames,

```julia
methods(reduce, DataFrames)

```

or perhaps more specifically, the `source` argument originates from `methods(vcat, DataFrames)` and is passed through reduce in this case.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 3, 2022, 3:04pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/4 "2022-03-03T15:04:04Z")

</div>

Performance isn’t ideal might be the use of a `Tuple` rather than a `Vector`. Large `Tuple`s don’t perform well… but that also might not be the issue.

I guess your problem is that there is a secialized method for `reduce` but not for `mapreduce`. One way to improve performance might be with a `Generator`.

You can also use an anonymous. They work just like in R

```julia
reduce((a, b) -> ..., dfs)

```

but it looks like this doesn’t work

```julia
julia> ps = [[1, 2], [3, 4], [5, 6]];

julia> mapreduce(f, (a, b) -> vcat(a, b, source = "id"), ps);
ERROR: ArgumentError: column(s) id are missing from argument(s) 2

```

Given this, i think maybe we should add a method for `mapreduce` in addition to `reduce`.

Another solution, along the lines of what I discussed yesterday, is to do more inside an anonymous function

```julia
julia> ps = [[1, 2], [3, 4], [5, 6]];

julia> mapreduce(vcat, eachindex(ps)) do i
           df = f(ps[i])
           df.source .= i
           df
       end

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [March 3, 2022, 3:04pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/5 "2022-03-03T15:04:38Z")

</div>

you can change your `f` function to create `:id` column and use `mapreduce`.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 3, 2022, 3:21pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/6 "2022-03-03T15:21:55Z")

</div>

something like this …

```julia

f(p) = DataFrame(x = collect(1:10), y = p[1]*rand(10), z = p[2]*rand(10), id=string(p[1])) 

mapreduce(f, (x,y)->vcat(x,y), ([1,2], [3,4],[5,6]))

```

ops … I arrived late

or something like that, if you really want to use library functions 😀

```julia

mapreduce(f, (x,y)->vcat(x,y, source=string("id",nrow(x)), cols=:union), ([1,2], [3,4],[5,6]))

```

---

<div class="post-metadata">

**Author:** ![baptnz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baptnz/32/24519_2.png) [@baptnz](https://discourse.julialang.org/u/baptnz)\
**Post date:** [March 3, 2022, 4:22pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/9 "2022-03-03T16:22:15Z")

</div>

> [@pdeffebach](#):
>
> You can also use an anonymous. They work just like in R
> 
> ```julia
> reduce((a, b) -> ..., dfs)
> 
> ```
> 
> but it looks like this doesn’t work

I’m also a bit puzzled by this; I thought I’d managed to get it to work with

```julia
mapreduce(f, (x,y) -> vcat(x,y, source="id", cols=:intersect), ps)

```

but [on my longer example](https://discourse.julialang.org/t/map-over-combinations-of-parameters-and-grouping-results-as-dataframe/77306/10) I get some unexpected `missing` values that I don’t understand at all.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 3, 2022, 4:52pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/10 "2022-03-03T16:52:51Z")

</div>

In this context, `vcat` _just_ knows about the two arguments it is given. Since each `vcat` takes two arguments in this example, `source` can only take the values of `1` or `2`. `cols = :intersect` tells `vcat` to just keep columns that are in _both_ data frames… I’m surprised that works tbh. I would think it would throw an error.

Also, maybe `source` isn’t the right move here. Why not just add the parameters directly to the data frame? In julia you can have vectors of vectors. No need to worry about mapping ids to parameters when you can just store the parameters.

---

<div class="post-metadata">

**Author:** ![baptnz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baptnz/32/24519_2.png) [@baptnz](https://discourse.julialang.org/u/baptnz)\
**Post date:** [March 3, 2022, 5:12pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/11 "2022-03-03T17:12:58Z")

</div>

> [@pdeffebach](#):
>
> I’m surprised that works tbh. I would think it would throw an error.

Yeah, I’m also a bit confused by my trial and error (:union fails, for example). Clearly this won’t work. I think I finally understand the relation between reduce and vcat for DataFrames, which I had mistakenly read the other way around in the source code. It’s the `vcat()` method that calls `reduce()` and passes it the `source` argument – because reduce has all the objects it can create the IDs, as you say, whereas `vcat()` on its own only gets passed x and y. It’s confusing because `vcat()` is defined as

```julia
Base.vcat(dfs::AbstractDataFrame...;
          cols::Union{Symbol, AbstractVector{Symbol},
                      AbstractVector{<:AbstractString}}=:setequal,
          source::Union{Nothing, SymbolOrString,
                           Pair{<:SymbolOrString, <:AbstractVector}}=nothing) =
    reduce(vcat, dfs; cols=cols, source=source)

```

suggesting (to my naive eyes) that `vcat` itself uses `source`, when in fact it just feeds it to the custom reduce.

---

<div class="post-metadata">

**Author:** ![baptnz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baptnz/32/24519_2.png) [@baptnz](https://discourse.julialang.org/u/baptnz)\
**Post date:** [March 3, 2022, 5:16pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/12 "2022-03-03T17:16:37Z")

</div>

> [@pdeffebach](#):
>
> Also, maybe `source` isn’t the right move here. Why not just add the parameters directly to the data frame? In julia you can have vectors of vectors. No need to worry about mapping ids to parameters when you can just store the parameters.

I liked the conciseness of it (no need to create an anonymous function to add the “id” (or all the parameters for that call directly, alternatively)), but that’s purely an aesthetic preference, which I wouldn’t even think about if I hadn’t been using `purrr::pmap_df()` for years (`plyr::mdply()` before that).

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 3, 2022, 5:30pm UTC](https://discourse.julialang.org/t/mapreduce-pass-extra-arguments-to-reduce-vcat-of-dataframes/77347/13 "2022-03-03T17:30:35Z")

</div>

in this way it is perhaps an acceptable compromise

```julia

f(p,id) = DataFrame(x = collect(1:10), y = p[1]*rand(10), z = p[2]*rand(10), id=id) 

mapreduce(t->f(t[2],t[1]), (x,y)->vcat(x,y), enumerate(([1,2], [3,4],[5,6],[7,8])))

```

this way you don’t have to change your f (p)

```julia

F(id,p)=hcat(f(p),DataFrame(id=fill(id,nrow(f(p)))))

mapreduce(t->F(t...), (x,y)->vcat(x,y), enumerate(([1,2], [3,4],[5,6])))

```
