# How Transducers.jl process table?

**URL:** <https://discourse.julialang.org/t/how-transducers-jl-process-table/80737>\
**Category:** General Usage\
**Tags:** question, transducers\
**Created:** [May 9, 2022, 7:06am UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737 "2022-05-09T07:06:09Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![zsz00](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zsz00/32/6314_2.png) [@zsz00](https://discourse.julialang.org/u/zsz00)\
**Post date:** [May 9, 2022, 7:06am UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/1 "2022-05-09T07:06:09Z")

</div>

How Transducers.jl process table?  
get each column mean ?

```julia
using Transducers, Table, Statistics
N = 1_000_000
a = [2randn(N÷2) .+ 6; randn(N÷2)]
b = [3randn(N÷2); 2randn(N÷2)]
c = randn(N)
d = c .+ 0.6randn(N)
table = (; a, b, c, d); # NTuple{4, Vector{Float64}}}

# table_df = DataFrame(table)
row_data = Tables.rows(table)

aa = Transducers.foldl(right, row_data |> Map(mean) |> collect)    

```

aa is a number such as 0.4  
I want to get like this:

```julia
a b c d
0.3, 0.4, 0.5 0.6

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [May 9, 2022, 7:41am UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/2 "2022-05-09T07:41:20Z")

</div>

I don’t know the functions of the Transducers package, but maybe to get what you are looking for, you should build the table as columns and stop at Map without letting foldl intervene

```julia
col_data = Tables.columns(table)

collect(Map(mean), col_data)

```

Indeed, probably not even the intervention of Tables is needed.

```julia
table = (; a, b, c, d); 
(; zip(keys(table),collect(Map(mean), table))...)

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [May 9, 2022, 8:10am UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/3 "2022-05-09T08:10:24Z")

</div>

Probably did not understand the question, but fwiw, doing simply:

```julia
map(mean, table)

# results in:
(a = 2.9984519617857526, b = 0.0008151385421462024, c = -0.00045218977614473727, d = -3.069694507042777e-5)

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [May 9, 2022, 12:59pm UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/4 "2022-05-09T12:59:59Z")

</div>

I understood that the request was related to how Map processes tables (in the form of namedtuples) …

---

<div class="post-metadata">

**Author:** ![zsz00](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zsz00/32/6314_2.png) [@zsz00](https://discourse.julialang.org/u/zsz00)\
**Post date:** [May 9, 2022, 1:39pm UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/5 "2022-05-09T13:39:50Z")

</div>

```julia
row_data = Tables.rows(table)
col_data = Tables.columns(table)

Transducers.foldl(right, col_data |> Map(mean) |> collect) # 0.002434656840427542
Transducers.collect(Map(mean), col_data) # is ok 
[2.998240292251592, 3.2144679602343106e-5, 0.0010958016920552937, 0.002434656840427542]

julia> map(mean, table) # is ok
| a | b | c | d |
| ------------------- | ----------------------- | ----------------------- | ---------------------- |
| 2.998240292251592 | 3.2144679602343106e-5 | 0.0010958016920552937 | 0.002434656840427542 |

```

I want to use Transducers.foldl do this,  
why foldl() don’t same to collect(Map(mean), col\_data) ??

My real intention is:  
Suppose my table is large and has many rows, I want to incrementally calculate the mean of each column .  
aa = Transducers.foldl(right, row\_data |\> Map(mean) |\> collect)

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [May 9, 2022, 1:59pm UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/6 "2022-05-09T13:59:09Z")

</div>

what you’re doing (calculating mean of each column) is completely columnar, you don’t want to iterater over rolls for both clarity and performance reasons. Don’t use Transducers.jl when you can simply do it in columnar fashion.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [May 9, 2022, 5:26pm UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/7 "2022-05-09T17:26:33Z")

</div>

Using foldl () you just have the average of the last column left.  
Maybe, if you really want to do this, accumulate () better suits your needs

```julia
aa = Transducers.foldl(right, col_data |> Map(mean) )    
aa = Transducers.accumulate(right, col_data |> Map(mean) )    

```

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [May 11, 2022, 10:19am UTC](https://discourse.julialang.org/t/how-transducers-jl-process-table/80737/8 "2022-05-11T10:19:03Z")

</div>

`xs |> Map(f) |> collect` and `collect(Map(f), xs)` are both equivalent to

```julia
ys = []
for x in xs
    push!(ys, f(x))
end
ys

```

So, `row_data |> Map(mean) |> collect` is computing

```julia
v = [
    mean((a[1], b[1], c[1], d[1])),
    mean((a[2], b[2], c[2], d[2])),
    ...
    mean((a[end], b[end], c[end], d[end])),
]

```

Then, since `foldl(right, v)` is equivalent to `v[end]`, you obtain a number `mean((a[end], b[end], c[end], d[end])),` from `foldl(right, row_data |> Map(mean) |> collect)`.

As others said, I think using “columnar” functions is the best way to do this. _That said,_ if you really want to do this in row-wise fashion (e.g., input does not fit in the memory), you can use [GitHub - JuliaFolds/DataTools.jl](https://github.com/JuliaFolds/DataTools.jl)

```julia
julia> using Transducers

julia> using DataTools: oncol, averaging

julia> foldxl(oncol(a = averaging, b = averaging, c = averaging, d = averaging), Tables.rowtable(table))
(a = 2.999957367033291, b = 0.00019357701607686137, c = 0.0009838445182253483, d = -0.00026117276633180807)

```
