# Functional style table processing

**URL:** <https://discourse.julialang.org/t/functional-style-table-processing/15336>\
**Category:** General Usage\
**Tags:** question\
**Created:** [September 22, 2018, 11:52am UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336 "2018-09-22T11:52:58Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrey.Borzunov](https://avatars.discourse-cdn.com/v4/letter/a/9f8e36/32.png) [@Andrey.Borzunov](https://discourse.julialang.org/u/Andrey.Borzunov)\
**Post date:** [September 22, 2018, 11:52am UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/1 "2018-09-22T11:52:58Z")

</div>

I have an Array of Tuples:

```julia
A = [
(1,10);
(1,20);
(1,30);
(2,15);
(2,25);
(2,35);
(3,17);
(3,27);
(3,37);
]

```

I want to get new Array of Tuples `(x,y)` where `x` is unique, and `y = mean(z)`, where `z` is all elements from previous table with corresponding `x`:

```julia
[
(1, 20);
(2, 25);
(3, 27)
]

```

I’m interested in beautiful functional style solution(not ugly nested loops, although i know, that it will be as efficient as one-line functional styled).  
Array of tuples can be changed at 2d Array.

Thx!

---

<div class="post-metadata">

**Author:** ![Andrey.Borzunov](https://avatars.discourse-cdn.com/v4/letter/a/9f8e36/32.png) [@Andrey.Borzunov](https://discourse.julialang.org/u/Andrey.Borzunov)\
**Post date:** [September 22, 2018, 12:42pm UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/2 "2018-09-22T12:42:03Z")

</div>

I managed to write only next solution, but it looks too long

```julia
# convert to 2d Array
B = zeros(length(A),2)
for x in LinearIndices(A)
    B[x,:] = [A[x][1], A[x][2]]
end

# Create unique array composed of 1st column values
C = unique(B[:,1])
out = zeros(length(C),2)

for i in LinearIndices(B[:,1])
    out[i,:] = [C[i], Statistics.mean( findall( C[i] .== B[:,2]) ) ]
end

```

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [September 22, 2018, 12:57pm UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/3 "2018-09-22T12:57:25Z")

</div>

Sounds like group by operation:

```julia
A = [
(1,10);
(1,20);
(1,30);
(2,15);
(2,25);
(2,35);
(3,17);
(3,27);
(3,37);
]

using DataFrames
using Statistics

df = DataFrame(Dict(:x => [a[1] for a in A], :y => [a[2] for a in A]))
by(df -> DataFrame(m = mean(df[:x])), df, :x)

```

Of course you can do the same thing manually using 2 dicts - for count and sum of elements, using loops or array/dict comprehension. E.g. using loops (which are clearer in this case IMO):

```julia
# make a dict with first element of each tuple as key and a list of corresponding second elements as value
d = Dict()
for (k, v) in A
    if !haskey(d, k)
        d[k] = []
    end
    push!(d[k], v)
end

# d now contains: 
# Dict{Any,Any} with 3 entries:
# 2 => Any[15, 25, 35]
# 3 => Any[17, 27, 37]
# 1 => Any[10, 20, 30]

# calculate means
dm = Dict(k => Statistics.mean(vs) for (k, vs) in d)

# which gives:
# Dict{Int64,Float64} with 3 entries:
# 2 => 25.0
# 3 => 27.0
# 1 => 20.0

```

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [September 22, 2018, 2:25pm UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/4 "2018-09-22T14:25:08Z")

</div>

In Julia 0.6 with IndexedTables:

```julia
using IndexedTables
groupby(mean, table(A), 1, select = 2)

```

Unfortunately is not ported to 1.0 yet, so this only works if you’re still on Julia 0.6

---

<div class="post-metadata">

**Author:** ![Seif\_Shebl](https://avatars.discourse-cdn.com/v4/letter/s/eada6e/32.png) [@Seif\_Shebl](https://discourse.julialang.org/u/Seif_Shebl)\
**Post date:** [September 22, 2018, 6:27pm UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/5 "2018-09-22T18:27:45Z")

</div>

This is quite head-spinning, but I hate reading the docs of a zillion packages just for doing a small task.

```julia
julia> A = [
       (1,10);
       (1,20);
       (1,30);
       (2,15);
       (2,25);
       (2,35);
       (3,17);
       (3,27);
       (3,37);
       ];

julia> B = unique( A[i][1] for i in 1:length(A) );

julia> C = [mean( i[2] for i in filter(M -> M[1] == j, A) ) for j in B ];

julia> [(B[i],C[i]) for i in 1:length(B)]
3-element Array{Tuple{Int64,Float64},1}:
 (1, 20.0)
 (2, 25.0)
 (3, 27.0)

```

---

<div class="post-metadata">

**Author:** ![Andrey.Borzunov](https://avatars.discourse-cdn.com/v4/letter/a/9f8e36/32.png) [@Andrey.Borzunov](https://discourse.julialang.org/u/Andrey.Borzunov)\
**Post date:** [September 22, 2018, 6:41pm UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/6 "2018-09-22T18:41:44Z")

</div>

Wow, so far, this is the closest solution.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [September 22, 2018, 7:36pm UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/7 "2018-09-22T19:36:02Z")

</div>

> [@Seif\_Shebl](#):
>
> This is quite head-spinning, but I hate reading the docs of a zillion packages just for doing a small task.

Note that your solution is O(N) + O(K \* N) + O(K) = O(K \* N) where N is length of array A, K - number of unique keys (i.e. length of B). At the same time solution with dict is only O(2N) = O(N), i.e. you see each element only twice instead of K times. Even more efficient algorithm would be to calculate averages on the fly, which would allow to make only 1 pass through each element in A.

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [September 22, 2018, 8:03pm UTC](https://discourse.julialang.org/t/functional-style-table-processing/15336/8 "2018-09-22T20:03:09Z")

</div>

There is a beautiful integration between IndexedTables and OnlineStats to compute statistics “on the fly” (with O(1) memory and doing only one pass). Here for example it’d be:

```julia
using IndexedTables, OnlineStats
groupreduce(Mean(), table(A), 1, select = 2)

```

It’s very useful for distributed datasets and OnlineStats has some sophisticated statistics (not only `Mean()` and co. but also linear models for example).
