# How to calcule the mean of values considering their tuples of another value:

**URL:** <https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974>\
**Category:** General Usage\
**Tags:** dataframes\
**Created:** [November 9, 2022, 11:40am UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974 "2022-11-09T11:40:20Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [November 9, 2022, 11:40am UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974/1 "2022-11-09T11:40:20Z")

</div>

Hi all, I have an issue that I can’t figure out how to solve it:

I have a df that stores two columns that look like this:

```julia
df.scores1 = [[1,2,2,3,5,6,1,2,9,2,1,6,4,2]]

df.normalized_len = [[0,0,0.1,0.1,0.2,0.3,0.4,0.5,0.5,0.6,0.7,0,8,0.9,1]]

```

And this goes on for many rows, all of them having the same length.

What I am trying to do is get the mean values of `df.scores1` that have the same value of `df.normalized_len`, so the result should look like this:

`df.mean_val_norm = [[1.5,2.5,5,6,1,5.5,2,1,6,4,2]] `

Any help is welcome!

Thanks a lot,  
Juan

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [November 9, 2022, 11:44am UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974/2 "2022-11-09T11:44:55Z")

</div>

I don’t understand the question - could you explain how the values in your desired output `mean_val_norm` are derived?

---

<div class="post-metadata">

**Author:** ![JorizovdZ](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jorizovdz/32/43064_2.png) [@JorizovdZ](https://discourse.julialang.org/u/JorizovdZ)\
**Post date:** [November 9, 2022, 12:07pm UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974/4 "2022-11-09T12:07:43Z")

</div>

You can also use the DataFrames documentation to form a DataFrame that is grouped with these values:

```julia
x = [1,2,2,3,5,6,1,2,9,2,1,6,4,2]
y = [0,0,0.1,0.1,0.2,0.3,0.4,0.5,0.5,0.6,0.7,0.8,0.9,1]
df = DataFrame(scores1=x, normalized_len=y)
gb=groupby(df, :normalized_len)
println(combine(gb, :scores1 => mean))

```

Output:

```julia
11×2 DataFrame
 Row │ normalized_len scores1_mean 
     │ Float64 Float64      
─────┼──────────────────────────────
   1 │ 0.0 1.5
   2 │ 0.1 2.5
   3 │ 0.2 5.0
   4 │ 0.3 6.0
   5 │ 0.4 1.0
   6 │ 0.5 5.5
   7 │ 0.6 2.0
   8 │ 0.7 1.0
   9 │ 0.8 6.0
  10 │ 0.9 4.0

```

---

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [November 9, 2022, 12:54pm UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974/5 "2022-11-09T12:54:05Z")

</div>

Sorry, I just realized my post was kinda vague.

I want to get the mean values of the items in `df.score1` that share the same value in `df.normalized_len`. So, for all the values that a `normalized_len` of 0 (the first two ones) would get summed and divided by two (because there are only two values that have that normalized len), and so on. Is it clearer now?

---

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [November 14, 2022, 11:16am UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974/6 "2022-11-14T11:16:54Z")

</div>

Hi, someone posted the correct answer but then deleted it! Just in case anyone has the same issue, this was it:

```julia
m = [mean(x[findall(==(u), y)]) for u in unique(y)]

```

In the end I used `@rtransform` like this and it worked great:

```julia
df_t = @rtransform df_t1 :mean_pos_rep = begin ## 
           [mean(:sum_total[findall(==(u), :norm_length)]) for u in unique(:norm_length)]##
end;

```

Cheers,  
Juan

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [November 14, 2022, 11:59am UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974/7 "2022-11-14T11:59:12Z")

</div>

But isn’t the correct answer the one proposed by @JorizovdZ?

This isn’t bad either, but maybe it’s not the first that comes to mind.

---

<div class="post-metadata">

**Author:** ![Juan\_Mac\_Donagh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan_mac_donagh/32/31798_2.png) [@Juan\_Mac\_Donagh](https://discourse.julialang.org/u/Juan_Mac_Donagh)\
**Post date:** [November 14, 2022, 12:02pm UTC](https://discourse.julialang.org/t/how-to-calcule-the-mean-of-values-considering-their-tuples-of-another-value/89974/8 "2022-11-14T12:02:18Z")

</div>

It works too, but the result that I intended was the one that got deleted. But, as the answer is still posted, I’ll mark it as the correcto solution.
