# Plotting difference from mean

**URL:** <https://discourse.julialang.org/t/plotting-difference-from-mean/58496>\
**Category:** Visualization\
**Created:** [April 3, 2021, 12:29pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496 "2021-04-03T12:29:00Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [April 3, 2021, 12:29pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/1 "2021-04-03T12:29:00Z")

</div>

If I have a CSV file with name and value pairs such as  
node01, 127097.3  
node02, 127263.5  
node03, 127132.3  
…

How would I make a plot of each values difference from the mean of the whole column?  
These are stream triad values from a benchmark, if anyone is interested.

I know that can be done in Excel but I would like to use a Pluto notebook or Queryverse  
I just know I have asked a dumb question and the answer is obvious.

(ps those are not real values - probably there are NDAs up the wazoo here)

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [April 3, 2021, 12:53pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/2 "2021-04-03T12:53:08Z")

</div>

Do you have multiple values for each node or the nodes are all different?

_PS: what are NDAs?_

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [April 3, 2021, 12:57pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/3 "2021-04-03T12:57:25Z")

</div>

Each node has a single value. This is the output of a streams memory benchmark, the Triad test  
There are copy, scale and add tests but I will play with them independently.

NDA = Non Disclosure Agreement  
These values are not particularly secret, but I work for a large company and dont want a rap over the knuckles.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [April 3, 2021, 1:03pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/4 "2021-04-03T13:03:52Z")

</div>

Not sure if this is what you need:

```julia
using CSV, DataFrames, Plots

df = DataFrame(CSV.File("difference_from_mean.csv"))
plot(df[:,1],df[:,2] .- mean(df[:,2]), ylabel="Difference from mean", legend=false)

```

![difference_from_mean](https://global.discourse-cdn.com/julialang/original/3X/3/6/36878a8a58d6118470a2f0bdcb72ffb869e23c9b.png)

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [April 3, 2021, 1:12pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/5 "2021-04-03T13:12:17Z")

</div>

Thankyou. Thats what I want. Would be nice to have histogram bars for each node.  
I can work on that though - thanks

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [April 3, 2021, 1:21pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/6 "2021-04-03T13:21:23Z")

</div>

What’s the structure of your data - do you have multiple observations per node?

If so you could have a look at [StatsPlots.jl](https://github.com/JuliaPlots/StatsPlots.jl) which has a bunch of grouped visualisations built in. It might be convenient to just demean the data ahead of any plotting, i.e. do

```julia
df[!, :data_demeaned] = df.data .- mean(df.data)

```

and then throw `df.data_demeaned` into the appropriate StatsPlots recipe

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [April 3, 2021, 1:23pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/7 "2021-04-03T13:23:08Z")

</div>

Thanks both. It is just a single value per node - the idea is to look for outliers, which indicate that node has something wrong with it.

Poor data. Being demeaned in code.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [April 3, 2021, 2:52pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/8 "2021-04-03T14:52:51Z")

</div>

Something like this then:

```julia
julia> using DataFrames, Statistics, Plots

julia> df = DataFrame(data = randn(50));

julia> bar(1:50, df.data .- mean(df.data), label = "Demeaned Data", linewidth = 0.01, xlabel = "Node", ylabel = "Deviation from mean");

julia> hline!([quantile(df.data .- mean(df.data), 0.1), quantile(df.data .- mean(df.data), 0.9)], label = "10th/90th percentile", linewidth = 2)

```

![image](https://global.discourse-cdn.com/julialang/original/3X/9/2/92e36287ca8fa7bbbc9f6ccb7ec1a3c9c1d60278.png)

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [April 3, 2021, 4:27pm UTC](https://discourse.julialang.org/t/plotting-difference-from-mean/58496/9 "2021-04-03T16:27:06Z")

</div>

That is perfect! Thankyou so much.
