# Peak finding from a distribution

**URL:** <https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485>\
**Category:** General Usage\
**Tags:** statistics, distributions, data\_science\
**Created:** [October 27, 2021, 2:48pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485 "2021-10-27T14:48:00Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![newtothis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/newtothis/32/32227_2.png) [@newtothis](https://discourse.julialang.org/u/newtothis)\
**Post date:** [October 27, 2021, 2:48pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/1 "2021-10-27T14:48:00Z")

</div>

I have a distribution as a vector in Julia. I can plot a histogram and/or a kde of for this data, but I want to find the peak values. The naive way in which I am doing it is to plot the histogram and then just retrieve the mode of the data in the vector used to plot the histogram. However, this output is sometimes not close to the peak I can see from the histogram. What do I do?

---

<div class="post-metadata">

**Author:** ![halleysfifthinc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/halleysfifthinc/32/206280_2.png) [@halleysfifthinc](https://discourse.julialang.org/u/halleysfifthinc)\
**Post date:** [October 27, 2021, 3:46pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/2 "2021-10-27T15:46:10Z")

</div>

If there is only a single peak of interest for the distribution vector, then `maximum`/`findmax` (in Base) will find it for you. For multiple peaks (local maxima), the [`Peaks.jl`](https://github.com/halleysfifthinc/Peaks.jl) package has relevant functions (e.g. [`findmaxima`](https://halleysfifthinc.github.io/Peaks.jl/stable/#Peaks.findmaxima)).

---

<div class="post-metadata">

**Author:** ![gustaphe](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gustaphe/32/18174_2.png) [@gustaphe](https://discourse.julialang.org/u/gustaphe)\
**Post date:** [October 27, 2021, 4:07pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/3 "2021-10-27T16:07:13Z")

</div>

Is it a vector of samples, or a probability density function/estimate?

If the former, you are looking for “Distribution Fitting” ([Distribution Fitting · Distributions.jl](https://juliastats.org/Distributions.jl/stable/fit/)). If the latter possibly [GitHub - JuliaNLSolvers/LsqFit.jl: Simple curve fitting in Julia](https://github.com/JuliaNLSolvers/LsqFit.jl)

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [October 27, 2021, 5:24pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/4 "2021-10-27T17:24:26Z")

</div>

If you can fit a distribution using Distributions, then compute the mode.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 27, 2021, 7:57pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/5 "2021-10-27T19:57:33Z")

</div>

> [@newtothis](#):
>
> The naive way in which I am doing it is to plot the histogram and then just retrieve the mode of the data in the vector used to plot the histogram. However, this output is sometimes not close to the peak I can see from the histogram. What do I do?

There’s a bug in your code if when you choose the middle of the histogram bin that has the highest count, it isn’t at the location that has the highest count on your histogram plot. Of course sometimes histograms are noisy, so the location with the highest count isn’t near the “peak” of the “smoothed histogram” you have in your brain. If that’s the issue, try using a kernel density estimate, and choose the sample with the highest estimated density.

```julia
using KernelDensity, Distributions,Random,StatsPlots

Random.seed!(1)
data = rand(Normal(3.0,3.0),200)

den = kde(data)
mval = findmax(den.density)
plot(den,xlim=(2,4))
println("Peak density occurs at $(den.x[mval[2]])")

```

---

<div class="post-metadata">

**Author:** ![newtothis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/newtothis/32/32227_2.png) [@newtothis](https://discourse.julialang.org/u/newtothis)\
**Post date:** [October 28, 2021, 1:12pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/7 "2021-10-28T13:12:55Z")

</div>

It’s a vector of samples. Essentially, I run a function that counts how many steps some process takes. I repeat the process over many realisations to find some sort of a sample distribution. I want to find what the most likely number of steps is, for this process to happen.

I’m sorry if this is badly phrased or if it’s a known problem. I’m quite new to both stats and julia.

---

<div class="post-metadata">

**Author:** ![newtothis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/newtothis/32/32227_2.png) [@newtothis](https://discourse.julialang.org/u/newtothis)\
**Post date:** [October 28, 2021, 1:13pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/8 "2021-10-28T13:13:35Z")

</div>

The kde from this code is just the line y = 0.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 28, 2021, 1:28pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/9 "2021-10-28T13:28:12Z")

</div>

Maybe something like this is what you want:

```julia
julia> nsteps() = rand(1:50) # the "number of steps"
nsteps (generic function with 1 method)

julia> counter = zeros(Int,50); # 50 is the maximum number of steps possible

julia> for i in 1:1000 # number of samples
           n = nsteps() # check the number of steps
           counter[n] +=1 # add to counter
       end

julia> findmax(counter) # find index and number of the maximum counter
(28, 7)

julia> counter[7]
28

```

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 28, 2021, 1:51pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/10 "2021-10-28T13:51:46Z")

</div>

> [@newtothis](#):
>
> kde from this code is just the line y = 0

No it’s not, but the code does zoom in to the region where the peak is and the curve is quite flat when zoomed.

Note, I was using Julia 1.7 so perhaps the RNG is different from your Julia version.

Note also, I edited the post to include `using Random,StatsPlots` because I had those already loaded in my julia session and didn’t notice you need to do that.

When I run the code above it finds a peak density at x = 2.85156… and graphs a fairly flat parabola like shape with peak value about 0.13589 at that x location.

---

<div class="post-metadata">

**Author:** ![newtothis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/newtothis/32/32227_2.png) [@newtothis](https://discourse.julialang.org/u/newtothis)\
**Post date:** [October 28, 2021, 6:35pm UTC](https://discourse.julialang.org/t/peak-finding-from-a-distribution/70485/11 "2021-10-28T18:35:09Z")

</div>

Thank you. It works now.
