# Finding parameter (via MLE) of distributons with binned data

**URL:** <https://discourse.julialang.org/t/finding-parameter-via-mle-of-distributons-with-binned-data/65060>\
**Category:** New to Julia\
**Tags:** question, statistics\
**Created:** [July 21, 2021, 7:15pm UTC](https://discourse.julialang.org/t/finding-parameter-via-mle-of-distributons-with-binned-data/65060 "2021-07-21T19:15:57Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 21, 2021, 7:15pm UTC](https://discourse.julialang.org/t/finding-parameter-via-mle-of-distributons-with-binned-data/65060/1 "2021-07-21T19:15:57Z")

</div>

We have many packages to fit distributions / mixture of them over `data` with MLE or Chi2 methods. But what tools to use right now if my `data` is already binned. I.e.

```julia
using Distributions, StatsBase
n = Normal()
u = Uniform(-1,1)
s() = rand() < 0.8 ? rand(n) : rand(u)
data = [s() for _=1:10^5]
binned = fit(Histogram, data)

```

Let’s say, the goal is to find `0.8` to be the most likely weight assuming we know the center of the `normal` and the range of the uniform.

One idea is of course using curve fitting such as `LsqFit` and treat the PDF as a curve and bin position-bin height as x,y. Or to sample the histogram to “recover” underlying data?

What would be a ready-made way in Julia’s ecosystem?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 21, 2021, 8:46pm UTC](https://discourse.julialang.org/t/finding-parameter-via-mle-of-distributons-with-binned-data/65060/2 "2021-07-21T20:46:31Z")

</div>

turns out a curve fit with LsqFit would kind of work:

```julia
using Distributions, LsqFit, FHist

n = Normal(0, 2)
u = Uniform(-1, 1)

data = rand(n, 8000) # 80% distributed as gaussian
data = vcat(data, rand(u, 2000)) # 20% as flat background

h = Hist1D(data, -3:0.3:3);
xs = FHist.bincenters(h)
ys = h.hist.weights;

# let's assume we know the center of normal distribution is 0 
# and we know uniform ranges from -1 to 1
# p[1] is sigma of Normal
# p[2] is strength for Normal, Uniform will have (1-p[2])
# p[3] is overall scale

mypdfN(x, sig) = pdf(Normal(0, sig), x)
mypdfU(x) = pdf(Uniform(-1, 1), x)

@. mypdf(x, p) = (mypdfN(x, p[1])*p[2] + (1-p[2]) * mypdfU(x))*p[3]

p0 = [3, 0.5, 10^4] # initial guess, 10^5 is basically our normalization
lb = [0.1, 0., 0.] # lower bounds because some parameter can't go negative
myfit = curve_fit(mypdf, xs, ys, p0; lower=lb);

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/1/7192aab2ea2a544dc4bc78976cf688141f92c2b4.png)
