# Using a (normalized) Histogram as a Distribution

**URL:** <https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682>\
**Category:** General Usage\
**Created:** [August 18, 2019, 4:02pm UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682 "2019-08-18T16:02:38Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 18, 2019, 4:02pm UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/1 "2019-08-18T16:02:38Z")

</div>

Do we have a wrapper somewhere that turns a `StatsBase.Histogram` into a `Distributions.Distribution?`

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [August 18, 2019, 7:00pm UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/3 "2019-08-18T19:00:10Z")

</div>

Sampling from a Histogram is basically just randomly sampling from the underlying data, no? Why not do that?

---

<div class="post-metadata">

**Author:** ![jkbest2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jkbest2/32/7350_2.png) [@jkbest2](https://discourse.julialang.org/u/jkbest2)\
**Post date:** [August 18, 2019, 7:01pm UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/4 "2019-08-18T19:01:06Z")

</div>

I’ve not used either, but `StatsBase` contains an `ecdf` (empirical CDF) function, and there is also [EmpiricalCDFs.jl](https://github.com/jlapeyre/EmpiricalCDFs.jl).

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [August 19, 2019, 1:43am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/5 "2019-08-19T01:43:51Z")

</div>

After reviewing `EmpiricalCDFs.jl`, I think the current implementation of random sampling is unnecessarily complicated. I think tbeason is correct.

> Sampling from a Histogram is basically just randomly sampling from the underlying data, no?

EDIT: I just did a benchmark and, complicated or no, the current implementation is 10% faster in one test than the straightforward and naive way: `data[rand(1:length(data))]`.

But, the OP may have other reasons for wanting to use the interface in `Distributions`.

`DiscreteNonParametric` does not do what the OP asks for.

---

<div class="post-metadata">

**Author:** ![pistacliffcho](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pistacliffcho/32/8518_2.png) [@pistacliffcho](https://discourse.julialang.org/u/pistacliffcho)\
**Post date:** [August 19, 2019, 4:09am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/6 "2019-08-19T04:09:40Z")

</div>

> [@tbeason](#):
>
> Sampling from a Histogram is basically just randomly sampling from the underlying data, no? Why not do that?

A good reason for not doing that is if you have a lot of data initially and you have a large number of ties in your data. Then a simple empirical distribution will greatly reduce what needs to saved and can greatly speed up various operations on the distribution.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [August 19, 2019, 4:51am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/7 "2019-08-19T04:51:08Z")

</div>

Before thinking about the implementation, you should tackle the conceptual question of what you want to do exactly, why you would want to go through a histogram, and what it means to sample from one. Eg one could imagine a histogram as a mixture distribution of uniforms (for each bin). Is this what you need?

Nonparametric methods all have various trade-offs you should consider, ie you get some nice theoretical properties when you can assume some regularity/smoothness about the distribution. Bootstrap (treating the sample as the distribution) assumes little but one can usually do better.

---

<div class="post-metadata">

**Author:** ![bernhard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bernhard/32/2619_2.png) [@bernhard](https://discourse.julialang.org/u/bernhard)\
**Post date:** [August 19, 2019, 4:51am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/8 "2019-08-19T04:51:48Z")

</div>

> [@jlapeyre](#):
>
> `DiscreteNonParametric` does not do what the OP asks for.

Thank you. I was looking for the ecdf function in the distribution package, but could not find it (as it is in StatsBase). I must have misinterpreted the meaning of DiscreteNonParametric

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 19, 2019, 7:59am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/9 "2019-08-19T07:59:02Z")

</div>

> Before thinking about the implementation, you should tackle the conceptual question of what you want to do exactly

What I want to do is using a histogram as a prior, in a context where priors are represented by `Distribution`. So I need a histogram with a `Distribution` API, mostly for `logpdf` and `rand`.

> why you would want to go through a histogram

As for why I’m using a histogram, it’s because in some cases that’s just the prior information we have. Of course having the original information, before histogramming, may be preferable - alas, that’s not always readily available.

> and what it means to sample from one.

Well, in this case it need to sample to draw starting values for MCMC chains according to my prior (histogram). But in general, interpreting a histogram as an ecdf and sampling from it - while maybe lacking in statistical purity - is quite a common use case, right?

> one could imagine a histogram as a mixture distribution of uniforms (for each bin). Is this what you need?

Basically, yes. It’s not hard to implement, of course - I was just curious if we already had that available in some package.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 19, 2019, 8:04am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/10 "2019-08-19T08:04:07Z")

</div>

> Sampling from a Histogram is basically just randomly sampling from the underlying data, no? Why not do that?

I usually encounter two reasons:

- The underlying data is not available

- The underlying data is very large (GB to TB) . Using a decently fine-binned histogram is basically just as precise, but comes with a very significant gain in performance.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [August 19, 2019, 8:08am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/11 "2019-08-19T08:08:17Z")

</div>

> [@oschulz](#):
>
> interpreting a histogram as an ecdf and sampling from it - while maybe lacking in statistical purity - is quite a common use case, right?

First, there is no reason to go through the ECDF, you can sample from a histogram directly (once you decide how you interpret it, eg a mixture of uniforms).

Second, statistical interpretation (not necessarily “purity”) is key here. If the procedure is _ad hoc_, it can of course be useful in some contexts, but I would not expect a readily available implementation, nor would I consider it a common use case. YMMV.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 19, 2019, 9:18am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/12 "2019-08-19T09:18:42Z")

</div>

> First, there is no reason to go through the ECDF, you can sample from a histogram directly

Yes, that was my intention. Sorry, my wording was a bit sloppy, of course a histogram isn’t really an ECDF. In any case, in many use cases I’ll probably have 2D (or sometimes even 3D) histograms, so probably best not to talk about this in term of CDFs.

> Second, statistical interpretation (not necessarily “purity”) is key here.

Well, the statistical interpretation would be that the histogram represents a step-wise approximation to the pdf of the underlying (prior) distribution. Though I agree that in some cases it may not be that simple - a step-wise function wouldn’t be very nice for HMC, for example, so maybe one my need to smoothe/approximate the histogram. That would be purely empirical, of course, but that doesn’t mean it wouldn’t work. 🙂

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [August 19, 2019, 9:22am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/13 "2019-08-19T09:22:25Z")

</div>

Perhaps it would be better if, instead of focusing on this specific question, you explained the underlying problem.

I don’t fully get the context here: why do you need to use data as a prior, is the issue Bayesian inference on large data (there are methods for that), what the model is, what kind of MCMC you want to use, etc.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 19, 2019, 9:47am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/14 "2019-08-19T09:47:20Z")

</div>

I think the discussion has veered a bit off-topic - maybe I wasn’t specific enough about what I was looking for in the first place. So, using the interpretation of a histogramm as a mixture distribution of uniforms (for each bin), as suggested by @Tamas_Papp: Do we have a `Distributions` compatible wrapper somewhere that does it already? If not, I’ll make one, I just wanted to avoid duplicating existing work.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 19, 2019, 9:59am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/15 "2019-08-19T09:59:22Z")

</div>

> [@Tamas\_Papp](#):
>
> why do you need to use data as a prior

The “data” may, e.g., be the marginalized posterior of a prior analysis, to be combined with new measurement data in a knowledge update process. You can of course also make it part of the likelihood (depending on your definition of prior and likelihood in the specific case). Or sometimes you do get non-marginalized MCMC output as a data source (e.g. cosmic parameter MC chains based on cosmic microwave background sets from Planck) - call that your prior or data - that you want to combine with other data, and one may decide to use histograms of the raw MC chains for performance reasons (e.g. to fit into CPU cache).

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [August 19, 2019, 11:37am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/16 "2019-08-19T11:37:45Z")

</div>

> [@oschulz](#):
>
> Do we have a `Distributions` compatible wrapper somewhere that does it already? If not, I’ll make one, I just wanted to avoid duplicating existing work.

FWIW, I am not aware of any. But this should not be that complicated, eg

```julia
using Distributions, StatsBase, IterTools, LinearAlgebra

function make_mixture(h)
    MixtureModel([Uniform(e...) for e in partition(h.edges[1], 2, 1)],
                 normalize(h.weights, 1))
end

```

should do it,

```julia
julia> h = fit(Histogram, randn(1000))
Histogram{Int64,1,Tuple{StepRangeLen{Float64,Base.TwicePrecision{Float64},Base.TwicePrecision{Float64}}}}
edges:
  -4.0:1.0:4.0
weights: [1, 17, 145, 346, 330, 143, 17, 1]
closed: left
isdensity: false

julia> m = make_mixture(h)
MixtureModel{Uniform{Float64}}(K = 8)
components[1] (prior = 0.0010): Uniform{Float64}(a=-4.0, b=-3.0)
components[2] (prior = 0.0170): Uniform{Float64}(a=-3.0, b=-2.0)
components[3] (prior = 0.1450): Uniform{Float64}(a=-2.0, b=-1.0)
components[4] (prior = 0.3460): Uniform{Float64}(a=-1.0, b=0.0)
components[5] (prior = 0.3300): Uniform{Float64}(a=0.0, b=1.0)
components[6] (prior = 0.1430): Uniform{Float64}(a=1.0, b=2.0)
components[7] (prior = 0.0170): Uniform{Float64}(a=2.0, b=3.0)
components[8] (prior = 0.0010): Uniform{Float64}(a=3.0, b=4.0)

julia> mean(m)
-0.010999999999999937

```

for univariate and it should be easy to extend to multivariate.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 19, 2019, 3:03pm UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/17 "2019-08-19T15:03:51Z")

</div>

> FWIW, I am not aware of any.

Good to know, thanks, Tamas!

> But this should not be that complicated, eg […]

Yep, something like that. I’ll code up something a bit more customized than using `MixtureModel` (`logpdf` would be too slow), but sure, it’s pretty straightforward.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [August 31, 2019, 8:37pm UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/18 "2019-08-31T20:37:52Z")

</div>

In case someone else is interested in this: Our Phd student @lmh91 implemented this now as part of BAT.jl: [https://github.com/bat/BAT.jl/blob/master/src/distributions/HistogramAsUvDistribution.jl](https://github.com/bat/BAT.jl/blob/master/src/distributions/HistogramAsUvDistribution.jl) .

We’ll let this mature and extend it further (n-dim histograms as multivariate distributions), then we’ll see if we can find a home for it in a more general package than BAT.jl.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 31, 2019, 10:57pm UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/19 "2019-08-31T22:57:37Z")

</div>

Taking this opportunity to brainstorm empirical CDFs in the Julia Stats ecosystem.

There are multiple implementations:

- StatsBase.jl: `ecdf`
- EmpiricalCDFs.jl: `EmpiricalCDF`
- GeoStatsBase.jl: `EmpiricalDistribution`

Would you like to join efforts to create a single high-performance, well-tested, widely-available implementation? Would it make sense to have it in Distributions.jl?

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [September 1, 2019, 6:58am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/20 "2019-09-01T06:58:24Z")

</div>

I’ve been thinking whether these things (ECDDF’s, histograms as distributions, etc.) could be combined in a new package “EmpiricalDistributions” or so.

---

<div class="post-metadata">

**Author:** ![oschulz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oschulz/32/2998_2.png) [@oschulz](https://discourse.julialang.org/u/oschulz)\
**Post date:** [September 1, 2019, 6:59am UTC](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682/21 "2019-09-01T06:59:26Z")

</div>

I think it would be nice if ECDFs would implement the `Distributions` API, in general (as far as possible).

[Next page](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682.md?page=2)
