# Different parametrizations of probability distributions in Distributions.jl

**URL:** <https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435>\
**Category:** Statistics\
**Created:** [December 7, 2018, 8:43pm UTC](https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435 "2018-12-07T20:43:03Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![alejandromerchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alejandromerchan/32/10500_2.png) [@alejandromerchan](https://discourse.julialang.org/u/alejandromerchan)\
**Post date:** [December 7, 2018, 8:43pm UTC](https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435/1 "2018-12-07T20:43:03Z")

</div>

In R there are several probability distributions that allow keywords for different parametrizations of the same distribution. For example, I use the Negative Binomial distribution a lot in my simulations of some biological processes and what particular case has me a little stumped.

I have some process that produces a vector with 950 Ints (no zeros) and I want to compare the actual distribution of the data with a theoretical NegBinom distribution. In R there is a parametrization that uses the mean, which I can calculate from my data, and the probability, which I can change to see how that affects the histogram. We don’t have access to that in Julia and while I’m sure there’s a way to get the parameter _r_ from the mean, I haven’t been able to figure it out.

What would be the best strategy to allow different parametrizations in different distributions? What would it take to change Negative Binomial, as an example, and I can probably try and do it as an exercise? I don’t know if this has been discussed before or if there are already plans in this direction, if so, I’ll appreciate a link to the relevant discussion.

---

<div class="post-metadata">

**Author:** ![mschauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mschauer/32/13946_2.png) [@mschauer](https://discourse.julialang.org/u/mschauer)\
**Post date:** [December 7, 2018, 9:39pm UTC](https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435/2 "2018-12-07T21:39:22Z")

</div>

For mean failure number `m` and success prob. `p`, it is `NegativeBinomial(m*(p/(1 - p)), p)`. Actually you would like to fit `p` and `r` directly without trying to guess `p` from the histogram, I do not know if that is implemented somewhere.

---

<div class="post-metadata">

**Author:** ![alejandromerchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alejandromerchan/32/10500_2.png) [@alejandromerchan](https://discourse.julialang.org/u/alejandromerchan)\
**Post date:** [December 7, 2018, 10:24pm UTC](https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435/3 "2018-12-07T22:24:52Z")

</div>

Thanks, that works! I think implementing different parametrizations would be a nice feature, but I really don’t know what would be the best way to do it.

And I’m not aware of any fitting function for Negative Binomial because there is no implementation of the sufficient statistics.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [December 8, 2018, 7:33am UTC](https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435/4 "2018-12-08T07:33:50Z")

</div>

> [@alejandromerchan](#):
>
> I think implementing different parametrizations would be a nice feature

I understand that different parametrizations are used in different communities/textbooks for some distributions, however I think that the best approach is to pick _one_, document it, and use it consistently in code.

Anything else will just lead to confusion and subtle bugs. IMO `Distributions` already has too much special-cased constructors, eg for `MvNormal`. Also cf [Distributions#584](https://github.com/JuliaStats/Distributions.jl/issues/584) and issues referenced there.

Usually the conversion to/from other parametrizations is trivial, so the documentation could mention that. So submitting PRs that extend docstrings would be my suggestion.

---

<div class="post-metadata">

**Author:** ![mschauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mschauer/32/13946_2.png) [@mschauer](https://discourse.julialang.org/u/mschauer)\
**Post date:** [December 8, 2018, 11:41am UTC](https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435/5 "2018-12-08T11:41:28Z")

</div>

I think that, considering that `mean(::NegativeBinomial)` is defined even if it is a trivial operation, the moral inverse  
`NegativeBinomial((mean=x, p=2))` (or `NegativeBinomial((mean=x, p=2))` to avoid conflict) could also be provided. Especially for the many distributions with `rate=` vs `scale=` parameters.

---

<div class="post-metadata">

**Author:** ![alejandromerchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alejandromerchan/32/10500_2.png) [@alejandromerchan](https://discourse.julialang.org/u/alejandromerchan)\
**Post date:** [December 10, 2018, 5:23pm UTC](https://discourse.julialang.org/t/different-parametrizations-of-probability-distributions-in-distributions-jl/18435/6 "2018-12-10T17:23:21Z")

</div>

Sure, documentation and examples are always good, however in the discussion you posted @simonbyrne talks about the possibility of using keyword arguments, with one canonical implementation, and I really believe that’s the way to go.
