# How do I fit a distribution to these data sets?!

**URL:** <https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243>\
**Category:** Statistics\
**Tags:** statistics\
**Created:** [August 31, 2019, 5:00pm UTC](https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243 "2019-08-31T17:00:31Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [August 31, 2019, 5:00pm UTC](https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243/1 "2019-08-31T17:00:31Z")

</div>

Let me start by saying that I’m very inexperienced in this area, so go easy on me @Tamas_Papp 😉. I have two datasets that are similar in shape and I’m trying to fit a distribution to each of them with the goal of being able to make probability statements about the processes that generated the data. For example, I want to be able to say that, “Under process A, the probability of a measurement being \> $10,000 is 0.1 while under process B, the probability of a measurement being \> $10,000 is 0.3 (or something like that).”

The problem is that there are a lot of zero values in the data and the data are bunched up around the lower end of the spectrum but the range is very wide, so I’m not sure what kind of distribution is appropriate. I’ve tried to deal with the zeros by doing log transformations (log.(data .+ 10), for example) , taking the square root, etc., but I’m not having any luck. I’m using the Distrbutions.jl package.

For one of the data sets, the summary stats look like this:

Summary Stats:  
Length: 30239  
Missing Count: 0  
Mean: 9011.465678  
Minimum: 0.000000  
1st Quartile: 0.000000  
Median: 250.400000  
3rd Quartile: 4129.670000  
Maximum: 3607200.690000  
Type: Float64

The 99th percentile is 138,688. I decided to lop the top 1% off with the hope that it would be easier to fit a distribution, but I’m still coming up short. Without the top 1% (so only including values \<= 138,688) a histogram of the data looks like this:

![image](https://global.discourse-cdn.com/julialang/original/3X/1/b/1b83016f9e026a597dee4b9742230aeff0e5b80f.png)

What I’ve tried is to fit basically every distribution possible form the Distributions.jl package, via the `fit` function, and then I do a `qqplot` (from StatsPlots) and the data never fit the distribution, no matter what I do.

My questions are:

1: Is there a distribution that would be a good natural fit for this kind of dataset?  
2: Should I exclude the zero-value data points?  
3: Should I exclude more of the upper-end values (maybe only keep the top 90% or 95%)?  
4: Should I explore more ways to transform the data?  
5: The distributions.jl docs talk about creating your own distribution, should I go down that road?

Any feedback is very much appreciated!

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [August 31, 2019, 5:46pm UTC](https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243/2 "2019-08-31T17:46:03Z")

</div>

Sorry, I have no idea why you pinged me above, except possibly if you want to fit a distribution using Bayesian methods. This would be able to answer all your questions, but requires some prior expertise. For this distribution in particular, a standard recommendation would be trying overdispersed families.

Regarding QQ plots: they are not necessarily good diagnostics for the very edge of tails, especially if you don’t have a lot of values there.

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [August 31, 2019, 6:47pm UTC](https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243/3 "2019-08-31T18:47:26Z")

</div>

This data seems to be truncated at 0, you can take natural log of it and then fit the log form data to certain distributions.

---

<div class="post-metadata">

**Author:** ![MatFi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matfi/32/10002_2.png) [@MatFi](https://discourse.julialang.org/u/MatFi)\
**Post date:** [August 31, 2019, 8:34pm UTC](https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243/4 "2019-08-31T20:34:22Z")

</div>

If I got you correctly, you just need a fitting distribution and there is no need that the parameters of the distribution have any meaning. So you are actually free to define your own distribution to fit the data.  
What speaks against a PDF of  
f = \frac{1}{n} \sum\_{i=1}^n \lambda\_i exp(-\lambda\_i x)   
or similar. Then then doing the rest (CDF) numerically.  
And for this the suggestion from @Yifan_Liu with doing the fit on a log scale is a good idea btw.

EDIT: forgot the norm  
the CDF should be somthing like  
f\_{CDF}= \frac{1}{n} \sum\_{i=1}^n (1- exp(-\lambda\_i x))

---

<div class="post-metadata">

**Author:** ![jkbest2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jkbest2/32/7350_2.png) [@jkbest2](https://discourse.julialang.org/u/jkbest2)\
**Post date:** [August 31, 2019, 9:30pm UTC](https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243/5 "2019-08-31T21:30:12Z")

</div>

Without more information about how these data were generated and your end goals, it’s going to be pretty difficult for anyone to help. That said, you probably want to look into zero-inflated models. One approach is to fit separate processes for the zeros and the positives. Here you’d have something like

```julia
p0 = fit(Bernoulli, y1 .== 0)
pos = fit(Exponential, y1[y1 .> 0]

```

You can then e.g. generate new data points as

```julia
rand(p0, 100) .* rand(pos, 100)

```

---

<div class="post-metadata">

**Author:** ![MatFi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/matfi/32/10002_2.png) [@MatFi](https://discourse.julialang.org/u/MatFi)\
**Post date:** [September 1, 2019, 5:59am UTC](https://discourse.julialang.org/t/how-do-i-fit-a-distribution-to-these-data-sets/28243/6 "2019-09-01T05:59:22Z")

</div>

I just have lerned from this [thread](https://discourse.julialang.org/t/using-a-normalized-histogram-as-a-distribution/27682) that there is [EmpiricalCDFs](https://github.com/jlapeyre/EmpiricalCDFs.jl) which should fulfill your purpose as well
