# Fitting of distribution histogram - axis-limit issues and problems with defining goodness of fit?

**URL:** <https://discourse.julialang.org/t/fitting-of-distribution-histogram-axis-limit-issues-and-problems-with-defining-goodness-of-fit/48617>\
**Category:** General Usage\
**Tags:** question, package, plotting\
**Created:** [October 19, 2020, 11:00am UTC](https://discourse.julialang.org/t/fitting-of-distribution-histogram-axis-limit-issues-and-problems-with-defining-goodness-of-fit/48617 "2020-10-19T11:00:52Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![MJMAG](https://avatars.discourse-cdn.com/v4/letter/m/f05b48/32.png) [@MJMAG](https://discourse.julialang.org/u/MJMAG)\
**Post date:** [October 19, 2020, 11:00am UTC](https://discourse.julialang.org/t/fitting-of-distribution-histogram-axis-limit-issues-and-problems-with-defining-goodness-of-fit/48617/1 "2020-10-19T11:00:53Z")

</div>

Hey Community,

I am facing the following “problem”:  
I am supposed to find a growth rate distribution in a certain data set. For this I am looking at new objects being created at a timestep. These counts I want to illustrate with a histogram, fit with an appropriate distribution and in the best case even evaluate the goodness of the fit.

After getting the different rates which are stored in the array “differences”, I am filtering only the values bigger than 10:

```julia
using StatsBase
using Distributions
using StatsPlots

daily_rates = []
pop = countmap(differences)
#only take the rates that are higher than zero 
filtered_pop = filter(tuple -> last(tuple) > 10, collect(pop))

#now we have a tuple and we only want to have the appearances
pop_vals = [filtered_pop[i][2] for i in 1:length(filtered_pop)]
sort!(pop_vals)
daily_rates=pop_vals

```

which gives me the following array:

```julia
[11, 11, 11, 12, 12, 12, 13, 14, 15, 15, 17, 18, 21, 24, 25, 25, 28, 38, 39, 43, 47, 54, 55, 64, 87, 99, 127, 187, 237, 306, 413, 611, 933, 1563, 3271]

```

This array I am fitting with the Distributions.jl package where I tried the Pareto distribution with the maximum likelihood method:

```julia
P = fit_mle(Pareto, daily_rates)

```

To plot the distribution as a straight line fit I also get the unique values of the array above, to have some x-values:

```julia
x = unique(daily_rates);

```

Now I am plotting the histogram and the fit with StatsPlots.jl and I chose logartihmic binning:

```julia
StatsPlots.histogram(daily_rates, bins = 10 .^range(0.0, length = 101, stop=log10(maximum(daily_rates))), fillalpha = 0.4,normalize=true, xaxis=:log, yaxis=:log, xlims = (10, maximum(daily_rates)),label=:data)
StatsPlots.plot!(P,x, xaxis=:log, yaxis=:log,label=:fit)

```

If I plot it, the histogram bars do not start at the very zero line. Which I don’t really understand.

![rate_distribution](https://global.discourse-cdn.com/julialang/original/3X/b/6/b69aae281f967a453b2ad1c82c9180c0b640894e.png)

This is my first issue and I would be happy for any hint what I can do to change it.

Moreover now it would be very useful the get a statement regarding the goodness of the fit. I see two options:  
Either I have to calculate for example the chi squared after pearson manually with:  
 ![Chi_sq](https://global.discourse-cdn.com/julialang/original/3X/a/3/a30c3c8ce8959a0c47754564687ecfeaa65770a1.png)  
But for this I would need my exact y-values of the histogram for which I didn’t find any function that could print me these values.  
Or I am using the HypothesisTests.jl package by which I am a bit overwhelmed such that I didn’t figure out yet how to use it in my case since I am quite unexperienced.

I don’t know if I am asking for too much help - but I would be very thankful for any reply.  
Thanks in advance! 🙂

---

<div class="post-metadata">

**Author:** ![mikkoku](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikkoku/32/16274_2.png) [@mikkoku](https://discourse.julialang.org/u/mikkoku)\
**Post date:** [October 19, 2020, 1:48pm UTC](https://discourse.julialang.org/t/fitting-of-distribution-histogram-axis-limit-issues-and-problems-with-defining-goodness-of-fit/48617/2 "2020-10-19T13:48:57Z")

</div>

The following gives histogram y values

```julia
using Distributions
using StatsBase
f = fit(Histogram, [0.3, 0.5, 0.7], 10 .^ [-1, -0.5, 0, 0.5])
f.weights

```

(I didn’t find a function to extract the weights and accessing them like this feels dirty.)
