# How to approximate a distribution function from an arbitrary list?

**URL:** <https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471>\
**Category:** General Usage\
**Tags:** question, package\
**Created:** [November 18, 2022, 7:08pm UTC](https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471 "2022-11-18T19:08:45Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![somranger](https://avatars.discourse-cdn.com/v4/letter/s/3ab097/32.png) [@somranger](https://discourse.julialang.org/u/somranger)\
**Post date:** [November 18, 2022, 7:08pm UTC](https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471/1 "2022-11-18T19:08:45Z")

</div>

I hope to obtain a approximated distribution function from list, counting exsiting times of each element. Is there any package can achieve this?  
In mathematica, function `SmoothKernelDistribution` (in Wolfram document, it is done by linear interpolation) could do this. For example:

```julia
data = RandomVariate[NormalDistribution[], 10^3];
Table[Plot[f[\[ScriptCapitalD], x], {x, -4, 4},PlotLabel -> f], {f, {PDF, CDF}}]

```

![mmexport1668796265060](https://global.discourse-cdn.com/julialang/original/3X/e/c/ec3680c7fed80a1c6f1053802351f062cfaca4b8.png)

I have tried `kde` and `pdf` in `KernelDensity`, but the result is not as good as mathematica. And I hope not only for normal distribution, but for any unknown type of distribution approximated by interpolation.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [November 18, 2022, 7:12pm UTC](https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471/2 "2022-11-18T19:12:41Z")

</div>

KernelDensity.jl should be using basically the exact same defaults as the defaults in that Mathematica function. I would be surprised if the results were noticeably any different for samples of that size (1000+ points).

---

<div class="post-metadata">

**Author:** ![somranger](https://avatars.discourse-cdn.com/v4/letter/s/3ab097/32.png) [@somranger](https://discourse.julialang.org/u/somranger)\
**Post date:** [November 18, 2022, 7:19pm UTC](https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471/3 "2022-11-18T19:19:04Z")

</div>

This is my code in Julia. I even generate more points.

```julia
using Distributions,KernelDensity,Plots
x = rand(Normal(), 100000)
f1 = kde(x);
plot(-10:10,x->pdf(f1,x))

```

 ![mmexport1668799128825](https://global.discourse-cdn.com/julialang/original/3X/f/b/fb452ac4ceb0549970313bcd690da47c21847790.png)

---

<div class="post-metadata">

**Author:** ![aramirezreyes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aramirezreyes/32/42573_2.png) [@aramirezreyes](https://discourse.julialang.org/u/aramirezreyes)\
**Post date:** [November 18, 2022, 7:27pm UTC](https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471/4 "2022-11-18T19:27:06Z")

</div>

I think your plot has too few points for your liking (you are only plotting on the 21 integers from -10 to 10).

E.g. it does not look that bad:

```julia
julia> using Distributions, KernelDensity, Plots

julia> f1 = kde(randn(10^3));

julia> x1 = range(-4,stop=4,length=1000)

julia> y1 = pdf.(Ref(f1),x1);

julia> plot(x1,y1)

```

Produces this:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/9/c/9cf859b4eb21ab812748180790452ebb73c7e73a.png)

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [November 18, 2022, 7:31pm UTC](https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471/5 "2022-11-18T19:31:54Z")

</div>

A few ideas:

1. Load the `StatsPlots` package and use the `density` function.
2. Generate more points:

```julia
plot(-10:.01:10, x -> pdf(f1, x))

```

3. I don’t know what your goal is, but you might also consider fitting a parametric distribution to your data. Something like:

```julia
d = fit(Normal, x)
plot(d)

```

---

<div class="post-metadata">

**Author:** ![somranger](https://avatars.discourse-cdn.com/v4/letter/s/3ab097/32.png) [@somranger](https://discourse.julialang.org/u/somranger)\
**Post date:** [November 18, 2022, 7:35pm UTC](https://discourse.julialang.org/t/how-to-approximate-a-distribution-function-from-an-arbitrary-list/90471/6 "2022-11-18T19:35:59Z")

</div>

You are right. I hope to fit a distribution from data. In this simple example, it works for normal distribution. I need a general fit for any type of data.
