# Changing the smoothness of density function

**URL:** <https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611>\
**Category:** Visualization\
**Tags:** statistics\
**Created:** [August 18, 2021, 4:14pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611 "2021-08-18T16:14:41Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 18, 2021, 4:14pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/1 "2021-08-18T16:14:41Z")

</div>

Hi, I wish to plot a density function (to compare to another density function I have plotted).  
Unfortunatley, all values are integers, typically between 0 and 5. This gives a very jagged, and ugly looking plot. Is there some option where I can tune this, making it smoother?

![density_problem](https://global.discourse-cdn.com/julialang/original/3X/2/8/28943702c467923b1208fcdecd93f1c6e2e0a74e.png)  
(Note, the axis labels are totally misgiving, not that it matter much, but still. x-axis should be molecule numbers and y-axis probably frequency)

---

<div class="post-metadata">

**Author:** ![jkbest2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jkbest2/32/7350_2.png) [@jkbest2](https://discourse.julialang.org/u/jkbest2)\
**Post date:** [August 18, 2021, 5:27pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/2 "2021-08-18T17:27:09Z")

</div>

I’m sure there’s a way to change the smoothness of the density estimate (but don’t know it off the top of my head sorry). But it seems like a histogram might be more appropriate here. You can still overlay your other density function for comparison.

---

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 18, 2021, 5:40pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/3 "2021-08-18T17:40:44Z")

</div>

How are you making that plot?

```julia
using StatsPlots

x=[0,0,1,1,1,1,2,2,2,2,2,3,3,3,4,5,6]
density(x)

```

![Capture](https://global.discourse-cdn.com/julialang/original/3X/b/c/bcf52c50566ba0986e33993155461eb10df109ce.png)

---

<div class="post-metadata">

**Author:** ![George9000](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/george9000/32/23619_2.png) [@George9000](https://discourse.julialang.org/u/George9000)\
**Post date:** [August 18, 2021, 6:52pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/4 "2021-08-18T18:52:49Z")

</div>

If this is a discrete distribution, perhaps a PMF plot is more appropriate? See  
[this draft, p. 93 - 96](https://statisticswithjulia.org/StatisticsWithJuliaDRAFT.pdf), and [this](https://github.com/h-Klok/StatsWithJuliaBook/blob/master/3_chapter/binomialCoinFlip.jl) corresponding to the PMF plot on [this page](https://statisticswithjulia.org/gallery.html).

---

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 19, 2021, 2:07pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/5 "2021-08-19T14:07:01Z")

</div>

Yes, but there are about 50,000 values in the array.

---

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 19, 2021, 2:07pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/6 "2021-08-19T14:07:37Z")

</div>

That’s a really extensive guide, thanks for sharing! Will look at it.

---

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 19, 2021, 2:22pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/7 "2021-08-19T14:22:45Z")

</div>

You could use KernelDensity.jl:

```julia
using StatsPlots, KernelDensity

x=rand(1:6,50000)
density(x, label="using StatsPlots")

f=kde(x,0:7) #specify the points at which to evaluate the kernel density estimation
plot!(f.x,f.density, label="using KernelDensity")

```

 ![Capture](https://global.discourse-cdn.com/julialang/original/3X/7/d/7d5adebd542bbe0582cb99a82024356a9f0a124c.png)

---

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 19, 2021, 2:39pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/8 "2021-08-19T14:39:04Z")

</div>

Thanks, that did work, but only partially. If I use an integer grid, I get a curve very similar to what I would have gotten if I would have plotted the values 0-\>10, against the count of each value in my data. It works, but looks jagged and rough (which I would rather want to avoid in this case, since I make a comparison to a smooth curve, and don’t want people distracted by what in this case is an artificial difference). If I try changing the grid, however, I get the same phenomena as when using density:

```julia
f=kde(no_stress_wt_vals,0:1:10) 
p1 = plot(f.x,f.density)
f=kde(no_stress_wt_vals,0:0.25:10) 
p2 = plot(f.x,f.density)
plot(p1,p2,size=(1000,350))

```

 ![kernel_density_curve](https://global.discourse-cdn.com/julialang/original/3X/d/5/d569080ef96cc42a21221d632d07ffb2515bbe81.png)

---

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 19, 2021, 2:41pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/9 "2021-08-19T14:41:07Z")

</div>

Yeah…

It’s probably best to plot a histogram where the bins are the integer values.  
Then you can plot other smooth curves on top of it.  
I don’t think density functions make much sense for discrete distributions since you are fitting the data to a continuous distribution (kernel).

---

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 19, 2021, 2:41pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/10 "2021-08-19T14:41:49Z")

</div>

Got it, thanks for the help everyone 🙂

---

<div class="post-metadata">

**Author:** ![genkuroki](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/genkuroki/32/18030_2.png) [@genkuroki](https://discourse.julialang.org/u/genkuroki)\
**Post date:** [August 20, 2021, 8:47am UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/11 "2021-08-20T08:47:12Z")

</div>

I agree with hdavid16-san.

Moreover, if you want to compare a sample of a discrete distribution with a continuous distribution, you can plot the ecdf of the sample and the cdf of the continuous distribution on top of each other.

The L^1-metric of cdf’s on \mathbb{R} coinsides with [the Wasserstein metric](https://en.wikipedia.org/wiki/Wasserstein_metric) for p=1.

Sample working code:

```julia
using Plots
using Distributions

"""empirical cumulative distribution function"""
ecdf(sample, x) = count(≤(x), sample)/length(sample)

λ = 3
dist_true = Poisson(λ)
n = 2^10
sample = rand(dist_true, n)
m, u² = mean(sample), var(sample)
gamma = Gamma(m^2/u², u²/m)
a, b = 0, 10

P = plot();
histogram!(sample; norm=true, alpha=0.3,
    bin=a-0.5:b+0.5, xtick=a:b,
    label="size-$n sample of Poisson($λ)");
plot!(x -> pdf(gamma, x), a, b;
    label="Gamma distribution approximation");

Q = plot(; legend=:bottomright, xtick=a:b);
plot!(x -> ecdf(sample, x), a-0.5, b+0.5;
    label="ecdf of size-$n sample of Poisson($λ)");
plot!(x -> cdf(gamma, x), a-0.5, b+0.5;
    label="cdf of Gamma distribution approximation");

plot(P, Q; size=(500, 500), layout=(2, 1))

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/6/f/6f0e6d82ee0a52b6ad5ecc9a2e7c218ed64da4aa.jpeg)

---

<div class="post-metadata">

**Author:** ![mschauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mschauer/32/13946_2.png) [@mschauer](https://discourse.julialang.org/u/mschauer)\
**Post date:** [August 20, 2021, 11:45am UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/12 "2021-08-20T11:45:14Z")

</div>

The key question is: Why are the numbers integers? Are they random, but rounded in some way? Or are they not random at all but set by the experimenter?

---

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 20, 2021, 11:50am UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/13 "2021-08-20T11:50:27Z")

</div>

Basically, I am comparing the presence of a protein in experiments vs simulations. The experiments use fluorescence as a proxy, which gives continuous data. By simulated data simulates the actual copy-numbers of how many proteins there are (discrete).

The units of the two categories don’t really correspond to each other, but I wanted to create an image to show that the shapes were similar. I have more details elsewhere, but in this case I figured if one case was obviously continuous and one discrete, that would be what people note first, while I want them to focus on the shapes.

What I ended up doing was to use Interpolations.jl to interpolate the data, and then plot it continuous (which I agree might seem like cheating a bit, but in this case I think it is OK).

---

<div class="post-metadata">

**Author:** ![mschauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mschauer/32/13946_2.png) [@mschauer](https://discourse.julialang.org/u/mschauer)\
**Post date:** [August 20, 2021, 5:37pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/14 "2021-08-20T17:37:58Z")

</div>

No, my question is why the time is in `n` minutes where `n` is an integer and not a real number (as expected if you actually measure time of physical processes).

---

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 20, 2021, 6:23pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/15 "2021-08-20T18:23:56Z")

</div>

Ahh, I’m sorry, my misstake.

I was using `default()` in the script I ran (because most plots are time plots). This is the only exception, but seem I forgot to clean the axises. The x-axis should be number of molecules, and the y-axis fraction or something.

---

<div class="post-metadata">

**Author:** ![cjdoris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cjdoris/32/213133_2.png) [@cjdoris](https://discourse.julialang.org/u/cjdoris)\
**Post date:** [August 20, 2021, 6:34pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/16 "2021-08-20T18:34:25Z")

</div>

One thing you can do if you have a distribution over the integers is to un-discretize your samples by adding uniform random noise in [0,1] to the location x of each sample.

This way, the KDE is actually estimating the histogram of your distribution, which is a bona-fide continuous PDF. That is, it is estimating the density with PDF f(x)=p(\lfloor x \rfloor) where p is the PMF of the discrete data distribution.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [August 20, 2021, 7:05pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/17 "2021-08-20T19:05:21Z")

</div>

what you want to do is specify the bandwidth for the density estimation in the density command.

```julia
x=rand(1:6,50000)
density(x,label="default")
density!(x,label="bandwidth=1",bandwidth=1)

```

![image](https://global.discourse-cdn.com/julialang/original/3X/0/c/0c816a71504fea04bfddd358a8ef4b9866eced94.png)

---

<div class="post-metadata">

**Author:** ![Torkel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torkel/32/5030_2.png) [@Torkel](https://discourse.julialang.org/u/Torkel)\
**Post date:** [August 20, 2021, 8:59pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/18 "2021-08-20T20:59:09Z")

</div>

Thanks, yes, something like that is what I was thinking about. Thanks 🙂

---

<div class="post-metadata">

**Author:** ![genkuroki](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/genkuroki/32/18030_2.png) [@genkuroki](https://discourse.julialang.org/u/genkuroki)\
**Post date:** [August 20, 2021, 9:48pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/19 "2021-08-20T21:48:32Z")

</div>

hdavid16-san is right:

> [@hdavid16](#):
>
> It’s probably best to plot a histogram where the bins are the integer values.  
> Then you can plot other smooth curves on top of it.  
> I don’t think density functions make much sense for discrete distributions since you are fitting the data to a continuous distribution (kernel).

Please don’t apply the kernel density estimation to a sample of a discrete distribution with values in 1, 2, 3, 4, 5, and 6. The kernel density estimation is a method for estimating the density function of the continuous distribution from a sample.

For a sample with values 1, 2, 3, 4, 5, and 6, both its histogram with integer bins and its ecdf keep the true sample information, but the kernel density estimation does not.

```julia
using StatsPlots
x = rand(1:6, 10^5)
histogram(x; norm=true, alpha=0.3, bin=0.5:6.5, label="histogram of true sample")
density!(x; label="kde with bandwidth=1", bandwidth=1, lw=2)
plot!(; ylim=(-0.005, 0.22), xtick=-10:10)

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/8/7820597419a4c1c22b0e4ce1ed2ed8419ffc1b87.jpeg)

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [August 20, 2021, 9:58pm UTC](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611/20 "2021-08-20T21:58:07Z")

</div>

There are plenty of good reasons to want to treat discrete distributions as if they were continuous. Whether it’s appropriate or not depends entirely on the application. I use KDE plots for discrete distributions all the time, particularly when the discreteness is basically uninteresting to the application.

[Next page](https://discourse.julialang.org/t/changing-the-smoothness-of-density-function/66611.md?page=2)
