# Get smooth CDF from OnlineStats's Quantile

**URL:** <https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886>\
**Category:** Statistics\
**Created:** [May 6, 2024, 8:50am UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886 "2024-05-06T08:50:13Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![disberd](https://avatars.discourse-cdn.com/v4/letter/d/8edcca/32.png) [@disberd](https://discourse.julialang.org/u/disberd)\
**Post date:** [May 6, 2024, 8:50am UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886/1 "2024-05-06T08:50:13Z")

</div>

Hello,

I am tracking statistics of some simulations using the `Quantile` object from the amazing OnlineStats.jl

I am also interested in being able to plot the CDF of the variable being observed with the `Quantile` but trying to do so directly with a code like so

```julia
using OnlineStats
q = fit!(Quantile(), randn(10^4))
# Get the CDF points
y = range(0,1; length = 101)
x = map(z -> value(q, z), y)

```

will create a CDF which is not smooth.

I could also `fit!` my data through an `OrderStats(100)` object but I have the impression that I should be able to reconstruct a smooth CDF from a `Quantile` (considering it holds an histogram with 500 bins) withouth the need of keeping a separate `OrdersStats` just for plotting purposes.

I have seen that I can get something more smooth by wrapping the `Quantile`’s histogram in the `Ash` object still from OnlineStats to get a smoothed PDF and integrate that, but is there maybe a better method to get the smoothed CDF from my `Quantile` object?

I am tagging @joshday as he is the main author of OnlineStats so he might already know the best approach 🙂

---

<div class="post-metadata">

**Author:** ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)\
**Post date:** [May 6, 2024, 11:35am UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886/2 "2024-05-06T11:35:58Z")

</div>

> [@disberd](#):
>
> ```julia
> using OnlineStats
> q = fit!(Quantile(), randn(10^4))
> # Get the CDF points
> y = range(0,1; length = 101)
> x = map(z -> value(q, z), y)
> 
> ```

Quantiles are estimated from a histogram, which will have jumps. I think the right way to plot it would be `Plots.plot(x, y, seriestype=:step)`. If you need a smoother estimate, you’ll need to add more bins in the histogram, e.g. `Quantile(b = 5000)` (default is 500).

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [May 6, 2024, 1:03pm UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886/3 "2024-05-06T13:03:59Z")

</div>

If you wanted it to be actually smooth, I’d use a kernel density first and then integrate it to get the CDF.

---

<div class="post-metadata">

**Author:** ![disberd](https://avatars.discourse-cdn.com/v4/letter/d/8edcca/32.png) [@disberd](https://discourse.julialang.org/u/disberd)\
**Post date:** [May 6, 2024, 1:27pm UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886/4 "2024-05-06T13:27:42Z")

</div>

Thanks @tbeason, I actually found the `Ash` function from this other somehow related post:

> [@On-the-flight kernel density estimation?](https://discourse.julialang.org/t/on-the-flight-kernel-density-estimation/95817):
>
> I have the following type of problem: imagine a Monte-Carlo simulation generating some data, and I want to estimate the density function for that simulation. I cannot store the data for each Monte-Carlos step, which would be prohibitive, but I want to do something better than simply constructing a histogram. To provide a concrete example: In theory a nice way to obtain the density would be: julia\> using Plots, KernelDensity julia\> data = randn(10^4); julia\> k = kde(data); julia\> plot(k.x,…

I understand that the AverageShiftedHistogram is somehow something that does similarly to the kernel density you mentioned, so using that and integrating its output was indeed the first thing I tried as mentioned in the original post.

Do you think this approach is actually different from what you are suggesting?

I was actually trying out the different options in a Pluto notebook:

> **Notebook Code**
>
> ```julia
> # ╔═╡ 21f03ec8-4e76-4e7c-8cac-3e67d9e79e44
> begin
> using OnlineStats
> using PlutoPlotly
> using Statistics
> end
> 
> # ╔═╡ dff0f3f2-6706-4f60-a607-1fcff9b3b314
> begin
> db2lin(x) = 10.0^(x/10)
> q = Quantile()
> o = OrderStats(100)
> a = db2lin.(randn(10^5) .* 2) # Let's create some lognormal variable
> # a = randn(10^4)
> fit!(q, a)
> fit!(o, a)
> end
> 
> # ╔═╡ d78b749b-a739-4fcf-a5c2-1c7d41fc11f0
> let
> d1 = let
> # We just build an Average Shifted Histogram for the smoothed pdf of the histogram used by Quantile
> ash = Ash(q.eh, 1)
> x, y = value(ash) # This extracts x and y for the smoothed pdf
> y = cumsum(y)
> # eltype to convert the step (which is TwicePrecision) to the actual type of the elements of y
> y *= eltype(y)(x.step) # We have to normalize by the step on x to have the sum to 1
> scatter(;x,y, name = "ASH")
> end
> d2 = let
> y = range(0, 1; length = 101)
> x = map(y) do y
> value(q, y)
> end
> scatter(;x, y, name = "Quantile", line_dash = :dash)
> end
> d3 = let
> y = range(0, 1; length = 101)
> x = map(y) do y
> quantile(o, y)
> end
> scatter(;x, y, name = "OrderStats", line_dash = :dot)
> end
> plot([d1, d2, d3], Layout(;
> template = "none",
> uirevision = 1,
> xaxis = attr(;
> title = "X"
> ),
> yaxis = attr(;
> title = "P{x <= X}"
> )
> ))
> end
> 
> ```

I see that Ash works for smoothing but the actual curve from OrderStats is “closer” to the Quantile curve (see below example zoom of the CDF plot of the notebook):  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/0/70693326f8da1943f1e2a1976ba88156b47fb1bd.png)

---

<div class="post-metadata">

**Author:** ![disberd](https://avatars.discourse-cdn.com/v4/letter/d/8edcca/32.png) [@disberd](https://discourse.julialang.org/u/disberd)\
**Post date:** [May 6, 2024, 1:30pm UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886/5 "2024-05-06T13:30:35Z")

</div>

Thanks a lot for the answer @joshday  
I get your point but I’d like to avoid increasing the number of bins in the quantile objects as I have potentially thousands of them and for everything else except “smoothness” of the ECDF plot I am more than fine with 500 bins.

I was just trying to find the _best_ approach to smooth out the cdf extraction from the quantile in post-processing just before plotting.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [May 6, 2024, 1:38pm UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886/6 "2024-05-06T13:38:30Z")

</div>

> [@disberd](#):
>
> Do you think this approach is actually different from what you are suggesting?

Effectively, no. That is basically what I was suggesting. I do not know why it appears to have some bias. I could speculate (I think the smoothing likely inflates the tails of the density), but if it is important to you then you need to decide how to approach the tradeoff.

---

<div class="post-metadata">

**Author:** ![disberd](https://avatars.discourse-cdn.com/v4/letter/d/8edcca/32.png) [@disberd](https://discourse.julialang.org/u/disberd)\
**Post date:** [May 6, 2024, 1:41pm UTC](https://discourse.julialang.org/t/get-smooth-cdf-from-onlinestatss-quantile/113886/7 "2024-05-06T13:41:44Z")

</div>

Thanks for confirming this. Indeed I probably do not need bigger precision than what the Ash method gives me, I just posted this question mostly to get inputs/suggestions on whether there were even better approaches 🙂
