# How to plot ~10 billion datapoint time series efficiently?

**URL:** <https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228>\
**Category:** Visualization\
**Tags:** question, plotting, diffeq\
**Created:** [May 18, 2022, 1:09am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228 "2022-05-18T01:09:30Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![hk69](https://avatars.discourse-cdn.com/v4/letter/h/c57346/32.png) [@hk69](https://discourse.julialang.org/u/hk69)\
**Post date:** [May 18, 2022, 1:09am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/1 "2022-05-18T01:09:30Z")

</div>

Hi, I am relatively new to Julia.  
I have a physics model that I implemented in Julia. I solve the differential equations to track the orbit of two blackholes over a period of about 5 years with very small time step. My output has ~10 billion elements which I want to plot. What is the best (fastest and memory efficient) way to do this in Julia?

I don’t need to zoom in or interact with my plot. I have tried gr() and inspectdr() backends for Plots.jl and I have also tried CairoMackie backend for Mackie but they just aren’t giving me the performance. pyplot() backend is the fastest that I have tested but that too is useful only up to 1e8 datapoints. I want to go to 1e10 datapoints.

Edit: I can’t use adaptive time-step to reduce the number of datapoints while solving my diff eq because that gives me uneven time steps. I have to take the FFT of my solution which doesn’t work with uneven time steps.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [May 18, 2022, 2:25am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/3 "2022-05-18T02:25:29Z")

</div>

It makes no sense to plot that many points. No one can see them. Compute a density distribution and plot that.

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [May 18, 2022, 3:10am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/4 "2022-05-18T03:10:44Z")

</div>

would someone show how that might be done?

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [May 18, 2022, 5:26am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/5 "2022-05-18T05:26:38Z")

</div>

Putting plotting aside for a moment, several solutions exist for handling uneven timesteps: you could use a nonuniform FFT as provided by [FINUFFT.jl](https://github.com/ludvigak/FINUFFT.jl) or [FastTransforms.jl](https://github.com/JuliaApproximation/FastTransforms.jl), or use the DiffEq.jl [solution interpolation](https://diffeq.sciml.ai/stable/basics/solution/#Interpolations-and-Calculating-Derivatives) interface to resample onto a uniform grid.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [May 18, 2022, 7:37am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/6 "2022-05-18T07:37:18Z")

</div>

That would occupy about 80GB of memory, so probably you can’t even load that trajectory in memory.

Can’t you just sample less frequently?

(You don’t need to save every times step of the simulation)

---

<div class="post-metadata">

**Author:** ![paulmelis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paulmelis/32/35063_2.png) [@paulmelis](https://discourse.julialang.org/u/paulmelis)\
**Post date:** [May 18, 2022, 7:43am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/7 "2022-05-18T07:43:53Z")

</div>

> [@hk69](#):
>
> I have a physics model that I implemented in Julia. I solve the differential equations to track the orbit of two blackholes over a period of about 5 years with very small time step. My output has ~10 billion elements which I want to plot. What is the best (fastest and memory efficient) way to do this in Julia?
> 
> I don’t need to zoom in or interact with my plot.

Putting aside the technical challenge here, what’s the goal for visualizing your data? Inspect the output for correctness? Get a feel for the patterns in the data? Prepare a publication? As you mention you don’t need interactivity I would assume it is not so much data exploration that you’re interested in, which makes me wonder why you would want to plot all of the data, given its size.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [May 18, 2022, 9:19am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/9 "2022-05-18T09:19:09Z")

</div>

Please check [this GR solution](https://discourse.julialang.org/t/is-there-a-way-to-do-digital-phosphor-type-of-plots-for-time-series-data/56490/12), using `shade()` to plot a large 1D time series.

Or [this other GR example](https://github.com/jheinen/GR.jl/blob/master/examples/shade_ex.jl) using `shadepoints()`. By using `Float32` , more than 1 billion points can be processed in seconds:

 ![GR_shadepoints_1billion_points](https://global.discourse-cdn.com/julialang/original/3X/e/7/e761ef225d3e7a29f1a464ac187a7dc6a41b7a39.jpeg)

---

<div class="post-metadata">

**Author:** ![wc4wc4wc4](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wc4wc4wc4/32/23038_2.png) [@wc4wc4wc4](https://discourse.julialang.org/u/wc4wc4wc4)\
**Post date:** [May 18, 2022, 10:55am UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/10 "2022-05-18T10:55:06Z")

</div>

Can you show a small example of the plot you want to improve?  
Do you plot trajectories (i.e. lines) or would a histogram (1D / 2D) make sense instead?

---

<div class="post-metadata">

**Author:** ![trilobit](https://avatars.discourse-cdn.com/v4/letter/t/a8b319/32.png) [@trilobit](https://discourse.julialang.org/u/trilobit)\
**Post date:** [May 18, 2022, 12:49pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/11 "2022-05-18T12:49:23Z")

</div>

Perhaps this helps:

> **[GitHub - joshday/OnlineStats.jl: ⚡ Single-pass algorithms for statistics](https://github.com/joshday/OnlineStats.jl)**
>
> ⚡ Single-pass algorithms for statistics. Contribute to joshday/OnlineStats.jl development by creating an account on GitHub.

[Online algorithm - Wikipedia](https://en.wikipedia.org/wiki/Online_algorithm)

---

<div class="post-metadata">

**Author:** ![scheidan1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scheidan1/32/24771_2.png) [@scheidan1](https://discourse.julialang.org/u/scheidan1)\
**Post date:** [May 18, 2022, 1:49pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/12 "2022-05-18T13:49:40Z")

</div>

You probably want to downsample your data before plotting.

The thesis _“Downsampling Time Series for Visual Representation”_ by Sveinn Steinarsson investigates different visaually pleasing subsampling strategies ([see here](http://skemman.is/stream/get/1946/15343/37285/3/SS_MSthesis.pdf)).  
If I remeber correctly the “Largest Triangle Three Bucket” (LTTB) algorithm is recommended.

~~Unfortunately, I could only find implementations in java and python~~

See [here](https://gist.github.com/jmert/4e1061bb42be80a4e517fc815b83f1bc) from @stevengj repley below.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [May 18, 2022, 1:50pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/13 "2022-05-18T13:50:03Z")

</div>

> [@joa-quim](#):
>
> It makes no sense to plot that many points.

There are a lot of different algorithms for downsampling huge timeseries datasets for visualization. See [this post](https://discourse.julialang.org/t/plotting-image-data-in-julia-is-much-slower-than-matlab/79660/18), for example (including some code).

If the downsampling algorithm is local (as is the case in the linked example above), you can process the dataset in chunks if the whole thing doesn’t fit into memory at once.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [May 18, 2022, 1:57pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/14 "2022-05-18T13:57:53Z")

</div>

Yep, I saw that post when it arrived and is on my todo list to adopt something like that in GMT.jl. But user mentioned orbits which makes think the problem might depend on x,y coordinates. What I had in mind is the use of the GMT module [blockmean](https://www.generic-mapping-tools.org/GMT.jl/dev/blockmean/) to compute a grid with some statistics of the orbits into a grid that could later be easily and quick turned into a plot.  
I’m not sure right now (would need to check the C code) but I think `blockmean` does a record-by-record reading so if data is in a disk file, it will take a wile but RAM memory should not be a problem (if it is, the procedure would have to be cut in chunks).

---

<div class="post-metadata">

**Author:** ![hk69](https://avatars.discourse-cdn.com/v4/letter/h/c57346/32.png) [@hk69](https://discourse.julialang.org/u/hk69)\
**Post date:** [May 18, 2022, 3:02pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/15 "2022-05-18T15:02:13Z")

</div>

I can’t sample less frequently because my orbit is spiraling inwards. It becomes smaller and smaller so if I sample less frequently I get very few data points per orbit which can’t really map the orbit accurately.  
I am aware of the memory problem so I have rented a virtual machine with lots of memory to run this code.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [May 18, 2022, 3:04pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/16 "2022-05-18T15:04:27Z")

</div>

> [@hk69](#):
>
> I can’t sample less frequently because my orbit is spiraling inwards. It becomes smaller and smaller so if I sample less frequently I get very few data points per orbit which can’t really map the orbit accurately.

That’s why downsampling strategies often need to be adaptive, so they sample more frequently when the data is changing more.

---

<div class="post-metadata">

**Author:** ![hk69](https://avatars.discourse-cdn.com/v4/letter/h/c57346/32.png) [@hk69](https://discourse.julialang.org/u/hk69)\
**Post date:** [May 18, 2022, 3:08pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/17 "2022-05-18T15:08:30Z")

</div>

It is for a publication. I don’t want to plot all of the data, I just can’t find a solution that helps me preserve the signal and create a publication-quality plot. (I also want to overlay other signals on the same plot)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [May 18, 2022, 3:29pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/18 "2022-05-18T15:29:58Z")

</div>

> [@stevengj](#):
>
> That’s why downsampling strategies often need to be adaptive, so they sample more frequently when the data is changing more.

For example, here are two recent adaptive downsampling algorithms; see also the references therein:

- Daniel et al, “[Adaptive resampling for data compression](https://doi.org/10.1016/j.array.2021.100076)” (2021)
- Gil et al “[Towards Smart Data Selection From Time Series Using Statistical Methods](https://doi.org/10.1109/ACCESS.2021.3066686)” (2021).

Not sure what free implementations are out there of this kind of technique, though; I couldn’t find anything even in Python, but maybe I was searching the wrong keywords. (e.g. `Pandas.resample` uses fixed bin sizes.)

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [May 18, 2022, 5:41pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/19 "2022-05-18T17:41:12Z")

</div>

If it is permissible, post a _short_ exerpt from the data file[s] you want to show along with (if you know, or just guess) any characterizing info (e.g. the largest, smallest and mean values in each data vector). It may help others to give more specific suggestions if you posted even a very coarse plot (one point each million) – or better yet, if there are other plots with a similar look that you could link.

---

<div class="post-metadata">

**Author:** ![hk69](https://avatars.discourse-cdn.com/v4/letter/h/c57346/32.png) [@hk69](https://discourse.julialang.org/u/hk69)\
**Post date:** [May 18, 2022, 8:33pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/20 "2022-05-18T20:33:42Z")

</div>

![hc_GW_DF_Orb240k_dp100_alpha_2point49_atol_1e14_rtol_1e12](https://global.discourse-cdn.com/julialang/original/3X/c/b/cbfc87b363842daae2a2ea89750d92abdcee299f.jpeg)  
This is the plot for a part of the orbit. The full plot will be very similar to it, the yellow curve will just be spanning a bit more of the x-axis

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [May 18, 2022, 8:50pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/21 "2022-05-18T20:50:05Z")

</div>

Is this correct?

And with your data, each of the two “curves” is given as a sequence of (x,y) coordinate pairs. And each “curve” may be determined by, say, 50 billion pairs; and the pairs are stored in some sensible sequence, so the (x,y) pair at index i1 bears some physical contiguity with the predecessor and successor pairs.

---

<div class="post-metadata">

**Author:** ![hk69](https://avatars.discourse-cdn.com/v4/letter/h/c57346/32.png) [@hk69](https://discourse.julialang.org/u/hk69)\
**Post date:** [May 18, 2022, 9:06pm UTC](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228/22 "2022-05-18T21:06:44Z")

</div>

Yes. The blue curve doesn’t have billions of data points. It’s just the noise curve for comparison. The yellow curve is theoretically calculated with a lot of data points but it is just a simple x,y line plot. The GR.shade example suggested by @rafael.guerra works well for this but it just doesn’t produce a publication quality plot.

[Next page](https://discourse.julialang.org/t/how-to-plot-10-billion-datapoint-time-series-efficiently/81228.md?page=2)
