# Exploring 24/7 Market Data with Julia: Volatility and Simple Visualizations

**URL:** <https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537>\
**Category:** General Usage\
**Tags:** plotting\
**Created:** [September 18, 2026, 7:01am UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537 "2026-09-18T07:01:47Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Raducanu\_Loisy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raducanu_loisy/32/224166_2.png) [@Raducanu\_Loisy](https://discourse.julialang.org/u/Raducanu_Loisy)\
**Post date:** [September 18, 2026, 7:01am UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/1 "2026-09-18T07:01:47Z")

</div>

Recently, I’ve been experimenting with Julia for analyzing 24/7 market data. I normally work with more general data tools, so this was partly an excuse to see how comfortable Julia feels for a small time-series project.

The dataset is pretty simple: timestamp, price, and trading volume. But continuously traded markets are interesting because there is no traditional market close. Instead of looking at “trading days,” I wanted to calculate returns over fixed intervals and see how volatility changes over time.

A simplified version looks like this:

```julia-auto

```

```julia-auto
using CSV
using DataFrames
using Statistics

df = CSV.read("market_data.csv", DataFrame)

df.return = [missing; diff(log.(df.price))]

returns = collect(skipmissing(df.return))

println("Mean return: ", mean(returns))
println("Std deviation: ", std(returns))

```

From there, I started experimenting with rolling volatility. For example, using a 24-observation window:

```julia-auto

```

```julia-auto
window = 24

rolling_vol = [
    i < window ? missing :
    std(df.return[i-window+1:i])
    for i in 1:nrow(df)
]

```

The visualization part is where it becomes much easier to spot what the summary statistics hide.

For market data, I’ve been looking at public datasets as well as exchange/API documentation, mainly to understand how price and timestamp data are structured before turning them into something usable for analysis.

One thing I’m still thinking about is the best Julia-native approach for larger datasets. With a few thousand rows, almost anything works. Once the dataset becomes millions of observations, though, repeatedly calculating rolling statistics this way obviously isn’t ideal.

I’m also curious about visualization. I’ve started with basic plotting, but Makie looks interesting for exploring larger time-series datasets.

For people who regularly use Julia for this kind of work, what would you recommend for rolling-window calculations and time-series visualization?

Would you keep everything in DataFrames, or use a more specialized package once the dataset gets larger?

---

<div class="post-metadata">

**Author:** ![discourse\_ai\_spam](https://avatars.discourse-cdn.com/v4/letter/d/c68b51/32.png) [@discourse\_ai\_spam](https://discourse.julialang.org/u/discourse_ai_spam)\
**Post date:** [September 18, 2026, 7:01am UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/2 "2026-09-18T07:01:51Z")

</div>



---

<div class="post-metadata">

**Author:** ![system](https://global.discourse-cdn.com/julialang/original/3X/1/2/12829a7ba92b924d4ce81099cbf99785bee9b405.png) [@system](https://discourse.julialang.org/u/system)\
**Post date:** [September 18, 2026, 8:03am UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/3 "2026-09-18T08:03:36Z")

</div>



---

<div class="post-metadata">

**Author:** ![dcelisgarza](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dcelisgarza/32/215951_2.png) [@dcelisgarza](https://discourse.julialang.org/u/dcelisgarza)\
**Post date:** [September 18, 2026, 11:08am UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/4 "2026-09-18T11:08:26Z")

</div>

Have a look at OnlineStats.jl for streaming updates of statistics. The JuliaIO org should also have something to work with streamed inputs so you don’t have to load everything into memory.

You may also be interested in my own package PortfolioOptimisers.jl, I’m working on online portfolio selection right now. It may be of interest to you.

---

<div class="post-metadata">

**Author:** ![mahmah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mahmah/32/214326_2.png) [@mahmah](https://discourse.julialang.org/u/mahmah)\
**Post date:** [September 18, 2026, 3:58pm UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/5 "2026-09-18T15:58:15Z")

</div>

Welcome! I guess if you keep the window size fairly short this won’t be a problem for quite some time. So do you want to keep track of everything at once? Maybe you could consider focusing on the last n days or so?

---

<div class="post-metadata">

**Author:** ![technocrat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/technocrat/32/220947_2.png) [@technocrat](https://discourse.julialang.org/u/technocrat)\
**Post date:** [September 18, 2026, 6:06pm UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/6 "2026-09-18T18:06:16Z")

</div>

Visualization is the easier part: Makie.jl has all the parts, and any of the frontier models will get you over the syntax hub. [Example](https://www.linkedin.com/pulse/i-call-bullshit-richard-careaga-vgkqc/)

For financial data, there’s the problem of microstructure noise if your iknterval is too fine. [See Luo and others](https://www.sciencedirect.com/science/article/pii/S2096232022000580) due to bid-ask bounce, rounding and discreteness. Two-scale realized volatility is supposed to correct for this. [See Zhang and others](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=468798).

---

<div class="post-metadata">

**Author:** ![adranka](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adranka/32/223829_2.png) [@adranka](https://discourse.julialang.org/u/adranka)\
**Post date:** [September 18, 2026, 10:40pm UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/7 "2026-09-18T22:40:43Z")

</div>

One quick optimization is instead of doing `std(df.return[i-window+1:i])`, wrap the array indexing in a view: `std(@view(df.return[i-window+1:i]))`.

When you index with a range, it actually creates a copy of the data. Whereas in your case you only need to read the values. So using `@view` avoids the allocation and associated copy. You can check the [performance tips in the julia documentation](https://docs.julialang.org/en/v1/manual/performance-tips/#man-performance-views) for more info.

I believe that’s the approach RollingFunctions.jl takes, which you could also use.

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [September 18, 2026, 11:28pm UTC](https://discourse.julialang.org/t/exploring-24-7-market-data-with-julia-volatility-and-simple-visualizations/139537/8 "2026-09-18T23:28:12Z")

</div>

> [@Raducanu\_Loisy](#):
>
> One thing I’m still thinking about is the best Julia-native approach for larger datasets. With a few thousand rows, almost anything works. Once the dataset becomes millions of observations, though, repeatedly calculating rolling statistics this way obviously isn’t ideal.

I mean, it depends on the row is and what’s “large” to you. Given 8-byte numbers (and you could do less), a million OHLCV points is 48MB, which is pretty manageable in this age. A billion at 48GB, much less affordable. You have to figure out how much memory and storage you have to spare, and use most of that memory because transferring data back and forth with storage only slows you down. Sometimes it’ll be more obvious how much data to work with at a time; a plot on a finite screen can only represent so many assets and data points, and rendering takes time to boot. I recommend processing that CSV to a database for efficient retrieval and to save some time parsing to the numeric types you need.
