# How do you save data in Monte Carlo simulations?

**URL:** <https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403>\
**Category:** Data\
**Tags:** question, data\
**Created:** [August 16, 2017, 3:53am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403 "2017-08-16T03:53:22Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 16, 2017, 3:53am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/1 "2017-08-16T03:53:23Z")

</div>

Let’s say I have a prior made of continuous and discrete variables:

```julia
using Distributions

a = Uniform()
b = Normal()
c = Binomial()
...

```

I generate random numbers from the prior, input them into an expensive model and the model gives back a list of arrays as a result. I repeat the process with another drawing of the input parameters and get a different set of arrays.

I have a routine that saves these “arrays” into disk, but the problem is that I want to track later on, what parameters from the prior have generated a particular array.

I am looking for a solution or file format that is portable enough, do you have recommendations? JLD, JLD2, HDF5, …

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [August 16, 2017, 4:01am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/2 "2017-08-16T04:01:56Z")

</div>

> [@juliohm](#):
>
> JLD, JLD2, HDF5

These all work. Or if you don’t care about version compatibility, you can just use serialization.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 16, 2017, 4:12am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/3 "2017-08-16T04:12:12Z")

</div>

Thanks @ChrisRackauckas, I will investigate better the pros and cons of each format…

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [August 16, 2017, 4:14am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/4 "2017-08-16T04:14:46Z")

</div>

JLD is pretty much standard. JLD2 can be faster but isn’t strictly compatible with JLD. Serialization is the fastest, but there’s no guarantee it will be the same when Julia changes versions. HDF5 is the standard for multiple programming languages so it can easily be opened in say MATLAB or Python, but I don’t think it can handle arbitrary Julia types? The others can just save your type, and when you read it back you’ll get that type back.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 16, 2017, 4:18am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/5 "2017-08-16T04:18:05Z")

</div>

That is a very good overview, thanks. I think I will go with JLD(2) for now, will see which one of the two exactly. I will very likely delete these files anyways, still in the phase of trying things out.

---

<div class="post-metadata">

**Author:** ![jkbest2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jkbest2/32/7350_2.png) [@jkbest2](https://discourse.julialang.org/u/jkbest2)\
**Post date:** [August 16, 2017, 4:53am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/6 "2017-08-16T04:53:46Z")

</div>

I’ve used HDF5 for this because it could memory map the arrays and write them to disk in the fly. This way if my run stopped early I still got partial results. This was useful for debugging. JLD(2) may be able to do this by now as well, I did this a year and a half ago. I was just storing arrays of MCMC samples though, no Julia types.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [August 16, 2017, 5:25am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/7 "2017-08-16T05:25:35Z")

</div>

Generally, I use posterior data from MCMC in two stages:

1. Convergence analysis, mostly Rhat (potential scale reduction factor) and effective sample size. For this I need scalars which may or may not correspond to variables (eg a scalar could be part of a symmetric matrix). I use vectors matrices (one for each chain). I save this using HDF5, keeping the whole sample.

2. For understanding the posterior, especially plotting and posterior predictive checks, I reconstitute the actual objects from the raw parameters (eg a matrix from a bunch of values). I frequently make a `struct` for this, otherwise a `Tuple` will do (if there are few values). I save the second half of each chain, if the mixing is good (mostly with a well-parametrized HMC) I save all values, otherwise I thin. This gives me a few thousand posterior draws. JLD handles this just fine.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [August 16, 2017, 5:48am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/8 "2017-08-16T05:48:59Z")

</div>

I have the impression JLD2 is meant to replace JLD and so be the more forward-compatible option. It should also do memory-mapping.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [August 16, 2017, 6:03am UTC](https://discourse.julialang.org/t/how-do-you-save-data-in-monte-carlo-simulations/5403/9 "2017-08-16T06:03:08Z")

</div>

> [@mkborregaard](#):
>
> I have the impression JLD2 is meant to replace JLD and so be the more forward-compatible option.

Seems like that might happen.

> <https://github.com/JuliaIO/JLD2.jl/issues/14>
>
> Since this package is now working on all platforms, I think it's about time to r…elease it. I'm not sure exactly how to do that. I see a few options:
> 
> \- Start with a separate package from JLD.jl, and merge once it's seen wider use. (The only potentially complication here is if/how the package should be integrated with FileIO.)
> \- Move code here to JLD.jl, but don't enable by default, since it hasn't received the same degree of testing. Once it sees wider use, enable it by default.
> \- Move code here to JLD.jl and immediately enable by default. I'm a little hesitant to do this, since there may be cases I'm not handling properly yet.
