# Saving and updating in-memory HDF5 files?

**URL:** <https://discourse.julialang.org/t/saving-and-updating-in-memory-hdf5-files/112395>\
**Category:** Performance\
**Tags:** hdf5\
**Created:** [April 2, 2024, 1:01am UTC](https://discourse.julialang.org/t/saving-and-updating-in-memory-hdf5-files/112395 "2024-04-02T01:01:14Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ahmed\_Salih](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ahmed_salih/32/206579_2.png) [@Ahmed\_Salih](https://discourse.julialang.org/u/Ahmed_Salih)\
**Post date:** [April 2, 2024, 1:01am UTC](https://discourse.julialang.org/t/saving-and-updating-in-memory-hdf5-files/112395/1 "2024-04-02T01:01:14Z")

</div>

Hello!

Reading the documentation of HDF5.jl ([Home · HDF5.jl](https://juliaio.github.io/HDF5.jl/stable/#In-memory-HDF5-files)) I see how it is possible to have an in-memory hdf5 file. I was wondering if;

1. Can I update a dataset in an hdf in place?
2. Can I save a copy of the current in-memory hdf5 file to disk?

I have a use case in which over N timesteps I have a fixed number of data points I want to save. I thought that HDF5 would be perfect for this, especially if one could do in-place.

Kind regards

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [April 2, 2024, 2:01am UTC](https://discourse.julialang.org/t/saving-and-updating-in-memory-hdf5-files/112395/2 "2024-04-02T02:01:24Z")

</div>

You should first familiarize yourself with the upstream documentation.

[https://docs.hdfgroup.org/hdf5/develop/\_h5\_f\_\_u\_g.html#subsubsec\_file\_alternate\_drivers\_mem](https://docs.hdfgroup.org/hdf5/develop/_h5_f__u_g.html#subsubsec_file_alternate_drivers_mem)

> **[HDF5 File Image Operations](https://portal.hdfgroup.org/documentation/hdf5-docs/advanced_topics/file_image_ops.html)**
>
> Ensuring long-term access and usability of HDF data and supporting users of HDF technologies

The example shows you how you can obtain a `Vector{UInt8}` of the file image. If you write that out to disk, you can open it just like a regular HDF5 file.

This is a highly specialized operation, and I’m not sure about your specific use case. Even for a HDF5 on disk, you can update a subsection of a dataset in-place.

Before we get into a XY problem situation, could you explain what you are trying to do or optimize?

---

<div class="post-metadata">

**Author:** ![Ahmed\_Salih](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ahmed_salih/32/206579_2.png) [@Ahmed\_Salih](https://discourse.julialang.org/u/Ahmed_Salih)\
**Post date:** [April 2, 2024, 2:16am UTC](https://discourse.julialang.org/t/saving-and-updating-in-memory-hdf5-files/112395/3 "2024-04-02T02:16:26Z")

</div>

Thanks!

What I want do in reality is saving datafiles at each output of my simulation. In my case since I know the number of data points before hand and that it never changes, I thought I could “preallocate” a HDF5 in memory, efficiently overwrite it in place, make a copy and save to disk. I thought by doing so that I could get an extra speed up / reduce allocations.

Kind regards

---

<div class="post-metadata">

**Author:** ![Ahmed\_Salih](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ahmed_salih/32/206579_2.png) [@Ahmed\_Salih](https://discourse.julialang.org/u/Ahmed_Salih)\
**Post date:** [April 2, 2024, 4:52pm UTC](https://discourse.julialang.org/t/saving-and-updating-in-memory-hdf5-files/112395/4 "2024-04-02T16:52:46Z")

</div>

I found that a simple solution for fast file writing using HDF5.jl, without resorting to the complexity above, is to save all data into one single file. In pseudo-code:

```julia
function SaveHDF5!(fid::HDF5.File, group_name, variable_names, args...)
    create_group(fid, group_name)
    if !isnothing(args)
        for i in eachindex(args)
            arg = args[i]
            var_name = variable_names[i]
            fid[group_name][var_name] = arg
        end
    end
end

```

Where I pass in a consistent file id, update the group name and the variable input as needed. This would give me the following timings:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/0/1/017da1d2837d160f39b68752903e60025dcfc1bf.png)

The reason for writing to one file is that surprisingly, atleast on Windows, `close` the HDF5.File is actually a bottle neck.

All in all, I am pretty pleased, writing to HDF5 in this way is about 10x faster than writing to individual `.vtp` files as I did in the past.

Just sharing my findings here, perhaps someone can benefit in the future.
