# Load HDF5 file larger than memory

**URL:** <https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412>\
**Category:** New to Julia\
**Tags:** hdf5\
**Created:** [December 11, 2023, 4:28am UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412 "2023-12-11T04:28:32Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![jisutich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jisutich/32/33342_2.png) [@jisutich](https://discourse.julialang.org/u/jisutich)\
**Post date:** [December 11, 2023, 4:28am UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412/1 "2023-12-11T04:28:32Z")

</div>

Hi,

I am trying to read data from a HDF5 file which is larger than the memory of my computer. I think every time I just need part of data in the file. Is there a way to get part of groups/dataset in the file without loading the whole file? Thanks

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [December 11, 2023, 4:30am UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412/2 "2023-12-11T04:30:49Z")

</div>

Yes. That’s how it HDF5.jl usually works.

Could you share some example code demonstrating your problem?

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [December 11, 2023, 4:51am UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412/3 "2023-12-11T04:51:38Z")

</div>

Here’s a demonstration of creating a 8 GB file and then retrieving a single element.

```julia-repl
julia> using HDF5

julia> h5open("bigfile.h5", "w") do h5f
           h5f["large_dataset"] = rand(1024, 1024, 1024)
       end;

julia> g() = h5open("bigfile.h5") do h5f
           h5f["large_dataset"][1024,512,256]
       end
g (generic function with 1 method)

julia> @time g()
  0.000527 seconds (51 allocations: 2.031 KiB)
0.5066790863746067

julia> @time g()
  0.001117 seconds (51 allocations: 2.031 KiB)
0.5066790863746067

```

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [December 11, 2023, 6:07am UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412/4 "2023-12-11T06:07:23Z")

</div>

You could also potentially memory map the file, look for `mmap`.

---

<div class="post-metadata">

**Author:** ![jisutich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jisutich/32/33342_2.png) [@jisutich](https://discourse.julialang.org/u/jisutich)\
**Post date:** [December 11, 2023, 7:10pm UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412/5 "2023-12-11T19:10:45Z")

</div>

Hi Mark:

Yes this works. Before I thought `h5open` will load the whole file which is too large (the file I am using is tens of GB) but I tried what you suggested and everything is good!

---

<div class="post-metadata">

**Author:** ![jisutich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jisutich/32/33342_2.png) [@jisutich](https://discourse.julialang.org/u/jisutich)\
**Post date:** [December 11, 2023, 9:24pm UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412/6 "2023-12-11T21:24:23Z")

</div>

Hi,

I have a further question about this. Suppose in my file I have many groups (like 10000) and I want just get 1000 random group each time. Is there a way to do this?Thanks.

For an array I know I can do `sample(data,1000,replace = true)` (the reason I want replace to be true is I am trying to do something like bootstrapping). But I don’t know how to do this for groups in a hdf5 file.

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [December 12, 2023, 7:27am UTC](https://discourse.julialang.org/t/load-hdf5-file-larger-than-memory/107412/7 "2023-12-12T07:27:27Z")

</div>

You could do something like this, but this is not lazy anymore.

```julia
julia> h5open("test.h5", "w") do h5f
           h5f["r/a"] = 1
           h5f["r/b"] = 2
           h5f["r/c"] = 3
           h5f["r/d"] = 4
       end
4

julia> h5open("test.h5", "r") do h5f
           _samples = sample(keys(h5f["r"]), 1000, replace = true)
           map(_samples) do _sample
               h5f["r"][_sample][]
           end
       end
1000-element Vector{Int64}:
 4
 2
 1
 1
 2
 3
 4
 1
 2
 2
 ⋮
 4
 3
 4
 1
 2
 3
 1
 3
 3

```
