# Reading a large array from an HDF5 file

**URL:** <https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983>\
**Category:** New to Julia\
**Tags:** hdf5\
**Created:** [May 8, 2019, 3:45am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983 "2019-05-08T03:45:33Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![huchiayu0517](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/huchiayu0517/32/8202_2.png) [@huchiayu0517](https://discourse.julialang.org/u/huchiayu0517)\
**Post date:** [May 8, 2019, 3:45am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/1 "2019-05-08T03:45:33Z")

</div>

I’m trying to read a large 3D array from an HDF5 file. If I just naively read a 3D array, I found that every time I try to access the element (e.g. in a for-loop), I see a lot of unexpected allocations and therefore the low speed (see the read1 function below). Instead, I have to first allocate an undef 3D array and then assign the value with `.=` (the read2 function) to avoid allocations when referencing the elements, which is a bit clumsy. The code example:

```julia
using BenchmarkTools
using HDF5
const N = 1000000
A = rand(3,3,N);
h5write("data_3darray.h5", "group", A);

function read1()
    data1 = h5read("data_3darray.h5", "group")
    for n in 1:N, i in 1:3, j in 1:3
        data1[j,i,n] += 0.0
    end
    return data1
end

function read2()
    data2 = Array{Float64,3}(undef, 3, 3, N)
    data2 .= h5read("data_3darray.h5", "group")
    for n in 1:N, i in 1:3, j in 1:3
        data2[j,i,n] += 0.0
    end
    return data2
end
@btime data=read1() #2.395 s (35990858 allocations: 617.84 MiB)
@btime data=read2() #88.712 ms (60 allocations: 137.33 MiB)

```

Why does it allocate so much memory during the loop in the first case? Is there a better way to read a large array from an HDF5 file?

---

<div class="post-metadata">

**Author:** ![tkoolen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkoolen/32/1603_2.png) [@tkoolen](https://discourse.julialang.org/u/tkoolen)\
**Post date:** [May 8, 2019, 4:13am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/2 "2019-05-08T04:13:12Z")

</div>

Yep, so `@code_warntype read1()` shows clearly that the inferred return type of `h5read` is `Any`, and so `data1` is also inferred to be of type `Any`. You can either explicitly annotate the type ([Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/#Annotate-values-taken-from-untyped-locations-1)),

```julia
data1::Array{Float64, 3} = h5read("data_3darray.h5", "group")

```

or

```julia
data1 = convert(Array{Float64, 3}, h5read("data_3darray.h5", "group"))

```

or introduce a function barrier ([Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/#kernel-functions-1)).

---

<div class="post-metadata">

**Author:** ![bicycle1885](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bicycle1885/32/107_2.png) [@bicycle1885](https://discourse.julialang.org/u/bicycle1885)\
**Post date:** [May 8, 2019, 4:24am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/3 "2019-05-08T04:24:28Z")

</div>

HD5 is not designed to read data element by element. Internally, it does data chunking and elements are read from/written to a file chunk-wise. See here for more explanation: [https://portal.hdfgroup.org/display/HDF5/Chunking+in+HDF5](https://portal.hdfgroup.org/display/HDF5/Chunking+in+HDF5).

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [May 8, 2019, 6:34am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/4 "2019-05-08T06:34:42Z")

</div>

If possible, I would suggest a different file format. We have spent quite a lot of time searching for a reliable and fast fileformat (you can find the discussions on this forum) and arrived to FlatBuffers. So if you can, I would suggest those. I have a good experience with them.

Tomas

---

<div class="post-metadata">

**Author:** ![tobias.knopp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tobias.knopp/32/7551_2.png) [@tobias.knopp](https://discourse.julialang.org/u/tobias.knopp)\
**Post date:** [May 8, 2019, 8:08am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/5 "2019-05-08T08:08:00Z")

</div>

HDF5 is fast and reliable. One can even `mmap` the array in question with `readmmap` as is outlined in this example:

[https://github.com/JuliaIO/HDF5.jl/blob/master/test/mmap.jl](https://github.com/JuliaIO/HDF5.jl/blob/master/test/mmap.jl)

It is clear that `read1` has two issues. Type instability and reading / writing scalar values from a file.

---

<div class="post-metadata">

**Author:** ![tkoolen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkoolen/32/1603_2.png) [@tkoolen](https://discourse.julialang.org/u/tkoolen)\
**Post date:** [May 8, 2019, 3:50pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/6 "2019-05-08T15:50:58Z")

</div>

> [@bicycle1885](#):
>
> HD5 is not designed to read data element by element.

> [@tobias.knopp](#):
>
> It is clear that `read1` has two issues. Type instability and reading / writing scalar values from a file.

Why do you both say this? The single `h5read` call in `read1` reads the whole array, I’d think?

---

<div class="post-metadata">

**Author:** ![JK\_loo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jk_loo/32/10822_2.png) [@JK\_loo](https://discourse.julialang.org/u/JK_loo)\
**Post date:** [March 4, 2020, 10:04pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/7 "2020-03-04T22:04:46Z")

</div>

Hi,  
i’m searching for a fast fileformt which can be converted from hdf5 file. Is it Flatbuffer? If yes, could you let me know the where in this Forum i can find about it?

JK

---

<div class="post-metadata">

**Author:** ![purplishrock](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/purplishrock/32/13451_2.png) [@purplishrock](https://discourse.julialang.org/u/purplishrock)\
**Post date:** [March 5, 2020, 12:34am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/8 "2020-03-05T00:34:53Z")

</div>

The search engine provides:

[https://github.com/JuliaData/FlatBuffers.jl](https://github.com/JuliaData/FlatBuffers.jl)

as to whether it’s faster than hdf5, i don’t think that question has been answered yet, most likely it depends on precisely what you are doing.

---

<div class="post-metadata">

**Author:** ![met-j](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/met-j/32/21919_2.png) [@met-j](https://discourse.julialang.org/u/met-j)\
**Post date:** [March 5, 2020, 1:02pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/9 "2020-03-05T13:02:16Z")

</div>

I am reading HDF5 files in the rage of 50 GB with out any issues. In fact I am reading 3d arrays of integers on the lines of:

```julia
using HDF5

function loadData(filename)
    fid = HDF5.h5open(filename, "r")
    obj = fid["key_lvl_1"]
    metadata = read(obj)
    close(fid)
    data = metadata["key_lvl_2"]["key_lvl_3"]["key_lvl_4"][:,:,:];
return data, metadata
end

@time (data,metadata) = loadData(filename); # 2.192142 seconds (586 allocations: 1018.752 MiB, 4.13% gc time)

```

Sure, the return of data & metadata is questionable. With a 54 GB file ready to use, I am very happy.

---

<div class="post-metadata">

**Author:** ![purplishrock](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/purplishrock/32/13451_2.png) [@purplishrock](https://discourse.julialang.org/u/purplishrock)\
**Post date:** [March 6, 2020, 11:16pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/10 "2020-03-06T23:16:45Z")

</div>

> [@met-j](#):
>
> Sure, the return of data & metadata is questionable.

Could you expand on what you mean by that ? My first reading of it make me think that the returned data was questionable, which would make HDF5 a poor choice indeed 😆

---

<div class="post-metadata">

**Author:** ![met-j](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/met-j/32/21919_2.png) [@met-j](https://discourse.julialang.org/u/met-j)\
**Post date:** [March 7, 2020, 6:53am UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/11 "2020-03-07T06:53:11Z")

</div>

All of `data` is contained in `metadata`.  
Would you really need both? Probably not. Anyhow, as it come at very little extra cost, I enjoy beeing able to choose if I just want my central `Array{Int, 3}` or all the metadata as well by simply choosing between:

```julia
data = loadData(filename)

```

and

```julia
(data, metadata) = loadData(filename)

```

---

<div class="post-metadata">

**Author:** ![mgiugliano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mgiugliano/32/5818_2.png) [@mgiugliano](https://discourse.julialang.org/u/mgiugliano)\
**Post date:** [February 6, 2023, 5:54pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/12 "2023-02-06T17:54:59Z")

</div>

hdf5 newbie here.

Is it correct that you are reading ALL the file into the memory? Is it desirable?  
How could one only load a chunk of data?

---

<div class="post-metadata">

**Author:** ![fft](https://avatars.discourse-cdn.com/v4/letter/f/e8c25b/32.png) [@fft](https://discourse.julialang.org/u/fft)\
**Post date:** [February 6, 2023, 8:22pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/13 "2023-02-06T20:22:13Z")

</div>

You can load all of the data or a subset of the data. Check out the documentation [Home · HDF5.jl](https://juliaio.github.io/HDF5.jl/stable/)

---

<div class="post-metadata">

**Author:** ![mgiugliano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mgiugliano/32/5818_2.png) [@mgiugliano](https://discourse.julialang.org/u/mgiugliano)\
**Post date:** [February 6, 2023, 9:26pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/14 "2023-02-06T21:26:06Z")

</div>

Thank you. I personally do find the documentation very poor, but it must be me, although people even pointed me to the documentation of the HDF5 Python library jus few weeks ago ☹

---

<div class="post-metadata">

**Author:** ![fft](https://avatars.discourse-cdn.com/v4/letter/f/e8c25b/32.png) [@fft](https://discourse.julialang.org/u/fft)\
**Post date:** [February 6, 2023, 9:37pm UTC](https://discourse.julialang.org/t/reading-a-large-array-from-an-hdf5-file/23983/15 "2023-02-06T21:37:14Z")

</div>

[https://juliaio.github.io/HDF5.jl/stable/#Reading-and-writing-data](https://juliaio.github.io/HDF5.jl/stable/#Reading-and-writing-data)

When you read the data just add add the slice for the subset of the data you want.

```julia

Asub = dset[2:3, 1:3]

```
