# Efficient way to read large array from binary file in slices?

**URL:** <https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847>\
**Category:** Performance\
**Tags:** question, binaryio, performance\
**Created:** [November 2, 2017, 7:20pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847 "2017-11-02T19:20:25Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![tlnagy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlnagy/32/5815_2.png) [@tlnagy](https://discourse.julialang.org/u/tlnagy)\
**Post date:** [November 2, 2017, 7:20pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/1 "2017-11-02T19:20:25Z")

</div>

I’ve written a TIFF parser ([GitHub - tlnagy/OMETIFF.jl: I/O operations for OME-TIFF files in Julia](https://github.com/tlnagy/OMETIFF.jl)) and there are still some inefficiencies that I would like to fix.

One problem is that I [allocate a large array](https://github.com/tlnagy/OMETIFF.jl/blob/8a9412c92fa7a6acde22a8e9ffae4348480374eb/src/loader.jl#L38) that will hold my multidimensional image data, but then have to also [allocate when I read in the slices of data](https://github.com/tlnagy/OMETIFF.jl/blob/8a9412c92fa7a6acde22a8e9ffae4348480374eb/src/loader.jl#L51-L52) and then copy the data from the latter into the former. I thought it would be to easy to fix by just passing a view of the larger array to `read!`, but that doesn’t work:

```julia
julia> s = open("julia_memory_blowup.tif")
IOStream(<file julia_memory_blowup.tif>)

julia> a = Array{Float64}(10, 10);

julia> read!(s, view(a, 1, :))
ERROR: MethodError: no method matching read!(::IOStream, ::SubArray{Float64,1,Array{Float64,2},Tuple{Int64,Base.Slice{Base.OneTo{Int64}}},true})
Closest candidates are:
  read!(::IO, ::BitArray) at bitarray.jl:2010
  read!(::AbstractString, ::Any) at io.jl:161
  read!(::IO, ::Array{UInt8,N} where N) at io.jl:387
  ...

```

Any thoughts of how to avoid allocating for each slice?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [November 2, 2017, 7:37pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/2 "2017-11-02T19:37:31Z")

</div>

I would just define a method for `read!` which works with subarrays. Possibly submit it as a PR. I came across a similar problem with `write` & bits types, and did that:  
[https://github.com/JuliaLang/julia/pull/24234](https://github.com/JuliaLang/julia/pull/24234)

---

<div class="post-metadata">

**Author:** ![ssfrr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ssfrr/32/3736_2.png) [@ssfrr](https://discourse.julialang.org/u/ssfrr)\
**Post date:** [November 2, 2017, 8:28pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/3 "2017-11-02T20:28:54Z")

</div>

you could also use `readbytes!`, which allows you to specify how many bytes you’d like to read. I’m not entirely sure why this is a separate function from `read!`, rather than just letting `read!` have an optional 3rd argument, but perhaps someone else knows.

---

<div class="post-metadata">

**Author:** ![ssfrr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ssfrr/32/3736_2.png) [@ssfrr](https://discourse.julialang.org/u/ssfrr)\
**Post date:** [November 2, 2017, 9:02pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/4 "2017-11-02T21:02:13Z")

</div>

btw, after doing some digging I found [here](https://github.com/JuliaLang/julia/pull/14660#issuecomment-171748241) that jeff has in the past supported merging the two functions, and also it looks like @samoconnor once [wrote a branch](https://github.com/JuliaLang/julia/compare/master...samoconnor:readbytes_branch) to do so.

---

<div class="post-metadata">

**Author:** ![tlnagy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlnagy/32/5815_2.png) [@tlnagy](https://discourse.julialang.org/u/tlnagy)\
**Post date:** [November 3, 2017, 6:50pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/5 "2017-11-03T18:50:42Z")

</div>

```julia
julia> readbytes!(s, view(a, 1, :))
ERROR: MethodError: no method matching readbytes!(::IOStream, ::SubArray{Float64,1,Array{Float64,2},Tuple{Int64,Base.Slice{Base.OneTo{Int64}}},true})
Closest candidates are:
  readbytes!(::IOStream, ::Array{UInt8,N} where N) at iostream.jl:278
  readbytes!(::IOStream, ::Array{UInt8,N} where N, ::Any; all) at iostream.jl:278
  readbytes!(::IO, ::AbstractArray{UInt8,N} where N) at io.jl:503
  ...

```

`readbytes!` similarly doesn’t work on this problem. `read!` already supports reading multiple bytes at a time. I was hoping there was an easier solution, but @Tamas_Papp’s might be best way forward. Not sure where to start though.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [November 5, 2017, 9:04am UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/6 "2017-11-05T09:04:34Z")

</div>

IMO writing performant code for `SubArray` will have to take [indexing](https://docs.julialang.org/en/latest/devdocs/subarrays/) into account.

---

<div class="post-metadata">

**Author:** ![y4lu](https://avatars.discourse-cdn.com/v4/letter/y/47e85d/32.png) [@y4lu](https://discourse.julialang.org/u/y4lu)\
**Post date:** [February 16, 2018, 7:10am UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/7 "2018-02-16T07:10:08Z")

</div>

How would mutating functions compare against non-mutating `a[1,:] = read(s, 10);`  
or possibly faster columnwise `a[:,1] = read(s, 10);` ?

---

<div class="post-metadata">

**Author:** ![henry2004y](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henry2004y/32/9284_2.png) [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Post date:** [July 26, 2019, 9:06pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/8 "2019-07-26T21:06:22Z")

</div>

For my problem in reading binary files using Julia 1.1, I still found that slices of array does not work for read!. For example,

```julia
w = Array{Float32,2}(undef,n1,nw)
read!(fileID, w[:,iw])

```

does not correctly get the values. As a workaround, I need to allocate another intermediate array to read in the correct ones.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 27, 2019, 5:50am UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/9 "2019-07-27T05:50:37Z")

</div>

> [@henry2004y](#):
>
> `read!(fileID, w[:,iw])`

You may want to try something like

```julia
read!(fileID, @view w[:,iw])

```

as your version just make a copy, reads into that, and then the copy is not accessible any more.

See [Arrays · The Julia Language](https://docs.julialang.org/en/v1/base/arrays/#Views-(SubArrays-and-other-view-types)-1)

---

<div class="post-metadata">

**Author:** ![mgkuhn](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mgkuhn/32/6276_2.png) [@mgkuhn](https://discourse.julialang.org/u/mgkuhn)\
**Post date:** [July 31, 2019, 12:49pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/10 "2019-07-31T12:49:06Z")

</div>

read!() still lacks a method for subarrays in Julia 1.1: [https://github.com/JuliaLang/julia/issues/32524](https://github.com/JuliaLang/julia/issues/32524)

---

<div class="post-metadata">

**Author:** ![tlnagy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlnagy/32/5815_2.png) [@tlnagy](https://discourse.julialang.org/u/tlnagy)\
**Post date:** [August 11, 2019, 5:41pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/11 "2019-08-11T17:41:58Z")

</div>

Thanks for creating that issue @mgkuhn. I can’t preallocate a single temporary array because the TIFF file type does not guarantee the same layout in memory for each slice\[1\]. I would really need a solution to read into a subarray where I could modify the shape of the array for each slice.

* * *

1. This is my reading of the spec. I doubt there are many TIFFs out there that would mix striped and non-striped images, but…you never know.

---

<div class="post-metadata">

**Author:** ![henry2004y](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henry2004y/32/9284_2.png) [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Post date:** [February 4, 2020, 9:30pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/12 "2020-02-04T21:30:05Z")

</div>

Will this be supported in the upcoming v1.4? I’m kind of confused by the discussion threads linked above.

---

<div class="post-metadata">

**Author:** ![tlnagy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tlnagy/32/5815_2.png) [@tlnagy](https://discourse.julialang.org/u/tlnagy)\
**Post date:** [February 4, 2020, 9:43pm UTC](https://discourse.julialang.org/t/efficient-way-to-read-large-array-from-binary-file-in-slices/6847/13 "2020-02-04T21:43:33Z")

</div>

Yup looks like it, the commit widening the signature for `read` is in [1.4.0-RC1](https://github.com/JuliaLang/julia/commit/1dbccbd110b06db8b4c17a0bed272075829a00cd):
