# Read binary data of arbitrary dims and type

**URL:** <https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560>\
**Category:** New to Julia\
**Tags:** binaryio\
**Created:** [September 9, 2019, 1:15pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560 "2019-09-09T13:15:56Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![David\_Mc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/david_mc/32/10191_2.png) [@David\_Mc](https://discourse.julialang.org/u/David_Mc)\
**Post date:** [September 9, 2019, 1:15pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/1 "2019-09-09T13:15:56Z")

</div>

Hi, I’m new to Julia (julia 1.1.1 &nbsp; ~~1.0.4~~ ). I’m trying to read binary data and struggling to understand how dimensions are handled. There are a few other posts about this but they only address 1-D cases. Typically our data is 2 or 3-D of any one type. From other answers it almost seems that reading the data as a 1-D vector and then reshaping it is the only way to go. But that seems so… inefficient that it’s more likely I’m just missing something.

Btw, I do not have access to other packages. This is on a stand-alone system. New packages will take months to get approval for.

Here’s stevengj’s nicely concise solution with the addition of type as a parameter (works well):

> read\_bin(filename, dims, T) = read!(filename, Vector{T}(undef, dims))  
> data = read\_bin(‘myfile.img’, (260 \* 251), Float32)  
> `>`65260-element Array(Float32,1)

I’ve tried passing a list to dims in various ways and changing Vector to Array… but to no avail.  
Ideas?

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [September 9, 2019, 1:33pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/2 "2019-09-09T13:33:02Z")

</div>

This works:

```julia
x = rand(Float32, 10, 10, 10)
write("test.bin", x)
y = Array{Float32}(undef, (10, 10, 10))
open("test.bin") do io
    read!(io, y)
end
y == x

```

---

<div class="post-metadata">

**Author:** ![anon92994695](https://avatars.discourse-cdn.com/v4/letter/a/ce7236/32.png) [@anon92994695](https://discourse.julialang.org/u/anon92994695)\
**Post date:** [September 9, 2019, 1:36pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/3 "2019-09-09T13:36:02Z")

</div>

I used to use a simple scheme for this. In the filename contain the type and parse it after reading. After this you can get the file size and determine the number of elements of that type there are by the bytes size.

As far as maintaining shape - well - not sure there. You could make a custom reading function to handle that?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 9, 2019, 1:41pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/4 "2019-09-09T13:41:35Z")

</div>

> [@David\_Mc](#):
>
> But that seems so… inefficient that it’s more likely I’m just missing something.

Did you benchmark this? `reshape` should be very efficient.

---

<div class="post-metadata">

**Author:** ![David\_Mc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/david_mc/32/10191_2.png) [@David\_Mc](https://discourse.julialang.org/u/David_Mc)\
**Post date:** [September 9, 2019, 1:53pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/5 "2019-09-09T13:53:48Z")

</div>

No benchmark, just conceptually reading data in as 1 shape then rearranging it _sounds_ inefficient. 🙂  
Most of our images are a few 100 MB to a couple of GB is size. This isn’t for production so speed isn’t crucial but it’s still a concern.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 9, 2019, 1:56pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/6 "2019-09-09T13:56:03Z")

</div>

Note that output from `reshape` shares data with the input (see `?reshape`), so it is rather fast.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 9, 2019, 2:09pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/7 "2019-09-09T14:09:22Z")

</div>

> [@Tamas\_Papp](#):
>
> Note that output from `reshape` shares data with the input (see `?reshape` ), so it is rather fast.

To expand upon this, realize that a multidimensional array is stored as a consecutive sequence of numbers (a “1d array”) in memory — there’s no such thing as “multidimensional memory” in standard CPUs. All `reshape` does is to _reinterpret_ the _same_ data as a different dimensionality. There is no physical rearrangement.

---

<div class="post-metadata">

**Author:** ![David\_Mc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/david_mc/32/10191_2.png) [@David\_Mc](https://discourse.julialang.org/u/David_Mc)\
**Post date:** [September 9, 2019, 2:24pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/8 "2019-09-09T14:24:16Z")

</div>

Sweet! Thank you. I was struggling for an embarrassing amount of time with this.

Since it’s a First Steps post, here’s the whole thing.

```julia
function read_dat(filename, dims, T)
    img = Array{T}(undef, (dims))
    open(filename) do io
        read!(io, img)
    end
end

> data = read_dat("myfilename.img", (251, 260), Float32)
251x260 Array{Float32,2}:...
```

Though I’m not sure how to pass the dims. This doesn’t work:

> bands=250; lines=260; samples=440;  
> fn = “myfile.img”;  
> dtype = Float32;  
> data = read\_dat(fn, (bands,lines,samples), dtype)

Tried a few variations; ([b,l,s]), ((b,s,l))… ???

---

<div class="post-metadata">

**Author:** ![David\_Mc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/david_mc/32/10191_2.png) [@David\_Mc](https://discourse.julialang.org/u/David_Mc)\
**Post date:** [September 9, 2019, 2:28pm UTC](https://discourse.julialang.org/t/read-binary-data-of-arbitrary-dims-and-type/28560/9 "2019-09-09T14:28:56Z")

</div>

Ohh, ok. That’s good to know, not what I pictured. Thanks.
