# Reading binary file to Julia

**URL:** <https://discourse.julialang.org/t/reading-binary-file-to-julia/103151>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [August 24, 2023, 12:23pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151 "2023-08-24T12:23:01Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![marianoarnaiz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marianoarnaiz/32/19377_2.png) [@marianoarnaiz](https://discourse.julialang.org/u/marianoarnaiz)\
**Post date:** [August 24, 2023, 12:23pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/1 "2023-08-24T12:23:01Z")

</div>

Hi all.  
I have a binary file that I need to load to julia. I have a python scrip that reads it, can anyone help me a bit to turn this to julia, it is not working for me

````julia
   # omit the first 4 values (header information) and reshape
    dtype = np.dtype([
        ("x", "<f4"),
        ("y", "<f4"),
        ("z", "<f4"),
        ("pdf", "<f4")])
    data = np.fromfile(filename, dtype=dtype)[4:]
    if coordinate_converter:
        data["x"], data["y"], data["z"] = coordinate_converter(
            data["x"], data["y"], data["z"])
    return data```
````

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [August 24, 2023, 1:02pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/2 "2023-08-24T13:02:21Z")

</div>

I don’t have some test data to be certain, but perhaps this could work:

```julia
struct Data{T}
    x::T
    y::T
    z::T
    pdf::T
end

function read_file(path; coordinate_conversion=nothing)
    data = Data{Float32}[]
    open(path, "r") do io
        
        # buffer of raw bytes for each item of data
        buffer = Vector{UInt8}(undef, sizeof(Data{Float32}))
        while !eof(io)
            readbytes!(io, buffer)
            push!(data, (reinterpret(Data{Float32}, buffer)[1]))
        end
    end
    data = @views data[5:end]
    if !isnothing(coordinate_conversion)
        data .= coordinate_conversion.(data) # convert in-place
    end

    return data
end

```

I believe `reinterpret` is little endian for the most part (see [this post](https://discourse.julialang.org/t/endian-of-reinterpret-t-array/5272)).

This isn’t the most performant implementation as it keeps resizing the array, but it should get you part of the way there.

---

<div class="post-metadata">

**Author:** ![marianoarnaiz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marianoarnaiz/32/19377_2.png) [@marianoarnaiz](https://discourse.julialang.org/u/marianoarnaiz)\
**Post date:** [August 24, 2023, 1:27pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/3 "2023-08-24T13:27:13Z")

</div>

I will try tonight.

Here is a test file if you get some time to check:

[https://github.com/marianoarnaiz/JULIA/blob/main/mine.20230703.195143.grid0.loc.scat](https://github.com/marianoarnaiz/JULIA/blob/main/mine.20230703.195143.grid0.loc.scat)

It should output 4 columns.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [August 24, 2023, 1:36pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/4 "2023-08-24T13:36:19Z")

</div>

Makes sense?

```julia
using GMT

julia> gmtconvert("mine.20230703.195143.grid0.loc.scat", binary="4f")
BoundingBox: [2.801335760031742e-41, 3.8999335765838623, 0.10017163306474686, 1.939615754657435e38, 0.0, 1.8999238014221191, 0.0, 107.7950668334961]
19992×4 GMTdataset{Float64, 2}
   Row │ col.1 col.2 col.3 col.4
       │ Float64 Float64 Float64 Float64
───────┼─────────────────────────────────────────────
     1 │ 2.80134e-41 1.93962e38 0.0 0.0
     2 │ 3.74212 0.491428 1.06699 106.243
     3 │ 3.75118 0.487598 1.07511 106.243
     4 │ 3.74868 0.499743 1.07752 106.243
     5 │ 3.74007 0.496765 1.06352 106.243
...

```

and add `header=16` to skip the text headers.

---

<div class="post-metadata">

**Author:** ![marianoarnaiz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marianoarnaiz/32/19377_2.png) [@marianoarnaiz](https://discourse.julialang.org/u/marianoarnaiz)\
**Post date:** [August 24, 2023, 2:18pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/5 "2023-08-24T14:18:21Z")

</div>

Hi Joaquim!. The last column should be between 0 and 1 if the manual es correct. Quote from the quotes manual:

> - **Scatter file** (_Binary_ , _FileExtension=_ `*.scat` )The Scatter file contains the _x,y,z_ locations and PDF value of each sample of the location [PDF](http://alomax.free.fr/nlloc/soft7.00/NLLoc.html#_inversion_). The number of samples to save is specified in the `LOCSEARCH` statement in the NLLoc Statements section of the Input Control File.Header: (_required_ ) one `integer and 3` `float` values`nSamples dummy dummy dummy` **Fields:**  
> `nSamples` (`integer` )  
> number of PDF samples in the following buffer  
> `dummy` (`float` )  
> unusedBuffer: (_required_ ) Sequence of four `float` values for each PDF sample`x(N), y(N), z(N), pdf(N) (N = 0, nSamples - 1)` **Fields:**  
> _`x(N), y(N), z(N)`_ (`float` )  
> _ **Non-GLOBAL** _: x, y and z location of the sample in kilometers relative to the geographic origin.  
> _ **GLOBAL:** _ longitude and latitude in degrees, and z location in kilometers of the sample.  
> `pdf(N)` (`float` )  
> PDF value of the sample, normalized so that the volume integral over the corresponding search grid of the PDF = 1.0

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [August 24, 2023, 3:04pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/6 "2023-08-24T15:04:49Z")

</div>

Not at computer right now. See the -bi documentation of GMT. It’s pretty flexible in column data type.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [August 24, 2023, 9:43pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/7 "2023-08-24T21:43:21Z")

</div>

Suspect all values shown are good. When reading binaries and use wrong data types all values are garbage, but in this case they are all reasonable.

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [August 24, 2023, 10:20pm UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/8 "2023-08-24T22:20:01Z")

</div>

> [@marianoarnaiz](#):
>
> PDF value of the sample, normalized so that the volume integral over the corresponding search grid of the PDF = 1.0

Doesn’t this imply the integral sums to 1? Which means that individual PDF values can be larger than one if the volume element is small.

My code produces nonsense output, so probably isn’t right. I have a feeling @joa-quim’s solution is the best one.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [August 25, 2023, 5:51am UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/9 "2023-08-25T05:51:45Z")

</div>

> [@marianoarnaiz](#):
>
> Hi Joaquim!. The last column should be between 0 and 1 if the manual es correct. Quote from the quotes manual:

Does it or does it not match the Python implementation you started off with?

---

<div class="post-metadata">

**Author:** ![marianoarnaiz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marianoarnaiz/32/19377_2.png) [@marianoarnaiz](https://discourse.julialang.org/u/marianoarnaiz)\
**Post date:** [August 25, 2023, 10:49am UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/10 "2023-08-25T10:49:09Z")

</div>

I´ve been trying to map @joa-quim values and they look right from this perspective.  
I will try something else this afternoon and report back

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [August 25, 2023, 10:53am UTC](https://discourse.julialang.org/t/reading-binary-file-to-julia/103151/11 "2023-08-25T10:53:09Z")

</div>

Also see StructIO.jl:

> **[GitHub - JuliaIO/StructIO.jl: Experimental new implementation of...](https://github.com/JuliaIO/StructIO.jl)**
>
> Experimental new implementation of StrPack.jl-like functionality - GitHub - JuliaIO/StructIO.jl: Experimental new implementation of StrPack.jl-like functionality
