# Reading binary data from raw PCM files

**URL:** <https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313>\
**Category:** General Usage\
**Tags:** binaryio\
**Created:** [January 31, 2019, 2:05pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313 "2019-01-31T14:05:38Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![srikar\_sastry](https://avatars.discourse-cdn.com/v4/letter/s/85f322/32.png) [@srikar\_sastry](https://discourse.julialang.org/u/srikar_sastry)\
**Post date:** [January 31, 2019, 2:05pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/1 "2019-01-31T14:05:38Z")

</div>

Hello,

I’m new to Julia and trying to figure out a way to read raw PCM samples of audio. When encoded in WAV, I can use the wavread function from the WAV package. But, when the file is simply raw PCM without encoder/decoder metadata, it is just binary data that needs to be read as int16 in little endian format (typically).

In python, I do this:

```julia
import numpy as np
f = open(infile, mode='r+')
y = np.fromfile(f,d)
if e == 'BIG':
    y.byteswap(True)
f.close()

```

[https://docs.scipy.org/doc/numpy/reference/generated/numpy.fromfile.html](https://docs.scipy.org/doc/numpy/reference/generated/numpy.fromfile.html)

Any suggestion on how to read binary data to specified data format would be helpful.

Thanks

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [January 31, 2019, 2:34pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/2 "2019-01-31T14:34:18Z")

</div>

Assuming that in your code `d` is the number of 16-bit values in the file and `e` is a string indicating endianness, this code will do it:

```julia
f = open(infile)
y = Vector{UInt16}(undef, d)
read!(f, y)
if e == "BIG"
    for (x, i) in enumerate(y)
        y[i] = bswap(x)
    end
end
close(f)

```

However, one thing about your Python code is unclear to me: how does it know to read `int16` code units?

---

<div class="post-metadata">

**Author:** ![dawbarton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dawbarton/32/215461_2.png) [@dawbarton](https://discourse.julialang.org/u/dawbarton)\
**Post date:** [January 31, 2019, 3:30pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/3 "2019-01-31T15:30:54Z")

</div>

From the numpy docs it appears that `d` is actually the data type. (`count` is the third argument and the default is to read all the data.) As such, you just need to change one line in @StefanKarpinski’s code

```julia
y = Vector{UInt16}(undef, stat(f).size ÷ sizeof(UInt16))

```

Where here `d` is `UInt16`.

---

<div class="post-metadata">

**Author:** ![srikar\_sastry](https://avatars.discourse-cdn.com/v4/letter/s/85f322/32.png) [@srikar\_sastry](https://discourse.julialang.org/u/srikar_sastry)\
**Post date:** [January 31, 2019, 4:19pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/4 "2019-01-31T16:19:20Z")

</div>

Thanks Stefan and dawbarton. Appreciate the quick help.

Yes, d is the data type which I can pass in as ‘int16’ (or whatever the type may be).

The code seems to work for a couple of vectors I tried. From the code, am I right to assume that the read!(f,y) function will do the type conversion and convert the binary data into the type of y?

Thanks

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [January 31, 2019, 5:02pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/5 "2019-01-31T17:02:38Z")

</div>

> [@srikar\_sastry](#):
>
> From the code, am I right to assume that the read!(f,y) function will do the type conversion and convert the binary data into the type of y?

There’s no conversion really, it just reads data from the file into the array’s memory.

---

<div class="post-metadata">

**Author:** ![srikar\_sastry](https://avatars.discourse-cdn.com/v4/letter/s/85f322/32.png) [@srikar\_sastry](https://discourse.julialang.org/u/srikar_sastry)\
**Post date:** [February 1, 2019, 2:49am UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/6 "2019-02-01T02:49:47Z")

</div>

Sorry, conversion was a poor descriptor.  
What I was trying to ask was that, the read function reads the binary data one byte at a time, but when placing it in the allocated memory, it smartly places two bytes in one slot since we have defined it to be Int16. Slightly less intuitive than python (MATLAB does it the same as python too), where we specify how to interpret the binary data in the reading operation, but not a real issue. I guess I have to just get used to the Julia way.

Thanks for the quick suggestions and explanations!!

---

<div class="post-metadata">

**Author:** ![c42f](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/c42f/32/52842_2.png) [@c42f](https://discourse.julialang.org/u/c42f)\
**Post date:** [February 1, 2019, 4:33am UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/7 "2019-02-01T04:33:29Z")

</div>

The `read(io, T)` function reads `sizeof(T)` bytes from the stream in “the canonical binary representation” of type `T`, which is the same as `fromfile` but for a single value of type `T` rather than an array. (For integers the “canonical binary representation” is the platform dependent binary format of the type in memory. For `UInt16` just a pair of bytes, little endian on the typical PC but can be big endian for embedded devices.)

Julia’s standard `read` doesn’t have a “read all” option equivalent of numpy’s `count=-1`. If you’re after pure convenience, you can read the whole file as bytes, reinterpret those (pairwise) as `UInt16` and swap the endianness all in one line:

```julia
y = bswap.(reinterpret(UInt16, read(filename)))

```

The solution using `read!` is more efficient when you know the size of the data but don’t know whether to do an endian swap (for type stability of the returned array at least). To combine the solutions above into a function:

```julia
function read_pcm(file_name, T; swap_endian=false)
    y = Vector{T}(undef, filesize(file_name) ÷ sizeof(T))
    read!(file_name, y)
    if swap_endian
        y .= bswap.(y) # In place broadcast of bswap over `y` is probably the neatest way to write this.
    end
    return y
end

```

Though I’m not sure whether the `T` makes sense here, or whether all PCM files are `UInt16` in practice?

Anyway, usage would be:

```julia
read_pcm("myfile.pcm", UInt16)
read_pcm("myfile.pcm", UInt8; swap_endian=true) # 8 bit audio ?!

```

---

<div class="post-metadata">

**Author:** ![srikar\_sastry](https://avatars.discourse-cdn.com/v4/letter/s/85f322/32.png) [@srikar\_sastry](https://discourse.julialang.org/u/srikar_sastry)\
**Post date:** [February 1, 2019, 6:56pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/8 "2019-02-01T18:56:57Z")

</div>

@c42f

Thanks for the detailed explanation. The one liner does seem to create an unfamiliar type:

```julia
Base.ReinterpretArray{Int16,1,UInt8,Array{UInt8,1}}

```

I’ll stick to the initial solution with the read! function and the “for loop” for the byte swap. The more I read up, the more my preferences adapt 🙂 I recently came across some Julia training that talked about for loops being more efficient than vector operations (as opposed to MATLAB). If that is true, then in the interest of speed and efficiency, the initial for loop approach by Stefan is probably the better way (which you already mentioned is more efficient).

Btw, the endian swap is a corner case. Most audio is in little endian format and we don’t need to do the swap. And audio is always signed integers - just adding the comment for completeness, all the above solutions are valid replacing UInt16 with Int16. The focus for me was the packing of multiple bytes and copying to one memory location.

Thanks

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [February 1, 2019, 8:16pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/9 "2019-02-01T20:16:52Z")

</div>

> [@srikar\_sastry](#):
>
> If that is true, then in the interest of speed and efficiency, the initial for loop approach by Stefan is probably the better way (which you already mentioned is more efficient).

The dotted version should be equivalent and is certainly slicker and more concise.

---

<div class="post-metadata">

**Author:** ![c42f](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/c42f/32/52842_2.png) [@c42f](https://discourse.julialang.org/u/c42f)\
**Post date:** [February 1, 2019, 8:47pm UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/10 "2019-02-01T20:47:08Z")

</div>

> [@srikar\_sastry](#):
>
> The one liner does seem to create an unfamiliar type

Yes, it’s the type which is the real reason for the comment about efficiency, not the “vectorized” broadcast notation. The compiler generates less efficient code when you have a structure like

```julia
x = Type1()
if some_runtime_condition()
    x = Type2()
end
# ... do something with `x`

```

People call this [“type stability”](https://docs.julialang.org/en/v1/manual/performance-tips/index.html#Write-%22type-stable%22-functions-1).

---

<div class="post-metadata">

**Author:** ![srikar\_sastry](https://avatars.discourse-cdn.com/v4/letter/s/85f322/32.png) [@srikar\_sastry](https://discourse.julialang.org/u/srikar_sastry)\
**Post date:** [February 2, 2019, 5:44am UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/11 "2019-02-02T05:44:38Z")

</div>

> [@StefanKarpinski](#):
>
> The dotted version should be equivalent and is certainly slicker and more concise.

Got it. I quickly checked and the bswap doesn’t cause an ambiguous type. The reinterpret function does.

Thanks!

---

<div class="post-metadata">

**Author:** ![srikar\_sastry](https://avatars.discourse-cdn.com/v4/letter/s/85f322/32.png) [@srikar\_sastry](https://discourse.julialang.org/u/srikar_sastry)\
**Post date:** [February 2, 2019, 5:45am UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/12 "2019-02-02T05:45:24Z")

</div>

> [@c42f](#):
>
> The compiler generates less efficient code when you have a structure like

Got it! Thank you so much for the explanation!

---

<div class="post-metadata">

**Author:** ![srikar\_sastry](https://avatars.discourse-cdn.com/v4/letter/s/85f322/32.png) [@srikar\_sastry](https://discourse.julialang.org/u/srikar_sastry)\
**Post date:** [February 2, 2019, 5:46am UTC](https://discourse.julialang.org/t/reading-binary-data-from-raw-pcm-files/20313/13 "2019-02-02T05:46:07Z")

</div>

Thank you everyone!  
You guys answered my questions and also helped my understand new concepts! Appreciate the help!
