# What are the benefits of using IOBuffer()?

**URL:** <https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693>\
**Category:** General Usage\
**Tags:** question, io\
**Created:** [September 12, 2019, 3:14pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693 "2019-09-12T15:14:25Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![chbian](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chbian/32/10260_2.png) [@chbian](https://discourse.julialang.org/u/chbian)\
**Post date:** [September 12, 2019, 3:14pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/1 "2019-09-12T15:14:25Z")

</div>

Are there potential benefits of writing to an `IOBuffer()` rather than printing to a file IO directly? For example, I see the following implementation,

```julia
function append(io::IO, data::SomeCompositeType)
# append a mutable struct to an IO
end

data_str = String(take!(append(IOBuffer(),element)))
open("data_file.txt", "w") do io
    write(io, data_str)
end

```

Is the `IOBuffer` step necessary in the above example? What does it buy us? Thanks.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 12, 2019, 3:27pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/2 "2019-09-12T15:27:13Z")

</div>

> [@chbian](#):
>
> Are there potential benefits of writing to an `IOBuffer()` rather than printing to a file IO directly?

Mainly, `IOBuffer` is used in circumstances where you _don’t_ just want to output to a file, e.g. if you want to output to a string, or you want to preprocess the data before it is written.

I don’t see much point in using an intermediate `IOBuffer` if you’re just going to dump it straight into a file. (Presumably file I/O is already buffered internally.)

---

<div class="post-metadata">

**Author:** ![chbian](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chbian/32/10260_2.png) [@chbian](https://discourse.julialang.org/u/chbian)\
**Post date:** [September 12, 2019, 3:33pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/3 "2019-09-12T15:33:01Z")

</div>

Thanks.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [September 13, 2019, 12:59am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/4 "2019-09-13T00:59:56Z")

</div>

> [@chbian](#):
>
> open(“data\_file.txt”, “w”) do io write(io, data\_str) end

Maybe something like

```julia
using BufferedStreams
BufferedOutputStream(open("data_file.txt", "w")) do io 
 # do lots of output
end

```

will be beneficial if you write lots of thigns to the file.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 13, 2019, 6:25am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/5 "2019-09-13T06:25:17Z")

</div>

I thought that the OS would take care of read-ahead and buffering pretty well when writing to a file directly.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [September 13, 2019, 6:43am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/6 "2019-09-13T06:43:04Z")

</div>

> [@xiaodai](#):
>
> will be beneficial if you write lots of thigns to the file.

Presumably you did some benchmarks to give this advice, do you mind sharing them?

---

<div class="post-metadata">

**Author:** ![fredrikekre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fredrikekre/32/1688_2.png) [@fredrikekre](https://discourse.julialang.org/u/fredrikekre)\
**Post date:** [September 13, 2019, 6:56am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/7 "2019-09-13T06:56:15Z")

</div>

> [@chbian](#):
>
> Are there potential benefits of writing to an `IOBuffer()` rather than printing to a file IO directly?

Another usecase might be that you don’t want to modify files on disk until you have all the content (printing might error and you don’t wanna leave a corrupted/half file on disk). E.g. here [https://github.com/JuliaLang/Pkg.jl/blob/191b7174cf955311d27820d2d1cb2cd870fb690b/src/project.jl#L174-L176](https://github.com/JuliaLang/Pkg.jl/blob/191b7174cf955311d27820d2d1cb2cd870fb690b/src/project.jl#L174-L176) we first print to an `IOBuffer`, and only if that is successfull do we open the file on disk and print to that.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [September 13, 2019, 9:17am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/8 "2019-09-13T09:17:49Z")

</div>

> [@kristoffer.carlsson](#):
>
> Presumably you did some benchmarks to give this advice, do you mind sharing them?

```julia
using BufferedStreams

x = "id".*string.(rand(UInt16,100_000_000))

fn(x) = begin
	io = BufferedOutputStream(open("c:/data/bin.bin", "w"))
	write.(Ref(io), x)
	close(io)
end

gn(x) = begin
	io = open("c:/data/bin2.bin", "w")
	write.(Ref(io), x)
	close(io)
end

using BenchmarkTools

@btime fn($x)
@btime gn($x)

```

there you go

![image](https://global.discourse-cdn.com/julialang/original/3X/e/0/e0ef1a18b5beda646e9dc062bb37d184f34d7387.png)

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [September 13, 2019, 9:21am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/9 "2019-09-13T09:21:50Z")

</div>

> [@Tamas\_Papp](#):
>
> I thought that the OS would take care of read-ahead and buffering pretty well when writing to a file directly.

See above. Hmmm, I heard from my Rust programmer friend that it’s still better to manager your own buffer with IO. He wrote this program for me that was 10x faster than anything on the market… so I think he knows his stuff

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 13, 2019, 9:34am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/10 "2019-09-13T09:34:48Z")

</div>

Thanks, I did not know this. I can reproduce your timings on Linux.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [September 13, 2019, 9:42am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/11 "2019-09-13T09:42:19Z")

</div>

Writing to files in Rust is AFAIU completely unbuffered so using a `BufWriter` is crucial, while Julia should use libuv for buffering. If Julia didn’t use any buffering there would be a much larger difference than a factor of 2 here. It is interesting to see that there is a difference at all though, perhaps the BufferedOutputStream buffering is more efficient than the libuv buffering in this case.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [September 13, 2019, 10:51am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/12 "2019-09-13T10:51:07Z")

</div>

I can’t reproduce this in WSL on W10.

```julia
julia> include("test.jl")
  3.322 s (5226 allocations: 763.30 MiB)
  3.160 s (11 allocations: 762.94 MiB)

shell> cat test.jl
using BufferedStreams, BenchmarkTools

x = "id".*string.(rand(UInt16, 100_000_000))

f(x) = begin
    io = BufferedOutputStream(open("tmp.bin", "w"))
    write.(Ref(io), x)
    close(io)
end

g(x) = begin
    io = open("tmp2.bin", "w")
    write.(Ref(io), x)
    close(io)
end

@btime f($x)
@btime g($x)

julia> versioninfo()
Julia Version 1.2.0
Commit c6da87ff4b (2019-08-20 00:03 UTC)
Platform Info:
  OS: Linux (x86_64-linux-gnu)
  CPU: Intel(R) Core(TM) i7-6600U CPU @ 2.60GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-6.0.1 (ORCJIT, skylake)
Environment:
  JULIA_NUM_THREADS = 4

```

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [September 13, 2019, 11:06am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/13 "2019-09-13T11:06:13Z")

</div>

> [@Sukera](#):
>
> I can’t reproduce this in WSL on W10.

I am on 1.3-rc1. I notice there are more allocations in 1.2 `3.322 s (5226 allocations: 763.30 MiB)`

I just thought BufferedOutput is always good but not so sure anymore…

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [September 13, 2019, 11:14am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/14 "2019-09-13T11:14:01Z")

</div>

> [@xiaodai](#):
>
> I notice there are more allocations in 1.2 `3.322 s (5226 allocations: 763.30 MiB)`

True, but does that matter when both timing and used memory are about the same?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [September 13, 2019, 11:22am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/15 "2019-09-13T11:22:09Z")

</div>

Just curious. Because in 1.3 there are much fewer allocations. I assume each allocation would incur some distinct cost

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [September 13, 2019, 11:29am UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/16 "2019-09-13T11:29:03Z")

</div>

Yeah, true! I’m more wondering why the additional buffered version is 3s faster on your machine than the regular version…

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [September 13, 2019, 12:23pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/17 "2019-09-13T12:23:26Z")

</div>

Ah, this is on 1.3… Small writes on 1.3 are slower because they are now thread safe (and thus needs to lock for every write). `BufferedStreams` is not thread safe so it avoids the overhead of locking.

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [September 13, 2019, 1:48pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/18 "2019-09-13T13:48:58Z")

</div>

> [@chbian](#):
>
> Is the `IOBuffer` step necessary in the above example? What does it buy us? Thanks.

In some circumstances, using the `IOBuffer` is faster and more efficient to build a string from an object. For example, consider constructing a string by concatenating a `String` objects instead, then you would be creating a lot of new instances of `String` in the process. With `IOBuffer` those constructors are skipped with the hypothetical `append` because you are streaming characters into a single constructor, instead of calling multiple constructors to build up whatever `String` you are trying to get from the `element` object.

---

<div class="post-metadata">

**Author:** ![wizofe](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wizofe/32/13056_2.png) [@wizofe](https://discourse.julialang.org/u/wizofe)\
**Post date:** [January 28, 2022, 3:20pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/19 "2022-01-28T15:20:52Z")

</div>

This is still valid 2 years later. Inserting it here just for reference 🙂 My system is a Macbook Pro 2018; 2.5GHz Intel i7 and 16GB 1600 DDR3 RAM running on Big Sur 11.6.3

![CleanShot 2022-01-28 at 17.20.34](https://global.discourse-cdn.com/julialang/original/3X/d/1/d10fb34ddca4adb3fe9ea58d41d2186116f91c51.png)

---

<div class="post-metadata">

**Author:** ![sem-geologist](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sem-geologist/32/36585_2.png) [@sem-geologist](https://discourse.julialang.org/u/sem-geologist)\
**Post date:** [May 25, 2022, 1:17pm UTC](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693/20 "2022-05-25T13:17:39Z")

</div>

I am very new to Julia, so please correct me if I am wrong or use this not as intended.  
I use `IOBuffer` when I need to concatenate larger binary strings for parsing. I come from `python` and I use it in Julia similarly to `BytesIO` from `python`. I am working with readers of proprietary data files where data can be chunked into pieces and pieces spread over whole file. Parsing of such file with direct IOStream (obtained with `open`) is painful, in particularly that it can be chunked with few more layers (chunked compression and encryption). Using IOBuffer allows to reduce complexity concatenating spread chunked data into continues data representation layer before going forward with next layer of parsing.

[Next page](https://discourse.julialang.org/t/what-are-the-benefits-of-using-iobuffer/28693.md?page=2)
