# Writing Arrow record batch requires a lot of RAM

**URL:** <https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313>\
**Category:** Data\
**Tags:** arrow\
**Created:** [September 17, 2021, 7:03am UTC](https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313 "2021-09-17T07:03:27Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![kobusherbst](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kobusherbst/32/19682_2.png) [@kobusherbst](https://discourse.julialang.org/u/kobusherbst)\
**Post date:** [September 17, 2021, 7:03am UTC](https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313/1 "2021-09-17T07:03:27Z")

</div>

I am using this function to save a batch Arrow files into a single Arrow file with record batches:

```julia
function combinebatches(path::String, file::String, batches)
    files = Array{String,1}()
    for i = 1:batches
        push!(files,joinpath(path, "$(file)$(i).arrow"))
    end
    arrow_parts = Tables.partitioner(Arrow.Table, files)
    open(joinpath(path, "$(file)_batched.arrow"), "w") do io
        Arrow.write(io, arrow_parts, compress=:zstd)
    end
    #delete chunks
    for i = 1:batches
        rm(joinpath(path, "$(file)$(i).arrow"))
    end
    return nothing
end #combinebatches

```

When saving about 25 uncompressed arrow files with an average size of 1.8 GB each, required a RAM size of around 50GB. As I am partitioning my data to save RAM, this is creating a problem with the high RAM usage (in this case \> 50GB). I was expecting that one would not require more than the maximum size of each single file (or perhaps number of threads x file size) to save the record batch.

`JULIA_NUM_THREADS` is set to 4.

I am using Arrow v1.6.2

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [September 17, 2021, 9:01am UTC](https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313/2 "2021-09-17T09:01:02Z")

</div>

CC @quinnj

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [September 18, 2021, 3:02am UTC](https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313/3 "2021-09-18T03:02:16Z")

</div>

This is certainly unexpected as you’ve described things. A few things that come to mind that we should look into:

- Although we’ve made _some_ efforts to ensure writing an `Arrow.Table` out avoids making copies, it’s possible we’re missing something here and things are getting unnecessarily converted or materialized
- We’ve seen a potentially related issue in [CSV.jl](https://github.com/JuliaData/CSV.jl/issues/850) where large chunks of memory seem to get “stuck” and not reclaimed by the GC. That bug is still under investigation.

---

<div class="post-metadata">

**Author:** ![kobusherbst](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kobusherbst/32/19682_2.png) [@kobusherbst](https://discourse.julialang.org/u/kobusherbst)\
**Post date:** [September 18, 2021, 4:55am UTC](https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313/4 "2021-09-18T04:55:20Z")

</div>

That would be great, although I am processing close to a billion records in my DataFrames-based workflow (longitudinal population studies - a record for each day in the life of some 150 000 individuals), I have managed to keep my RAM usage generally below 64GB using the great partition features of Arrow, coupled with DataFrame’s amazing performance, but reducing the time from 8.5 days using a Java-based workflow to 1 day using Julia.

Happy to know that I am not doing something stupid in my code, also happy to test any change on my workflow if that would help.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [September 18, 2021, 10:13am UTC](https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313/5 "2021-09-18T10:13:38Z")

</div>

> [@kobusherbst](#):
>
> using a Java-based workflow

do you happen to use Spark?

---

<div class="post-metadata">

**Author:** ![kobusherbst](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kobusherbst/32/19682_2.png) [@kobusherbst](https://discourse.julialang.org/u/kobusherbst)\
**Post date:** [September 19, 2021, 4:31am UTC](https://discourse.julialang.org/t/writing-arrow-record-batch-requires-a-lot-of-ram/68313/6 "2021-09-19T04:31:25Z")

</div>

No, Pentaho.
