# Writing Arrow files by column

**URL:** <https://discourse.julialang.org/t/writing-arrow-files-by-column/113955>\
**Category:** Performance\
**Created:** [May 7, 2024, 4:22pm UTC](https://discourse.julialang.org/t/writing-arrow-files-by-column/113955 "2024-05-07T16:22:14Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![djholiver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djholiver/32/50470_2.png) [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Post date:** [May 7, 2024, 4:22pm UTC](https://discourse.julialang.org/t/writing-arrow-files-by-column/113955/1 "2024-05-07T16:22:14Z")

</div>

Hi,

It appears (and I have observed) that Arrow.jl requires the entire data to be held in [memory](https://github.com/apache/arrow-julia/blob/main/src/write.jl) before flushing to disk. This is prohibitively big for my use case (500GB+) and the use of Chained Vectors for batch writing makes reading quite slow.

I am not comfortable enough with Julia IO to interpret how to write in column batches, but assume this would be possible - I’ll attempt to do so if someone could provide hints on how to start (or confirm it to be impossible).

Regards,

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [May 8, 2024, 2:54pm UTC](https://discourse.julialang.org/t/writing-arrow-files-by-column/113955/2 "2024-05-08T14:54:04Z")

</div>

Does [`Arrow.append`](https://arrow.apache.org/julia/stable/reference/#Arrow.append) do what you want?

I have tried using `Arrow.append` but, so far, without success. I’m reading data from a couple of hundred separate csv files all with the same columns and trying to consolidate them into a single Arrow file. However, differences between files in the data held in each column (in some files, some columns contain entirely `missing` data, for example) means I always get an error relating to inconsistent Arrow schemas. I haven’t yet had time to work out how to address this.
