# Write data to Arrow file row by row

**URL:** <https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165>\
**Category:** General Usage\
**Tags:** arrow\
**Created:** [April 13, 2021, 7:40am UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165 "2021-04-13T07:40:12Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [April 13, 2021, 7:40am UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/1 "2021-04-13T07:40:12Z")

</div>

I have a large amount of data I am loading from an API and I would like to write the output of each call to a “row” of a file in order not to keep it in RAM/ in case the program errors. I have been using the Arrow.jl package for data reading/writing needs recently and really like it. However, I have not found a way to append rows to an Arrow file and was wondering if that is possible?

Thanks!

Here is my current example that does not work:

```julia
using Arrow, Tables

dat = [(a=1, b=2), (a=3,b=4)]
open("test.feather", "a+") do io
    for row in dat
        Arrow.write(io, [row])
    end
end

Arrow.Table("test.feather") |> Tables.rowtable

```

---

<div class="post-metadata">

**Author:** ![davibarreira](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davibarreira/32/20270_2.png) [@davibarreira](https://discourse.julialang.org/u/davibarreira)\
**Post date:** [May 30, 2022, 8:38pm UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/2 "2022-05-30T20:38:47Z")

</div>

Did you find an answer for this?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [May 31, 2022, 4:33am UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/3 "2022-05-31T04:33:29Z")

</div>

[https://stackoverflow.com/questions/66388141/how-to-append-a-dataframe-to-an-existing-apache-arrow-file-on-disk](https://stackoverflow.com/questions/66388141/how-to-append-a-dataframe-to-an-existing-apache-arrow-file-on-disk)

I think the answer is no

---

<div class="post-metadata">

**Author:** ![danielw2904](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielw2904/32/10890_2.png) [@danielw2904](https://discourse.julialang.org/u/danielw2904)\
**Post date:** [May 31, 2022, 4:53am UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/4 "2022-05-31T04:53:07Z")

</div>

I ended up using the jsonlines format instead. If the file ends up too large I use it with GZip.jl

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [May 31, 2022, 10:42am UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/5 "2022-05-31T10:42:04Z")

</div>

I don’t see it in the docs, but

- [https://github.com/apache/arrow-julia/pull/160](https://github.com/apache/arrow-julia/pull/160) and
- [refactor Arrow.write to support incremental writes by baumgold · Pull Request #277 · apache/arrow-julia · GitHub](https://github.com/apache/arrow-julia/pull/277)

added some support for this.

---

<div class="post-metadata">

**Author:** ![phantom](https://avatars.discourse-cdn.com/v4/letter/p/e0b2c6/32.png) [@phantom](https://discourse.julialang.org/u/phantom)\
**Post date:** [April 4, 2023, 1:04am UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/6 "2023-04-04T01:04:26Z")

</div>

Just in case it is helpful to anyone. As noted above I think this can be done now if you put the named tuples in a vector.

so something like

```julia
dat = [[(a=1, b=2)], [(a=3,b=4)]]
for d = dat
     Arrow.append("newfilepath",d)
end

```

bkamins has a good example and explanation of `Arrow.append`[here](https://discourse.julialang.org/t/arrow-stream-usage-clarification/89508/6).

Also if you are appending to a pre-existing `Arrow` file, it must be written like `Arrow.write(filename::String, tbl; file=false)` with the `file` keyword argument set to `false` for `append` to work. See [API Reference · Arrow.jl](https://arrow.juliadata.org/stable/reference/#Arrow.append)

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [April 4, 2023, 1:33am UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/7 "2023-04-04T01:33:29Z")

</div>

> [@phantom](#):
>
> `Arrow.append("newfilepath",d)`

I think this will produce a file with excessive metadata. You should minimize the time `append` happens because each time a new “batch” is written to the file which implies more metadata

---

<div class="post-metadata">

**Author:** ![phantom](https://avatars.discourse-cdn.com/v4/letter/p/e0b2c6/32.png) [@phantom](https://discourse.julialang.org/u/phantom)\
**Post date:** [April 7, 2023, 9:31pm UTC](https://discourse.julialang.org/t/write-data-to-arrow-file-row-by-row/59165/8 "2023-04-07T21:31:06Z")

</div>

Thanks so much for pointing this out! Sorry been out with a stomach bug so just saw this.
