# Streaming Tables to CSV

**URL:** <https://discourse.julialang.org/t/streaming-tables-to-csv/80849>\
**Category:** New to Julia\
**Tags:** csv, tables\
**Created:** [May 10, 2022, 10:47pm UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849 "2022-05-10T22:47:57Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![oconnore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oconnore/32/26654_2.png) [@oconnore](https://discourse.julialang.org/u/oconnore)\
**Post date:** [May 10, 2022, 10:47pm UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849/1 "2022-05-10T22:47:57Z")

</div>

How can I stream some rows to a CSV file (i.e. I have an iterator that creates rows, and I want those rows to be written to a file without requiring enough memory for all of the rows).

It seems like the CSV.jl library is designed to write a stream of data from an in-memory table, but it is not capable of streaming data from a streaming table:

```julia
CSV.RowWriter(table; kwargs...)

```

This table must implement Tables.jl interface, which seems like a low level interface library primarily implemented by DataFrames.jl. I can’t find a way to make a DataFrame that is generated dynamically row by row, such as when the data size is too large to fit in memory.

[https://csv.juliadata.org/stable/writing.html](https://csv.juliadata.org/stable/writing.html)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [May 10, 2022, 11:25pm UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849/2 "2022-05-10T23:25:22Z")

</div>

Is a loop that calls `print` too slow?

My impression is that a lot of the optimizations in the CSV package would be inapplicable to writing row by row, so another alternative would be to write the data in chunks: save a few thousand rows to an array or a DataTable and call CSV.write, then move on to the next chunk (re-using/over-writing the same memory buffer for every chunk).

---

<div class="post-metadata">

**Author:** ![oconnore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oconnore/32/26654_2.png) [@oconnore](https://discourse.julialang.org/u/oconnore)\
**Post date:** [May 10, 2022, 11:32pm UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849/3 "2022-05-10T23:32:34Z")

</div>

Calling print isn’t too slow — I just want to leverage the CSV library so I don’t have to write my own escaping/delimiter logic.

I guess I can implement chunking, it just seems like it increases the complexity of what I’m doing.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [May 11, 2022, 12:02am UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849/4 "2022-05-11T00:02:13Z")

</div>

> [@oconnore](#):
>
> Calling print isn’t too slow — I just want to leverage the CSV library so I don’t have to write my own escaping/delimiter logic.

The [`writedlm` function](https://docs.julialang.org/en/v1/stdlib/DelimitedFiles/#DelimitedFiles.writedlm) in the DelimitedFiles standard library accepts “an iterable collection of iterable rows”.

The main reasons to use CSV.jl instead of DelimitedFiles are performance and Tables.jl support. Since it seems you need neither of these, I would just use DelimitedFiles.

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [May 11, 2022, 12:20am UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849/5 "2022-05-11T00:20:45Z")

</div>

> [@oconnore](#):
>
> This table must implement Tables.jl interface, which seems like a low level interface library primarily implemented by DataFrames.jl.

An iterator of named tuples would be a valid (and lazy) Tables.jl table:

```julia
get_row(n) = (; col1="hi $n", col2=n^2)
table = (get_row(n) for n=1:100)

```

Definitely not used just by DataFrames!

---

<div class="post-metadata">

**Author:** ![oconnore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oconnore/32/26654_2.png) [@oconnore](https://discourse.julialang.org/u/oconnore)\
**Post date:** [May 11, 2022, 7:05pm UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849/6 "2022-05-11T19:05:10Z")

</div>

Thanks! I think @stevengj 's suggestion to use writedlm definitely work for what I’m doing, but this is exactly what I was asking about.

---

<div class="post-metadata">

**Author:** ![alvarodiaz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alvarodiaz/32/14629_2.png) [@alvarodiaz](https://discourse.julialang.org/u/alvarodiaz)\
**Post date:** [April 25, 2023, 1:15pm UTC](https://discourse.julialang.org/t/streaming-tables-to-csv/80849/7 "2023-04-25T13:15:24Z")

</div>

A better solution is to pass the dataframe to be written to the function eachrow() in the argument of either writedlm() or CSV.write():

```julia
# first with writedlm from DelimitedFiles
writedlm("path_to_save_folder/file_name.csv", eachrow(dataframe))
# now with CSV.write 
CSV.write("path_to_save_folder/file_name.csv", eachrow(dataframe))

```

Both work well but I experienced a sudden overshoot in memory for writedlm(), which did not happen for CSV.write().
