# Questions about csv（How to write to csv faster）

**URL:** <https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247>\
**Category:** General Usage\
**Tags:** question, csv\
**Created:** [October 25, 2022, 1:12pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247 "2022-10-25T13:12:58Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![zhangchunyong](https://avatars.discourse-cdn.com/v4/letter/z/d9b06d/32.png) [@zhangchunyong](https://discourse.julialang.org/u/zhangchunyong)\
**Post date:** [October 25, 2022, 1:12pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/1 "2022-10-25T13:12:58Z")

</div>

I have a huge dataframe which has 26M rows.How do I write it to a csv file faster and comsume fewer memories?

```julia
julia> pdpdpd=vcat(pd,pd,pd)
26590302×6 DataFrame
      Row │ chr position pdr discordant sum allsum
          │ String Int64 Float64 Int64 Int64 Int64
──────────┼───────────────────────────────────────────────────────
        1 │ chr1 10542 0.0 0 1 1
        2 │ chr1 10563 0.0 0 1 1
        3 │ chr1 10571 0.0 0 1 1
        4 │ chr1 10577 0.0 0 1 1
        5 │ chr1 10579 0.0 0 1 1
        6 │ chr1 10589 0.0 0 1 1
        7 │ chr1 10609 1.0 1 1 1
        8 │ chr1 10617 1.0 1 1 1
        9 │ chr1 10620 1.0 1 1 1
       10 │ chr1 10633 1.0 1 1 1
       11 │ chr1 10636 1.0 1 1 1
       12 │ chr1 10638 1.0 1 1 1
       13 │ chr1 10641 1.0 1 1 1
       14 │ chr1 10644 1.0 1 1 1
       15 │ chr1 10650 1.0 1 1 1
       16 │ chr1 10660 1.0 1 1 1
       17 │ chr1 10662 1.0 1 1 1
       18 │ chr1 10665 1.0 1 1 1
       19 │ chr1 10667 1.0 1 1 1
       20 │ chr1 10670 1.0 1 1 1
       21 │ chr1 13303 NaN 0 0 1
       22 │ chr1 13668 NaN 0 0 1
       23 │ chr1 13694 NaN 0 0 1
    ⋮ │ ⋮ ⋮ ⋮ ⋮ ⋮ ⋮
 26590281 │ chr1 248944853 1.0 1 1 1
 26590282 │ chr1 248944864 1.0 1 1 1
 26590283 │ chr1 248944868 1.0 1 1 1
 26590284 │ chr1 248944875 1.0 1 1 1
 26590285 │ chr1 248944889 1.0 1 1 1
 26590286 │ chr1 248944897 1.0 1 1 1
 26590287 │ chr1 248944900 1.0 1 1 1

julia> @time CSV.write("a.csv",pdpdpd)
 40.550157 seconds (718.14 M allocations: 17.845 GiB, 11.81% gc time, 0.28% compilation time)
"a.csv"

```

It took a lot of time,and cosumes large memeories.What should i do?Thanks.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 25, 2022, 2:03pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/2 "2022-10-25T14:03:13Z")

</div>

Does it have to be CSV (because you send it somewhere that needs a CSV) or do you just want to write the data and read it quicker? If it can be something else consider writing an Arrow file using Arrow.jl

---

<div class="post-metadata">

**Author:** ![zhangchunyong](https://avatars.discourse-cdn.com/v4/letter/z/d9b06d/32.png) [@zhangchunyong](https://discourse.julialang.org/u/zhangchunyong)\
**Post date:** [October 26, 2022, 11:57am UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/4 "2022-10-26T11:57:56Z")

</div>

> [@dlakelan](#):
>
> Does it have to be CSV (because you send it somewhere that needs a CSV) or do you just want to write the data and read it quicker? If it can be something else consider writing an Arrow file using Arrow.jl

Sorry，I just want to generate a csv.How to be faster?

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 26, 2022, 2:14pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/5 "2022-10-26T14:14:10Z")

</div>

> [@zhangchunyong](#):
>
> `, 11.81% gc time,`

So allocations and GC were not a huge bottleneck here. 12% isn’t nothing but it’s not a very attractive target for optimization.

I just back of enveloped that 26M rows at 100 bytes per row over 40 seconds gives you around 65MB/s which depending on your hard drive speed could mean hardware is the bottleneck… Probably not if you have a SSD but spinning drives might top out in that range. Or at least maybe at 2x that speed. So I think you should benchmark your hardware first before trying to make things faster.

Once you know your hardware limits, then compare your throughout to that.

A useful test is to use `dd` on Linux.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 26, 2022, 2:15pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/6 "2022-10-26T14:15:50Z")

</div>

This post was temporarily hidden by the community for possibly being off-topic, unfocused, inappropriate, or spammy.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 26, 2022, 2:22pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/7 "2022-10-26T14:22:02Z")

</div>

If you need to produce a CSV because that’s what the consumer needs… Then you have no choice. That’s why I asked about that.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [October 26, 2022, 2:39pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/8 "2022-10-26T14:39:13Z")

</div>

You CAN read an individual CSV file using parallelism (I believe it’s done already), but I’m not sure you can write them in parallel in general (or at least that Julia packages do).

Does Julia at least do that with Arrow.jl (or some other formats)? I believe some formats are meant for that and just googling, while CSV wasn’t made for that, I think it could be done (at least for your file, looking at its data, assuming it’s typical): [how to write content of df into csv file using multiprocessing in python - splunktool](https://splunktool.com/how-to-write-content-of-df-into-csv-file-using-multiprocessing-in-python)  
[Writing to CSV-file from multiple threads](https://gunnarpeipman.com/write-csv-from-multiple-threads/)  
[CSV / File writer conflict during parallel execution - KNIME Analytics Platform - KNIME Community Forum](https://forum.knime.com/t/csv-file-writer-conflict-during-parallel-execution/24637)

If you have an SSD then writing (or reading) in parallel could help, but for a spinning disk will likely kill performance.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 26, 2022, 5:00pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/9 "2022-10-26T17:00:04Z")

</div>

> [@dlakelan](#):
>
> because that’s what the consumer needs

how are they gonna use it, Excel??

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 26, 2022, 5:07pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/10 "2022-10-26T17:07:15Z")

</div>

R Python Stata SAS some custom C code, COBOL legacy systems whatever

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 26, 2022, 6:17pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/11 "2022-10-26T18:17:50Z")

</div>

except SAS and COBOL I’m sure all of them can read .arrow file.

Also, consider not adding Julia into this mix, this sounds terrible LOL

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [October 26, 2022, 6:21pm UTC](https://discourse.julialang.org/t/questions-about-csv-how-to-write-to-csv-faster/89247/12 "2022-10-26T18:21:32Z")

</div>

When you control both the writer and the reader, you can choose format. When the reader is out of your control… like it’s a different department and they don’t want to change their code, or it’s a customer or whatever… you provide what the reader needs.

I agree with you that Arrow files are better for fast reading and writing. But it’s not always an option to change formats.
