# Any general ideas about reducing GC time involving DataFrames?

**URL:** <https://discourse.julialang.org/t/any-general-ideas-about-reducing-gc-time-involving-dataframes/104740>\
**Category:** Performance\
**Tags:** dataframes\
**Created:** [October 9, 2023, 1:35am UTC](https://discourse.julialang.org/t/any-general-ideas-about-reducing-gc-time-involving-dataframes/104740 "2023-10-09T01:35:12Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![liuyxpp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liuyxpp/32/9870_2.png) [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Post date:** [October 9, 2023, 1:35am UTC](https://discourse.julialang.org/t/any-general-ideas-about-reducing-gc-time-involving-dataframes/104740/1 "2023-10-09T01:35:12Z")

</div>

Say I have thousands of CSV files to be processed. For each, I will read the file into a Dataframe, process it, and then store the results in another DataFrame. During this process, it seems two main allocations (one for reading and one for storing) are unavoidable. I use multi-threads to process these files and the result df for each file is pushed to the Channel. I then collected all result dfs and vcat them using `take!`. It seems almost time (70%+) is spent on GC. How can I improve this situation?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [October 9, 2023, 1:41am UTC](https://discourse.julialang.org/t/any-general-ideas-about-reducing-gc-time-involving-dataframes/104740/2 "2023-10-09T01:41:04Z")

</div>

One posibility: do you need to `vcat` them all together? This will be problem dependent, but you often can do your analysis on the chunks separately. The other thing that might help is julia 1.10 (currently in beta) which adds multi-threaded GC.

---

<div class="post-metadata">

**Author:** ![liuyxpp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liuyxpp/32/9870_2.png) [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Post date:** [October 9, 2023, 1:57am UTC](https://discourse.julialang.org/t/any-general-ideas-about-reducing-gc-time-involving-dataframes/104740/3 "2023-10-09T01:57:40Z")

</div>

I have already used Julia 1.10.

Sometimes I don’t need to `vcat` them: each chunk is just a single training sample. I will try it. Thanks!
