# Why DataFrames v.0.21.2 (julia v1.4.2) requires more memory than the previous version

**URL:** <https://discourse.julialang.org/t/why-dataframes-v-0-21-2-julia-v1-4-2-requires-more-memory-than-the-previous-version/41171>\
**Category:** Performance\
**Tags:** dataframes\
**Created:** [June 10, 2020, 11:55pm UTC](https://discourse.julialang.org/t/why-dataframes-v-0-21-2-julia-v1-4-2-requires-more-memory-than-the-previous-version/41171 "2020-06-10T23:55:42Z")\
**Posts on this page:** 3\
**Page:** 2

<div class="post-metadata">

**Author:** ![tamasgal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamasgal/32/27946_2.png) [@tamasgal](https://discourse.julialang.org/u/tamasgal)\
**Post date:** [June 15, 2020, 9:40pm UTC](https://discourse.julialang.org/t/why-dataframes-v-0-21-2-julia-v1-4-2-requires-more-memory-than-the-previous-version/41171/21 "2020-06-15T21:40:29Z")

</div>

If it’s only 0s and 1s, which might end up being interpreted as `Int64`, then it’s no wonder that the size in memory blows up.

Such a row in your CSV looks like this (if I understood correctly):

`0,1,0,1,1,0,...`

Which means you have ~2 bytes per value. `Int64` has a size of 8 bytes, so you will occupy 4x more “space” in ram.

I’d recommend using some kind of in-place data conversion to `Bool` or `UInt8`. You need to help CSV/DataFrames and tell them the exact types. They cannot guess it unless they read the whole data, but that’s already too late…

---

<div class="post-metadata">

**Author:** ![freeman](https://avatars.discourse-cdn.com/v4/letter/f/ec9cab/32.png) [@freeman](https://discourse.julialang.org/u/freeman)\
**Post date:** [June 16, 2020, 7:07pm UTC](https://discourse.julialang.org/t/why-dataframes-v-0-21-2-julia-v1-4-2-requires-more-memory-than-the-previous-version/41171/22 "2020-06-16T19:07:43Z")

</div>

How about `BitArray`?

[https://web.mit.edu/julia\_v0.6.0/julia/share/doc/julia/html/en/stdlib/arrays.html#BitArrays-1](https://web.mit.edu/julia_v0.6.0/julia/share/doc/julia/html/en/stdlib/arrays.html#BitArrays-1)

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [June 29, 2020, 2:45am UTC](https://discourse.julialang.org/t/why-dataframes-v-0-21-2-julia-v1-4-2-requires-more-memory-than-the-previous-version/41171/23 "2020-06-29T02:45:50Z")

</div>

Note that CSV.jl’s memory footprint has been fixed in the latest 0.7 release.

One idea others have mentioned is reading the data in as `Bool`, which you could do by passing `truestrings=["1"], falsestrings=["0"]`, in case you wanted to go that route.

[Previous page](https://discourse.julialang.org/t/why-dataframes-v-0-21-2-julia-v1-4-2-requires-more-memory-than-the-previous-version/41171.md?page=1)
