# Reading large CSV files

**URL:** <https://discourse.julialang.org/t/reading-large-csv-files/139061>\
**Category:** New to Julia\
**Tags:** dataframes, csv\
**Created:** [August 27, 2026, 10:56am UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061 "2026-08-27T10:56:06Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![mahmah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mahmah/32/214326_2.png) [@mahmah](https://discourse.julialang.org/u/mahmah)\
**Post date:** [August 27, 2026, 10:56am UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/1 "2026-08-27T10:56:06Z")

</div>

Hi,

is there any way to read a 88 GB csv file into a DataFrame? It kills julia, even with the -t auto option.

Thanks!

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [August 27, 2026, 11:12am UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/2 "2026-08-27T11:12:30Z")

</div>

How much RAM does your machine have ie `Sys.total_physical_memory()/(2^30) # GiB` ?

---

<div class="post-metadata">

**Author:** ![chrisnekarda](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisnekarda/32/207277_2.png) [@chrisnekarda](https://discourse.julialang.org/u/chrisnekarda)\
**Post date:** [August 27, 2026, 11:21am UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/3 "2026-08-27T11:21:35Z")

</div>

If you don’t need to materialize the entire dataset in memory, consider tyring with DuckDB—either directly using [DuckDB](https://duckdb.org/docs/lts/clients/julia) or using the excellent frontend [QuackIO.jl](https://github.com/JuliaAPlavin/QuackIO.jl). DuckDB allows you to load the data lazily, materializing only the subset of the data you need.

---

<div class="post-metadata">

**Author:** ![mahmah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mahmah/32/214326_2.png) [@mahmah](https://discourse.julialang.org/u/mahmah)\
**Post date:** [August 27, 2026, 11:41am UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/4 "2026-08-27T11:41:49Z")

</div>

62.449…

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [August 27, 2026, 12:11pm UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/5 "2026-08-27T12:11:56Z")

</div>

Definitely not enough RAM for typical in-memory table libraries to completely read a file of that size. Look into chrisnekarda’s suggestion.

---

<div class="post-metadata">

**Author:** ![croberts](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/croberts/32/9465_2.png) [@croberts](https://discourse.julialang.org/u/croberts)\
**Post date:** [August 27, 2026, 1:55pm UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/6 "2026-08-27T13:55:56Z")

</div>

I’ll add that the size of the CSV in memory depends on variable types.

For example, you might have a column of strings, but there are only 1000 unique strings. If these strings are long, (say 20 characters), they are going to take up an enormous amount of memory stored as strings. In contrast, if you store them as categorical arrays, then you have a length 1000 String array (which takes up approximately zero memory) and you have a huge possibly UInt16 refs array (default is UInt32 for CategoricalArrays but you can compress to UInt16). So moving from strings to UINTs reduces your memory footprint by a factor of 80 in this example.

Of course this does not solve your immediate problem – if you can change variable types to reduce memory only after loading the data with CSV.read, then you still run out of memory before loading the data.

One thing I have done is process a large number of smaller CSV files into strongly typed HDF5 files, then join the columns together to form a larger HDF5 file. This larger HDF5 gives me a smaller memory footprint (and also allows me to rapidly load the specific variables I need for an analysis).

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [August 27, 2026, 2:23pm UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/7 "2026-08-27T14:23:19Z")

</div>

CSV was not designed for data of this size.  
Allowing an 88 GB CSV file to exist is environmentally unfriendly.

---

<div class="post-metadata">

**Author:** ![stephancb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stephancb/32/14243_2.png) [@stephancb](https://discourse.julialang.org/u/stephancb)\
**Post date:** [August 28, 2026, 12:49pm UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/8 "2026-08-28T12:49:54Z")

</div>

Julia DatafFrames need to fit into memory, therefore are not very suitable for this amount of data. Available memory could be enlarged by using swap space, and this would require full control over the system, reboot, details depend on the system (Linux, Windows). The `mmap` function maps memory into file space. Afaik, Julia DataFrames do not support mmap. The csv file itself could be `mmap`ed. With an index array to the byte offsets of each line, this would be something like a “poor man’s data frame.”

SQLite can convert csv to a database table. This would not change things fundamentally, instead of millions of lines, there are then millions of rows. But SQL has tools and window functions which make working with such data more convenient. These would be available via the SQLite.jl package.

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [August 30, 2026, 3:27pm UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/9 "2026-08-30T15:27:55Z")

</div>

Another possibility is to convert the CSV file to [Arrow](https://arrow.apache.org) format, perhaps through DuckDB or [Polars](https://pola.rs). The Arrow package for Julia, which presents a `columntable` interface to an Arrow file, memory-maps the file and is able to process files larger than physical memory.

Arrow also supports categorical data in what they call `dict-encoded` columns. Often CSV files are very large because they have thousands or millions of repetitions of the same strings in a particular column.

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [August 31, 2026, 11:56am UTC](https://discourse.julialang.org/t/reading-large-csv-files/139061/10 "2026-08-31T11:56:40Z")

</div>

> [@stephancb](#):
>
> Afaik, Julia DataFrames do not support mmap.

I think the individual columns of the dataframe can be `AbstractVector`s that are saved on disk and “loaded” via `mmap`, so `DataFrame`s can work for larger than RAM data. Saving the data in the `.arrow` format (as suggested above) and loading with [Arrow.Table](https://arrow.apache.org/julia/stable/manual/#Arrow.Table), either as `Arrow.Table(file)` or `DataFrame(Arrow.Table(file))`, gives you a memory mapped table in Julia. I might personally still prefer the plain `Arrow.Table(file)`, as I’m not sure how many DataFrames operations are optimized for memory mapped columns, whereas the `Arrow.Table` is clearly designed for this use case.

You could also try `Arrow.Stream(file)` (see [here](https://arrow.apache.org/julia/stable/manual/#Arrow.Stream)) and then get an iterator of tables via `Tables.partitions` (Arrow files are subdivided into many smaller tables that can be streamed one by one, see the `.arrows` stream format).

To convert from `.csv` to `.arrow` (or actually `.arrows`, the streaming format), DuckDB is a great option, but you can do it also in julia (see [User Manual · Arrow.jl](https://arrow.apache.org/julia/stable/manual/#Arrow.write)) though I haven’t tested this latter option on large datasets.
