# Read subset of columns from a CSV?

**URL:** <https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040>\
**Category:** General Usage\
**Tags:** csv\
**Created:** [July 9, 2026, 6:33pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040 "2026-07-09T18:33:20Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 9, 2026, 6:33pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/1 "2026-07-09T18:33:20Z")

</div>

I have a dataset that contains several thousand gzipped CSV files each with a few hundred thousand lines. Using CSV.jl and `CSV.File` takes about 4 seconds, which is really good, but for the number of files till means ~3 hours just to read everything.

However, I really only need 4-5 out of the 152 columns, so I am wondering if ignoring the data I don’t need will speed this up. Is there a way to use CSV.jl (or other package) to only read certain columns of data? Or do I need to homebrew a solution?

Additional info: The columns are not fixed width, and `null` is used in place of any data with no values

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [July 9, 2026, 6:58pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/2 "2026-07-09T18:58:26Z")

</div>

Check out these `select` and `drop` [Examples · CSV.jl](https://csv.juliadata.org/stable/examples.html#select_example)

I don’t know about its impact on efficiency

---

<div class="post-metadata">

**Author:** ![technocrat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/technocrat/32/220947_2.png) [@technocrat](https://discourse.julialang.org/u/technocrat)\
**Post date:** [July 9, 2026, 7:04pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/3 "2026-07-09T19:04:22Z")

</div>

You’ve tried this?

```julia
using CSV
source = CSV.Rows("your.csv", select=[1, 3])
df = DataFrames(source)

```

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 9, 2026, 7:18pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/4 "2026-07-09T19:18:22Z")

</div>

oooh, wow, I don’t know how I missed that in the documentation! Selecting the 5 columns knocks off about 0.75 seconds, which would be about 40 minutes when I’m ready to read it all! It also really reduces the amount of memory too.

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 9, 2026, 7:28pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/5 "2026-07-09T19:28:43Z")

</div>

I had not tried `CSV.Rows`, but it is giving me weird results compared to `CSV.File`. I imagine the speed is comparable though assuming that they are using the same code to do the reading.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [July 9, 2026, 9:46pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/6 "2026-07-09T21:46:37Z")

</div>

Shameless plug: QuackIO.jl can do arbitrary column selectors or filtering with simple julian syntax.

---

<div class="post-metadata">

**Author:** ![technocrat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/technocrat/32/220947_2.png) [@technocrat](https://discourse.julialang.org/u/technocrat)\
**Post date:** [July 9, 2026, 9:54pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/7 "2026-07-09T21:54:53Z")

</div>

Another alternative is to preprocess through a `cut` or `awk` commands in the CLI

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 10, 2026, 12:16pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/8 "2026-07-10T12:16:37Z")

</div>

oooh, thanks. I’ll take a look at this too.

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 10, 2026, 12:18pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/9 "2026-07-10T12:18:07Z")

</div>

Yeah, saving out a data subset to disk crossed my mind too. I’ll have to take a look at the `man` pages for those functions.

---

<div class="post-metadata">

**Author:** ![mihalybaci](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mihalybaci/32/13528_2.png) [@mihalybaci](https://discourse.julialang.org/u/mihalybaci)\
**Post date:** [July 10, 2026, 1:01pm UTC](https://discourse.julialang.org/t/read-subset-of-columns-from-a-csv/138040/10 "2026-07-10T13:01:51Z")

</div>

ha! I just found your EnhancedCSV.jl package, which is exactly what I needed!
