# Processing csv's in parallel

**URL:** <https://discourse.julialang.org/t/processing-csvs-in-parallel/8791>\
**Category:** General Usage\
**Tags:** question\
**Created:** [February 3, 2018, 7:31am UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791 "2018-02-03T07:31:21Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![chasecb](https://avatars.discourse-cdn.com/v4/letter/c/c67d28/32.png) [@chasecb](https://discourse.julialang.org/u/chasecb)\
**Post date:** [February 3, 2018, 7:31am UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/1 "2018-02-03T07:31:21Z")

</div>

I’m basically trying to serialize csvs in parallel. The parallel stuff seems to have changed a bit from Julia’s more sophomoric (to be as polite as I can) versions but I could never get it working back then either. Let’s say I had a list of files [1.csv, 2.csv,3.csv] where I just wanted to do a readable() and a writeable() using DataFrames what would be the most efficient way to achieve that (seems pmap or @parallel could both work so wondering what approach is better and why)? Are there any good resources for this type of thing? Many many thanks in advance.  
Chase CB

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [February 3, 2018, 8:09pm UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/2 "2018-02-03T20:09:40Z")

</div>

Whenever you’re performing operations that include I/O to/from a single disk, the bottleneck is the I/O, and that can’t be sped up by parallelizing.

If this is a limiting factor in a production application, use faster disks (PCIe SSDs and/or RAID).

 ![](https://global.discourse-cdn.com/julialang/original/3X/d/d/ddf92e087e0ab915b7e2fb7a3f5c051b53e2eede.png)

---

<div class="post-metadata">

**Author:** ![chasecb](https://avatars.discourse-cdn.com/v4/letter/c/c67d28/32.png) [@chasecb](https://discourse.julialang.org/u/chasecb)\
**Post date:** [February 3, 2018, 8:46pm UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/3 "2018-02-03T20:46:27Z")

</div>

Sorry should have been clearer there is an intermediary data processing  
step where the paralleization might pay some dividends but I get what you  
are saying about I/O.

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [February 3, 2018, 9:09pm UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/4 "2018-02-03T21:09:59Z")

</div>

Have you profiled your script’s execution yet? Unless the time spent on data processing is on the same order as the I/O time, parallelizing will only buy you a minuscule speed-up.

If it’s on the same order, I would read the files into memory serially, process in parallel, and write the results back serially, but the speed-up will only be ~25-50% at best.

---

<div class="post-metadata">

**Author:** ![chasecb](https://avatars.discourse-cdn.com/v4/letter/c/c67d28/32.png) [@chasecb](https://discourse.julialang.org/u/chasecb)\
**Post date:** [February 3, 2018, 9:27pm UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/5 "2018-02-03T21:27:19Z")

</div>

Yes I did use the profiler actually and the data processing takes roughly  
12x the I/O time.

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [February 3, 2018, 11:26pm UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/6 "2018-02-03T23:26:52Z")

</div>

Can you post a MWE of your current code? If processing takes that long, parallelizing it will likely help, but your approach will depend on the size of the files you’re working with and the type of operations you’re performing on the data.

---

<div class="post-metadata">

**Author:** ![jwu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jwu/32/3775_2.png) [@jwu](https://discourse.julialang.org/u/jwu)\
**Post date:** [February 3, 2018, 11:52pm UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/7 "2018-02-03T23:52:41Z")

</div>

I used pmap to do a similar stuff, it works great.  
The IO goes through network though.

---

<div class="post-metadata">

**Author:** ![chasecb](https://avatars.discourse-cdn.com/v4/letter/c/c67d28/32.png) [@chasecb](https://discourse.julialang.org/u/chasecb)\
**Post date:** [February 4, 2018, 12:10am UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/8 "2018-02-04T00:10:51Z")

</div>

The files are roughly 100mb and there are many of them.

---

<div class="post-metadata">

**Author:** ![chasecb](https://avatars.discourse-cdn.com/v4/letter/c/c67d28/32.png) [@chasecb](https://discourse.julialang.org/u/chasecb)\
**Post date:** [February 4, 2018, 12:12am UTC](https://discourse.julialang.org/t/processing-csvs-in-parallel/8791/9 "2018-02-04T00:12:26Z")

</div>

```julia
using DataFrames

files = ["2000a.csv", "2000b.csv", "2000c.csv", "2000d.csv"]

function process(df::DataFrame)
    const conversions = Dict(
                :date => c -> Date( c, "mm/dd/yyyy"),
                :value => c -> [float(s) for s in c],
                :date2 => c -> Date( c, "yyyy-mm-dd" ),
                :value2 => c -> [float(s) for s in c],
                 )
    for (name,col) in eachcol(df)
        if haskey(conversions, name )
            df[name] = conversions[name](col)
        end
    end
    return df
end

for file in files ## this loop is what I would like to make parallel
    df = readtable(file)
    df = process(df)
    writetable(df, "path")
end

```
