# Reading and processing Data files concurrently

**URL:** <https://discourse.julialang.org/t/reading-and-processing-data-files-concurrently/5972>\
**Category:** Data\
**Tags:** parallel\
**Created:** [September 19, 2017, 2:43pm UTC](https://discourse.julialang.org/t/reading-and-processing-data-files-concurrently/5972 "2017-09-19T14:43:44Z")\
**Posts on this page:** 1\
**Showing post:** 18

<div class="post-metadata">

**Author:** ![amit.murthy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amit.murthy/32/47_2.png) [@amit.murthy](https://discourse.julialang.org/u/amit.murthy)\
**Post date:** [September 20, 2017, 4:37pm UTC](https://discourse.julialang.org/t/reading-and-processing-data-files-concurrently/5972/18 "2017-09-20T16:37:54Z")

</div>

You can try both these approaches and see which fits your use case better.

1. Distributed, multi-process - use `addprocs(N)` and `pmap`

pseudocode:

```julia
addprocs()
@everywhere begin
   using DataFrames
   function process_file(fname)
   .......
   end
end

results = pmap(process_file, list_of_files)

```

1. Multi-threaded, single process. Process in sets of N. Psuedocode would be something like this:

```julia
const N = 4
const data_chnl = Channel{Any}(N)
@schedule begin
  @sync for f in list_of_files
    @async put!(data_chnl, readTable(f, sep, h))
  end
  close(data_chnl)
end

data=[]
for d in data_chnl
  push!(data, d)
  if length(data) == N
    Threads.@threads for d2 in data
       process_read_file(d2)
    end
    empty!(data)
  end 
end

if length(data) > 0
 Threads.@threads for d2 in data
   process_read_file(d2)
 end
 empty!(data)
end 

```

---

_[View the full topic](https://discourse.julialang.org/t/reading-and-processing-data-files-concurrently/5972)._
