# Parallel deployment in DataFrames and CSV operations

**URL:** <https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070>\
**Category:** General Usage\
**Tags:** question\
**Created:** [March 18, 2022, 8:29am UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070 "2022-03-18T08:29:16Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![AlanAylmer](https://avatars.discourse-cdn.com/v4/letter/a/b19c9b/32.png) [@AlanAylmer](https://discourse.julialang.org/u/AlanAylmer)\
**Post date:** [March 18, 2022, 8:29am UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070/1 "2022-03-18T08:29:16Z")

</div>

Hi. Using Julia happily for 3 years.

Just checking before I do some potentially destructive file handling. I have huge amounts of monthly data and I want to clean-reformat it and I have a working code and an 8 core laptop.

In the code, I check if a substring is in a file name and then deletes the ones that are irrelevant and copies the keepers with a flag. Then later I delete everything except keepers.

I tested and did proof of concept for Jan. In Feb I optimised for speed with type specs and considering loop invariants blah blah.

All the action happens in a one discrete parent/working folder, per month, though things get copied and removed a little between folders within that.

I propose to open five terminals for five months at time for the remaining 10 months, and run a single thread Julia instance in each.

The question is more abstracted than the code. The code is unremarkable. It’s a question about Julia’s behaviour in parallel teriminals in Linux.

Here is the question.

Ubuntu box on standard x86-64. If I CTRL+ALT+T and bring up five parallel 1 thread instances can I safely run five months in parallel? Is it thread safe in that regard?

Thanks

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [March 18, 2022, 8:53am UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070/2 "2022-03-18T08:53:56Z")

</div>

Instances of Julia do not interfere with each other.

Although I have never run it for 5 months at a time.

---

<div class="post-metadata">

**Author:** ![AlanAylmer](https://avatars.discourse-cdn.com/v4/letter/a/b19c9b/32.png) [@AlanAylmer](https://discourse.julialang.org/u/AlanAylmer)\
**Post date:** [March 18, 2022, 11:35am UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070/3 "2022-03-18T11:35:36Z")

</div>

Good  
Thanks  
Solved

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [March 18, 2022, 11:53am UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070/4 "2022-03-18T11:53:06Z")

</div>

You might want to look into Distributed

which will spawn independent processes for you

then you won’t need 5 terminals

[https://docs.julialang.org/en/v1/stdlib/Distributed/](https://docs.julialang.org/en/v1/stdlib/Distributed/)

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [March 18, 2022, 12:07pm UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070/5 "2022-03-18T12:07:37Z")

</div>

The only concern here that I would have is to make sure that the 5 processes are working on disjoint sets of files. I think it is also important to point out that if your code is very IO heavy then using 5 workers may not help much because IO does not parallelize well.

---

<div class="post-metadata">

**Author:** ![AlanAylmer](https://avatars.discourse-cdn.com/v4/letter/a/b19c9b/32.png) [@AlanAylmer](https://discourse.julialang.org/u/AlanAylmer)\
**Post date:** [March 18, 2022, 4:39pm UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070/6 "2022-03-18T16:39:45Z")

</div>

Thank you @lawless-m  
I am just doing it once so I’ll do it across terminals  
I will look into this Pkg anyway for my edu

---

<div class="post-metadata">

**Author:** ![AlanAylmer](https://avatars.discourse-cdn.com/v4/letter/a/b19c9b/32.png) [@AlanAylmer](https://discourse.julialang.org/u/AlanAylmer)\
**Post date:** [March 18, 2022, 4:41pm UTC](https://discourse.julialang.org/t/parallel-deployment-in-dataframes-and-csv-operations/78070/7 "2022-03-18T16:41:39Z")

</div>

Thank you @tbeason  
I do not need superefficiency, just a degree of parallelism is qualitatively different from none  
I take your point re IO being especially constraining though
