# Distributed computing for functions in scripts inside a local module?

**URL:** <https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729>\
**Category:** Julia at Scale\
**Tags:** question, parallel, distributed\
**Created:** [March 28, 2023, 11:11pm UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729 "2023-03-28T23:11:32Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![math\_opt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/math_opt/32/25317_2.png) [@math\_opt](https://discourse.julialang.org/u/math_opt)\
**Post date:** [March 28, 2023, 11:11pm UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/1 "2023-03-28T23:11:32Z")

</div>

I have a local module, say `MyProject`, which has a folder `src` with two scripts `script1.jl` and `script2.jl`.

```julia
module MyProject

using Package1
using Package2
using Package3

include("script1.jl")
include("script2.jl")

export function1_in_script1,
    function1_in_script2

```

Now, there is a `for` loop in `function1_in_script2` that I want to parallelize, which depends on the content in the `script1.jl` and `Package1` and `Package2`.

If I don’t need parallelization, I would generally run the following code:

```julia
#Activating the environment
using Pkg
if isfile("Project.toml") && isfile("Manifest.toml")
    Pkg.activate(".")
end

using MyProject
using AnyOtherRequiredPackage

# Call the function
function1_in_script2(...)

```

Now, how do I modify the above code to make sure that the for loop in the `function1_in_script2` is available on all worker processes? I am using `pmap` to parallelize the for loop. I understand, I can do something like:

```julia
using Distributed
addprocs(5)

#Now, I am confused about where to use the @everywhere macro
#Should it be on the entire module, i.e., @everywhere using MyProject 
#Or should I use it inside the MyProject module?

```

Thanks!

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [March 28, 2023, 11:22pm UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/2 "2023-03-28T23:22:01Z")

</div>

Check the ultimate guide to distributed computing:

> **[GitHub - Arpeggeo/julia-distributed-computing: The ultimate guide to...](https://github.com/Arpeggeo/julia-distributed-computing)**
>
> The ultimate guide to distributed computing in Julia - GitHub - Arpeggeo/julia-distributed-computing: The ultimate guide to distributed computing in Julia

---

<div class="post-metadata">

**Author:** ![math\_opt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/math_opt/32/25317_2.png) [@math\_opt](https://discourse.julialang.org/u/math_opt)\
**Post date:** [March 29, 2023, 12:06am UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/3 "2023-03-29T00:06:59Z")

</div>

Thanks for pointing to the GitHub repository. I went through it, and I understand the parallelization part described there; however, in my case, I want to parallelize something that’s part of the module (corresponding to which the environment is instantiated) and not something that’s defined and called after the module is loaded post-instantiating the environment.

May be it’s a trivial extension of what’s in the GitHub repo, but I just wanted to make sure that I am doing it correctly.

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [March 29, 2023, 12:35am UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/4 "2023-03-29T00:35:34Z")

</div>

You can use `remotecall_eval` to load the module on a worker (assuming it has access to the code), even if the worker is being spun up from a function defined within the module. I learned this from @dave.f.kleinschmidt, possibly he has a code example handy somewhere.

---

<div class="post-metadata">

**Author:** ![dave.f.kleinschmidt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dave.f.kleinschmidt/32/55_2.png) [@dave.f.kleinschmidt](https://discourse.julialang.org/u/dave.f.kleinschmidt)\
**Post date:** [March 29, 2023, 2:59pm UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/5 "2023-03-29T14:59:10Z")

</div>

Nothing I can share publicly unfortunately, but here’s a lightly sanitized version:

```julia
module MyStuff

using Distributed
using Pkg

# async create n_workers worker processes
function create_worker_processes(n_workers, manager=Disributed.LocalManager(); revise=false)
    tasks = map(1:n_workers) do n
        @async create_worker_process(manager; revise, n)
    end

    return @async map(fetch, tasks)
end

function create_worker_process(manager; revise=false, n=nothing)
    n_str = n === nothing ? "" : " $(n)"
    worker_str = string("worker", n_str)
    @info "Requesting $(worker_str)..."
    pid = only(addprocs(manager))
    # make sure we activate the ACTUAL PROJECT that's active on the manager,
    # which may be different than `@.` during e.g. CI runs
    project = Pkg.project().path
    Distributed.remotecall_eval(Main, pid,
                                :(using Pkg; Pkg.activate($(project))))
    if revise
        @info "Loading Revise on $(worker_str)..."
        Distributed.remotecall_eval(Main, pid, :(using Revise))
    end
    @info "Loading MyStuff on $(worker_str)..."
    Distributed.remotecall_eval(Main, pid, :(using MyStuff))
    @info "$(worker_str) ready, PID $(pid)"
    return pid
end

end # module

```

passing around the manager is a bit of extra cognitive overhead but v. useflu when you’re juggling, say, different kinds of K8s resources (GPU-equipped pods for training, CPU-only for batching etc.). in that case, we usually have another layer like

```julia
provision_workers(config)
    train_workers = create_worker_processes(train_manager(config), config.n_train_workers)
    batch_workers = create_worker_processes(batch_manager(config), config.n_batch_workers)
    return (; train_workers = fetch(train_workers), batch_workers = fetch(batch_workers), config)
end

```

then we pass this “harness” to the functions that actually do the training/batching work so they know what workers to use. that is, there’s no _magical distributed_ execution: the user/driver script has to specify which resources to use; this just makes it a bit more convenient to set up those resources so that they can be used effectively so the user does not need to do `@everywhere using MyStuff`, set the project, etc.

---

<div class="post-metadata">

**Author:** ![math\_opt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/math_opt/32/25317_2.png) [@math\_opt](https://discourse.julialang.org/u/math_opt)\
**Post date:** [March 29, 2023, 6:28pm UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/6 "2023-03-29T18:28:04Z")

</div>

Hi @dave.f.kleinschmidt,

Thanks for the detailed explanation. I probably need some time to understand all parts of your code; however, yesterday, I was playing around with my code, and I noticed that simply requesting (via a slurm script) a certain number of cores (say 24) on a particular node on a cluster, and then activating my project environment (shown below) on all cores automatically parallelized the `pmap` function in my module.

```julia
using Distributed
addprocs(24)

#activating the MyProject environment on all workers
@everywhere begin
    using Pkg
    if isfile("Project.toml") && isfile("Manifest.toml")
        Pkg.activate(".")
    end
end

using MyProject # the pmap function lies somewhere inside one of the scripts in this module

#and then whatever I need to call from the MyProject module

```

I was not expecting it to parallelize but in the output log I can see that the function called inside `pmap` ran on different workers. I was assuming that I would at least need to do `@everywhere using MyProject` but that seems to be not required. Any comments on why it works, or may be I am missing something here?

Thanks for the help.

---

<div class="post-metadata">

**Author:** ![dave.f.kleinschmidt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dave.f.kleinschmidt/32/55_2.png) [@dave.f.kleinschmidt](https://discourse.julialang.org/u/dave.f.kleinschmidt)\
**Post date:** [March 29, 2023, 10:56pm UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/7 "2023-03-29T22:56:08Z")

</div>

Mysterious! I’ve always found that unless you’ve actually run the code that defines the work functions on the workers, then you’ll get a undefvar error. But it’s possible that the work you’re trying to do (i.e. the function that you’re `pmap`ing) does not require anything in `MyProject`, otherwise it’s a mystery to me

---

<div class="post-metadata">

**Author:** ![kdheepak](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kdheepak/32/10881_2.png) [@kdheepak](https://discourse.julialang.org/u/kdheepak)\
**Post date:** [May 5, 2023, 7:33pm UTC](https://discourse.julialang.org/t/distributed-computing-for-functions-in-scripts-inside-a-local-module/96729/8 "2023-05-05T19:33:23Z")

</div>

For my own reference more than anything else, here’s the function I’m using:

```julia
function create_worker_processes(; n_workers = Sys.CPU_THREADS, revise=false)
    N = last(sort(workers()))
    asyncmap(1:n_workers) do i
      worker_number = N + i
      worker_str = "Worker $worker_number"
      @info "Requesting Worker $(worker_number)..."
      pid = only(addprocs(1))
      project = Pkg.project().path
      Distributed.remotecall_eval(Main, pid,
                                  :(using Pkg; Pkg.activate($(project))))
      if revise
          @info "Loading Revise on $(worker_str)..."
          Distributed.remotecall_eval(Main, pid, :(using Revise))
      end
      @info "Loading PackageName on $(worker_str)..."
      Distributed.remotecall_eval(Main, pid, :(import PackageName))
      @info "$(worker_str) ready, PID $(pid)"
      pid
    end
end

```

Thanks to @dave.f.kleinschmidt for the initial implementation and @ericphanson for suggesting checking out `asyncmap`. Adding this function to my package, I’m able to write other functions that call `@spawn` or `@spawnat` and that allows me to make including `Distributed` seamless for users of my package.
