# How to use distributed and pmap across GPU cores

**URL:** https://discourse.julialang.org/t/how-to-use-distributed-and-pmap-across-gpu-cores/78852
**Category:** GPU
**Tags:** question, cuda, distributed, pmap
**Created:** [April 1, 2022, 8:16am UTC](https://discourse.julialang.org/t/how-to-use-distributed-and-pmap-across-gpu-cores/78852 "2022-04-01T08:16:16Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![Nadou407](https://avatars.discourse-cdn.com/v4/letter/n/f07891/32.png) [@Nadou407](https://discourse.julialang.org/u/Nadou407)
#### Post date: [April 1, 2022, 8:16am UTC](https://discourse.julialang.org/t/how-to-use-distributed-and-pmap-across-gpu-cores/78852/1 "2022-04-01T08:16:17Z")

</div>

Hi all,

I’m still new to Julia GPU and Julia in general so got quite confused about how to do parallel computation using a GPU. Any help would be greatly appreciated!

My current code (a simple example) using 20 CPU cores is the following. Basically, I use distributed package to create CPU workers, send the needed data to each worker, use -pmap- to divide the whole job into batched pieces and distribute them to each worker, then collect them back.

I would like to take advantage of the large number of cores in a GPU (my V100 GPU should have 4000+ cores) but didn’t find enough references on how to do the transition…I would guess I need to make some CuArray somewhere but really confused about where to start, so thank you so much for any help and guidance!

```julia
using Distributed
N_worker = 20
addprocs(N_worker)
using DelimitedFiles
@everywhere using Distributions, Random, ParallelDataTransfer

Data = rand(10000,25) #in real application, this will be imported from a CSV file
sendto(workers(), Data = Data)

@everywhere function F(i)
    eps = rand(1)
    out = Data[i,1] + eps #this is silly but just an illustration of the real calculation (which will take some time per worker)
    return out
end

function parallel(NN)
    pool = CachingPool(workers())
    f_obj = pmap(F, pool, 1:NN, batch_size = Int(ceil(NN/N_worker)))
    f_obj = hcat(f_obj...)
    return f_obj
end

result = parallel(10000)
writedlm("result.txt", [result])
rmprocs(workers())

```

---

<div class="post-metadata">

### Author: ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)
#### Post date: [April 1, 2022, 9:02am UTC](https://discourse.julialang.org/t/how-to-use-distributed-and-pmap-across-gpu-cores/78852/2 "2022-04-01T09:02:39Z")

</div>

While I am sure you can access the GPU from different processes, it will be considerably easier to work your problem out using a single CPU thread to initiate the computation on the multiple GPU cores.

So while learning the GPU side, I would abandon using Distributed

I have found KernelAbstractions.jl simple to get started

[https://juliagpu.github.io/KernelAbstractions.jl/stable](https://juliagpu.github.io/KernelAbstractions.jl/stable)

---

<div class="post-metadata">

### Author: ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)
#### Post date: [April 1, 2022, 1:04pm UTC](https://discourse.julialang.org/t/how-to-use-distributed-and-pmap-across-gpu-cores/78852/3 "2022-04-01T13:04:25Z")

</div>

See the introductory tutorial: [Introduction · CUDA.jl](https://cuda.juliagpu.org/stable/tutorials/introduction/). If possible, use array abstractions, and if you need to you can write custom kernels (either with CUDA.jl directly or using KernelAbstractions.jl).
