# CUDA performing scalar indexing when used along with Distributed

**URL:** <https://discourse.julialang.org/t/cuda-performing-scalar-indexing-when-used-along-with-distributed/119470>\
**Category:** GPU\
**Tags:** cuda, distributed\
**Created:** [September 16, 2024, 6:07pm UTC](https://discourse.julialang.org/t/cuda-performing-scalar-indexing-when-used-along-with-distributed/119470 "2024-09-16T18:07:15Z")\
**Posts on this page:** 1\
**Showing post:** 5

<div class="post-metadata">

**Author:** ![eldee](https://avatars.discourse-cdn.com/v4/letter/e/b5a626/32.png) [@eldee](https://discourse.julialang.org/u/eldee)\
**Post date:** [September 19, 2024, 7:37pm UTC](https://discourse.julialang.org/t/cuda-performing-scalar-indexing-when-used-along-with-distributed/119470/5 "2024-09-19T19:37:44Z")

</div>

Hi Paulo, you’re welcome.

Could you add some more information to the code: declarations of `input1` etc., a (dummy) implementation of `_myFunction`, …? See also point 4 in [this PSA](https://discourse.julialang.org/t/please-read-make-it-easier-to-help-you/14757). It’s hard to see how something can be improved, when you’re not sure what is concretely going on 🙂 .

> [@Paulo\_Refosco](#):
>
> ```julia
> x = Array{Any}(undef, (3, 1)...)
> for i in 1:3
> x[i] = [input1, input2, input3]
> end
> 
> ```

I’m also not sure what the intent is here. Why is `size(x) == (3, 1)`, instead of just `(3,)`? By the way, note that you don’t need the splatting `...`: `(undef, (3, 1))` and `(undef, 3, 1)` (what the splatting results in) are equivalent. Is it intended that `x[1] == x[2] == x[3]`? Why `Array{Any}` and not something a bit more concrete like `Vector{NTuple{3, CuArray}}`?

> [@Paulo\_Refosco](#):
>
> The problem is that since pmap only takes one group of elements (as per my understanding), in order to input many variables I first need to place them into an array

The second argument to `(p)map` does not need to be a `Vector`, but could also be e.g. a `Tuple` or generator, which might help with the inefficient memory usage you mention. For example:

```julia
using Distributed
addprocs(2)
@everywhere begin
    using CUDA 
    using Statistics: mean
end

pmap(mean, (CUDA.rand(2) .+ myid() for i = 1:3))
#=
3-element Vector{Float32}:
 1.9031491
 1.707502
 1.1796367
=#

pmap(x -> myid() + mean(x), (CUDA.rand(2) for i = 1:3))
#=
3-element Vector{Float32}:
 3.316328
 2.5876007
 2.284222
=#

```

Note also that this example shows that the CUDA data is generated here by the master process (`myid() == 1`) and sent over to the other processes without any issues.

---

_[View the full topic](https://discourse.julialang.org/t/cuda-performing-scalar-indexing-when-used-along-with-distributed/119470)._
