# Faster small CUDA memory transfers (UnifiedMem?)

**URL:** <https://discourse.julialang.org/t/faster-small-cuda-memory-transfers-unifiedmem/45850>\
**Category:** GPU\
**Tags:** cuda\
**Created:** [August 31, 2020, 5:18pm UTC](https://discourse.julialang.org/t/faster-small-cuda-memory-transfers-unifiedmem/45850 "2020-08-31T17:18:06Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alex\_Ellison](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alex_ellison/32/5749_2.png) [@Alex\_Ellison](https://discourse.julialang.org/u/Alex_Ellison)\
**Post date:** [August 31, 2020, 5:18pm UTC](https://discourse.julialang.org/t/faster-small-cuda-memory-transfers-unifiedmem/45850/1 "2020-08-31T17:18:06Z")

</div>

Hey,

I’m trying to write a program which launches many kernels, but the launches depend on results of earlier kernels. Sometimes the result means a kernel should not be launched at all. They can be distributed across many streams.

Currently, I store the results of a kernel in a `CuArray` called `c`, copy it to the host with `a = Array(c)`, and can work directly with the value(s) in `a`. Lots of kernels means lots of memory transfers, and these all happen in the default stream. This is a bottleneck.

It looks like Unified Memory might help because it is accessible from host and device, and you can do async transfers in any stream. This [example](https://developer.nvidia.com/blog/unified-memory-cuda-beginners/) in C++ makes it look like standard array syntax works on host and device. The [interface](https://juliagpu.gitlab.io/CUDA.jl/lib/driver/#Memory-Management) in CUDA.jl is less transparent to me.

I can create a buffer and write to it like:

```julia
julia> begin
       using CUDA
       unified = CUDA.Mem.alloc(CUDA.Mem.Unified, 4)
       cuptr = convert(CuPtr{UInt32}, unified)
       CUDA.Mem.set!(cuptr, UInt32(17), UInt32(1))
       ptr = convert(Ptr{UInt32}, unified)
       @info "Value: $(unsafe_load(ptr))"
       end
[ Info: Value: 17
julia> 

```

but `CUDA.Mem.set!` doesn’t seem to be kernel friendly. If this use case isn’t totally insane, any tips for how to use unified memory in a kernel?

Thanks

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [August 31, 2020, 7:15pm UTC](https://discourse.julialang.org/t/faster-small-cuda-memory-transfers-unifiedmem/45850/2 "2020-08-31T19:15:40Z")

</div>

There was [https://github.com/JuliaGPU/CuArrays.jl/pull/336](https://github.com/JuliaGPU/CuArrays.jl/pull/336) which would’ve made the interface much friendlier, but it doesn’t seem to have been moved to CUDA.jl…

---

<div class="post-metadata">

**Author:** ![Alex\_Ellison](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alex_ellison/32/5749_2.png) [@Alex\_Ellison](https://discourse.julialang.org/u/Alex_Ellison)\
**Post date:** [August 31, 2020, 8:14pm UTC](https://discourse.julialang.org/t/faster-small-cuda-memory-transfers-unifiedmem/45850/3 "2020-08-31T20:14:21Z")

</div>

`unsafe_wrap` gives me the rest of the functionality I need I think. TBD if the performance is what I’m looking for

```julia
using CUDA

function kernel(cuarray)
    cuarray[1] += 3
    return nothing
end

unified = CUDA.Mem.alloc(CUDA.Mem.Unified, 4)
cuptr = convert(CuPtr{UInt32}, unified)
ptr = convert(Ptr{UInt32}, unified)
cuarray = unsafe_wrap(CuArray{UInt32}, cuptr, 1)
array = unsafe_wrap(Array{UInt32}, ptr, 1)
array[1] = 17
@cuda kernel(cuarray)
synchronize()
println(map(Int, array))

```

Gives us

```julia
[Info: [20]
julia> 

```
