# CUDA.jl - When to synchronize

**URL:** <https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576>\
**Category:** General Usage\
**Tags:** cuda\
**Created:** [June 12, 2024, 9:20pm UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576 "2024-06-12T21:20:30Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jason\_Meziere](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jason_meziere/32/207966_2.png) [@Jason\_Meziere](https://discourse.julialang.org/u/Jason_Meziere)\
**Post date:** [June 12, 2024, 9:20pm UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/1 "2024-06-12T21:20:30Z")

</div>

The project I’m working on is run entirely on a GPU. What I mean by that is that there is no copying back and forth between the GPU and CPU until the very end. Along the way, I go through 5 main types of operations:

Array programming  
Kernels (3 functions I couldn’t figure out with array programming)  
mapreducedim-type function calls  
Calls to cufinufft in python (gpu version of flatiron institute’s NUFFT code)  
Calls to a GPU compiled LAMMPS through LAMMPS.jl (molecular dynamics code written in C++)

I’m not quite sure where I need to put explicit synchronize commands in the code. I think that array programming calls it for me, so I’m thinking I may need to call it before and after the kernels, the call to the cufinufft library in python, and the call to LAMMPS. Is this correct?

Is there a general rule that I should follow, or is it case dependent?

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [June 12, 2024, 9:43pm UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/2 "2024-06-12T21:43:37Z")

</div>

Everytime you are passing data to be processed on another stream you need to synchronize beforehand.

So as an example if you are passing a CuArray to a C++ library that launches CUDA operations internally, you will need to synchronize before the ccall (and at the end of the C++ code) to make sure that all operations created by Julia are finished before C++ operates on the memory and vice-versa.

---

<div class="post-metadata">

**Author:** ![Jason\_Meziere](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jason_meziere/32/207966_2.png) [@Jason\_Meziere](https://discourse.julialang.org/u/Jason_Meziere)\
**Post date:** [June 12, 2024, 10:45pm UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/3 "2024-06-12T22:45:27Z")

</div>

Awesome, thank you. But the same does not apply to kernels compiled with @cuda, correct? These would not be on another stream.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [June 13, 2024, 10:56am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/4 "2024-06-13T10:56:47Z")

</div>

> [@Jason\_Meziere](#):
>
> But the same does not apply to kernels compiled with @cuda, correct? These would not be on another stream.

Correct. In fact, with the latest version of CUDA.jl it’s not strictly required anymore to synchronize when performing operations on other streams, as CUDA.jl will synchronize for you: [CUDA.jl 5.4: Memory management mayhem ⋅ JuliaGPU](https://juliagpu.org/post/2024-05-28-cuda_5.4/#tracked_memory_allocations). This of course does not hold when calling out to non-CUDA.jl code.

---

<div class="post-metadata">

**Author:** ![ww1g11](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ww1g11/32/14205_2.png) [@ww1g11](https://discourse.julialang.org/u/ww1g11)\
**Post date:** [November 2, 2024, 11:23am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/5 "2024-11-02T11:23:55Z")

</div>

Hi, I have a very basic question related to synchronize using KernelAbstractions.jl. In the example, such as [Matmul · KernelAbstractions.jl](https://juliagpu.github.io/KernelAbstractions.jl/dev/examples/matmul/), there is a call to `KernelAbstractions.synchronize(backend)`. I ran a quick test using the following code, and it appears that removing `sync1`, `sync2`, and `sync3` still produces the correct results. Are all three synchronizations unnecessary?

```julia
using KernelAbstractions, Test
using CUDA
using BenchmarkTools

# Increase x by 1
@kernel function kernel1!(x)
    i = @index(Global)
    x[i] += 1
end

# y = x^2
@kernel function kernel2!(y,x)
    i = @index(Global)
    y[i] = x[i]^2
end

function test_fun(y, x)
    backend = KernelAbstractions.get_backend(x)
    kernel1!(backend)(x; ndrange=length(x))
    
    KernelAbstractions.synchronize(backend) # sync1 

    kernel2!(backend)(y, x; ndrange=length(x))
    
    return nothing
end

N = 10000
x = CUDA.zeros(N)
y = CUDA.zeros(N)
y_cpu = zeros(N)
for i in 1:100
    test_fun(y, x)

    backend = KernelAbstractions.get_backend(x)
    KernelAbstractions.synchronize(backend) # sync2
end

backend = KernelAbstractions.get_backend(x)
KernelAbstractions.synchronize(backend) # sync3
copyto!(y_cpu, y)

@test all(y .== 100^2)
@test all(y_cpu .== 100^2)

```

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [November 4, 2024, 4:27pm UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/6 "2024-11-04T16:27:28Z")

</div>

> [@ww1g11](#):
>
> `copyto!(y_cpu, y)`

Copying to CPU memory automatically synchronizes.

> [@ww1g11](#):
>
> ```julia
> kernel1!(backend)(x; ndrange=length(x))
>     
> KernelAbstractions.synchronize(backend) # sync1 
> 
> kernel2!(backend)(y, x; ndrange=length(x))
> 
> ```

Kernel execution is ordered on the task-local stream, so there’s no need to synchronize in between.

---

<div class="post-metadata">

**Author:** ![ww1g11](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ww1g11/32/14205_2.png) [@ww1g11](https://discourse.julialang.org/u/ww1g11)\
**Post date:** [November 5, 2024, 4:29am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/7 "2024-11-05T04:29:38Z")

</div>

> [@CUDA.jl - When to synchronize](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/6):
>
> Copying to CPU memory automatically synchronizes. Kernel execution is ordered on the task-local stream, so there’s no need to synchronize in between.

Many thanks, so all three synchronizations are unnecessary? Does this conclusion also hold for other GPUs supported by KernelAbstractions.jl?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [November 5, 2024, 5:40am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/8 "2024-11-05T05:40:30Z")

</div>

Yes, they are all unnecessary. That should hold for all our back-ends.

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [November 5, 2024, 10:40am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/9 "2024-11-05T10:40:53Z")

</div>

Yeah, this is mostly an artifact from a previous version of KernelAbstractions.jl where one had to to be more explicit.

---

<div class="post-metadata">

**Author:** ![Yassin\_ElBedwihy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yassin_elbedwihy/32/208506_2.png) [@Yassin\_ElBedwihy](https://discourse.julialang.org/u/Yassin_ElBedwihy)\
**Post date:** [March 5, 2025, 9:37am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/10 "2025-03-05T09:37:46Z")

</div>

I suppose then that the multiple warnings to synchronize and await kernel launches in the [julia con 2021 workshop](https://www.youtube.com/watch?v=Hz9IMJuW5hU&t=292s) is outdated correct?

---

<div class="post-metadata">

**Author:** ![Yassin\_ElBedwihy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yassin_elbedwihy/32/208506_2.png) [@Yassin\_ElBedwihy](https://discourse.julialang.org/u/Yassin_ElBedwihy)\
**Post date:** [March 5, 2025, 9:49am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/11 "2025-03-05T09:49:06Z")

</div>

For example, the code at timestamp 2:35:47

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [March 6, 2025, 6:53am UTC](https://discourse.julialang.org/t/cuda-jl-when-to-synchronize/115576/12 "2025-03-06T06:53:52Z")

</div>

That is correct.
