# Understanding Shuffle in CUDA

**URL:** <https://discourse.julialang.org/t/understanding-shuffle-in-cuda/90235>\
**Category:** Performance\
**Tags:** cuda\
**Created:** [November 14, 2022, 1:33pm UTC](https://discourse.julialang.org/t/understanding-shuffle-in-cuda/90235 "2022-11-14T13:33:05Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![nethaji8](https://avatars.discourse-cdn.com/v4/letter/n/58f4c7/32.png) [@nethaji8](https://discourse.julialang.org/u/nethaji8)\
**Post date:** [November 14, 2022, 1:33pm UTC](https://discourse.julialang.org/t/understanding-shuffle-in-cuda/90235/1 "2022-11-14T13:33:05Z")

</div>

Hi everyone,  
I am trying to use shuffle in a Cuda kernel, can someone explain to me how to use CUDA.shfl\_sync, CUDA.shfl\_up\_sync, CUDA.shfl\_down\_sync, CUDA.shfl\_xor\_sync  
with a simple example…? It’s not clear in the documentation.

---

<div class="post-metadata">

**Author:** ![Gnimuc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gnimuc/32/2194_2.png) [@Gnimuc](https://discourse.julialang.org/u/Gnimuc)\
**Post date:** [November 14, 2022, 1:43pm UTC](https://discourse.julialang.org/t/understanding-shuffle-in-cuda/90235/2 "2022-11-14T13:43:43Z")

</div>

> **[Using CUDA Warp-Level Primitives | NVIDIA Technical Blog](https://developer.nvidia.com/blog/using-cuda-warp-level-primitives/)**
>
> NVIDIA GPUs execute groups of threads known as warps in SIMT (Single Instruction, Multiple Thread) fashion. Many CUDA programs achieve high performance by taking advantage of warp execution.
