# \#cuda

**URL:** https://discourse.julialang.org/tag/cuda/104.md

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

---

## [Convolutions does not work on CUDA](https://discourse.julialang.org/t/convolutions-does-not-work-on-cuda/138703)

<div class="topic-metadata">

**Author:** [@Matej\_Zorek](https://discourse.julialang.org/u/Matej_Zorek)\
**Replies:** 0\
**Last updated:** [August 9, 2026, 12:49pm UTC](https://discourse.julialang.org/t/convolutions-does-not-work-on-cuda/138703 "2026-08-09T12:49:18Z")

</div>

Hi, I have a problem with Convolutional layers on gpu. It simply does not work. I have installed CUDA and cuDNN. GPU works for “simple” layers like Dense, LayerNorm etc, but when i try evaluate Conv it gives me an error…

---

## [CUDA.jl only supports NVIDIA drivers for CUDA 11.x or 12.x (yours is for CUDA 13.0.0)](https://discourse.julialang.org/t/cuda-jl-only-supports-nvidia-drivers-for-cuda-11-x-or-12-x-yours-is-for-cuda-13-0-0/137803)

<div class="topic-metadata">

**Author:** [@Cyan](https://discourse.julialang.org/u/Cyan)\
**Replies:** 5\
**Last updated:** [July 16, 2026, 1:53am UTC](https://discourse.julialang.org/t/cuda-jl-only-supports-nvidia-drivers-for-cuda-11-x-or-12-x-yours-is-for-cuda-13-0-0/137803 "2026-07-16T01:53:37Z")

</div>

Hi, all. I have upgraded the CUDA several times, but still got: julia\> using CUDA ┌ Error: This version of CUDA.jl only supports NVIDIA drivers for CUDA 11.x or 12.x (yours is for CUDA 13.0.0) └ @ CUDA ~/.julia/packag…

---

## [KernelForge.jl — High-performance portable GPU primitives for arbitrary types and operators](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780)

<div class="topic-metadata">

**Author:** [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Replies:** 13\
**Last updated:** [June 16, 2026, 3:31pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780 "2026-06-16T15:31:07Z")

</div>

I’m happy to announce two related packages for high-performance GPU computing in Julia: KernelForge.jl — high-level GPU primitives (mapreduce, scan, matvec, search, vectorized copy) with performance competitive with ve…

---

## [GPU (CUDA) vs CPU QR decomposition performance of dense wide complex matrix](https://discourse.julialang.org/t/gpu-cuda-vs-cpu-qr-decomposition-performance-of-dense-wide-complex-matrix/137448)

<div class="topic-metadata">

**Author:** [@BambOoxX](https://discourse.julialang.org/u/BambOoxX)\
**Replies:** 0\
**Last updated:** [June 4, 2026, 3:54pm UTC](https://discourse.julialang.org/t/gpu-cuda-vs-cpu-qr-decomposition-performance-of-dense-wide-complex-matrix/137448 "2026-06-04T15:54:00Z")

</div>

Hello all, I have an algorithm that requires to compute the QR decomposition (though I’m only interested in R) of a wide complex double precision matrix with a typical size of (400 rows x 16000 columns). I have squeeze…

---

## [Reactant.jl CUDA Version](https://discourse.julialang.org/t/reactant-jl-cuda-version/137211)

<div class="topic-metadata">

**Author:** [@csvance](https://discourse.julialang.org/u/csvance)\
**Replies:** 0\
**Last updated:** [May 20, 2026, 11:45am UTC](https://discourse.julialang.org/t/reactant-jl-cuda-version/137211 "2026-05-20T11:45:16Z")

</div>

Is there a recommended CUDA driver version + version of CUDA.jl I should be using with Reactant.jl? It looks like there are conflicting versions of CUDA somewhere. I0000 00:00:1779236186.222218 3515013 service.cc:178\] X…

---

## [\[ANN\] cuTile.jl v0.3 + webinar](https://discourse.julialang.org/t/ann-cutile-jl-v0-3-webinar/136988)

<div class="topic-metadata">

**Author:** [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Replies:** 0\
**Last updated:** [May 5, 2026, 2:04pm UTC](https://discourse.julialang.org/t/ann-cutile-jl-v0-3-webinar/136988 "2026-05-05T14:04:03Z")

</div>

I’ve just tagged cuTile.jl v0.3, featuring: CUDA.jl integration. Launching a cuTile kernel is now just @cuda backend=cuTile .... Better performance. We now match or outperform NVIDIA’s cuTile Python on every benchmark …

---

## [Could not instantiate \`CUDA\` in container](https://discourse.julialang.org/t/could-not-instantiate-cuda-in-container/136964)

<div class="topic-metadata">

**Author:** [@Chrysoberyl](https://discourse.julialang.org/u/Chrysoberyl)\
**Replies:** 2\
**Last updated:** [May 3, 2026, 7:07pm UTC](https://discourse.julialang.org/t/could-not-instantiate-cuda-in-container/136964 "2026-05-03T19:07:33Z")

</div>

I have a project manifest for Julia 1.12.4 containing \[\[deps.CUDA\]\] deps = \["AbstractFFTs", "Adapt", "BFloat16s", "CEnum", "CUDA\_Compiler\_jll", "CUDA\_Driver\_jll", "CUDA\_Runtime\_Discovery", "CUDA\_Runtime\_jll", "Crayons",…

---

## [CUDA.jl with @threads causing memory leak?](https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853)

<div class="topic-metadata">

**Author:** [@noetheriankoala](https://discourse.julialang.org/u/noetheriankoala)\
**Replies:** 3\
**Last updated:** [April 27, 2026, 10:23pm UTC](https://discourse.julialang.org/t/cuda-jl-with-threads-causing-memory-leak/136853 "2026-04-27T22:23:37Z")

</div>

Hello. First of all, I am using Julia v1.12.6, CUDA v5.11.2, and FLoops v0.2.2. Consider the following code and comments. using CUDA using Base.Threads function func1() A = CuArray{Float64}(undef, 1500, 1500, 1000)…

---

## [Argmax mapreduce on GPU](https://discourse.julialang.org/t/argmax-mapreduce-on-gpu/134971)

<div class="topic-metadata">

**Author:** [@noetheriankoala](https://discourse.julialang.org/u/noetheriankoala)\
**Replies:** 6\
**Last updated:** [April 21, 2026, 9:35am UTC](https://discourse.julialang.org/t/argmax-mapreduce-on-gpu/134971 "2026-04-21T09:35:03Z")

</div>

Hello! I am trying to quickly compute \\text{argmax}\_{\\substack{1 \\leq s \\leq k\\\\ k+1 \\leq t \\leq n}} A\_{s,t} + (1-\\ell\_s)(1 + \\ell\_t) I do this on the CPU with the following code. f = ((i, j),) -\> (i, j, A\[i, j\]^2 + …

---

## [\[ANN\] FastCUDASSIM.jl](https://discourse.julialang.org/t/ann-fastcudassim-jl/136382)

<div class="topic-metadata">

**Author:** [@eldee](https://discourse.julialang.org/u/eldee)\
**Replies:** 2\
**Last updated:** [April 12, 2026, 12:59pm UTC](https://discourse.julialang.org/t/ann-fastcudassim-jl/136382 "2026-04-12T12:59:07Z")

</div>

Hi everyone, I’m happy to announce my first package: FastCUDASSIM.jl. It does what it says on the tin: efficiently compute the Structural Similarity Index Measure (SSIM) via CUDA.jl. Additionally we also (optionally) ou…

---

## [Choosing between KernelAbstractions, AcceleratedKernels, ParallelStencils, or just CUDA.jl](https://discourse.julialang.org/t/choosing-between-kernelabstractions-acceleratedkernels-parallelstencils-or-just-cuda-jl/135833)

<div class="topic-metadata">

**Author:** [@Ludovic\_Dumoulin](https://discourse.julialang.org/u/Ludovic_Dumoulin)\
**Replies:** 18\
**Last updated:** [March 18, 2026, 3:38pm UTC](https://discourse.julialang.org/t/choosing-between-kernelabstractions-acceleratedkernels-parallelstencils-or-just-cuda-jl/135833 "2026-03-18T15:38:30Z")

</div>

Hello, Until today, I have used mainly CUDA.jl to run simulations. I mostly do hydrodynamics and solve various PDEs. I would like to make my code available for other users in my lab, to be able to run simulations also …

---

## [\[ANN\] cuTile.jl: Tile-based GPU programming for CUDA GPUs](https://discourse.julialang.org/t/ann-cutile-jl-tile-based-gpu-programming-for-cuda-gpus/136011)

<div class="topic-metadata">

**Author:** [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Replies:** 4\
**Last updated:** [March 4, 2026, 4:21pm UTC](https://discourse.julialang.org/t/ann-cutile-jl-tile-based-gpu-programming-for-cuda-gpus/136011 "2026-03-04T16:21:45Z")

</div>

I’m happy to announce an initial release of cuTile.jl, a new JuliaGPU package that makes it possible to program (Blackwell) NVIDIA GPUs using a tile-based abstraction by NVIDIA. This simplifies writing kernels, because y…

---

## [\[ANN\] Breeze.jl: GPU-based high-res atmospheric modeling based on Oceananigans.jl](https://discourse.julialang.org/t/ann-breeze-jl-gpu-based-high-res-atmospheric-modeling-based-on-oceananigans-jl/135347)

<div class="topic-metadata">

**Author:** [@glwagner](https://discourse.julialang.org/u/glwagner)\
**Replies:** 10\
**Last updated:** [February 6, 2026, 10:06pm UTC](https://discourse.julialang.org/t/ann-breeze-jl-gpu-based-high-res-atmospheric-modeling-based-on-oceananigans-jl/135347 "2026-02-06T22:06:45Z")

</div>

Hello Julia community! We’re building GPU-first, finite volume, pure Julia software for atmospheric modeling called “Breeze.jl”, based on Oceananigans.jl. Breeze looks and smells like Oceananigans and reuses grids, fiel…

---

## [ElasticWave2D.jl - 2D seismic wave simulation on your laptop](https://discourse.julialang.org/t/elasticwave2d-jl-2d-seismic-wave-simulation-on-your-laptop/135349)

<div class="topic-metadata">

**Author:** [@Wuheng10086](https://discourse.julialang.org/u/Wuheng10086)\
**Replies:** 0\
**Last updated:** [January 30, 2026, 2:32pm UTC](https://discourse.julialang.org/t/elasticwave2d-jl-2d-seismic-wave-simulation-on-your-laptop/135349 "2026-01-30T14:32:30Z")

</div>

Hi, everyone! I made ElasticWave2D.jl — GPU-accelerated 2D elastic wave simulation in Julia. I wanted to explore Julia’s potential for seismic simulation — easy install, GPU support, and readable code. Features: C…

---

## [SparseArrays vs CUSPARSE](https://discourse.julialang.org/t/sparsearrays-vs-cusparse/135320)

<div class="topic-metadata">

**Author:** [@rcalxrc08](https://discourse.julialang.org/u/rcalxrc08)\
**Replies:** 0\
**Last updated:** [January 29, 2026, 10:55am UTC](https://discourse.julialang.org/t/sparsearrays-vs-cusparse/135320 "2026-01-29T10:55:13Z")

</div>

I have no idea if this is intentional or not, at the moment there is an inconstency in the julia and or CUDA stack when dense and sparse vectors are interacting: using SparseArrays, CUDA N=10 x\_cpu=randn(N); x\_sparse=sp…

---

## [CUDA package having issues due to different versions of CUDA toolkit and NVIDIA driver CUDA version](https://discourse.julialang.org/t/cuda-package-having-issues-due-to-different-versions-of-cuda-toolkit-and-nvidia-driver-cuda-version/134755)

<div class="topic-metadata">

**Author:** [@Aditya747S](https://discourse.julialang.org/u/Aditya747S)\
**Replies:** 11\
**Last updated:** [January 24, 2026, 10:20pm UTC](https://discourse.julialang.org/t/cuda-package-having-issues-due-to-different-versions-of-cuda-toolkit-and-nvidia-driver-cuda-version/134755 "2026-01-24T22:20:48Z")

</div>

I have CUDA toolkit for CUDA 13.0 and my drivers have CUDA 13.1 which is why I am getting this error in julia Error: You are using CUDA 13.1.0, but CUDA.jl was precompiled for CUDA 13.0.0. │ │ This is unexpected; pleas…

---

## [Multi-GPU inference in Flux.jl](https://discourse.julialang.org/t/multi-gpu-inference-in-flux-jl/135210)

<div class="topic-metadata">

**Author:** [@Chrysoberyl](https://discourse.julialang.org/u/Chrysoberyl)\
**Replies:** 2\
**Last updated:** [January 24, 2026, 12:07am UTC](https://discourse.julialang.org/t/multi-gpu-inference-in-flux-jl/135210 "2026-01-24T00:07:26Z")

</div>

Flux allows for parallel training with multiple GPUs: GPU Support · Flux In my use case, I need to run inference of a single model on multiple GPUs. Is there a way to load balance inference calls to the model so all GPU…

---

## [CUDA Warning about freeing DeviceMemory](https://discourse.julialang.org/t/cuda-warning-about-freeing-devicememory/135042)

<div class="topic-metadata">

**Author:** [@Chrysoberyl](https://discourse.julialang.org/u/Chrysoberyl)\
**Replies:** 1\
**Last updated:** [January 15, 2026, 7:37am UTC](https://discourse.julialang.org/t/cuda-warning-about-freeing-devicememory/135042 "2026-01-15T07:37:09Z")

</div>

I’m using Flux.jl for machine learning. When I run training, I got this error (seems to be generated from compiled code in Zygote) ERROR: LoadError: CUDA error: an illegal memory access was encountered (code 700, ERROR\_…

---

## [Cartesian Indices Sequence on the GPU](https://discourse.julialang.org/t/cartesian-indices-sequence-on-the-gpu/134998)

<div class="topic-metadata">

**Author:** [@Chrysoberyl](https://discourse.julialang.org/u/Chrysoberyl)\
**Replies:** 0\
**Last updated:** [January 12, 2026, 8:24am UTC](https://discourse.julialang.org/t/cartesian-indices-sequence-on-the-gpu/134998 "2026-01-12T08:24:05Z")

</div>

I have an array of integers i, e.g. \[3,1,2\], and I want to map them to Cartesian indices \[(1, 3), (2, 1), (3, 2)\]. On the CPU, the easy way of creating this array is CartesianIndex.(enumerate(i)). However this does not …

---

## [Feedback wanted: GPU-accelerated 2D elastic wave simulation (staggered-grid FD) in Julia](https://discourse.julialang.org/t/feedback-wanted-gpu-accelerated-2d-elastic-wave-simulation-staggered-grid-fd-in-julia/134869)

<div class="topic-metadata">

**Author:** [@Wuheng10086](https://discourse.julialang.org/u/Wuheng10086)\
**Replies:** 10\
**Last updated:** [January 10, 2026, 3:16pm UTC](https://discourse.julialang.org/t/feedback-wanted-gpu-accelerated-2d-elastic-wave-simulation-staggered-grid-fd-in-julia/134869 "2026-01-10T15:16:09Z")

</div>

Hi everyone, I’m currently working on a GPU-accelerated 2D elastic wave simulation code in Julia, based on staggered-grid finite-difference discretization. The original motivation is seismic forward modeling, but I’m …

---

## [Failed to precompile CUDA](https://discourse.julialang.org/t/failed-to-precompile-cuda/134255)

<div class="topic-metadata">

**Author:** [@WuSiren](https://discourse.julialang.org/u/WuSiren)\
**Replies:** 14\
**Last updated:** [December 16, 2025, 10:59pm UTC](https://discourse.julialang.org/t/failed-to-precompile-cuda/134255 "2025-12-16T22:59:21Z")

</div>

julia\> using CUDA ┌ Warning: Circular dependency detected. Precompilation will be skipped for: │ SparseArraysExt \[85068d23-b5fb-53f1-8204-05c2aba6942f\] │ AtomixCUDAExt \[13011619-4c7c-5ef0-948f-5fc81565cd05\] │ Linea…

---

## [KernelAbstractions + CUDA + Reactant - how to get minimal working example](https://discourse.julialang.org/t/kernelabstractions-cuda-reactant-how-to-get-minimal-working-example/134164)

<div class="topic-metadata">

**Author:** [@Jakub\_Mitura](https://discourse.julialang.org/u/Jakub_Mitura)\
**Replies:** 12\
**Last updated:** [November 30, 2025, 6:34am UTC](https://discourse.julialang.org/t/kernelabstractions-cuda-reactant-how-to-get-minimal-working-example/134164 "2025-11-30T06:34:06Z")

</div>

Hello I tried to get minimal working example for reactant powered kernel abstractions plus Enzyme like using Pkg Pkg.activate(".") using Reactant using KernelAbstractions using CUDA using Enzyme using Test # Simple squ…

---

## [Error in testset gpuarrays/linalg/core](https://discourse.julialang.org/t/error-in-testset-gpuarrays-linalg-core/134091)

<div class="topic-metadata">

**Author:** [@Roger\_Powell](https://discourse.julialang.org/u/Roger_Powell)\
**Replies:** 2\
**Last updated:** [November 25, 2025, 11:48am UTC](https://discourse.julialang.org/t/error-in-testset-gpuarrays-linalg-core/134091 "2025-11-25T11:48:08Z")

</div>

First time user of CUDA.jl. My Julia versioninfo() is Julia Version 1.12.2 Commit ca9b6662be4 (2025-11-20 16:25 UTC) Build Info: Official https://julialang.org release Platform Info: OS: Linux (x86\_64-linux-gnu) C…

---

## [InvalidIRError when running AcceleratedKernels.sum on a GPU SubArray (CuArray view)](https://discourse.julialang.org/t/invalidirerror-when-running-acceleratedkernels-sum-on-a-gpu-subarray-cuarray-view/133994)

<div class="topic-metadata">

**Author:** [@0samuraiE](https://discourse.julialang.org/u/0samuraiE)\
**Replies:** 2\
**Last updated:** [November 20, 2025, 9:20am UTC](https://discourse.julialang.org/t/invalidirerror-when-running-acceleratedkernels-sum-on-a-gpu-subarray-cuarray-view/133994 "2025-11-20T09:20:41Z")

</div>

I’m running into an issue when trying to use AcceleratedKernels.sum on a GPU SubArray (a view into a CuArray). Calling AK.sum on the full array works fine, but calling it on the view throws an InvalidIRError. Question …

---

## [How to initialize/fix the RNG seed on the GPU?](https://discourse.julialang.org/t/how-to-initialize-fix-the-rng-seed-on-the-gpu/133831)

<div class="topic-metadata">

**Author:** [@jwtkeeble](https://discourse.julialang.org/u/jwtkeeble)\
**Replies:** 4\
**Last updated:** [November 13, 2025, 7:27pm UTC](https://discourse.julialang.org/t/how-to-initialize-fix-the-rng-seed-on-the-gpu/133831 "2025-11-13T19:27:07Z")

</div>

Hi All, I have a question regarding initializing the random seed within CUDA.jl. I saw in this thread for the CURAND library, however, it seems the CURAND library is now deprecated for CUDA.jl. I do see there’s CUDA.de…

---

## [Wrapping CUDA.jl with juliacall](https://discourse.julialang.org/t/wrapping-cuda-jl-with-juliacall/133554)

<div class="topic-metadata">

**Author:** [@chrissm23](https://discourse.julialang.org/u/chrissm23)\
**Replies:** 4\
**Last updated:** [November 7, 2025, 9:09am UTC](https://discourse.julialang.org/t/wrapping-cuda-jl-with-juliacall/133554 "2025-11-07T09:09:08Z")

</div>

Hello everyone! I’m trying to make a Python wrapper using PythonCall.jl for a Julia package that uses CUDA.jl. The package initializes random integers from a UnitRange using the CPU’s Random.rand into a Matrix and then …

---

## [Improving performance of CUDA GPU kernel: LU factorization](https://discourse.julialang.org/t/improving-performance-of-cuda-gpu-kernel-lu-factorization/132971)

<div class="topic-metadata">

**Author:** [@Leo\_I](https://discourse.julialang.org/u/Leo_I)\
**Replies:** 17\
**Last updated:** [October 28, 2025, 10:37am UTC](https://discourse.julialang.org/t/improving-performance-of-cuda-gpu-kernel-lu-factorization/132971 "2025-10-28T10:37:38Z")

</div>

I am learning GPU programming via CUDA.jl. I wish to implement an efficient LU factorization for matrices over a finite field (I implemented it as UInt8/UInt16/... with operations + mod and \* mod). For the sake of simpl…

---

## [Slow matrix multiplication in CUBLAS.gemm\_strided\_batched with ComplexF64](https://discourse.julialang.org/t/slow-matrix-multiplication-in-cublas-gemm-strided-batched-with-complexf64/132826)

<div class="topic-metadata">

**Author:** [@JackC](https://discourse.julialang.org/u/JackC)\
**Replies:** 1\
**Last updated:** [October 7, 2025, 6:25am UTC](https://discourse.julialang.org/t/slow-matrix-multiplication-in-cublas-gemm-strided-batched-with-complexf64/132826 "2025-10-07T06:25:05Z")

</div>

I’ve been developing some GPU code that seems does a lot of matrix multiplications with complex numbers. I figured I’d use the CUBLAS implementations that are available in the CUDA.jl package. However, I’ve found that us…

---

## [Failure to download artifact: CUDA\_Compiler](https://discourse.julialang.org/t/failure-to-download-artifact-cuda-compiler/131580)

<div class="topic-metadata">

**Author:** [@mukund-gupta](https://discourse.julialang.org/u/mukund-gupta)\
**Replies:** 1\
**Last updated:** [September 1, 2025, 11:19am UTC](https://discourse.julialang.org/t/failure-to-download-artifact-cuda-compiler/131580 "2025-09-01T11:19:14Z")

</div>

Hi everyone, I am new to CUDA.jl and am receiving the following error issues when performing using CUDA on a GPU node of a Linux cluster. Precompiling CUDA... 22064.9 ms ✓ InvertedIndices 22158.5 ms ✓ LaTeXString…

---

## [Custom (NumPy style) broadcasting rule that avoids iterating over elements (for GPU-acceleration)](https://discourse.julialang.org/t/custom-numpy-style-broadcasting-rule-that-avoids-iterating-over-elements-for-gpu-acceleration/131515)

<div class="topic-metadata">

**Author:** [@TimHargreaves](https://discourse.julialang.org/u/TimHargreaves)\
**Replies:** 10\
**Last updated:** [August 24, 2025, 7:48pm UTC](https://discourse.julialang.org/t/custom-numpy-style-broadcasting-rule-that-avoids-iterating-over-elements-for-gpu-acceleration/131515 "2025-08-24T19:48:57Z")

</div>

Minimal Example As a minimal example, suppose that I have the following two structs defined, struct Foo{T, M\<:AbstractMatrix{T}} A::M B::M end struct Bar{T, M\<:AbstractMatrix{T}} C::M foo::Foo{T, M} end…

[Next page](https://discourse.julialang.org/tag/cuda/104.md?match_all_tags=true&page=1&tags%5B%5D=cuda)
