# GPU

**URL:** https://discourse.julialang.org/c/domain/gpu/11.md?page=3

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 4

---

## [Calculate associated Legendre polynomials on the GPU](https://discourse.julialang.org/t/calculate-associated-legendre-polynomials-on-the-gpu/127418)

<div class="topic-metadata">

**Author:** [@PetrosStefanou](https://discourse.julialang.org/u/PetrosStefanou)\
**Replies:** 3\
**Last updated:** [March 27, 2025, 7:15pm UTC](https://discourse.julialang.org/t/calculate-associated-legendre-polynomials-on-the-gpu/127418 "2025-03-27T19:15:55Z")

</div>

Hi everyone, I would like to evaluate the associated Legendre polynomials P\_l^m (\\cos{\\theta}) on the GPU. I tried the packages AssociatedLegendrePolynomials.jl and LegendrePolynomials.jl, both of which work as intended…

---

## [Lightweight dependency for GPU programming](https://discourse.julialang.org/t/lightweight-dependency-for-gpu-programming/127386)

<div class="topic-metadata">

**Author:** [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Replies:** 7\
**Last updated:** [March 27, 2025, 5:22pm UTC](https://discourse.julialang.org/t/lightweight-dependency-for-gpu-programming/127386 "2025-03-27T17:22:24Z")

</div>

As a package author is it possible to define KernelAbstractions.jl in my package without adding heavy dependencies? I see that the package currently depends on \[deps\] Adapt = "79e6a3ab-5dfb-504d-930d-738a2a938a0e" Atomi…

---

## [@inbounds slower](https://discourse.julialang.org/t/inbounds-slower/126634)

<div class="topic-metadata">

**Author:** [@eldee](https://discourse.julialang.org/u/eldee)\
**Replies:** 8\
**Last updated:** [March 25, 2025, 1:52am UTC](https://discourse.julialang.org/t/inbounds-slower/126634 "2025-03-25T01:52:00Z")

</div>

Hi, I’m encountering some situations in my GPU code where adding @inbounds increases execution time by 10%-20%. Of course it would be great if we would never need @inbounds, but on the other hand there are certainly als…

---

## [I32 indexing](https://discourse.julialang.org/t/i32-indexing/126702)

<div class="topic-metadata">

**Author:** [@MrBelette](https://discourse.julialang.org/u/MrBelette)\
**Replies:** 8\
**Last updated:** [March 24, 2025, 3:04pm UTC](https://discourse.julialang.org/t/i32-indexing/126702 "2025-03-24T15:04:46Z")

</div>

I’m currently exploring using different precision (fp32, fp64, doublefloat) calculations in the financial domain using RTX 4090 gpus. These gpus have great fp32 performance but fp64 performance isn’t so hot due to there …

---

## [Floating point exceptions on the gpu](https://discourse.julialang.org/t/floating-point-exceptions-on-the-gpu/127292)

<div class="topic-metadata">

**Author:** [@MrBelette](https://discourse.julialang.org/u/MrBelette)\
**Replies:** 1\
**Last updated:** [March 24, 2025, 8:16am UTC](https://discourse.julialang.org/t/floating-point-exceptions-on-the-gpu/127292 "2025-03-24T08:16:08Z")

</div>

I’m wondering how people treat floating point exceptions on the gpu, such as division by zero or overflow? I’m not sure about other gpus, but I believe that Nvidia gpus don’t have hardware capability to detect IEEE fp e…

---

## [Unable to use AMDGPU.jl on RX6600](https://discourse.julialang.org/t/unable-to-use-amdgpu-jl-on-rx6600/126889)

<div class="topic-metadata">

**Author:** [@bpsomu](https://discourse.julialang.org/u/bpsomu)\
**Replies:** 13\
**Last updated:** [March 19, 2025, 10:40pm UTC](https://discourse.julialang.org/t/unable-to-use-amdgpu-jl-on-rx6600/126889 "2025-03-19T22:40:15Z")

</div>

I am trying to use AMDGPU.jl but it installs with several errors and does not work. Operating Systems: Fedora and Endeavor OS Julia version: 1.11.4 installed via the curl command on the Julia website julia\> using AMDG…

---

## [Adapt BroadcastStyle for CUDA](https://discourse.julialang.org/t/adapt-broadcaststyle-for-cuda/126861)

<div class="topic-metadata">

**Author:** [@RainerHeintzmann](https://discourse.julialang.org/u/RainerHeintzmann)\
**Replies:** 1\
**Last updated:** [March 18, 2025, 5:50pm UTC](https://discourse.julialang.org/t/adapt-broadcaststyle-for-cuda/126861 "2025-03-18T17:50:40Z")

</div>

I am trying to get CUDA.jl to work seamlessly with a MutableShiftedArray . By writing a CUDASupportExt.jl class: module CUDASupportExt using CUDA using Adapt using MutableShiftedArrays indexing CUDA error # lets do t…

---

## [Moving ahead with CUDA support](https://discourse.julialang.org/t/moving-ahead-with-cuda-support/126208)

<div class="topic-metadata">

**Author:** [@RainerHeintzmann](https://discourse.julialang.org/u/RainerHeintzmann)\
**Replies:** 2\
**Last updated:** [March 17, 2025, 4:13pm UTC](https://discourse.julialang.org/t/moving-ahead-with-cuda-support/126208 "2025-03-17T16:13:48Z")

</div>

For a number of applications CUDA (and AD) support of various packages (e.g. FourierTools.jl) would be great. One road-block is the support of a form of shifted and padded images. Sometimes a requirement is a shifted an…

---

## [How to benchmark a function that uses KernelAbstractions kernels?](https://discourse.julialang.org/t/how-to-benchmark-a-function-that-uses-kernelabstractions-kernels/127026)

<div class="topic-metadata">

**Author:** [@pitsianis](https://discourse.julialang.org/u/pitsianis)\
**Replies:** 4\
**Last updated:** [March 17, 2025, 1:27pm UTC](https://discourse.julialang.org/t/how-to-benchmark-a-function-that-uses-kernelabstractions-kernels/127026 "2025-03-17T13:27:19Z")

</div>

What is the correct way to benchmark a function that uses KernelAbstractions.jl (KA) kernels? I noticed that KA/CUDA i.e. KA with the CUDA.jl backend requires the @synchronize(device) macro. Consider this @btime beg…

---

## [Occasional long delays in CUDA.jl](https://discourse.julialang.org/t/occasional-long-delays-in-cuda-jl/81545)

<div class="topic-metadata">

**Author:** [@jpdoane](https://discourse.julialang.org/u/jpdoane)\
**Replies:** 17\
**Last updated:** [March 15, 2025, 11:06am UTC](https://discourse.julialang.org/t/occasional-long-delays-in-cuda-jl/81545 "2025-03-15T11:06:00Z")

</div>

We have a real-time CUDA.jl application that must maintain a given time budget in order to keep up with a streaming input signal. Although the steady state performance meets timing, we occasional see fairly long delays t…

---

## [Profiling CUDA kernels on the Jetson](https://discourse.julialang.org/t/profiling-cuda-kernels-on-the-jetson/126120)

<div class="topic-metadata">

**Author:** [@CaG21](https://discourse.julialang.org/u/CaG21)\
**Replies:** 3\
**Last updated:** [March 3, 2025, 11:13am UTC](https://discourse.julialang.org/t/profiling-cuda-kernels-on-the-jetson/126120 "2025-03-03T11:13:57Z")

</div>

I am trying to profile a CUDA kernel (using CUDA.jl) on a Jetson AGX Orin DevKit. However, when I try to do so, for example by running using CUDA a = CUDA.rand(1024,1024,1024); CUDA.@profile sin.(a) I get the error ERR…

---

## [Code snippet for multiGPU fft](https://discourse.julialang.org/t/code-snippet-for-multigpu-fft/26035)

<div class="topic-metadata">

**Author:** [@rveltz](https://discourse.julialang.org/u/rveltz)\
**Replies:** 8\
**Last updated:** [March 3, 2025, 7:33am UTC](https://discourse.julialang.org/t/code-snippet-for-multigpu-fft/26035 "2025-03-03T07:33:25Z")

</div>

Hi, I would like to perform a 3d FFT on multi GPUs. It can be done in cuda as explained here. However the functions required do not seem ported in CuArrays. I was wondering if anybody has ever done this in Julia? A sec…

---

## [Bad interaction of Metal.jl and PyPlot on julia 1.11.2](https://discourse.julialang.org/t/bad-interaction-of-metal-jl-and-pyplot-on-julia-1-11-2/123759)

<div class="topic-metadata">

**Author:** [@Andrea\_Pagnani](https://discourse.julialang.org/u/Andrea_Pagnani)\
**Replies:** 1\
**Last updated:** [February 26, 2025, 6:01pm UTC](https://discourse.julialang.org/t/bad-interaction-of-metal-jl-and-pyplot-on-julia-1-11-2/123759 "2025-02-26T18:01:09Z")

</div>

Dearests, Still fighting to test the GPU on my new macbook pro M4. Fresh new environment on version 1.11.2 (TEST\_PYPLOT) pkg\> add Metal, PyPlot (TEST\_PYPLOT) pkg\> add Metal, PyPlot Resolving package versions... …

---

## [Is there anything like vmap to vectorize a computation](https://discourse.julialang.org/t/is-there-anything-like-vmap-to-vectorize-a-computation/126288)

<div class="topic-metadata">

**Author:** [@Leander](https://discourse.julialang.org/u/Leander)\
**Replies:** 10\
**Last updated:** [February 25, 2025, 6:48pm UTC](https://discourse.julialang.org/t/is-there-anything-like-vmap-to-vectorize-a-computation/126288 "2025-02-25T18:48:42Z")

</div>

I want to evaluate using julia for a project and so far it was straightforward. The setting is the following, I have a bunch of functions (moment-generating functions over different parameters), for which I want to compu…

---

## [CUDNN in Julia](https://discourse.julialang.org/t/cudnn-in-julia/84330)

<div class="topic-metadata">

**Author:** [@Jonas208](https://discourse.julialang.org/u/Jonas208)\
**Replies:** 6\
**Last updated:** [February 25, 2025, 5:05pm UTC](https://discourse.julialang.org/t/cudnn-in-julia/84330 "2025-02-25T17:05:17Z")

</div>

Is there any way - or what is the best way - to use CUDNN in Julia? Details: I am interested in that topic because I want do add GPU-acceleration (Nvidia) to my own little Deep Learning module (it is not available anyw…

---

## [How does a kernel function in KernelAbstractions.jl work when the backend is a CPU?](https://discourse.julialang.org/t/how-does-a-kernel-function-in-kernelabstractions-jl-work-when-the-backend-is-a-cpu/126173)

<div class="topic-metadata">

**Author:** [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Replies:** 1\
**Last updated:** [February 22, 2025, 4:59pm UTC](https://discourse.julialang.org/t/how-does-a-kernel-function-in-kernelabstractions-jl-work-when-the-backend-is-a-cpu/126173 "2025-02-22T16:59:50Z")

</div>

I’m trying to add GPU-based parallelization to my code. I found KernelAbstractions.jl and liked its backend (vendor) agnostic approach, especially the fall-back support for CPU multi-threading. However, I didn’t find de…

---

## [How to perform a sparse matrix dense matrix product with addition (cuda library style)](https://discourse.julialang.org/t/how-to-perform-a-sparse-matrix-dense-matrix-product-with-addition-cuda-library-style/126093)

<div class="topic-metadata">

**Author:** [@victor\_vhrn](https://discourse.julialang.org/u/victor_vhrn)\
**Replies:** 1\
**Last updated:** [February 20, 2025, 7:33am UTC](https://discourse.julialang.org/t/how-to-perform-a-sparse-matrix-dense-matrix-product-with-addition-cuda-library-style/126093 "2025-02-20T07:33:09Z")

</div>

I want to compute the following result = -1 \* (e \* (e' \* U)) + M \* U Where e = ones(n,1), M is a n x n CuSparseMatrixCSR and U is a n x r dense CuMatrix. Since it is my understanding that the above operation invokes 4…

---

## [I get a warning when i use Upsample layer with AMDGPU](https://discourse.julialang.org/t/i-get-a-warning-when-i-use-upsample-layer-with-amdgpu/112728)

<div class="topic-metadata">

**Author:** [@Radu\_Mihai\_Diaconu](https://discourse.julialang.org/u/Radu_Mihai_Diaconu)\
**Replies:** 1\
**Last updated:** [February 18, 2025, 11:00am UTC](https://discourse.julialang.org/t/i-get-a-warning-when-i-use-upsample-layer-with-amdgpu/112728 "2025-02-18T11:00:35Z")

</div>

Hello guys so i have this code using Flux: Upsample noise = 0.5f0 + 0.001f0 .+ AMDGPU.randn(Float32, 32, 32, 1, 1) layer1 = Upsample(:bilinear, size=(64, 64, 1, 1)) |\> gpu layer1(noise) any reason for me getting thi…

---

## [Correct utilisation of CUDA kernel for simulations](https://discourse.julialang.org/t/correct-utilisation-of-cuda-kernel-for-simulations/109882)

<div class="topic-metadata">

**Author:** [@Ludovic\_Dumoulin](https://discourse.julialang.org/u/Ludovic_Dumoulin)\
**Replies:** 16\
**Last updated:** [February 13, 2025, 9:53am UTC](https://discourse.julialang.org/t/correct-utilisation-of-cuda-kernel-for-simulations/109882 "2025-02-13T09:53:24Z")

</div>

Hello, I am new to julia and GPU computing, before today I didn’t care about optimisation (it was fast enough for me). I use this kind of CUDA kernel (example for 2D diffusion using finite difference with euler forward…

---

## [Is it possible to use CuStaticSharedArray(T, n) with n const?](https://discourse.julialang.org/t/is-it-possible-to-use-custaticsharedarray-t-n-with-n-const/125756)

<div class="topic-metadata">

**Author:** [@Ludovic\_Dumoulin](https://discourse.julialang.org/u/Ludovic_Dumoulin)\
**Replies:** 2\
**Last updated:** [February 11, 2025, 8:59am UTC](https://discourse.julialang.org/t/is-it-possible-to-use-custaticsharedarray-t-n-with-n-const/125756 "2025-02-11T08:59:19Z")

</div>

Hello, I would like to compare the size of different block to see how the calculation speed-up or not. I use SharedArray of size CuStaticSharedArray(Float64, (B+2,B+2)) with B from 16 to 32, and now I am changing B by …

---

## [How to use CLArray with OpenCL 0.10](https://discourse.julialang.org/t/how-to-use-clarray-with-opencl-0-10/125757)

<div class="topic-metadata">

**Author:** [@jheinen](https://discourse.julialang.org/u/jheinen)\
**Replies:** 1\
**Last updated:** [February 10, 2025, 7:47pm UTC](https://discourse.julialang.org/t/how-to-use-clarray-with-opencl-0-10/125757 "2025-02-10T19:47:44Z")

</div>

I migrated the GR Mandelbrot example to OpenCL 0.10 (based on this notebook). Unfortunately I could not use the access parameter in CLArray(). Does this feature require special package versions?

---

## [Another freezing test CUDA](https://discourse.julialang.org/t/another-freezing-test-cuda/125667)

<div class="topic-metadata">

**Author:** [@togo59](https://discourse.julialang.org/u/togo59)\
**Replies:** 4\
**Last updated:** [February 10, 2025, 11:19am UTC](https://discourse.julialang.org/t/another-freezing-test-cuda/125667 "2025-02-10T11:19:04Z")

</div>

I get no errors at all that I can find. But when I run Pkg.test("CUDA"; test\_args=–verbose --jobs=1) after a while my computer freezes up. I have 6-core Intel CPU with 12GB and I have just installed an Intel A2000, which…

---

## [Using cuBLASDx in Julia](https://discourse.julialang.org/t/using-cublasdx-in-julia/125527)

<div class="topic-metadata">

**Author:** [@CaG21](https://discourse.julialang.org/u/CaG21)\
**Replies:** 6\
**Last updated:** [February 9, 2025, 2:54pm UTC](https://discourse.julialang.org/t/using-cublasdx-in-julia/125527 "2025-02-09T14:54:39Z")

</div>

I am really interested in using cuBLASDx, that is the CUDA API set to perform BLAS calculations inside CUDA kernels. However, since it is not provided in the CUDA Toolkit, but it should be downloaded separately, it seems…

---

## [How to Use Native FP4 and FP8 for Computation in the Julia Environment with CUDA.jl](https://discourse.julialang.org/t/how-to-use-native-fp4-and-fp8-for-computation-in-the-julia-environment-with-cuda-jl/125475)

<div class="topic-metadata">

**Author:** [@ajsagit](https://discourse.julialang.org/u/ajsagit)\
**Replies:** 0\
**Last updated:** [February 2, 2025, 3:15pm UTC](https://discourse.julialang.org/t/how-to-use-native-fp4-and-fp8-for-computation-in-the-julia-environment-with-cuda-jl/125475 "2025-02-02T15:15:47Z")

</div>

How to Use Native FP4 and FP8 for Computation in the Julia Environment with CUDA.jl TIA

---

## [Why the Floating-Point Calculation Efficiency of CUDA.jl Does Not Reach the Official Theoretical Value](https://discourse.julialang.org/t/why-the-floating-point-calculation-efficiency-of-cuda-jl-does-not-reach-the-official-theoretical-value/125474)

<div class="topic-metadata">

**Author:** [@ajsagit](https://discourse.julialang.org/u/ajsagit)\
**Replies:** 1\
**Last updated:** [February 2, 2025, 3:55pm UTC](https://discourse.julialang.org/t/why-the-floating-point-calculation-efficiency-of-cuda-jl-does-not-reach-the-official-theoretical-value/125474 "2025-02-02T15:55:10Z")

</div>

using CUDA using BenchmarkTools using Printf function benchmark\_floating\_point(T::Type, size=2048) aligned\_size = div(size, 8) \* 8 A = CUDA.rand(T, (aligned\_size, aligned\_size)) B = CUDA.rand(T, (aligned\_siz…

---

## [How to develop code in Vulkan using Julia?](https://discourse.julialang.org/t/how-to-develop-code-in-vulkan-using-julia/125456)

<div class="topic-metadata">

**Author:** [@mj2984](https://discourse.julialang.org/u/mj2984)\
**Replies:** 1\
**Last updated:** [February 1, 2025, 2:33pm UTC](https://discourse.julialang.org/t/how-to-develop-code-in-vulkan-using-julia/125456 "2025-02-01T14:33:52Z")

</div>

I am looking to see if it is possible to generate SPIR-V code from Julia that can be used in Vulkan (Compute). Or something that can transform Julia code to Vulkan for compute purposes. The intention is to write a code …

---

## [Why is CUDA.FFT slow only when performed over the second dimension of a 3D array?](https://discourse.julialang.org/t/why-is-cuda-fft-slow-only-when-performed-over-the-second-dimension-of-a-3d-array/125358)

<div class="topic-metadata">

**Author:** [@marcsgil](https://discourse.julialang.org/u/marcsgil)\
**Replies:** 0\
**Last updated:** [January 29, 2025, 6:49pm UTC](https://discourse.julialang.org/t/why-is-cuda-fft-slow-only-when-performed-over-the-second-dimension-of-a-3d-array/125358 "2025-01-29T18:49:19Z")

</div>

Consider the following test on the CPU: using CUDA, FFTW, BenchmarkTools x = randn(ComplexF32, 2^7, 2^7, 2^7) for dim ∈ 1:ndims(x) @info "FFT along dimension $dim" display(@benchmark CUDA.@sync fft($x, $dim)) …

---

## [AMDGPU.versioninfo() trips an assertion in AMD's code](https://discourse.julialang.org/t/amdgpu-versioninfo-trips-an-assertion-in-amds-code/125203)

<div class="topic-metadata">

**Author:** [@nsajko](https://discourse.julialang.org/u/nsajko)\
**Replies:** 1\
**Last updated:** [January 26, 2025, 8:51am UTC](https://discourse.julialang.org/t/amdgpu-versioninfo-trips-an-assertion-in-amds-code/125203 "2025-01-26T08:51:27Z")

</div>

System and Julia info, ./julia-1.10.8/bin/julia -g2 -e 'using InteractiveUtils; versioninfo()': Julia Version 1.10.8 Commit 4c16ff44be8 (2025-01-22 10:06 UTC) Build Info: Official https://julialang.org/ release Platfo…

---

## [Unexpected coalesced group behaviour in CUDA.jl](https://discourse.julialang.org/t/unexpected-coalesced-group-behaviour-in-cuda-jl/125109)

<div class="topic-metadata">

**Author:** [@eldee](https://discourse.julialang.org/u/eldee)\
**Replies:** 3\
**Last updated:** [January 25, 2025, 8:07am UTC](https://discourse.julialang.org/t/unexpected-coalesced-group-behaviour-in-cuda-jl/125109 "2025-01-25T08:07:31Z")

</div>

Hi, When experimenting a bit with cooperative groups and in particular coalesced groups, I’ve encountered some strange behaviour which I cannot explain. As a simple example, consider using CUDA using CUDA: i32 functio…

---

## [MLX and Apple silicon](https://discourse.julialang.org/t/mlx-and-apple-silicon/125136)

<div class="topic-metadata">

**Author:** [@Hareruya](https://discourse.julialang.org/u/Hareruya)\
**Replies:** 4\
**Last updated:** [January 24, 2025, 2:44am UTC](https://discourse.julialang.org/t/mlx-and-apple-silicon/125136 "2025-01-24T02:44:02Z")

</div>

Good evening. I have recently bought an M4 MacbookPro Max and have been using it to do some programming in my free time, messing around with image processing on the side. The machine is very powerful and energy efficien…

[Previous page](https://discourse.julialang.org/c/domain/gpu/11.md?page=2)

[Next page](https://discourse.julialang.org/c/domain/gpu/11.md?page=4)
