# \#cudanative

**URL:** https://discourse.julialang.org/tag/cudanative/50.md

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

---

## [How to run ptx code on CUDA from julia?](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306)

<div class="topic-metadata">

**Author:** [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Replies:** 11\
**Last updated:** [January 26, 2024, 8:42pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306 "2024-01-26T20:42:17Z")

</div>

Hi I am trying to create a NN library that directly output a very optimized version of ptx code, also would be open to output even lower level code if possible. But of course I see I have to go step by step as it is get…

---

## [CUDA fft wrapper problem](https://discourse.julialang.org/t/cuda-fft-wrapper-problem/29540)

<div class="topic-metadata">

**Author:** [@Michael\_Benton](https://discourse.julialang.org/u/Michael_Benton)\
**Replies:** 1\
**Last updated:** [August 24, 2023, 12:56pm UTC](https://discourse.julialang.org/t/cuda-fft-wrapper-problem/29540 "2023-08-24T12:56:29Z")

</div>

A 1d fft across the 2nd dimension of 3 dimensional CuArray is not enabled by the wrapper (ERROR: ArgumentError: batching dims must be sequential) to reproduce: dim = 2 data = CuArrays.rand(ComplexF32,512,512,512); my…

---

## [Cuda-memcheck reports over 1300 errors with 4 lines of julia code with CUDA.jl](https://discourse.julialang.org/t/cuda-memcheck-reports-over-1300-errors-with-4-lines-of-julia-code-with-cuda-jl/84516)

<div class="topic-metadata">

**Author:** [@Lian\_Yunlong](https://discourse.julialang.org/u/Lian_Yunlong)\
**Replies:** 2\
**Last updated:** [July 20, 2022, 11:39pm UTC](https://discourse.julialang.org/t/cuda-memcheck-reports-over-1300-errors-with-4-lines-of-julia-code-with-cuda-jl/84516 "2022-07-20T23:39:00Z")

</div>

I have a file test.jl of four lines of Julia code using CUDA A = CuArray(rand(1000,1000)); B = CuArray(rand(1000,1000)); C = A \* B; After running the command cuda-memcheck julia test.jl \>\> test.jl.log I have got so m…

---

## [Is it possible to create a virtual environment using CUDA?](https://discourse.julialang.org/t/is-it-possible-to-create-a-virtual-environment-using-cuda/84185)

<div class="topic-metadata">

**Author:** [@Fathya.Salih](https://discourse.julialang.org/u/Fathya.Salih)\
**Replies:** 3\
**Last updated:** [July 19, 2022, 3:25pm UTC](https://discourse.julialang.org/t/is-it-possible-to-create-a-virtual-environment-using-cuda/84185 "2022-07-19T15:25:03Z")

</div>

Hello! I am trying to use julia 1.2.0 on a remote GPU system and I am having trouble creating a virtual environment where I can add the packages I need. These are the packages available in the remote machine: (v1.2) pk…

---

## [How to reset GPU after launch failure](https://discourse.julialang.org/t/how-to-reset-gpu-after-launch-failure/28768)

<div class="topic-metadata">

**Author:** [@Michael\_Benton](https://discourse.julialang.org/u/Michael_Benton)\
**Replies:** 8\
**Last updated:** [March 18, 2022, 1:49pm UTC](https://discourse.julialang.org/t/how-to-reset-gpu-after-launch-failure/28768 "2022-03-18T13:49:16Z")

</div>

If I write a kernel that fails for whatever reason, julia becomes unable to use the GPU for the remainder of the session example : julia\>cu(\[1.,2.\]) CUDA error: unspecified launch failure (code #719, ERROR\_LAUNCH\_FAIL…

---

## [The most general way to estimate the optimal arguments for @cuda macro](https://discourse.julialang.org/t/the-most-general-way-to-estimate-the-optimal-arguments-for-cuda-macro/39342)

<div class="topic-metadata">

**Author:** [@fedoroff](https://discourse.julialang.org/u/fedoroff)\
**Replies:** 6\
**Last updated:** [April 6, 2021, 2:12pm UTC](https://discourse.julialang.org/t/the-most-general-way-to-estimate-the-optimal-arguments-for-cuda-macro/39342 "2021-04-06T14:12:17Z")

</div>

How one can choose the optimal arguments (threads, blocks, etc.) to launch CUDA kernels using @cuda macro? In the following example I consider a simple kernel which does vector addition of two matrices: using CuArrays u…

---

## [Julia Cuda  Matrix multiplication](https://discourse.julialang.org/t/julia-cuda-matrix-multiplication/21741)

<div class="topic-metadata">

**Author:** [@Noobie76](https://discourse.julialang.org/u/Noobie76)\
**Replies:** 3\
**Last updated:** [February 24, 2021, 6:52pm UTC](https://discourse.julialang.org/t/julia-cuda-matrix-multiplication/21741 "2021-02-24T18:52:38Z")

</div>

Hi, I’m relatively new to Julia and want to implement a numerical method using the CUDA libraries for Julia. I worked myself through the introduction files on GitHub and gained all the basic knowledge to write my own co…

---

## [Flux Pendulum DDPG example fails on GPU](https://discourse.julialang.org/t/flux-pendulum-ddpg-example-fails-on-gpu/49975)

<div class="topic-metadata">

**Author:** [@lilanger](https://discourse.julialang.org/u/lilanger)\
**Replies:** 2\
**Last updated:** [November 12, 2020, 9:52am UTC](https://discourse.julialang.org/t/flux-pendulum-ddpg-example-fails-on-gpu/49975 "2020-11-12T09:52:55Z")

</div>

I am pretty new to Flux and try to get the Pendulum environment (using a slightly modified Reinforce.jl environment: Pendulum update) running with DDPG with newer package versions (based on Flux model-zoo). It seems to …

---

## [Column-wise reduction on a CUDA.CuArray matrix](https://discourse.julialang.org/t/column-wise-reduction-on-a-cuda-cuarray-matrix/45069)

<div class="topic-metadata">

**Author:** [@mkarikom](https://discourse.julialang.org/u/mkarikom)\
**Replies:** 0\
**Last updated:** [August 16, 2020, 10:33pm UTC](https://discourse.julialang.org/t/column-wise-reduction-on-a-cuda-cuarray-matrix/45069 "2020-08-16T22:33:44Z")

</div>

As suggested here, I am trying to create a version of this findfirst kernel that operates over dimension 2 of input matrix xs and returns the first match for each vector along dimension 1. I think a good approach is a c…

---

## [PTX JIT compilation failed with Flux and CUDAnative while using Flux](https://discourse.julialang.org/t/ptx-jit-compilation-failed-with-flux-and-cudanative-while-using-flux/41034)

<div class="topic-metadata">

**Author:** [@bafonso](https://discourse.julialang.org/u/bafonso)\
**Replies:** 3\
**Last updated:** [June 9, 2020, 2:37pm UTC](https://discourse.julialang.org/t/ptx-jit-compilation-failed-with-flux-and-cudanative-while-using-flux/41034 "2020-06-09T14:37:55Z")

</div>

This one is a bit weird, not sure it should be here or in Juno? Anyway, this code runs fine via reply from terminal but I get an error running this in Juno once it is using the GPU. The second part errors out. using Flu…

---

## [CUDAnative out of resources, but only when run from Atom/Juno?](https://discourse.julialang.org/t/cudanative-out-of-resources-but-only-when-run-from-atom-juno/40428)

<div class="topic-metadata">

**Author:** [@Alex\_Ellison](https://discourse.julialang.org/u/Alex_Ellison)\
**Replies:** 2\
**Last updated:** [May 29, 2020, 10:20pm UTC](https://discourse.julialang.org/t/cudanative-out-of-resources-but-only-when-run-from-atom-juno/40428 "2020-05-29T22:20:16Z")

</div>

Hey GPU Gang, I have a kernel I’m trying to execute with CUDAnative. I’ve tested it with Nvidia Visual Profiler and it uses 46 registers/thread, and I can achieve the theoretical max of 1024 threads/block when I run it …

---

## [Crashes and high utilization while training with Flux with GPU](https://discourse.julialang.org/t/crashes-and-high-utilization-while-training-with-flux-with-gpu/39626)

<div class="topic-metadata">

**Author:** [@PeterKeffer](https://discourse.julialang.org/u/PeterKeffer)\
**Replies:** 2\
**Last updated:** [May 17, 2020, 12:40pm UTC](https://discourse.julialang.org/t/crashes-and-high-utilization-while-training-with-flux-with-gpu/39626 "2020-05-17T12:40:27Z")

</div>

Hey there lovely Julia community! I’m quite new to Julia, but I’m so excited about all the awesome packages and the wonderful community. I’ve already used PyTorch a lot, also with GPU on several machines, and got it alw…

---

## [Virtual GPU office hours](https://discourse.julialang.org/t/virtual-gpu-office-hours/38385)

<div class="topic-metadata">

**Author:** [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Replies:** 0\
**Last updated:** [April 28, 2020, 5:09pm UTC](https://discourse.julialang.org/t/virtual-gpu-office-hours/38385 "2020-04-28T17:09:41Z")

</div>

We have decided to experiment with a new format and decided to hold virtual office hours for GPU programming in Julia. So everything that is under the JuliaGPU · GitHub umbrella. Feel free to drop by and to discuss your…

---

## [Use GPU subfunction in a bigger-function?](https://discourse.julialang.org/t/use-gpu-subfunction-in-a-bigger-function/37768)

<div class="topic-metadata">

**Author:** [@Ahmed\_Salih](https://discourse.julialang.org/u/Ahmed_Salih)\
**Replies:** 5\
**Last updated:** [April 19, 2020, 6:18pm UTC](https://discourse.julialang.org/t/use-gpu-subfunction-in-a-bigger-function/37768 "2020-04-19T18:18:59Z")

</div>

Hello! My question title is properly worded poorly, but here goes. I have made a “subfunction”, called gpu\_ParticlePackStep! which calculates properties for one particle. If I call this function as such in Julia: gpu\_P…

---

## [CUDAnative dynamic allocation](https://discourse.julialang.org/t/cudanative-dynamic-allocation/35435)

<div class="topic-metadata">

**Author:** [@Michael\_Benton](https://discourse.julialang.org/u/Michael_Benton)\
**Replies:** 5\
**Last updated:** [March 4, 2020, 6:29pm UTC](https://discourse.julialang.org/t/cudanative-dynamic-allocation/35435 "2020-03-04T18:29:52Z")

</div>

I need to be able to dynamically allocate in kernel. val = MArray{Tuple{4}, Float32}(undef) is a way to get static allocation with StaticArrays.jl inside a CUDA kernel I figured something like Array{Float32}(undef, 4) …

---

## [Improving cache behavior with ldg()](https://discourse.julialang.org/t/improving-cache-behavior-with-ldg/33628)

<div class="topic-metadata">

**Author:** [@Michael\_Benton](https://discourse.julialang.org/u/Michael_Benton)\
**Replies:** 6\
**Last updated:** [February 6, 2020, 11:56pm UTC](https://discourse.julialang.org/t/improving-cache-behavior-with-ldg/33628 "2020-02-06T23:56:02Z")

</div>

What exactly is meant by https://juliagpu.github.io/CUDAnative.jl/stable/lib/device/array.html that ldg “loads the value through the read-only texture cache for improved cache behavior” ? The following is a memory band…

---

## [Debugging in CUDAnative](https://discourse.julialang.org/t/debugging-in-cudanative/34193)

<div class="topic-metadata">

**Author:** [@Michael\_Benton](https://discourse.julialang.org/u/Michael_Benton)\
**Replies:** 1\
**Last updated:** [February 5, 2020, 7:01am UTC](https://discourse.julialang.org/t/debugging-in-cudanative/34193 "2020-02-05T07:01:51Z")

</div>

struggling to find the source of problems in some kernels I write due to I don’t know of a julia CUDA debugger, I’ve found some for C. I get: CUDA error: unspecified launch failure (code 719, ERROR\_LAUNCH\_FAILED) and …

---

## [Passing array of pointers to CUDA kernel](https://discourse.julialang.org/t/passing-array-of-pointers-to-cuda-kernel/33570)

<div class="topic-metadata">

**Author:** [@Michael\_Benton](https://discourse.julialang.org/u/Michael_Benton)\
**Replies:** 2\
**Last updated:** [January 20, 2020, 5:00pm UTC](https://discourse.julialang.org/t/passing-array-of-pointers-to-cuda-kernel/33570 "2020-01-20T17:00:09Z")

</div>

I need to be able to stack arbitrarily sized cuarrays and CuTextures as a single input to a kernel, so I can do operations in the kernel that loop over the data like so: (get GitHub - cdsousa/CuTextures.jl: \[DEPRECATED,…

---

## [Faster read only memory](https://discourse.julialang.org/t/faster-read-only-memory/32457)

<div class="topic-metadata">

**Author:** [@Michael\_Benton](https://discourse.julialang.org/u/Michael_Benton)\
**Replies:** 5\
**Last updated:** [January 8, 2020, 6:50am UTC](https://discourse.julialang.org/t/faster-read-only-memory/32457 "2020-01-08T06:50:57Z")

</div>

I’ve got something that works, but it needs to work faster. I’ve tested different methods outside of Julia that can do this much faster, and by comparison with peakflops.jl and a rudimentary calculation of how many opera…

---

## [Composite types array, and composite type with array fields](https://discourse.julialang.org/t/composite-types-array-and-composite-type-with-array-fields/32466)

<div class="topic-metadata">

**Author:** [@garychow](https://discourse.julialang.org/u/garychow)\
**Replies:** 5\
**Last updated:** [December 20, 2019, 3:43am UTC](https://discourse.julialang.org/t/composite-types-array-and-composite-type-with-array-fields/32466 "2019-12-20T03:43:18Z")

</div>

Hi, I am using CUDAnative now and wonder if anyone could point me to a correct direction such that these operations is possible. Is there anyone else found that it will be easier if these 2 functions are supported? I un…

---

## [Problem with GPU programming](https://discourse.julialang.org/t/problem-with-gpu-programming/28704)

<div class="topic-metadata">

**Author:** [@sergevic](https://discourse.julialang.org/u/sergevic)\
**Replies:** 4\
**Last updated:** [September 13, 2019, 4:09pm UTC](https://discourse.julialang.org/t/problem-with-gpu-programming/28704 "2019-09-13T16:09:20Z")

</div>

I have the following code that I can run on Julia v.1.1: using LinearAlgebra L=Symmetric(rand(Float32,10000,10000)) C=zeros(Float32,size(L,1),size(L,2)) D=zeros(Float32,size(L,1),size(L,2)) D\[1,1\]=1 k=750 function Lap…

---

## [CUDAnative , performance drop after several timesteps](https://discourse.julialang.org/t/cudanative-performance-drop-after-several-timesteps/28495)

<div class="topic-metadata">

**Author:** [@YJC](https://discourse.julialang.org/u/YJC)\
**Replies:** 5\
**Last updated:** [September 8, 2019, 7:43am UTC](https://discourse.julialang.org/t/cudanative-performance-drop-after-several-timesteps/28495 "2019-09-08T07:43:33Z")

</div>

Hi all, I am building up a time marching simulation via using CUDAnative, I found that after several time steps , the performance drop significantly. If I “synchronize” all the thread at some point, the performance was…

---

## [ERROR: InvalidIRError: compiling reduce\_kernel with \`eigen\` command](https://discourse.julialang.org/t/error-invalidirerror-compiling-reduce-kernel-with-eigen-command/28307)

<div class="topic-metadata">

**Author:** [@sergevic](https://discourse.julialang.org/u/sergevic)\
**Replies:** 4\
**Last updated:** [September 3, 2019, 12:31am UTC](https://discourse.julialang.org/t/error-invalidirerror-compiling-reduce-kernel-with-eigen-command/28307 "2019-09-03T00:31:34Z")

</div>

I’m trying to convert my Julia code with CUDA programming, because I get a OutOfMemory() error when I run the original code on Julia v.1.2. My converted code starts with the following lines: using CuArrays using Dista…

---

## [ERROR: Package CuArrays errored during testing](https://discourse.julialang.org/t/error-package-cuarrays-errored-during-testing/28006)

<div class="topic-metadata">

**Author:** [@sergevic](https://discourse.julialang.org/u/sergevic)\
**Replies:** 2\
**Last updated:** [August 27, 2019, 4:31pm UTC](https://discourse.julialang.org/t/error-package-cuarrays-errored-during-testing/28006 "2019-08-27T16:31:59Z")

</div>

I followed all the steps in https://juliagpu.gitlab.io/CuArrays.jl/tutorials/generated/intro/ to install Cuda for using it in Julia v. 1.1. I also installed the CUDA Toolkit v. 10.1. After the installation, I run test Cu…

---

## [Initializing CuArrays](https://discourse.julialang.org/t/initializing-cuarrays/27976)

<div class="topic-metadata">

**Author:** [@wiktorkujawa](https://discourse.julialang.org/u/wiktorkujawa)\
**Replies:** 0\
**Last updated:** [August 25, 2019, 11:24pm UTC](https://discourse.julialang.org/t/initializing-cuarrays/27976 "2019-08-25T23:24:39Z")

</div>

How to initialize this type of array: ArrayName=cu(fill(\[0.0f0,0.0f0\],(20,20,5))), I don’t want to use Tuples because there are too many problems with math operations doing on them(they can’t for example increment and …

---

## [Using BenchmarkTools with CUDAnative and CuArrays and running out of CPU or GPU memory](https://discourse.julialang.org/t/using-benchmarktools-with-cudanative-and-cuarrays-and-running-out-of-cpu-or-gpu-memory/27278)

<div class="topic-metadata">

**Author:** [@Josiah\_Slack](https://discourse.julialang.org/u/Josiah_Slack)\
**Replies:** 0\
**Last updated:** [August 7, 2019, 5:41pm UTC](https://discourse.julialang.org/t/using-benchmarktools-with-cudanative-and-cuarrays-and-running-out-of-cpu-or-gpu-memory/27278 "2019-08-07T17:41:54Z")

</div>

I’m doing some simple “get acquainted” experimenting with convolutions using DSP, CUDAnative, CUDAdrv and CuArrays. I’m creating random 3-d arrays - rand(Float32, N, N, N). I then create “device” versions of the arrays b…

---

## [QuadGK don't work in CUDA kernel](https://discourse.julialang.org/t/quadgk-dont-work-in-cuda-kernel/26754)

<div class="topic-metadata">

**Author:** [@wiktorkujawa](https://discourse.julialang.org/u/wiktorkujawa)\
**Replies:** 6\
**Last updated:** [July 30, 2019, 9:05pm UTC](https://discourse.julialang.org/t/quadgk-dont-work-in-cuda-kernel/26754 "2019-07-30T21:05:37Z")

</div>

I have problem with executing Julia function: function biotSavartCalculation(x,y,z,Segment,I,SegmentsOnElement,segmentlength,Bx,By,Bz) xIndex=(blockIdx().x-1) \* blockDim().x + threadIdx().x yIndex=(blockIdx().y-1) \*…

---

## [CUDAdrv fails to build](https://discourse.julialang.org/t/cudadrv-fails-to-build/26476)

<div class="topic-metadata">

**Author:** [@knavely](https://discourse.julialang.org/u/knavely)\
**Replies:** 0\
**Last updated:** [July 17, 2019, 6:47pm UTC](https://discourse.julialang.org/t/cudadrv-fails-to-build/26476 "2019-07-17T18:47:33Z")

</div>

Got it working but not quite sure exactly how so im just deleting this

---

## [Error in Cuda function : ERROR: LoadError: CuError(1, nothing)](https://discourse.julialang.org/t/error-in-cuda-function-error-loaderror-cuerror-1-nothing/26808)

<div class="topic-metadata">

**Author:** [@wiktorkujawa](https://discourse.julialang.org/u/wiktorkujawa)\
**Replies:** 8\
**Last updated:** [July 26, 2019, 10:11pm UTC](https://discourse.julialang.org/t/error-in-cuda-function-error-loaderror-cuerror-1-nothing/26808 "2019-07-26T22:11:04Z")

</div>

I have this error while running Cuda ERROR: LoadError: CuError(1, nothing) Stacktrace: \[1\] (::getfield(CUDAdrv, Symbol("##25#26")){Bool,Int64,CuStream,CuFunction})(::Array{Ptr{Nothing},1}) at C:\\Users\\Wiktor\\.julia\\pac…

---

## [Cuda - run code with list](https://discourse.julialang.org/t/cuda-run-code-with-list/26441)

<div class="topic-metadata">

**Author:** [@wiktorkujawa](https://discourse.julialang.org/u/wiktorkujawa)\
**Replies:** 0\
**Last updated:** [July 16, 2019, 8:54pm UTC](https://discourse.julialang.org/t/cuda-run-code-with-list/26441 "2019-07-16T20:54:05Z")

</div>

I wanted to get line coordinates from startpoint to endpoint, but it didn’t work. I guess that the problem is in data type of Segment list(and overally a lists). In my program I will use lists to load data. using CUDAd…

[Next page](https://discourse.julialang.org/tag/cudanative/50.md?match_all_tags=true&page=1&tags%5B%5D=cudanative)
