# How to run ptx code on CUDA from julia?

**URL:** <https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306>\
**Category:** General Usage\
**Tags:** gpu, cudanative, cuda\
**Created:** [January 26, 2024, 1:31pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306 "2024-01-26T13:31:35Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 1:31pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/1 "2024-01-26T13:31:35Z")

</div>

Hi

I am trying to create a NN library that directly output a very optimized version of ptx code, also would be open to output even lower level code if possible. But of course I see I have to go step by step as it is getting exponentially more and more complex to go lower and have to understand times more details.

Can anyone help me on how to run, ptx or lower level of code or even binary on the CUDA gpu?

I saw `CUDA.jl` has `CuDeviceFunction` and `cufunction` that could be used, but couldn’t understand how could I use this to run low level codes.

Thank you for any help!

---

<div class="post-metadata">

**Author:** ![Pangoraw](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pangoraw/32/24719_2.png) [@Pangoraw](https://discourse.julialang.org/u/Pangoraw)\
**Post date:** [January 26, 2024, 1:52pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/2 "2024-01-26T13:52:43Z")

</div>

You can use the lower-level `CuModule` to load a cubin binary and `CuFunction(mod name)` to create a cufunction. ptxas can be used to get cubin from the ptx.

```julia

using CUDA

N, D = 1, 1

a = cu(randn(Float32,N,D));
b = cu(randn(Float32,N,D));
out = cu(zeros(Float32,N,N));

ptx_path = "output.ptx"
cubin_path = "output.cubin"

run(`$(CUDA.ptxas()) --gpu-name sm_75 $ptx_path --output-file $cubin_path --verbose`)

mod = CuModule(read(cubin_path))
func = CuFunction(mod, "pairwise_l2_kernel_0d1d2d3c4c5c")

CUDA.@sync CUDA.cudacall(
    func, (CuPtr{Float32}, CuPtr{Float32}, CuPtr{Float32}),
    a, b, out;
    blocks=1, threads=1, shmem=512,
)

```

---

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 2:14pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/3 "2024-01-26T14:14:22Z")

</div>

Wow! This is crazy if I can run `.cubin` files basically!

---

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 2:51pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/4 "2024-01-26T14:51:41Z")

</div>

Any idea on why I get error at this line:

```julia
mod = CuModule(read(cubin_path))

```

`ERROR: CUDA error: device kernel image is invalid (code 300, ERROR_INVALID_SOURCE)`

For you this code works perfectly do I assume it correctly? I tried multiple different working ptx files but always the same message, even tried to change the sm\_75 to sm\_52.

The code `module.jl` where the error comes. So I guess res is something “INVALID\_SOURCE”

 ![image](https://global.discourse-cdn.com/julialang/original/3X/b/9/b9ce2b8b2c4402ad0d9fc66bd6f5804a3a86a48b.png)

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [January 26, 2024, 3:02pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/5 "2024-01-26T15:02:02Z")

</div>

It’s not strictly required to compile PTX to a CUBIN; you can invoke `CuModule` with PTX code too, and have the driver JIT-compile it to native code.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [January 26, 2024, 3:04pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/6 "2024-01-26T15:04:59Z")

</div>

Apart from that though, `cuModuleLoadDataEx` should support both CUBIN and PTX input. Maybe verify that the cubin is valid (e.g. pass it to `cuobjdump` or `nvdisasm` or so) and check that the buffer is NULL terminated?

---

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 3:07pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/7 "2024-01-26T15:07:11Z")

</div>

With that way, it worked. No idea what could be the problem with my explicit build. I will look into your ptx build process too.

Also this way I stuck at the next line:

```julia
func = CuFunction(mod, "pairwise_l2_kernel_0d1d2d3c4c5c")

```

Where is this functionname comes from? “pairwise\_l2\_kernel\_0d1d2d3c4c5c”

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [January 26, 2024, 3:08pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/8 "2024-01-26T15:08:17Z")

</div>

That’s the kernel you define in your PTX input.

---

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 3:11pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/9 "2024-01-26T15:11:11Z")

</div>

> [@maleadt](#):
>
> pass it to `cuobjdump` or `nvdisasm` or so

`cuobjdump` return with nothing.  
`nvdisasm` prints out the asm.  
I guess this will be good then.

---

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 3:14pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/10 "2024-01-26T15:14:46Z")

</div>

I could use the “test\_sum” in this case.

Now I crashed my gpu by running the code. So there must be some mistake with the ptrs. 😃  
`ERROR: CUDA error: an illegal memory access was encountered (code 700, ERROR_ILLEGAL_ADDRESS)`

I think I made it based on your recommendations, just this crash has to be solved! Will be back if everything works perfectly.

---

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 3:27pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/11 "2024-01-26T15:27:44Z")

</div>

I modified this `.cu` file and regenerated the `.ptx` file to be sure that it doesn’t access the edge of the arrays just to make sure. No success.

Tried everything I could and now I made it work with running with `block=1`, `threads=1`:

```julia
CUDA.@sync CUDA.cudacall(
    func, 
    (CuPtr{Float32}, CuPtr{Float32}, CuPtr{Float32}),
    a, b, out;
    blocks=1, threads=1, shmem=0,
)

```

After this I can change threads to anything.  
Really strange behaviour.

And now I tried it multiple times already this way. It always succeeds. What can this problem be.

---

<div class="post-metadata">

**Author:** ![Marcell\_Havlik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marcell_havlik/32/37424_2.png) [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Post date:** [January 26, 2024, 8:42pm UTC](https://discourse.julialang.org/t/how-to-run-ptx-code-on-cuda-from-julia/109306/12 "2024-01-26T20:42:17Z")

</div>

I know it is weekend. But want to share the problem was with my `.cu` kernel. Really. 😃 In the most basic kernel.

The problem was that in the cu kernel I used this:  
`int I = ((blockIdx.x - 1) * blockDim.x) + threadIdx.x+1;`  
instead of:  
`int i = blockIdx.x*blockDim.x + threadIdx.x;`

Thank you for the help for everyone, this is already extraordinary! 🥰
