# Presentation on effective use of CUDAnative/CuArrays

**URL:** https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919
**Category:** GPU
**Created:** [December 22, 2018, 8:13am UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919 "2018-12-22T08:13:57Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)
#### Post date: [December 22, 2018, 8:13am UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/1 "2018-12-22T08:13:57Z")

</div>

Since our GPU stack is pretty ill-documented, I figured it can’t hurt to cross-post any content there is. Here’s a recent presentation of mine that explains how the relevant packages work, with demos, as well as tips/tricks and tools on how to do so effectively: [Julia BeNeLux 2018-12 - GPU Tutorial - Google Slides](https://docs.google.com/presentation/d/1l-BuAtyKgoVYakJSijaSqaTL3friESDyTOnU2OLqGoA/)

---

<div class="post-metadata">

### Author: ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)
#### Post date: [December 23, 2018, 9:56pm UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/2 "2018-12-23T21:56:27Z")

</div>

I don’t suppose there’s a video of that? Looks extremely useful, but for me a little hard to follow

---

<div class="post-metadata">

### Author: ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)
#### Post date: [December 24, 2018, 7:23am UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/3 "2018-12-24T07:23:08Z")

</div>

Sadly, no. I did add presenter notes though, and I could add some more if certain parts are unclear.

---

<div class="post-metadata">

### Author: ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)
#### Post date: [December 26, 2018, 11:22am UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/4 "2018-12-26T11:22:49Z")

</div>

This is beautiful! Thanks for sharing! We should have more of those… 🙂

---

<div class="post-metadata">

### Author: ![floswald](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/floswald/32/195_2.png) [@floswald](https://discourse.julialang.org/u/floswald)
#### Post date: [December 30, 2018, 2:26pm UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/5 "2018-12-30T14:26:10Z")

</div>

hi @maleadt, thanks a lot for sharing those. can i ask a quick question please? I don’t really understand what’s going on with this example here on slide 22:

```julia
function diff_y(a, b)
   a .= @views b[:, 2:end] .- b[:, 1:end-1]
end
# vs
function diff_y(a, b)
   s = size(a)
   for j = 1:s[2]
       @inbounds a[:,j] .= b[:,j+1] - b[:,j]
   end
end

```

so, in the second case, upon doing `@cuda diff_y(a,b)` we would effectively generate one GPU kernel for each `j`, whereas in the first case it’s one unique kernel?

More in general: there is nothing wrong per se to split tasks on the GPU into functions, right? I mean, I could have 2 kernels, where kernel 1 calls kernel 2, instead of stuffing all tasks into one big function? (kernel 2 is written as a regular julia function, i.e. I don’t need `@cuda` in front of it, correct?)  
thanks!

---

<div class="post-metadata">

### Author: ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)
#### Post date: [January 3, 2019, 6:36am UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/6 "2019-01-03T06:36:28Z")

</div>

> [@floswald](#):
>
> in the second case, upon doing `@cuda diff_y(a,b)` we would effectively generate one GPU kernel for each `j` , whereas in the first case it’s one unique kernel?

No, we’re never doing an explicit `@cuda` in these examples, but relying on the array abstractions by CuArrays.jl to call `@cuda` in the implementations of these abstractions. And that’s exactly why the second example is better: only a single fused call to broadcast is executed, resulting in a single `@cuda`, whereas the original version puts that in a loop causing multiple calls to broadcast and consequently `@cuda`.

> [@floswald](#):
>
> More in general: there is nothing wrong per se to split tasks on the GPU into functions, right? I mean, I could have 2 kernels, where kernel 1 calls kernel 2, instead of stuffing all tasks into one big function? (kernel 2 is written as a regular julia function, i.e. I don’t need `@cuda` in front of it, correct?)

Correct, but we don’t call kernel 2 a kernel then, just an ordinary (device) function. Only when launching multiple kernels (ie. multiple calls to `@cuda`, either explicitly or as part of array abstractions from CuArrays.jl), and when those kernels are sufficiently small not to saturate the GPU easily, then fusion makes sense. This typically happens with short operations as the ones you end up with when doing broadcast. When writing your own kernels, it is much easier to saturate the GPU. But again, profile to be sure (see the screenshots at the end of my talk).

---

<div class="post-metadata">

### Author: ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)
#### Post date: [January 3, 2019, 7:24am UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/7 "2019-01-03T07:24:50Z")

</div>

Does that mean that on the GPU we’re back to trying to make code vectorized?

---

<div class="post-metadata">

### Author: ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)
#### Post date: [January 3, 2019, 7:28am UTC](https://discourse.julialang.org/t/presentation-on-effective-use-of-cudanative-cuarrays/18919/8 "2019-01-03T07:28:17Z")

</div>

> [@mkborregaard](#):
>
> Does that mean that on the GPU we’re back to trying to make code vectorized?

Pretty much, although the definition of “vectorized” is much more broad nowadays (with broadcast fusion). Here’s hoping for similar expressibility improvements for other operators like reduce.
