# Auto-diff Friendly GPU Stencils

**URL:** <https://discourse.julialang.org/t/auto-diff-friendly-gpu-stencils/92126>\
**Category:** GPU\
**Created:** [December 26, 2022, 6:53am UTC](https://discourse.julialang.org/t/auto-diff-friendly-gpu-stencils/92126 "2022-12-26T06:53:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dreycen\_Foiles](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dreycen_foiles/32/44744_2.png) [@Dreycen\_Foiles](https://discourse.julialang.org/u/Dreycen_Foiles)\
**Post date:** [December 26, 2022, 6:53am UTC](https://discourse.julialang.org/t/auto-diff-friendly-gpu-stencils/92126/1 "2022-12-26T06:53:56Z")

</div>

I’m solving a semi-discretized PDE on my GPU and one of the steps involves performing the Laplacian via a 5-point stencil. My current method does this with a CUDA kernel, but I also want to perform automatic differentiation so that I can use a stiff ODE solver. It seems pretty clear to me that a ForwardDiff dual will not work with a CUDA kernel because it’s not a bitstype.

Are there any recommended ways to either get automatic differentiation to work with a generic CUDA kernel or a better way to implement a stencil on a GPU that is compatible with ForwardDiff?

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [December 26, 2022, 12:27pm UTC](https://discourse.julialang.org/t/auto-diff-friendly-gpu-stencils/92126/2 "2022-12-26T12:27:56Z")

</div>

Handwritten CUDA.jl kernels will recompile with different number types so that’s fine. Though stencils are linear operators so directly defining their derivatives isn’t hard: I’d just define the derivative overload for NNLib.jl’s `conv` and use it.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [December 26, 2022, 3:21pm UTC](https://discourse.julialang.org/t/auto-diff-friendly-gpu-stencils/92126/3 "2022-12-26T15:21:05Z")

</div>

You could try things like `@tullio y[i] := -x[i] + 2x[i+1] - x[i+2]`, I believe such cases ought to be fairly efficient (including GPU & derivatives).

(You may also be interested in [ParallelStencil.jl](https://github.com/omlins/ParallelStencil.jl) but not sure this will help with derivatives.)

---

<div class="post-metadata">

**Author:** ![luraess](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/luraess/32/16189_2.png) [@luraess](https://discourse.julialang.org/u/luraess)\
**Post date:** [January 2, 2023, 9:16pm UTC](https://discourse.julialang.org/t/auto-diff-friendly-gpu-stencils/92126/4 "2023-01-02T21:16:34Z")

</div>

Not using ForwardDiff though, you could have a look at using [Enzyme.jl](https://enzyme.mit.edu/julia/) for differentiating GPU kernels. An example can be found here [GitHub - PTsolvers/PT-AD: Pseudo-transient auto-diff playground](https://github.com/PTsolvers/PT-AD). There we use Enzyme to get the Jacobian vector product and gradient of the cost function from AD in the scope of an adjoint-based gradient descent inversion procedure having a nonlinear diffusion equation as forward problem.
