# ODE solver algorithms that exploit massive parallelism (eg GPU)?

**URL:** <https://discourse.julialang.org/t/ode-solver-algorithms-that-exploit-massive-parallelism-eg-gpu/82726>\
**Category:** Modelling & Simulations\
**Tags:** gpu, ode\
**Created:** [June 14, 2022, 7:09am UTC](https://discourse.julialang.org/t/ode-solver-algorithms-that-exploit-massive-parallelism-eg-gpu/82726 "2022-06-14T07:09:11Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![marius311](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marius311/32/3953_2.png) [@marius311](https://discourse.julialang.org/u/marius311)\
**Post date:** [June 14, 2022, 7:09am UTC](https://discourse.julialang.org/t/ode-solver-algorithms-that-exploit-massive-parallelism-eg-gpu/82726/1 "2022-06-14T07:09:11Z")

</div>

My scenario is that I’m solving some stiff ODE of O(100) variables with an ensemble of O(200) parameters. KenCarp4 or Rodas5 work the best that we’ve found. The problem is mildly sparse, but not super disconnected, and on CPU at least using sparse jacobian is slower due to overhead.

We’re working on getting this on GPU, but even using something like `EnsembleProblem` to parallelize across the ensemble of 200 parameters leaves \>95% of a modern GPU’s threads idle in a given timestep, so it seems like there’s huge potential for further speedup.

Are there stiff algorithms that further exploit this parallelism available? I know Runge-Kutta (not that we’re using it here) has parallel tableaus, but even those only really give opportunity for small factors speedups. Is there anything better? Given that at every time step for every parameter I can pretty much evaluate my ODE function at 100 extra values of `t` and `u` for free on a modern GPU, seems I must be able to reduce the number of time-steps needed? Or are there any other strategies one could use to saturate a GPU and speed this up? Thanks for any advice.

---

<div class="post-metadata">

**Author:** ![rejuvyesh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rejuvyesh/32/91_2.png) [@rejuvyesh](https://discourse.julialang.org/u/rejuvyesh)\
**Post date:** [June 14, 2022, 7:58pm UTC](https://discourse.julialang.org/t/ode-solver-algorithms-that-exploit-massive-parallelism-eg-gpu/82726/2 "2022-06-14T19:58:37Z")

</div>

[https://github.com/SciML/DiffEqGPU.jl/pull/148](https://github.com/SciML/DiffEqGPU.jl/pull/148) might be of interest? Although not a stiff solver.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [June 14, 2022, 11:52pm UTC](https://discourse.julialang.org/t/ode-solver-algorithms-that-exploit-massive-parallelism-eg-gpu/82726/3 "2022-06-14T23:52:45Z")

</div>

Yeah that’s the direction we’re heading. It’s an improvement of over 100x for the non-stiff case. We need to do it for the stiff ODE case too.

---

<div class="post-metadata">

**Author:** ![marius311](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marius311/32/3953_2.png) [@marius311](https://discourse.julialang.org/u/marius311)\
**Post date:** [June 15, 2022, 1:52am UTC](https://discourse.julialang.org/t/ode-solver-algorithms-that-exploit-massive-parallelism-eg-gpu/82726/4 "2022-06-15T01:52:45Z")

</div>

Thanks for the replies. Is there any reference for what these GPU versions of these methods do exactly?

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [June 15, 2022, 6:00am UTC](https://discourse.julialang.org/t/ode-solver-algorithms-that-exploit-massive-parallelism-eg-gpu/82726/5 "2022-06-15T06:00:03Z")

</div>

Not right now. They are pre pre paper
