# Are forward-mode derivatives calculated in parallel?

**URL:** <https://discourse.julialang.org/t/are-forward-mode-derivatives-calculated-in-parallel/128870>\
**Category:** Machine Learning\
**Tags:** forwarddiff, autodiff\
**Created:** [May 9, 2025, 10:43am UTC](https://discourse.julialang.org/t/are-forward-mode-derivatives-calculated-in-parallel/128870 "2025-05-09T10:43:50Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bizzi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bizzi/32/51484_2.png) [@Bizzi](https://discourse.julialang.org/u/Bizzi)\
**Post date:** [May 9, 2025, 10:43am UTC](https://discourse.julialang.org/t/are-forward-mode-derivatives-calculated-in-parallel/128870/1 "2025-05-09T10:43:50Z")

</div>

When I calculate the output of a function and a directional derivative using forward-mode autodiff (ForwardDiff, TaylorDiff etc), is the calculation of the derivatives carried out in parallel?

By parallel I mean: assuming both my function and its derivatives are gpu-friendly (i.e. a Neural Network), should calculating its derivatives along with the result make overall evaluation slower? Or could I hope for “free” derivatives given enough memory?

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [May 9, 2025, 10:57am UTC](https://discourse.julialang.org/t/are-forward-mode-derivatives-calculated-in-parallel/128870/2 "2025-05-09T10:57:47Z")

</div>

Just to clarify, I think there are three ways to interpret your question:

1. If the function runs on the GPU (parallelizing over input), does the derivative run on the GPU too?
2. Assuming the derivative runs on the GPU (parallelizing over input), how much will one derivative slow down the primal program?
3. Assuming the derivative runs on the GPU (parallelizing over input), how many derivatives are computed simultaneously (parallelizing over directions/tangents)?

* * *

As far as ForwardDiff.jl is concerned:

1. Depends on the operator. If I remember correctly, `ForwardDiff.derivative` will run fine on GPU arrays, but `ForwardDiff.gradient` will fail due to scalar indexing.
2. In forward mode, autodiff theory tells us that evaluation shouldn’t be slowed down too much when derivatives are propagated alongside the primals… but that’s not true in practice. For instance, if the primal function takes an optimized code path for `Matrix{Float64}` (like a BLAS call), the derivative requires working with `Matrix{Dual{Float64}}`, which is much slower because it is pure Julia code.
3. This is determined by the so-called [chunk size](https://juliadiff.org/ForwardDiff.jl/stable/user/advanced/#Configuring-Chunk-Size), which determines how many [partials](https://juliadiff.org/ForwardDiff.jl/stable/dev/how_it_works/) are stored in the derivative tuple.
