# Automatic Differentiation and Parallelism

**URL:** <https://discourse.julialang.org/t/automatic-differentiation-and-parallelism/9478>\
**Category:** Optimization (Mathematical)\
**Tags:** parallel, optimization\
**Created:** [March 3, 2018, 7:15pm UTC](https://discourse.julialang.org/t/automatic-differentiation-and-parallelism/9478 "2018-03-03T19:15:26Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juser](https://avatars.discourse-cdn.com/v4/letter/j/34f0e0/32.png) [@Juser](https://discourse.julialang.org/u/Juser)\
**Post date:** [March 3, 2018, 7:15pm UTC](https://discourse.julialang.org/t/automatic-differentiation-and-parallelism/9478/1 "2018-03-03T19:15:26Z")

</div>

I am trying to optimize a function that is expensive to evaluate but has many component parts that can be computed in parallel. However, DualNumbers are not bits-types and therefore cannot be stored in SharedArrays. Is there anything that can be done analogously to using SharedArrays to avoid reallocating and distributing memory between many processors for intermediate steps that contain DualNumbers?

As an example, let’s say we wanted to minimize f(X,theta) with respect to theta, where X is a large SharedArray not modified by f. (We use a SharedArray to avoid the overhead of passing X to all of the processors many times.)

As an example: let f(X,theta)= sum\_k h(x\_k,g(X,theta)), where x\_k is the k-th row of X. It is clear that once g=g(X,theta) is computed that f(X,theta) = sum\_k h(x\_k,g) is embarrassingly parallelizeable, so the “obvious” way to evaluate f(X,theta) is to just compute g once and parallelize the computation of sum\_k h(x\_k,g) along k. However, during optimization using automatic differentiation, g will take on a dual-number. That means that we can’t just stick g in a SharedArray. Instead we have to pass g to each processor, even though g isn’t modified at all during the parallel computation of sum\_k h(x\_k,g).

Does anyone have a suggestion for a good way to avoid reallocating and passing g around many times in such a case, especially since it isn’t even modified during the parallel computation.

Thank you!

---

<div class="post-metadata">

**Author:** ![tkoolen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkoolen/32/1603_2.png) [@tkoolen](https://discourse.julialang.org/u/tkoolen)\
**Post date:** [March 3, 2018, 7:34pm UTC](https://discourse.julialang.org/t/automatic-differentiation-and-parallelism/9478/2 "2018-03-03T19:34:07Z")

</div>

> DualNumbers are not bits-types

Assuming you’re talking about `DualNumber`s from DualNumbers.jl, they are as long as `T` is a bitstype (you can check with the `isbits` function).

> <https://github.com/JuliaDiff/DualNumbers.jl/blob/c82597b4e1ef4ee15b10ac6ee85a6d6bcc5f70e3/src/dual.jl#L3-L6>

---

<div class="post-metadata">

**Author:** ![Juser](https://avatars.discourse-cdn.com/v4/letter/j/34f0e0/32.png) [@Juser](https://discourse.julialang.org/u/Juser)\
**Post date:** [March 3, 2018, 8:30pm UTC](https://discourse.julialang.org/t/automatic-differentiation-and-parallelism/9478/3 "2018-03-03T20:30:33Z")

</div>

@tkoolen, thanks for letting me know. I did not know that and just verified that one can, in fact, generate a SharedArray with element-type Dual.

Given this, my current thinking is to implement g as a SharedArray of DualNumbers.Dual{Float64} and then just fill that same vector with either Floats or Duals depending on whether I’m on a gradient or evaluation step. I imagine this would add some unnecessary overhead in doing computations for the “epsilon” part of the Duals even when they are all zero (i.e. in the function evaluation step). However, there may be clever ways to fill a SharedArray{Float64} in function calls and a SharedArray{Dual{Float64}} in gradient calls.

Thanks for the help!
