# How to do gradient clipping in Julia for large for loops

**URL:** <https://discourse.julialang.org/t/how-to-do-gradient-clipping-in-julia-for-large-for-loops/108402>\
**Category:** General Usage\
**Tags:** ad, autodiff\
**Created:** [January 5, 2024, 6:26pm UTC](https://discourse.julialang.org/t/how-to-do-gradient-clipping-in-julia-for-large-for-loops/108402 "2024-01-05T18:26:45Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Devetak](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/devetak/32/50611_2.png) [@Devetak](https://discourse.julialang.org/u/Devetak)\
**Post date:** [January 5, 2024, 6:26pm UTC](https://discourse.julialang.org/t/how-to-do-gradient-clipping-in-julia-for-large-for-loops/108402/1 "2024-01-05T18:26:45Z")

</div>

Hello.

I have a function that has many for loops. For example:

```julia
function f(x)
           res = 1.0
           for i in 1:1000000
               res = res * x^2
           end
           return res / x
           end

```

of course then

```julia
ForwardDiff.derivative(f, 2.0) == Inf

```

how do I do gradient clipping? I am not sure it is called like that. What I want to do is to normalize the gradient after every iteration of the for loop so that the final answer is a direction which makes sense.

In particular my input is an array and I want to preserve the relative importance of the gradients. Say `x[1]` is very impactful so it has grarient `0.999`and `x[2]` is not very impactful and it has gradient `0.00001` and so on…

I am not sure on how to do it/how to actually make this work or MORE IMPORTANTLY if there are better things to do in this case.

PS: Maybe this is not the best example as `f(2.0) == Inf`. In practice my gradient comes out to be `Nan`, which is less informative that `Inf`.

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [January 5, 2024, 8:08pm UTC](https://discourse.julialang.org/t/how-to-do-gradient-clipping-in-julia-for-large-for-loops/108402/2 "2024-01-05T20:08:33Z")

</div>

Hi,  
Can you tell us a little more about the context, why you need AD, why your function might diverge, etc?

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [January 6, 2024, 12:40am UTC](https://discourse.julialang.org/t/how-to-do-gradient-clipping-in-julia-for-large-for-loops/108402/3 "2024-01-06T00:40:51Z")

</div>

> [@Devetak](#):
>
> Maybe this is not the best example as `f(2.0) == Inf`. In practice my gradient comes out to be `Nan`, which is less informative that `Inf`.

Do you have an example of a function that is finite and differentiable but the gradient still comes out to be `NaN`?

If your function is theoretically finite but is overflowing to `Inf` due to the finite floating-point precision, maybe you should instead compute the logarithm or similar. For example,

```julia
function logf(x)
    logres = 0.0
    for i in 1:1000000
        logres += 2 * log(x)
    end
    return logres - log(x)
end

```

computes the logarithm of your function `f(x)` above, but without overflowing, and both `logf` and its derivative work fine:

```julia
julia> logf(2.0)
1.3862936679852684e6

julia> ForwardDiff.derivative(logf, 2.0)
999999.5

```

---

<div class="post-metadata">

**Author:** ![Devetak](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/devetak/32/50611_2.png) [@Devetak](https://discourse.julialang.org/u/Devetak)\
**Post date:** [January 18, 2024, 10:10am UTC](https://discourse.julialang.org/t/how-to-do-gradient-clipping-in-julia-for-large-for-loops/108402/4 "2024-01-18T10:10:53Z")

</div>

Hi. Sorry for the radio silence I was sick.

The NaNs were part of an unrelated bug. Nevertheless to answer @gdalle’s question.

I just have a very “long” function (ie. many loops) that I want to take a derivative of. Its quite similar to applying a RNN multiple times. Of course then the gradients explode or go to zero.

I read that in that case gradient clipping helps. I wanted to implement the same in Forwardiff(and/or Zygote) but could not find any resource.

To make it clearer. What I want is after every iteration of the for loop to renormalize the gradient so that it will never blow up.

I tought this could be a nice post to have for the community.
