# Understanding the results from Optim.jl for a black box function

**URL:** <https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507>\
**Category:** Optimization (Mathematical)\
**Tags:** optim, optimization\
**Created:** [June 24, 2021, 4:38pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507 "2021-06-24T16:38:55Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 24, 2021, 4:38pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/1 "2021-06-24T16:38:56Z")

</div>

I’m using `Optim.jl` to solve an unconstrained minimization problem.  
In this particular problem I have a _black-box_ function, which can take a long time on a single function evaluation. I don’t have access to gradient information, and even though I have tried to use automatic differentiation, there are some parts of the code that the differentiator cannot handle and throws some errors.

Nevertheless, I’m using the L-BFGS implementation from `Optim.jl` with the finite differences approximation for the gradient information, and it seems to work fine. Unfortunately, I cannot show the full code here because it is not my property, but I can at least show a snippet of the code that does the optimization call

```julia
closure(x) = sqerror_model(x; phi=phi, rho=rho, model=model) # The objective function, a black-box
initial_x = randn(dim)
optim_res = Optim.optimize(
    closure, initial_x, Optim.LBFGS(; m=15), Optim.Options(iterations=200, show_trace=true)
)

```

With this snippet I’m trying to point out that there are 200 iterations and that I want to see the trace for each iteration. I also changed to `m` parameter for the optimization method to get a little bit more precision.

What I get in return is the following

 ![opt-results](https://global.discourse-cdn.com/julialang/original/3X/3/c/3c8d19c16529590d00b7b85ece6b33777137e4f9.png)

Here, the `dimension=3001` is the size of my parameter vector. This is the expected length of my minimizer.

I have some questions from this.

1. What does the gradient norm equal to zero mean?
2. I specified 200 iterations on the optimization call, why do I only see one iteration, namely the zeroth iteration?
3. Why do I see zero on all the convergence measures? And also, what does the NaN mean in this convergence measures?
4. At the end, it says that only one call to both the function and the gradient were done. Does this mean that only one iteration was done?

---

<div class="post-metadata">

**Author:** ![cvanaret](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cvanaret/32/11594_2.png) [@cvanaret](https://discourse.julialang.org/u/cvanaret)\
**Post date:** [June 24, 2021, 6:35pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/2 "2021-06-24T18:35:06Z")

</div>

That looks fishy. Your screenshot indeed suggests that only the initial point is evaluated, then the method terminates because the gradient is 0 (= it is a stationary point).  
Since you’re picking a random initial point, looks like the gradient is not evaluated correctly (if at all). NaN (not a number) probably indicates that, since you stopped at the first iteration, there’s nothing to compare to.

To see a trace at each iteration, you could put some printing instructions in your “closure” function, ie something like:

```julia
function closure(x)
  evaluation = sqerror_model(x; phi=phi, rho=rho, model=model)
  println("Current x: ", x)
  println("Current evaluation: ", evaluation)
  return evaluation
end

```

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 24, 2021, 9:25pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/3 "2021-06-24T21:25:55Z")

</div>

Thanks for the reply.  
I used your suggested code, although not with the printing of the current value of `x` because it is very large, but with the current evaluation I get the feeling that the function is being evaluated a lot of times, I guess that’s because of the gradient estimation that’s taking place.

Anyway, the value is always the same as seen here

```julia
Current evaluation: 0.13872807013353833
Current evaluation: 0.13872807013353833
Current evaluation: 0.13872807013353833
Current evaluation: 0.13872807013353833
Current evaluation: 0.13872807013353833
Current evaluation: 0.13872807013353833
Current evaluation: 0.13872807013353833
Current evaluation: 0.13872807013353833

```

I have now used an initial value of zero everywhere, instead of a random initial point.  
Maybe this does mean that this is a stationary point and no optimization is being done.  
I’ll have to try with a different optimizer to check if the current evaluation changes.

---

<div class="post-metadata">

**Author:** ![cvanaret](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cvanaret/32/11594_2.png) [@cvanaret](https://discourse.julialang.org/u/cvanaret)\
**Post date:** [June 24, 2021, 10:30pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/4 "2021-06-24T22:30:56Z")

</div>

This looks like a line search: BFGS computes a descent direction based on the gradient and tries to determine what fraction along this direction it should move to. But since the gradient is 0, it just stays where it is.

Does BFGS use finite differences by default to approximate the gradient?  
Have you tried finite differences on your function to see what comes out of it? For example, you could try evaluating \frac{f(x + h e\_1) - f(x)}{h} where e\_1 = (1, 0, \ldots, 0) is the first unit basis vector, and h is around 10^{-3}. Do you get something nonzero?

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 25, 2021, 5:41pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/5 "2021-06-25T17:41:33Z")

</div>

> [@cvanaret](#):
>
> Does BFGS use finite differences by default to approximate the gradient?

Yes, `Optim.jl` uses finite differences to approximate the gradient. I’m using this as well because the automatic differentation does not work on my function.

> [@cvanaret](#):
>
> Have you tried finite differences on your function to see what comes out of it?

I’ll test this out soon and let you know, thanks for all the help!

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 25, 2021, 6:27pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/6 "2021-06-25T18:27:12Z")

</div>

I implemented the following solution, very slow and improvised

```julia
function gradient_test(phi, rho, model, dim)
    function closure(x)
        evaluation = sqerror_model(x; phi=phi, rho=rho, model=model)
        println("Current evaluation: ", evaluation)

        return evaluation
    end

    h_val = 1e-5
    for i in eachcol(I(dim))
        init_value = randn(dim) .- 50.0 # Change the evaluation every time!
        gradient_finite = closure(init_value .+ (h_val .* i)) - closure(init_value)
        gradient_finite /= h_val
        display(gradient_finite)
    end

    return nothing
end

```

and obtained the following

```julia
Current evaluation: 75.8385823602473
Current evaluation: 75.8385823602473
0.0
Current evaluation: 75.8385823602473
Current evaluation: 75.8385823602473
0.0
Current evaluation: 75.8385823602473
Current evaluation: 75.8385823602473
0.0
Current evaluation: 75.8385823602473
Current evaluation: 75.8385823602473
0.0

```

Here are some of my thoughts:

- Even if I change the input vector `init_value` and use random values in each iteration, the function evaluation returns the same value, which is the reason why the gradient is zero. I don’t think this is intended, nor I believe it is right. Both evaluations, f(x + h e) and f(x) should be different, at least to some extent.
- Indeed, the gradient is zero and that’s the reason why the optimizer was not doing anything at all.

I really don’t know how or why this is happening, but at least I got some useful insights that need further investigation.

---

<div class="post-metadata">

**Author:** ![cvanaret](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cvanaret/32/11594_2.png) [@cvanaret](https://discourse.julialang.org/u/cvanaret)\
**Post date:** [June 25, 2021, 6:56pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/7 "2021-06-25T18:56:14Z")

</div>

What does `h_val .* i` do? Is it really the i th unit basis vector?

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 25, 2021, 8:15pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/8 "2021-06-25T20:15:05Z")

</div>

Yeah, if I loop over each column from `I` I get the boolean version of an identity matrix, which I believe is equivalent to the i th unit basis vector.

```julia
julia> using LinearAlgebra

julia> for i in eachcol(I(5))
       println(i)
       end
Bool[1, 0, 0, 0, 0]
Bool[0, 1, 0, 0, 0]
Bool[0, 0, 1, 0, 0]
Bool[0, 0, 0, 1, 0]
Bool[0, 0, 0, 0, 1]

```

and I could convert that to Float64 with the following

```julia
julia> for i in eachcol(I(5))
       println(Float64.(i))
       end
[1.0, 0.0, 0.0, 0.0, 0.0]
[0.0, 1.0, 0.0, 0.0, 0.0]
[0.0, 0.0, 1.0, 0.0, 0.0]
[0.0, 0.0, 0.0, 1.0, 0.0]
[0.0, 0.0, 0.0, 0.0, 1.0]

```

and when I do `h_val .* i` in the previous post, I get the value of he\_i in the expression f(x + he\_i), where e\_i is the i th unit basis vector.  
Well, hopefully I didn’t make a mistake somewhere, but I think it looks right.  
Anyway, both versions give me the same results.

---

<div class="post-metadata">

**Author:** ![cvanaret](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cvanaret/32/11594_2.png) [@cvanaret](https://discourse.julialang.org/u/cvanaret)\
**Post date:** [June 25, 2021, 8:27pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/9 "2021-06-25T20:27:10Z")

</div>

Cool, looks good.

You get evaluations different from 0.13872807013353833 (in your second message), so the function is not totally constant. But it’s weird that its value changes so little. Can you try with larger values of h?

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 25, 2021, 9:25pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/10 "2021-06-25T21:25:54Z")

</div>

> [@cvanaret](#):
>
> Can you try with larger values of h ?

I plugged in h=0.1 and, unfortunately, the result is almost the same.

```julia
Current evaluation: 75.6876240505038
Current evaluation: 75.6876240505038
0.0
Current evaluation: 75.6876240505038
Current evaluation: 75.6876240505038
0.0
Current evaluation: 75.6876240505038
Current evaluation: 75.6876240505038
0.0
Current evaluation: 75.6876240505038
Current evaluation: 75.6876240505038
0.0
Current evaluation: 75.6876240505038
Current evaluation: 75.6876240505038
0.0
Current evaluation: 75.6876240505038
Current evaluation: 75.6876240505038
0.0
Current evaluation: 75.6876240505038
Current evaluation: 75.6876240505038
0.0

```

> [@cvanaret](#):
>
> You get evaluations different from 0.13872807013353833 (in your second message)

Yeah, you’re totally right, very sorry about that. That’s because I had to change some other parameters within the code. If I stick with the same ones, I get something similar to those values obtained before

```julia
Current evaluation: 0.12281985076905867
Current evaluation: 0.12281985076905867
Current evaluation: 0.12281985076905867

```

I was changing the parameters because I need this procedure to be quite general. Those values only represent that some parameters give better estimations than others, but the problem is still the same, it keeps evaluating the same value, regardless of initial condition or any other thing, which is really troublesome. But that’s part of the numerical procedure itself, I don’t think it has something to do with the optimization part.

---

<div class="post-metadata">

**Author:** ![cvanaret](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cvanaret/32/11594_2.png) [@cvanaret](https://discourse.julialang.org/u/cvanaret)\
**Post date:** [June 25, 2021, 9:32pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/11 "2021-06-25T21:32:50Z")

</div>

Got it.  
Can you explain what the x values represent for your function? For example, are they parameters for a model that you need to fit against data points?  
Can you check in the body of the function how these x values are used (and maybe post a small snippet to illustrate it)?

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 25, 2021, 10:04pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/12 "2021-06-25T22:04:22Z")

</div>

Unfortunately, I can’t post any more code.

I can say, though, that those x values are actually neural network weights, and the idea is to try to find the weights that give the minimum value for the black box function.  
The way I’m trying to accomplish this is by evaluating some weights, and adjusting them using the L-BFGS optimization method. But if the gradient is zero, I’m getting nowhere, as you can already tell.

I guess the real problem lies within the black-box function, because regardless of the vector that I pass to the function I always get the same values in return.

---

<div class="post-metadata">

**Author:** ![cvanaret](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cvanaret/32/11594_2.png) [@cvanaret](https://discourse.julialang.org/u/cvanaret)\
**Post date:** [June 25, 2021, 10:13pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/13 "2021-06-25T22:13:51Z")

</div>

So check that your function actually updates the weights of the NN from the x vector passed as parameter. Something in `sqerror_model` must be wrong.

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 25, 2021, 10:15pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/14 "2021-06-25T22:15:15Z")

</div>

Yeah, there’s something very wrong there, I’ll make sure to check it out thoroughly.  
Anyways, thanks for all the help and time, I really appreciate it!

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [June 25, 2021, 10:18pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/15 "2021-06-25T22:18:39Z")

</div>

Without derivatives probably you should try a derivative-free method, like [bobyqa](https://nlopt.readthedocs.io/en/latest/NLopt_Algorithms/#bobyqa). Numerical derivatives can be very expensive and very unstable.

Or, from Optim.jl, [Nelder Mead - Optim.jl](http://julianlsolvers.github.io/Optim.jl/v0.9.3/algo/nelder_mead/)

---

<div class="post-metadata">

**Author:** ![edwinb-ai](https://avatars.discourse-cdn.com/v4/letter/e/bb73d2/32.png) [@edwinb-ai](https://discourse.julialang.org/u/edwinb-ai)\
**Post date:** [June 25, 2021, 10:19pm UTC](https://discourse.julialang.org/t/understanding-the-results-from-optim-jl-for-a-black-box-function/63507/16 "2021-06-25T22:19:37Z")

</div>

Thanks, I’ll try that out!
