# New Theoretical Analysis for ADVI / BBVI

**URL:** <https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527>\
**Category:** Probabilistic Programming\
**Created:** [May 28, 2023, 7:31pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527 "2023-05-28T19:31:45Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Red-Portal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/red-portal/32/9102_2.png) [@Red-Portal](https://discourse.julialang.org/u/Red-Portal)\
**Post date:** [May 28, 2023, 7:31pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/1 "2023-05-28T19:31:45Z")

</div>

Hi all,

We’ve recently uploaded to ArXiv a theoretical analysis for ADVI (although the title says black-box VI), which I believe is the first formal full convergence proof for any black-box VI-type algorithm that uses SGD:

[[2305.15349] Black-Box Variational Inference Converges (arxiv.org) ](https://arxiv.org/abs/2305.15349).

What’s interesting is that the analysis reveals some very unexpected properties of the covariance parameterization that we use in practice. In particular, if we use non-linear transformations for the diagonal elements (as done for the mean-field parameterization in Stan and Turing), such as

L\_{ii} = \exp(\ell\_i),

where \mathbf{L} = \mathrm{diag}\left(L\_{11}, \ldots, L\_{dd}\right) is the (diagonal) Cholesky factor for the variational approximation, we provably lose speed. One could have achieved an \mathcal{O}(1/T) converge rate for nice posteriors but only gets \mathcal{O}(1/\sqrt{T}) instead. And if one does the same for _full-rank parameterizations_, as PyMC3, but not Stan, the ELBO might not even be convex even if the posterior is log-concave!

~~In the context of Julia’s probabilistic programming ecosystem, this means using `positive_definite` from `ParameterHandling.jl` to construct variational approximations is suboptimal.~~

In our experiments, we indeed observe that the mean-field parameterization without any `exp` or `softplus` transformation, for enforcing the scale to be positive, converges the fastest. To me, this is one of those rare occasions where optimization theory precisely tells you what happens in practice, which is not that common, unfortunately.

Please let me know if you have any comments or questions.

---

<div class="post-metadata">

**Author:** ![simsurace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simsurace/32/30216_2.png) [@simsurace](https://discourse.julialang.org/u/simsurace)\
**Post date:** [May 29, 2023, 8:54pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/2 "2023-05-29T20:54:11Z")

</div>

Interesting work, thanks for posting that here.

From a [quick glance](https://github.com/JuliaGaussianProcesses/ParameterHandling.jl/blob/e16589bc7e37aac793f351022e2207a98b44f682/src/parameters_matrix.jl#L41-L91) it seems that the diagonal components of the Cholesky factor are linear in the underlying parameter in ParameterHandling.jl, which if I understand the paper correctly is what seems to achieve the fastests convergence. But maybe I’m missing the point. Could you point out what is suboptimal in the code or maybe open an issue?

---

<div class="post-metadata">

**Author:** ![Red-Portal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/red-portal/32/9102_2.png) [@Red-Portal](https://discourse.julialang.org/u/Red-Portal)\
**Post date:** [May 29, 2023, 9:33pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/3 "2023-05-29T21:33:27Z")

</div>

Oh that’s true. I have no idea why I thought that it enforced positive-definiteness. I think I may have mixed it up with `positive`. Thanks for pointing this out!

But, wouldn’t the fact that it does _not_ enforce positive-definiteness be misleading? Maybe a name like `lower_triangular` would be more appropriate.

---

<div class="post-metadata">

**Author:** ![simsurace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simsurace/32/30216_2.png) [@simsurace](https://discourse.julialang.org/u/simsurace)\
**Post date:** [May 30, 2023, 2:29pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/4 "2023-05-30T14:29:00Z")

</div>

I think this is something that could be discussed. Maybe adding an optional second argument `positive_definite(X, strict=false)` is a non-breaking way of going about this.  
But I’m thinking that given your results would there be any benefit in using a non-linearity for the diagonal elements in practice? In my applications (GPs) we always add some jitter anyway.

---

<div class="post-metadata">

**Author:** ![mohamed82008](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mohamed82008/32/18171_2.png) [@mohamed82008](https://discourse.julialang.org/u/mohamed82008)\
**Post date:** [May 30, 2023, 8:53pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/5 "2023-05-30T20:53:50Z")

</div>

Instead of proximal SGD, could be worth trying stochastic Frank Wolfe here [GitHub - ZIB-IOL/FrankWolfe.jl: Julia implementation for various Frank-Wolfe and Conditional Gradient variants](https://github.com/ZIB-IOL/FrankWolfe.jl). CC: @mbesancon

---

<div class="post-metadata">

**Author:** ![Red-Portal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/red-portal/32/9102_2.png) [@Red-Portal](https://discourse.julialang.org/u/Red-Portal)\
**Post date:** [May 30, 2023, 10:39pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/6 "2023-05-30T22:39:21Z")

</div>

Indeed, our paper shows that it shouldn’t be necessary for VI. But I can imagine people trying to use `positive_definite` for other applications where PD-ness must be enforced and get confused.

---

<div class="post-metadata">

**Author:** ![Red-Portal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/red-portal/32/9102_2.png) [@Red-Portal](https://discourse.julialang.org/u/Red-Portal)\
**Post date:** [May 30, 2023, 10:40pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/7 "2023-05-30T22:40:38Z")

</div>

Thanks for the suggestions! Didn’t know Frank-Wolfe could be used to deal with non-smoothness. But would it have theoretical/practical benefits for this application? Excuse my lack of knowledge of Frank-Wolfe.

---

<div class="post-metadata">

**Author:** ![mohamed82008](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mohamed82008/32/18171_2.png) [@mohamed82008](https://discourse.julialang.org/u/mohamed82008)\
**Post date:** [May 30, 2023, 10:55pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/8 "2023-05-30T22:55:42Z")

</div>

I don’t think the non-smoothness would be needed when using FW. If I understood your paper correctly (only briefly skimmed through), you are modelling the constraints using a non-smooth objective instead of using domain transformation. FW gives you a third choice where you can ensure the optimiser always remains in the feasible domain without bringing the constraints up in the objective as a non-smooth term. Modelling the problem as a conic-constrained (stochastic) nonlinear optimization problem will allow (stochastic) FW to respect these conic constraints, e.g. non-negativity or positive semi-definite constraints, without introducing non-smooth objectives.

---

<div class="post-metadata">

**Author:** ![Red-Portal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/red-portal/32/9102_2.png) [@Red-Portal](https://discourse.julialang.org/u/Red-Portal)\
**Post date:** [May 30, 2023, 11:19pm UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/9 "2023-05-30T23:19:42Z")

</div>

Oh, I see where you’re coming from. Yes, your understanding of our paper is spot on. The problem with constraining the domain is that choosing the constraints is not straightforward. One needs to set a lower bound on the variance of the variational approximation, which means we need to have a good guess of the posterior variance.

---

<div class="post-metadata">

**Author:** ![mohamed82008](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mohamed82008/32/18171_2.png) [@mohamed82008](https://discourse.julialang.org/u/mohamed82008)\
**Post date:** [May 31, 2023, 12:30am UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/10 "2023-05-31T00:30:02Z")

</div>

> [@Red-Portal](#):
>
> One needs to set a lower bound on the variance of the variational approximation

0? Or in the dense case, positive definite.

---

<div class="post-metadata">

**Author:** ![Red-Portal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/red-portal/32/9102_2.png) [@Red-Portal](https://discourse.julialang.org/u/Red-Portal)\
**Post date:** [May 31, 2023, 1:37am UTC](https://discourse.julialang.org/t/new-theoretical-analysis-for-advi-bbvi/99527/11 "2023-05-31T01:37:38Z")

</div>

The smoothness of the entropy term scales as 1/\sigma. So to get a smooth loss function you actually need a strictly positive lower bound. The problem is how small must that positive constant be.
