# Amortized Hierarchical Variational Model

**URL:** <https://discourse.julialang.org/t/amortized-hierarchical-variational-model/46665>\
**Category:** Probabilistic Programming\
**Tags:** flux, turing, machine-learning\
**Created:** [September 15, 2020, 9:22pm UTC](https://discourse.julialang.org/t/amortized-hierarchical-variational-model/46665 "2020-09-15T21:22:17Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![darleybarreto](https://avatars.discourse-cdn.com/v4/letter/d/a9a28c/32.png) [@darleybarreto](https://discourse.julialang.org/u/darleybarreto)\
**Post date:** [September 15, 2020, 9:22pm UTC](https://discourse.julialang.org/t/amortized-hierarchical-variational-model/46665/1 "2020-09-15T21:22:17Z")

</div>

Suppose I have a complex neural network giving four parameters `mu, precision, alpha, beta`, which may be fixed for all training, learned during training or varying for each batch, such as:

```julia
X,Y = next_batch(data)
(mu,precision,alpha,beta) = ComplexNN(X)

```

Then I have `Parent := NormalInverseGamma(mu,precision,alpha,beta)`, `Q := Normal(mean, variance)`, and `P := Normal(mean_prior, variance_prior)`. Where `(mean, variance) ~ Parent` and `target ~ Q`. The `target` samples will be used by some other task which has it’s own loss. How could I perform the training with `L = LogLikelihood(Q,Y) - DKL(Q,P) + other_task_loss`?

---

<div class="post-metadata">

**Author:** ![mohamed82008](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mohamed82008/32/18171_2.png) [@mohamed82008](https://discourse.julialang.org/u/mohamed82008)\
**Post date:** [September 16, 2020, 1:33am UTC](https://discourse.julialang.org/t/amortized-hierarchical-variational-model/46665/2 "2020-09-16T01:33:43Z")

</div>

You can increment the log joint probability manually in a Turing model using `acclogp!(_varinfo, lp)` where `lp` is your new “loss”. Technically, it would be negative the “loss” added because in an optimization context we would maximize the log probability so its negative is the “loss” minimized.

---

<div class="post-metadata">

**Author:** ![darleybarreto](https://avatars.discourse-cdn.com/v4/letter/d/a9a28c/32.png) [@darleybarreto](https://discourse.julialang.org/u/darleybarreto)\
**Post date:** [September 16, 2020, 12:25pm UTC](https://discourse.julialang.org/t/amortized-hierarchical-variational-model/46665/3 "2020-09-16T12:25:04Z")

</div>

Thank you for your reply. I have a few questions about it. So `lp = -(LogLikelihood(Q,Y) - DKL(Q,P) + other_task_loss)`, right? Also, what would `_varinfo` be? One more thing, will ADVI update the parameters of my (Flux) NN out-of-the-box, or do I need to register them somehow?

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![mohamed82008](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mohamed82008/32/18171_2.png) [@mohamed82008](https://discourse.julialang.org/u/mohamed82008)\
**Post date:** [September 16, 2020, 8:08pm UTC](https://discourse.julialang.org/t/amortized-hierarchical-variational-model/46665/4 "2020-09-16T20:08:24Z")

</div>

> [@darleybarreto](#):
>
> `lp = -(LogLikelihood(Q,Y) - DKL(Q,P) + other_task_loss)`

I think you may not need the negative sign here. I was mostly talking about `other_task_loss` term which I don’t know what it refers to in this context. But the term “loss” means that we are interested in its low values while we are interested in the high values of log likelihood for example. Could be just a terminology confusion thing and you didn’t mean to imply “loss”.

Turing treats all parameters equally, NN parameters or not, that’s irrelevant. There is a tutorial in the docs on Bayesian NN. I don’t know if we have an ADVI one. @torfjelde might know.

`_varinfo` is a reserved term only available inside the `@model` body that lets you access the internal data structure we use to track random variables and log probabilities.

---

<div class="post-metadata">

**Author:** ![darleybarreto](https://avatars.discourse-cdn.com/v4/letter/d/a9a28c/32.png) [@darleybarreto](https://discourse.julialang.org/u/darleybarreto)\
**Post date:** [September 17, 2020, 2:49pm UTC](https://discourse.julialang.org/t/amortized-hierarchical-variational-model/46665/5 "2020-09-17T14:49:16Z")

</div>

Thank you for you answers!

> I was mostly talking about `other_task_loss` term which I don’t know what it refers to in this context

It depends on the final task, but it can be Cross Entropy.

I have a question regarding the ADVI part, do you have any idea on how efficient this is in practice (wrt training time) compared to amortized VI?
