# \[ANN\] Announcing AlphaZero.jl

**URL:** <https://discourse.julialang.org/t/ann-announcing-alphazero-jl/36877>\
**Category:** Package Announcements\
**Tags:** machine-learning\
**Created:** [April 1, 2020, 9:32pm UTC](https://discourse.julialang.org/t/ann-announcing-alphazero-jl/36877 "2020-04-01T21:32:36Z")\
**Posts on this page:** 1\
**Showing post:** 29

<div class="post-metadata">

**Author:** ![jonathan-laurent](https://avatars.discourse-cdn.com/v4/letter/j/ecae2f/32.png) [@jonathan-laurent](https://discourse.julialang.org/u/jonathan-laurent)\
**Post date:** [April 5, 2020, 8:12pm UTC](https://discourse.julialang.org/t/ann-announcing-alphazero-jl/36877/29 "2020-04-05T20:12:02Z")

</div>

> [@Oscar\_Smith](#):
>
> Lc0 has a really active discord, and some of these questions, I’m not the perfect person to answer, so I encourage you to ask around a bit here [https://discord.gg/bZvDNk](https://discord.gg/bZvDNk).

I just had an interesting conversation on the Lc0 discord, which answered my question about _distinguishing two different uses of the move selection temperature_. For people following this thread, I am summarizing my conclusions here.

- After running N MCTS iterations to plan a move, let’s write N\_a the number of times action a is explored. We have N = \sum\_a N\_a.
- The resulting game policy is to play action a with probability \pi\_a := (N\_a/N)^{1/\tau} with \tau the move selection parameter. However, the policy target that should be used to update the neural network is (N\_a/N)\_a and **not** \pi.
- I think the AlphaGo Zero paper is misleading here as it uses notation \pi to denote both the policy to follow during self-play and the target update, suggesting these should be the same.

Also:

- In Lc0, two temperature parameters are introduced. The first one is the _move selection temperature_, which corresponds to the \tau parameter described above. The second one (which does not appear in the AlphaGo Zero paper) is called the _policy temperature_ and it is applied to the softmax output of the neural network to form the prior probabilities used by MCTS.
- Typically, the policy temperature should be greather than 1 and the move selection temperature should be less than 1.

I am going to update AlphaZero.jl accordingly. I expect it should result in a significant improvement of the connect four agent.

---

_[View the full topic](https://discourse.julialang.org/t/ann-announcing-alphazero-jl/36877)._
