Google's MuZero

StatisticalMouse · December 25, 2020, 10:42am

This looks like a big deal. It uses model-based reinforcement learning and can learn to play both chess and Pac-Man. Abstract from the Science article:

" Constructing agents with planning capabilities has long been one of the main challenges in the pursuit of artificial intelligence. Tree-based planning methods have enjoyed huge success in challenging domains, such as chess1 and Go2, where a perfect simulator is available. However, in real-world problems, the dynamics governing the environment are often complex and unknown. Here we present the MuZero algorithm, which, by combining a tree-based search with a learned model, achieves superhuman performance in a range of challenging and visually complex domains, without any knowledge of their underlying dynamics. The MuZero algorithm learns an iterable model that produces predictions relevant to planning: the action-selection policy, the value function and the reward. When evaluated on 57 different Atari games3—the canonical video game environment for testing artificial intelligence techniques, in which model-based planning approaches have historically struggled4—the MuZero algorithm achieved state-of-the-art performance. When evaluated on Go, chess and shogi—canonical environments for high-performance planning—the MuZero algorithm matched, without any knowledge of the game dynamics, the superhuman performance of the AlphaZero algorithm5 that was supplied with the rules of the game."

The Science article, as referred to by above, is here.

http://dx.doi.org/10.1038/s41586-020-03051-4

Topic		Replies	Views
Sokoban and Julia Machine Learning	2	1181	April 9, 2020
New Alpha zero replica Machine Learning announcement	6	1311	June 19, 2020
[ANN] Announcing AlphaZero.jl Package Announcements machine-learning	36	6175	July 3, 2021
Reinforcement learning with A* and a deep heuristic Machine Learning	0	537	November 13, 2018
It is too easy to beat AlphaGo.jl Offtopic	13	2759	June 30, 2019

Google's MuZero

Related topics