# Sokoban and Julia

**URL:** <https://discourse.julialang.org/t/sokoban-and-julia/37227>\
**Category:** Machine Learning\
**Created:** [April 8, 2020, 9:52am UTC](https://discourse.julialang.org/t/sokoban-and-julia/37227 "2020-04-08T09:52:16Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [April 8, 2020, 9:52am UTC](https://discourse.julialang.org/t/sokoban-and-julia/37227/1 "2020-04-08T09:52:16Z")

</div>

Hi,

just for a curiosity, did played with reinforcement learning to solve sokoban? Is there any repo, which would allow to play with it?

Thanks for response,  
Tomas

---

<div class="post-metadata">

**Author:** ![jonathan-laurent](https://avatars.discourse-cdn.com/v4/letter/j/ecae2f/32.png) [@jonathan-laurent](https://discourse.julialang.org/u/jonathan-laurent)\
**Post date:** [April 8, 2020, 2:26pm UTC](https://discourse.julialang.org/t/sokoban-and-julia/37227/2 "2020-04-08T14:26:22Z")

</div>

I am not aware of any existing Julia repo but here are some thoughts.

I think that learning a good sokoban player is a pretty hard problem given the current state of the art in RL. I would not expect an approach based purely on learning (DQN, policy gradient) to be successful unless you throw some Deepmind’s level of computing power at it.

Instead, I would really bet on a combination of learning and search, as done in AlphaZero. (Note that I may be a bit biased here as I just [released](https://discourse.julialang.org/t/announcing-alphazero-jl/36877) the first version of [AlphaZero.jl](https://github.com/jonathan-laurent/AlphaZero.jl)).

With a few modifications, I think AlphaZero could work pretty well. However:

- Sokoban is a one player game and you cannot rely on self-play as easily as in Connect Four, Chess or Go. However, I would expect [curriculum learning](https://lilianweng.github.io/lil-log/2020/01/29/curriculum-for-reinforcement-learning.html) to work pretty well in the case of Sokoban (in particular, see the section on assymetric self-play).
- To make search more efficient, you might want to expose more high-level actions than just {Left, Right, Up, Down}. For example, it might be interesting to have a macro action such as “go there”, which is implemented using a hard-coded path finding algorithm.

Another big challenge in Sokoban is that some moves are irreversible and the game is never-ending from an impossible position. Therefore, you may want the agent to “learn when to give up” (or just put a limit on the length of episodes).

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [April 9, 2020, 8:47am UTC](https://discourse.julialang.org/t/sokoban-and-julia/37227/3 "2020-04-09T08:47:08Z")

</div>

Hi Jonathan,

I am aware of the complexity of Sokoban and thats why I have choosen that. My students mainly use python and I wanted to know, if there is something in python. We are indeed interested in combination of RL a search method. Sad, that we cannot easily test AlphaZero.jl.

Tomas
