# \#rl

**URL:** https://discourse.julialang.org/tag/rl/1314.md

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

---

## [Juliacall with SB3 for RL in WSL bings to segmetation fault problems](https://discourse.julialang.org/t/juliacall-with-sb3-for-rl-in-wsl-bings-to-segmetation-fault-problems/122989)

<div class="topic-metadata">

**Author:** [@Claudio](https://discourse.julialang.org/u/Claudio)\
**Replies:** 0\
**Last updated:** [November 23, 2024, 10:41am UTC](https://discourse.julialang.org/t/juliacall-with-sb3-for-rl-in-wsl-bings-to-segmetation-fault-problems/122989 "2024-11-23T10:41:49Z")

</div>

I am trying to train a RL agent using SB3, torch, and Gymnasium. I use a linux environment through wsl2 (Ubuntu-22.04), and VisualStudio Code. To speed up the step phase in my environment, I invoke Julia passing through …

---

## [Trouble setting up PPO with reinforcementlearning](https://discourse.julialang.org/t/trouble-setting-up-ppo-with-reinforcementlearning/116653)

<div class="topic-metadata">

**Author:** [@clempe](https://discourse.julialang.org/u/clempe)\
**Replies:** 0\
**Last updated:** [July 5, 2024, 10:01am UTC](https://discourse.julialang.org/t/trouble-setting-up-ppo-with-reinforcementlearning/116653 "2024-07-05T10:01:13Z")

</div>

I’m new to julia reinforcementlearning, trying to run some experiments with PPO. Tried a lot of things, somehow can’t get ActorCritic to work. Maybe someone could give me a tip? Julia Version 1.10.4 Commit 48d4fd4843 (2…

---

## [How to solve a Markov decision process with randomness](https://discourse.julialang.org/t/how-to-solve-a-markov-decision-process-with-randomness/96426)

<div class="topic-metadata">

**Author:** [@Jian\_ZUO](https://discourse.julialang.org/u/Jian_ZUO)\
**Replies:** 0\
**Last updated:** [March 22, 2023, 12:24am UTC](https://discourse.julialang.org/t/how-to-solve-a-markov-decision-process-with-randomness/96426 "2023-03-22T00:24:21Z")

</div>

Hi, I have a Markov decision process as follows: states: S1, S2; the increment of S1 follows a Gamma law, i.e. S1(t+d)-S1(t) ~ gamma(alpha, beta), S1 has a fixed range of zero to a failure threshold FT (thus the transi…
