# GPU and Thread-Parallel Support for Turing.jl

**URL:** <https://discourse.julialang.org/t/gpu-and-thread-parallel-support-for-turing-jl/58958>\
**Category:** Probabilistic Programming\
**Created:** [April 10, 2021, 12:30am UTC](https://discourse.julialang.org/t/gpu-and-thread-parallel-support-for-turing-jl/58958 "2021-04-10T00:30:48Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Koshroy](https://avatars.discourse-cdn.com/v4/letter/k/91b2a8/32.png) [@Koshroy](https://discourse.julialang.org/u/Koshroy)\
**Post date:** [April 10, 2021, 12:30am UTC](https://discourse.julialang.org/t/gpu-and-thread-parallel-support-for-turing-jl/58958/1 "2021-04-10T00:30:48Z")

</div>

Hey all,

I’m a bit new to the Julia ecosystem and had a couple questions that some searching around didn’t answer:

1. Is there a way to use the GPU to sample for Turing.jl models? And if there isn’t, is there a way to try to use a GPU to sample from probability distributions in general?
2. Is there a source of documentation on ways to distribute out sampling to threads? Right now I’m running the multi-threaded chain sampler which is sampling independently on each thread, but I was wondering if there was more fine-grained parallelism I could get.

Thanks

---

<div class="post-metadata">

**Author:** ![EvoArt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/evoart/32/25357_2.png) [@EvoArt](https://discourse.julialang.org/u/EvoArt)\
**Post date:** [April 13, 2021, 10:03pm UTC](https://discourse.julialang.org/t/gpu-and-thread-parallel-support-for-turing-jl/58958/2 "2021-04-13T22:03:20Z")

</div>

Have a look at this reply

> [@Parallelism within Turing.jl model](https://discourse.julialang.org/t/parallelism-within-turing-jl-model/54064):
>
> Hi there, just wondering how safe it is to use Threads.@threads for loops within turing models e.g. @model function my\_func(Y) alpha ~ Normal(0,1) sigma ~ Normal(0,1) Threads.@threads for i in 1:size(Y)[2] Y[:,j] .~ Normal(alpha,sigma) end end This seems to work on my laptop. But I don’t currently have more than 1 thread available to me, so I can’t test it out properly. Is there any reason to avoid this in Turing? Would also be nice to get a general view of how nice…
