# \[ANN\] PhiloxRNG.jl: Generate random numbers on CPU and GPU using the Philox4x32 counter-based RNG

**URL:** <https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224>\
**Category:** Package Announcements\
**Created:** [March 16, 2026, 4:26pm UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224 "2026-03-16T16:26:33Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [March 16, 2026, 4:26pm UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/1 "2026-03-16T16:26:33Z")

</div>

> **[GitHub - medyan-dev/PhiloxRNG.jl: Generate random numbers on CPU and GPU using the...](https://github.com/medyan-dev/PhiloxRNG.jl)**
>
> Generate random numbers on CPU and GPU using the Philox4x32 counter-based RNG.

PhiloxRNG.jl is a package for generating random numbers on both the CPU and GPU. It will generate similar numbers on all devices (results are not exactly the same when sampling floating point distributions due to fast math differences).

The underlying algorithm is a 10 round [Philox4x32](https://doi.org/10.1145/2063384.2063405) combined with a fast boxmuller transformation for sampling the normal distribution.

Performance is currently significantly better than `randn!` `Float32`, but slower in other cases.

## Benchmarks

Julia 1.12.5, CUDA 5.11.0, AMD Ryzen 7 9800X3D, NVIDIA GeForce RTX 3080.

### CPU (ns/value, N = 100,000,000)

| Function | PhiloxRNG.jl | Random.jl |
| --- | --- | --- |
| `rand` F32 | 0.791 | 0.522 |
| `rand` F64 | 1.997 | 1.052 |
| `randn` F32 | 1.009 | 2.114 |
| `randn` F64 | 3.098 | 1.795 |

### GPU (ns/value, N = 100,000,000)

| Function | PhiloxRNG.jl | CUDA.jl |
| --- | --- | --- |
| `rand` F32 | 0.006 | 0.006 |
| `randn` F32 | 0.007 | 0.032 |

Random123.jl also implements the philox family of RNGs. The main difference is the API. Random123.jl uses a mutable struct `AbstractRNG` interface, while PhiloxRNG.jl uses pure functions. This can make PhiloxRNG.jl easier to use in some situations.

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [March 17, 2026, 3:20am UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/2 "2026-03-17T03:20:56Z")

</div>

By replacing a `sqrt_fast` with a `sqrt_llvm` and manually unrolling a loop, the rng functions now have all clean effects! [Improve effects by nhz2 · Pull Request #1 · medyan-dev/PhiloxRNG.jl · GitHub](https://github.com/medyan-dev/PhiloxRNG.jl/pull/1)

Performance is now reasonably close to Random.jl on the CPU when generating large batches.

### CPU (ns/value, N = 100,000,000)

| Function | PhiloxRNG.jl | Random.jl |
| --- | --- | --- |
| `rand` F32 | 0.679 | 0.528 |
| `rand` F64 | 1.371 | 1.074 |
| `randn` F32 | 0.898 | 2.103 |
| `randn` F64 | 2.009 | 1.801 |

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [March 17, 2026, 3:26am UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/3 "2026-03-17T03:26:54Z")

</div>

Very odd that Random.jl stdlib is slower for `randn(::Float32)` than `randn(::Float64)`. Seems worth investigating.

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [March 18, 2026, 12:59am UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/4 "2026-03-18T00:59:03Z")

</div>

From what I can tell Random.jl uses the Float64 method, and converts to Float32 at the end. [julia/stdlib/Random/src/normal.jl at 942c262fa4002fad29e36e19ae1e052401680f2f · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/blob/942c262fa4002fad29e36e19ae1e052401680f2f/stdlib/Random/src/normal.jl#L224)

This is a more correct way of doing things and probably has more accurate tails. PhiloxRNG.jl uses exactly 32 bits of rng output per normal Float32, which means the absolute value of the distribution gets truncated to between 1E-14 and 7.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [March 18, 2026, 1:23am UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/5 "2026-03-18T01:23:53Z")

</div>

For the vectorized usecase, it does look like we could gain some perf by using box-muller like philoxrng.jl. The lack of branch divergence is quite nice.

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [March 18, 2026, 6:01pm UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/6 "2026-03-18T18:01:53Z")

</div>

Something else to try is using box-muller with [UInt64 rng inputs and Float32 math](https://github.com/medyan-dev/PhiloxRNG.jl/pull/2) to increase the maximum output to 9.5, this might be overkill for Float32.

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [April 21, 2026, 5:53am UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/7 "2026-04-21T05:53:02Z")

</div>

Version 1.1.1 greatly improves Float64 performance on GPUs.

It turns out CUDA.jl’s compiler has problems constant propagating `::Float64 ^ ::Int64` and is happier with `::Float64 ^ ::Int32`.

### GPU — NVIDIA GeForce RTX 3080 (ns/value, N = 100,000,000)

| Function | PhiloxRNG | CUDA.jl |
| --- | --- | --- |
| `rand` F32 | 0.0056 | 0.0056 |
| `randn` F32 | 0.0058 | 0.0318 |
| `rand` F64 | 0.0111 | 0.0195 |
| `randn` F64 | 0.1332 | 0.2743 |

### GPU — NVIDIA A100-SXM4-40GB (ns/value, N = 100,000,000)

| Function | PhiloxRNG | CUDA.jl |
| --- | --- | --- |
| `rand` F32 | 0.0028 | 0.0064 |
| `randn` F32 | 0.0049 | 0.0359 |
| `rand` F64 | 0.0053 | 0.0163 |
| `randn` F64 | 0.0096 | 0.0762 |

---

<div class="post-metadata">

**Author:** ![nhz2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nhz2/32/44428_2.png) [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Post date:** [October 4, 2026, 9:22pm UTC](https://discourse.julialang.org/t/ann-philoxrng-jl-generate-random-numbers-on-cpu-and-gpu-using-the-philox4x32-counter-based-rng/136224/8 "2026-10-04T21:22:34Z")

</div>

GPUArrays.jl now uses the functions from PhiloxRNG.

Passing `CUDA.default_rng()` as the first argument to `rand!` and `randn!` should now match PhiloxRNG speeds, and it even supports part of the `AbstractRNG` interface so it works with custom samplers (assuming the sampling code is GPU compatible).

Ref: [Replace Tausworthe+LCG RNG with stateless Philox4x32-10 - Pull Request #707 - JuliaGPU/GPUArrays.jl - GitHub](https://github.com/JuliaGPU/GPUArrays.jl/pull/707) [Update PhiloxRNG.jl to v1.1.1 and fix counter overflow with UInt64 - Pull Request #717 - JuliaGPU/GPUArrays.jl - GitHub](https://github.com/JuliaGPU/GPUArrays.jl/pull/717) [Avoid extra philox calls with ElementRNG - Pull Request #723 - JuliaGPU/GPUArrays.jl - GitHub](https://github.com/JuliaGPU/GPUArrays.jl/pull/723)
