# KernelForge.jl — High-performance portable GPU primitives for arbitrary types and operators

**URL:** <https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780>\
**Category:** Package Announcements\
**Tags:** gpu, cuda, kernelabstractions, high-performance\
**Created:** [February 22, 2026, 8:27pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780 "2026-02-22T20:27:48Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [February 22, 2026, 8:27pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/1 "2026-02-22T20:27:48Z")

</div>

I’m happy to announce two related packages for high-performance GPU computing in Julia:

- [**KernelForge.jl**](https://epilliat.github.io/KernelForge.jl/) — high-level GPU primitives (mapreduce, scan, matvec, search, vectorized copy) with performance competitive with vendor-optimized libraries
- [**KernelIntrinsics.jl**](https://github.com/...) — the low-level foundation, providing warp shuffles, vectorized memory access, and memory ordering primitives on top of KernelAbstractions.jl

## Motivation

Julia already has excellent options for GPU computing:

- **CUDA.jl** is mature and highly optimized, but is CUDA-specific and relies on vendor libraries (cuBLAS, CUB) for performance-critical primitives
- **AcceleratedKernels.jl** takes a cross-architecture approach via KernelAbstractions.jl, trading some performance against vendor libraries for portability

**KernelForge.jl** aims to combine the best of both: **cross-architecture portability** (once the underlying intrinsics are available on other backends) with **performance matching or exceeding vendor libraries** (at least on a (NVIDIA RTX 1000)), through pure Julia implementations using memory ordering and flags, warp-level reductions, and vectorized memory stores/loads.

## What makes this different?

Two complementary goals drive the design:

- **Generality** : support for arbitrary `isbitstype` structs, custom operators, and non-contiguous views — not just `Float32` with `+`
- **Performance** : match cuBLAS and CUB on concrete types like `Float32`

Both goals are achieved simultaneously, at least on a laptop GPU (NVIDIA RTX 1000) — see the [performance section](https://epilliat.github.io/KernelForge.jl/stable/performances/) for benchmarks.

## Available primitives

| Primitive | Description |
| --- | --- |
| `mapreduce` | General reduction over any dimension(s), arbitrary types and operators |
| `scan!` | 1D prefix scan (like `accumulate` in Julia Base), supports non-commutative operators |
| `vcopy!` | Vectorized copy with configurable load/store widths |
| `matvec`, `vecmat` | Matrix-vector and vector-matrix products with custom operators |
| `argmax`, `argmin` | Index of extremum |
| `findfirst`, `findlast` | Search on GPU arrays |

Moreover, for all these primitives, we do not need to any `init` or `neutral` argument as CUDA.jl or AcceleratedKernel.jl !

```julia
using KernelForge as KF
using CUDA

src = CUDA.rand(Float32, 10^6)
dst = similar(src)

# Full reduction with custom function and operator
total = KF.mapreduce(abs2, +, src)

# Reduction over specific dimensions
B = CUDA.rand(Float32, 4, 8, 16)
result = KF.mapreduce(identity, +, B; dims=(1, 3)) # shape: (1, 8, 1)

# Views are fully supported
v = view(src, 1:2:10^6)
KF.scan!(+, dst, v)

# Search
i = KF.findfirst(>(0.99f0), src)
j = KF.argmax(src)

```

## Roadmap

Time permitting, the next planned primitives are **sort** and **matrix-matrix product** and/or some graph algorithms

## Architecture

```julia-auto
KernelForge.jl ← GPU primitives (mapreduce, scan, matvec, search, ...)
       │
KernelIntrinsics.jl ← Low-level intrinsics (warp shuffle, vectorized load/store, memory fences)
       │
KernelAbstractions.jl ← Backend dispatch
       │
   CUDA.jl ← (AMD/Intel planned)

```

`KernelIntrinsics.jl` is similar in scope to `GPUArraysCore.jl` — a building block for library developers rather than end users. Extending support to AMD or Intel GPUs would primarily require work in `KernelIntrinsics.jl`; `KernelForge.jl` itself is designed to need minimal adaptation.

## Current status

Both packages are **experimental** and currently **CUDA-only** , though an extensive test suite is already in place. Note that vectorized loads and stores in `KernelIntrinsics.jl` do not perform bounds checking. Correctness and performance have been validated on an NVIDIA RTX 1000.

Feedback, bug reports, and contributions are very welcome — especially from anyone with access to AMD or Intel hardware!

---

<div class="post-metadata">

**Author:** ![yolhan\_mannes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yolhan_mannes/32/220485_2.png) [@yolhan\_mannes](https://discourse.julialang.org/u/yolhan_mannes)\
**Post date:** [February 23, 2026, 4:31am UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/2 "2026-02-23T04:31:06Z")

</div>

Awesome packages, not sure about the name of the second since there is already an KernelIntrinsics module within KernelAbstraction

> <https://github.com/JuliaGPU/KernelAbstractions.jl/blob/main/src/intrinsics.jl>

---

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [February 23, 2026, 8:17am UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/3 "2026-02-23T08:17:24Z")

</div>

Oh I didn’t know about that! I already changed the name once since it was first “MemoryAccess”, but that felt too general. Maybe something like KernelTools? Or it might be worth integrating the features into an existing package rather than maintaining a separate one. It can be quite heavy though already since it already comes with several features, each requiring an extension per backend.

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [February 23, 2026, 9:20am UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/4 "2026-02-23T09:20:56Z")

</div>

I am also happy to take PRs to KernelAbstractions.jl or Atomix.jl/UnsafeAtomics.jl.

For KernelAbstraction, we currently maintain the 0.9 release branch and the 0.10 development branch. The only “challenge” for KA is that functionality should work on all backends, or at least there should be a query function.

---

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [February 23, 2026, 10:28am UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/5 "2026-02-23T10:28:18Z")

</div>

Thanks! I agree that warp shuffle operations and vectorized loads/stores would be a natural fit for KernelAbstractions.jl (although I wouldn’t quite know how to handle them on CPU — maybe with warp size equal to 1?), and that memory ordering primitives would belong in UnsafeAtomics.jl. I went the separate package route to get a quick proof of concept out before committing to full multi-backend support — and honestly, writing my own package is also a great way to learn quickly. I’d be happy to contribute the relevant pieces upstream once they’re more mature!

---

<div class="post-metadata">

**Author:** ![caleb-allen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/caleb-allen/32/14054_2.png) [@caleb-allen](https://discourse.julialang.org/u/caleb-allen)\
**Post date:** [February 25, 2026, 2:10pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/6 "2026-02-25T14:10:32Z")

</div>

This is a very interesting package, thank you for putting this together!

Do I understand correctly that one could use KenelForge.jl to, for example, run integer operations on a GPU?

---

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [February 25, 2026, 2:15pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/7 "2026-02-25T14:15:31Z")

</div>

Yes exactly. Actually you can already do this on CUDA.jl and AcceleratedKernels.jl.  
But here the difference is performance matching CuBlas and Cub (see the benchmarks), (even for copy operation!).

Also I put a lot of example showing KernelForge works also on more complex structures (check UnitFloat8 type, or Quaternions for scan function).

Types must be bitstypes though but it’s a general GPU requirement.

---

<div class="post-metadata">

**Author:** ![jgreener64](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jgreener64/32/2483_2.png) [@jgreener64](https://discourse.julialang.org/u/jgreener64)\
**Post date:** [March 14, 2026, 7:12pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/8 "2026-03-14T19:12:51Z")

</div>

Nice work. Backend-agnostic warp operations would be useful for us in Molly.jl, where the force kernel relies on warp operations and we want to work on different GPUs.

There is a KernelAbstractions.jl issue open about warp operations: [exposing warp-level semantics · Issue #420 · JuliaGPU/KernelAbstractions.jl · GitHub](https://github.com/JuliaGPU/KernelAbstractions.jl/issues/420).

---

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [March 15, 2026, 8:08pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/9 "2026-03-15T20:08:04Z")

</div>

Thanks! Great to hear this could be useful for Molly.jl.

For context on the current state of KernelForge.jl: I just landed ROCm backend support in KernelIntrinsics.jl and updated KernelForge.jl accordingly. Both packages now work on CUDA and ROCm, with all tests passing on both backends — this includes warp shuffles, vectorized memory accesses, and memory fences/ordering primitives. Performance is very close to vendor libraries (CUB/rocPRIM) on several primitives, tested on RTX 1000 and A40 (CUDA) and MI300X (ROCm).

Regarding #420: I wanted to move quickly without waiting on a KA PR. I also think warp operations are fundamental enough that they belong at this level — every GPU backend supports them, and one can always define a warp as a single thread for degenerate cases like CPUs.

Beyond just exposing backend warp ops though, I wanted something more abstract. At the hardware level, warp shuffles operate on 32-bit registers, and backends handle more complex types by recursing down to primitives — but they stop at complex numbers (quite surprisingly). Arbitrary structs and tuples aren’t supported. KernelIntrinsics.jl extends this recursion to handle general structs and tuples, which I think is important for use cases like Molly.jl where you’re naturally shuffling composite types (e.g. force/position vectors) across threads.

---

<div class="post-metadata">

**Author:** ![jgreener64](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jgreener64/32/2483_2.png) [@jgreener64](https://discourse.julialang.org/u/jgreener64)\
**Post date:** [March 16, 2026, 11:53am UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/10 "2026-03-16T11:53:41Z")

</div>

> [@epilliat](#):
>
> [KernelIntrinsics.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/KernelIntrinsics) extends this recursion to handle general structs and tuples, which I think is important for use cases like [Molly.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/Molly) where you’re naturally shuffling composite types (e.g. force/position vectors) across threads.

Yes that’s important, currently we do some sketchy struct dis-assembly and re-assembly to allow warp shuffles with arbitrary atom types. This might provide a better way to do that.

---

<div class="post-metadata">

**Author:** ![WilliBee](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/willibee/32/218038_2.png) [@WilliBee](https://discourse.julialang.org/u/WilliBee)\
**Post date:** [March 20, 2026, 11:52am UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/11 "2026-03-20T11:52:18Z")

</div>

Hello Emmanuel, thank you very much for this awesome package, it provides exactly the kind of warp level control and vector load I needed to port the RadiK GPU TopK algo ([paper](https://arxiv.org/abs/2501.14336), CUDA ++ [repo](https://github.com/leefige/radik)) to backend-agnostic Julia (see [BitonicSort.jl](https://github.com/WilliBee/BitonicSort.jl) and [RadiK.jl](https://github.com/WilliBee/RadiK.jl)).

I also forked your package to implement some Metal functionality, I could open a PR if you’re interested !

Cheers !

---

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [March 22, 2026, 8:14pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/12 "2026-03-22T20:14:01Z")

</div>

Thanks !

Yes open a PR ! It would be interesting to have metal backend also 🙂 Ideally it would have to pass some tests like the ones already presents (mostly generated by claude)  
If orderings in memory access are not as developped as for CUDA, you can fall back to strong thread fence that should already exist.

By the way I have also planed to implement a sorting algo looking more like quicksort (using only one kernel). I don’t know if my idea is good but I guess it is worth trying !

Emmanuel

---

<div class="post-metadata">

**Author:** ![WilliBee](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/willibee/32/218038_2.png) [@WilliBee](https://discourse.julialang.org/u/WilliBee)\
**Post date:** [April 16, 2026, 4:44pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/13 "2026-04-16T16:44:34Z")

</div>

Hi Emmanuel, hope all is well on your side !

I (finally) opened two PRs :

- [Add Metal backend support by WilliBee · Pull Request #5 · epilliat/KernelIntrinsics.jl · GitHub](https://github.com/epilliat/KernelIntrinsics.jl/pull/5)
- [Add litmus tests for memory ordering operations by WilliBee · Pull Request #6 · epilliat/KernelIntrinsics.jl · GitHub](https://github.com/epilliat/KernelIntrinsics.jl/pull/6) : this one was a bit of a detour, it implements behavioral tests for memory reordering of atomic operations (as you’ll see in the first PR, Acquire-Load and Release-Store needed to be implemented as a two-step fence+atomic operation, I wanted to make sure that worked)

All tests are passing on my Macbook; I tried on Google Colab for CUDA and all pass as well.

Also great that you’re trying the quicksort-like algo ! How is this going ?

Talk to you soon on here or on github ! Cheers !

William

---

<div class="post-metadata">

**Author:** ![epilliat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/epilliat/32/219946_2.png) [@epilliat](https://discourse.julialang.org/u/epilliat)\
**Post date:** [June 16, 2026, 3:31pm UTC](https://discourse.julialang.org/t/kernelforge-jl-high-performance-portable-gpu-primitives-for-arbitrary-types-and-operators/135780/14 "2026-06-16T15:31:07Z")

</div>

Hi everyone,

Happy to announce I’ve just released **v0.2.0** 🎉

The headline of this release is a GPU `sort` (plus `sortperm` and a batched per-column sort): on an **RTX 1000 (Turing)** it reaches **CUB radix-sort performance**. There’s also a counter-based RNG (`Random` submodule) and a good amount of perf work across scan and map-reduce.  
There’s also a batch sort function if you want to sort many vectors efficiently.

Two honest caveats:

- **Sort is only benchmarked on the RTX 1000 so far** — it’s currently **unbenchmarked on AMD** , and I’d love numbers on other NVIDIA cards too. But I don’t have infinite money, so I can only measure on the hardware I actually own (and on AMD dropplets for which i have little credit).

- That constraint is exactly why I’m now putting a lot of effort into **autotuning**. The direction is simple: **anyone can autotune.** The tuner runs on _your_ GPU and emits a JSON of the best parameters for that architecture; drop it into the library’s data folder and it’s picked up automatically — so one tuning run on a given GPU becomes reusable for everyone with that card.

Long-term, the idea is that people run the autotuner on their card and contribute the resulting JSON back say by email or a PR (not sure how yet though), so the library gradually accumulates tuned configs for more and more architectures.

One note for transparency: I’m the sole developer on this, and this new version was largely **coded with Claude code**. It’s now able to write very efficient GPU kernels compared to six months ago… I helped it sometimes but it can find a lot of fine optimization by itself now.

Links:

- **KernelForge.jl** : [GitHub - epilliat/KernelForge.jl · GitHub](https://github.com/epilliat/KernelForge.jl)
- **KernelIntrinsics.jl** : [GitHub - epilliat/KernelIntrinsics.jl: Julia functions for invoking GPU-specific memory access instructions · GitHub](https://github.com/epilliat/KernelIntrinsics.jl)
