# \[ANN\] ReactantNitro.jl: Reactant-first training framework

**URL:** <https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500>\
**Category:** Package Announcements\
**Tags:** machine-learning, lux, reactant\
**Created:** [September 16, 2026, 3:19pm UTC](https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500 "2026-09-16T15:19:25Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [September 16, 2026, 3:19pm UTC](https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500/1 "2026-09-16T15:19:25Z")

</div>

I’m excited to finally release ReactantNitro.jl; a Reactant-first training framework inspired by PyTorch Lightning, but specifically designed around making working with Reactant.jl easy while still allowing for a large amount of flexibility without sacrificing performance. While Reactant and Enzyme have some truly incredible potential to speed up machine learning workflows, there are a number of stumbling blocks researchers and ML practitioners often hit coming from other ML ecosystems.

We developed ReactantNitro.jl at [Medical Metrics Inc.](https://medicalmetrics.com) to explicitly avoid these sorts of problems while keeping the massive 2-10x speedup we were seeing for our models compared to PyTorch eager mode and `torch.compile`. It’s been in the works for around half a year now and has been used internally to successfully train many different models.

> **[GitHub - EnzymeAD/ReactantNitro.jl: Reactant-first training for Lux models: declare...](https://github.com/EnzymeAD/ReactantNitro.jl)**
>
> Reactant-first training for Lux models: declare the experiment, compile once, train without boilerplate.

## Experiment DSL designed specifically for Reactant

One of the core features of the framework is an Experiment DSL which lets you categorize your models parameters in a way that the framework can ensure correct and optimal behavior:

| Marker | Reaches traced code as | In the compile key? | Changing the value |
| --- | --- | --- | --- |
| `GraphConst{T}` | a baked literal | yes | recompiles, correctly: a different value is a different program |
| `Device{T}` | a device-resident traced input | no, by construction | never recompiles: sweep it, schedule it, rewrite it live |
| unmarked, i.e. `Host{T}` | not at all | no | never recompiles: driver-only, invisible to the tracer |

The frameworks callback hooks are keyed on your `@experiment` struct:

```julia
@experiment struct MyExp
    "Structural: changes the emitted graph, so it bakes and is part of the compile key."
    width::GraphConst{Int} = 128

    "A traced input: sweep it or schedule it without recompiling."
    smoothing::Device{Float32} = 0.05f0

    "Unmarked, therefore Host: driver-only and invisible to the tracer."
    max_epochs::Int = 20
end

```

Structuring things this way eliminates entire categories of pitfalls:

- The tracer will never trace over your dataset (`Host` by default); for large datasets, this can massively increase Reactant’s compile time
- One of the first principles of the framework is reducing wasted compile time. Changes to `GraphConst` values are the only thing beyond code changes that need a recompile, which allows the framework to know when we should invalidate a cached program. `Host` values are guaranteed to be invisible to the compiled program, so they cannot change it. Changes to `Device` values do not need a recompile unless the shape or type of it changes.
- The framework knows exactly when `Device` values should be transferred to the device. For scheduled values it happens every step, for constant buffers it’s done a single time at the start of training. More on flexible scheduling later.

## Eliminate Boilerplate

We take inspiration from PyTorch Lightning when it comes to eliminating boilerplate by providing hooks that are dispatched on your experiment type. You handle your forwards pass, loss function, and metrics; the framework handles the boilerplate glue that connects it all while aiming to still provide the user with sufficient flexibility. A manual optimization mode is also provided which can be used to train models like GANs.

```julia
ReactantNitro.forward(::MnistMLP, model, ps, st; img) = Lux.apply(model, img, ps, st)

function ReactantNitro.loss(e::MnistMLP, logits; label)
    smoothed = (1f0 - e.smoothing) .* label .+ e.smoothing / 10f0
    return -sum(smoothed .* logsoftmax(logits; dims = 1)) / size(label, 2)
end

```

These hooks are used to compile three different programs: training forwards/backwards, training optimizer step, and validation/test/infer. The reason we don’t roll the optimizer step into the forwards/backwards is to provide flexibility to do things like gradient accumulation and allow you to switch optimisers without recompiling the entire program.

## REPL Driven Workflow + Kaimon MCP Extension

Every part of the interface was designed to be used in a REPL making use of Revise.jl to allow you to quickly iterate and test new ideas without unnecessary recompile. The invalidation system is Revise aware and tracks changes that would invalidate compiled programs, so compile only happens when its actually needed.

ReactantNitro.jl also has a Kaimon.jl extension which lets your agents drive training jobs entirely through MCP and a full set of skills for your agents to use. Thanks to @kahliburke for all the amazing work he has done on making this possible in the Julia ecosystem! We have experimented some with integrating the framework into KaimonSlate.jl as well, some follow up work is likely needed on my end to make it entirely seamless.

## Flexible Metrics

Metrics can either be handled on the device side for maximum performance or host side for increased flexibility and avoiding re-compile after changing the metrics. This can be defined separately between for train and validation/test. Reductions are structured in a normalization aware way for validation/test, you return a tuple each step with `(scalar, n)` and the n is used in `ReactantNitro.finalize_metrics`. You can also return `(arr, nothing)` at each step and handle the reduction yourself for metrics like AUROC or R^2 that need to be done over all of the individual samples.

## Parameter Schedules

You can schedule pretty much any kind of `Device` (including optimizer parameters like learning rate, weight decay, etc)

```julia
@experiment struct Seq2SeqExp
    teacher_forcing::Device{Float32} = 1f0
    max_epochs::Int = 20
end

ReactantNitro.schedules(::Seq2SeqExp) = (;
    device = (; teacher_forcing = total -> t -> max(0f0, 1f0 - 2f0 * t / total)),
    opt = (; eta = total -> OneCycle(total, 1f-3)), # ParameterSchedulers.jl, no dependency
)

```

This even works with more complicated configurations such as multiple optimisers keyed by the individual optimizer name.

## Export Models to ReactantServer.jl

A ReactantServerExport.jl extension is provided which makes it straightforward to export your nitro models to ReactantServer.jl for serving in production. In the future we are also looking at different export paths such as tensorflow/jax which would allow for deployment in existing serving stacks.

## Status

I plan on registering the package shortly. The API should remain stable as I’m fairly confident it can handle most common use cases, but eventually there will be a 1.0.0 release where we can make improvements to the interface based on feedback.

Work is ongoing to provide end-to-end examples, but the framework is simple enough that getting something off of the ground shouldn’t be terribly difficult.

### AI Disclaimer

- **This post was written entirely by me with no AI assistance.**
- Coding agents were used heavily during the development of ReactantNitro.jl, but the design is entirely mine down to the last detail. The framework has been used to produce many models that had their predictions independently verified during internal testing. Of course as with all software there may be bugs; if you encounter any issues we would love an issue report or a PR addressing them.
- Documentation is largely AI generated but reviewed / refined by me; the process of improving it is ongoing.

---

<div class="post-metadata">

**Author:** ![Dale\_James\_Black](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dale_james_black/32/11595_2.png) [@Dale\_James\_Black](https://discourse.julialang.org/u/Dale_James_Black)\
**Post date:** [September 16, 2026, 4:19pm UTC](https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500/2 "2026-09-16T16:19:21Z")

</div>

Very cool!

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [September 17, 2026, 6:55pm UTC](https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500/3 "2026-09-17T18:55:21Z")

</div>

Before registering, I decided to refine the experience of using ReactantNitro in a REPL with PrettyTables.jl integration (thanks @Ronis_BR!) and a compile aware ProgressMeter.jl so you don’t have to ask whether something hung or if it’s just compiling.

 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/1/71674d8ed62dc095451b01e9380a053a6d4f541c.png)

I also added an extension for TensorBoardLogger.jl and I’m planning to add one for both DearDiary.jl and a CometML logger which I have yet to register. PRs adding other loggers would be much appreciated if you use them!

Finally I wrote an MLUtils.jl extension to ensure that your DataLoader behaves exactly how you expect and warns / raises errors when using combinations of options that do no map into ReactantNitro cleanly. In general DataLoader is the “golden path” where where things should just work, and as such its the suggested way to setup your `build_data` function. You can technically return any iterable, but having a more complete contract lets the framework do a better job at ensuring optimal behavior.

The only thing outstanding is doing a bit of (human) work with the README.md and then we should be good to register!

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [September 21, 2026, 4:53pm UTC](https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500/4 "2026-09-21T16:53:29Z")

</div>

v0.1.0 is out on general 🥳

ReactantNitro.jl now integrates with notebook environments via HTML tables, progress bars, and some basic plotting via Makie without requiring a logger. Shown below is KaimonSlate.jl:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/3/732ff29080d40ba18bacf9794810d22f382ad15f.png)

 ![image](https://global.discourse-cdn.com/julialang/original/3X/5/b/5ba73bcd008278dbb5fbb5991cfbcb1d373e832a.png)

 ![image](https://global.discourse-cdn.com/julialang/original/3X/0/c/0ce5d5f91361490ea0072921cf22e61556ec85ae.png)

We internally keep track of each validation, but not each training step. The principle is provide things to the user which provide the most value while taking the least resources and let experiment logger backends handle the rest. In the future we could handle this through an expanded experiment logger interface to pull the needed data.

One caveat with reactivity; neither Pluto.jl or KaimonSlate.jl understands that when you update a method that dispatches on your experiment type, that makes the cells that instantiate a `Nitro` with an instance of the experiment type stale. You need to make the dependency structure explicit:

```julia-auto
# Cell A
compile_hook = begin
  ReactantNitro.forward(::MnistMLP, model, ps, st; img) = Lux.apply(model, img, ps, st)
  
  function ReactantNitro.loss(e::MnistMLP, logits; label)
      smoothed = (1f0 - e.smoothing) .* label .+ e.smoothing / 10f0
      return -sum(smoothed .* logsoftmax(logits; dims = 1)) / size(label, 2)
  end
end

# Cell B
data_hook = begin
  ...
end

# Cell C
host_hook = begin
  ...
end

# Cell D
begin
  # Explicit dependency on A, B, C
  compile_hook, host_hook, data_hook
  n = Nitro(...)
end

```

A standard Jupyter notebook style model works fine, and of course REPL + Revise.jl is completely seamless.

---

<div class="post-metadata">

**Author:** ![madppiper](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/madppiper/32/220029_2.png) [@madppiper](https://discourse.julialang.org/u/madppiper)\
**Post date:** [September 22, 2026, 7:47am UTC](https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500/5 "2026-09-22T07:47:15Z")

</div>

This is great! I want to play around with it in a private project soon. Really cool!

This is the kind of project I think the Julia space needs more of!

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [October 1, 2026, 6:31pm UTC](https://discourse.julialang.org/t/ann-reactantnitro-jl-reactant-first-training-framework/139500/6 "2026-10-01T18:31:08Z")

</div>

I accidentally started a run today for 50 epochs (default) when I should have done far fewer. Stopped the run and started the new one, no re-compile required 😎 No need to think, when you skip the compile it’s because it wasn’t needed and when you don’t its because something was invalidated.

[v0.1.3](https://github.com/EnzymeAD/ReactantNitro.jl/releases/tag/v0.1.3) has been released with a ton of improvements to working with checkpoints and resuming / loading weights / device buffers. When working in the same experiment directory over many experiments, you can do things like :

- Load the best weights from the most recent experiment: `Nitro(e; weights=:latest => :best)`
- Load the best weights from all previous experiments: `Nitro(e; weights=:all => :best)`
- Load the latest weights from the most recent experiment: `Nitro(e; weights=:all => :latest)`
  - You can do the same thing with the `resume` kwarg; `resume` was formerly `checkpoint` which still works but is deprecated. This makes the users intent clear: they intend to resume the training run which is more involved than just loading weights and optimizer state.

- If you use an experiment directory per experiment, there are no behavior changes and everything continues to work the way it did before. There was a bug when re-using the same experiment directory using `topk`, previously it did `topk` over all previous experiments and the current. Now you can control the behavior by choosing `scope=:run` (default) or `scope=:all`. Only checkpoints with a matching metric key are considered.

More improvements:

- You can now also disable always keeping the most recent checkpoint with `keep_latest=false`.
- `restore_optimizer` flag is provided for when you use `weights=` incase for some reason you also want the optimizer state (default `false`)
  - A logical next step would be making this do something for `resume=`, it would default to `true` there of course.

- `restore_best` option added for `train!` so you can ensure the best checkpoint gets loaded into your `Nitro` at the end of training.
