# \[ANN\] Luximm.jl: Lux ports of timm image backbones, with HuggingFace pretrained weights

**URL:** <https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153>\
**Category:** Package Announcements\
**Tags:** package, lux, computer-vision\
**Created:** [May 18, 2026, 4:51am UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153 "2026-05-18T04:51:33Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [May 18, 2026, 4:51am UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/1 "2026-05-18T04:51:34Z")

</div>

I’m happy to share a first public look at **Luximm.jl** , a Julia package that ports image-classification backbones from Ross Wightman’s [`timm`](https://github.com/huggingface/pytorch-image-models) (PyTorch Image Models) to [Lux.jl](https://lux.csail.mit.edu/). Pretrained weights load directly from the HuggingFace Hub in `.safetensors` format, sharing the same on-disk cache `timm` and `huggingface_hub` use.

**Repo:** [GitHub - csvance/Luximm.jl: Lux.jl Image Models · GitHub](https://github.com/csvance/Luximm.jl) _(see the README for install, quickstart, the full variant table, and the porting workflow; this post is just the elevator pitch.)_

## Why this exists

The motivation was concrete: I needed Julia’s SciML ecosystem together with modern pretrained vision backbones, and Python doesn’t have a peer for SciML. The original stack was vision encoders feeding `torchdiffeq` in PyTorch, which works but leaves much to be desired. Moving the DiffEQ side to Julia meant the vision side had to come too. Jimm started as a one-off port of a single backbone for that internal use case and snowballed from there. If your work also lives at that intersection of pretrained vision encoders and the rest of the SciML stack, the hope is that Jimm makes Julia a more complete option for that workload.

## What you get today

A small but useful set of modern CNN backbones for research and practitioner use, with pretrained weights: **BiT ResNetV2** (15 variants, Apache 2.0), **ConvNeXt v1** (19 Facebook AI variants from the original 2022 paper, Apache 2.0, plus 4 DINOv3 encoders under Meta’s DINOv3 License), and **ConvNeXt V2** (26 FCMAE variants, CC BY-NC 4.0). 64 checkpoints in total. Pretrained weights load by passing the variant key to a one-line `load_<family>_pretrained` call. ViT, EfficientNet, Swin, and the rest of the `timm` catalog are open targets for contribution.

## What Jimm is, and isn’t

It is a strict port: same architectures, same hyperparameters, same weight init, same `state_dict` key layout, so any `timm/<variant>` checkpoint on HuggingFace loads without manual rewiring and the forward pass matches `timm` to within float32 round-off. It is **not** a Julia-native reimagining, a general CV toolkit, or a training framework, and **it is not at 1:1 parity with the full `timm` catalog, nor is it likely to ever be.** `timm` has hundreds of architectures and thousands of checkpoints; Jimm tracks the subset its contributors actually use. Backbones land via PR.

## Correctness gate

Every registered variant has a parity test that downloads the real `.safetensors` from HuggingFace, loads them through the Lux model, runs the forward pass, and asserts max-abs-diff against `timm`’s output on the same input is under `1e-3` (most variants land closer to `1e-4`). That single test covers both the architecture port and the weight loader: if the safetensors loader misroutes a tensor or applies the wrong axis permutation, the forward output diverges and the test fails. CI is still a work in progress (the full sweep is expensive and we are figuring out how to run it with the resources we have), but contributors can scope runs to a single variant via `JIMM_TEST_VARIANTS`.

## How this code was produced (and a caveat)

Most of Jimm was written by AI agents driving the porting workflow encoded in `.claude/skills/timm-to-lux/`, with human review at each phase and the parity tests as the correctness backstop. The code is already being used in real projects, so it works, but **expect bugs and rough edges** , especially around features that the parity tests do not exercise (anything past forward inference with the released weights). File issues; we will fix them.

## Porting new backbones with Claude Code

If you have Claude Code and tokens to spare, the practical path to a new backbone is: open the repo, ask Claude to port `timm/<your_model>`, and follow the skill at `.claude/skills/timm-to-lux/`. It produces a working PR for most `timm` architectures in a single session taking into account the differences between the two frameworks.

**Contributions are welcome and encouraged** , with or without the skill. See the README’s “Contributing a new backbone” section for the acceptance criteria. Bug reports, PRs for new variants of registered families, and PRs for entirely new families are all in scope.

Thanks to Ross Wightman for `timm`, the Julia ML ecosystem maintainers, and to my employer [Medical Metrics Inc.](https://medicalmetrics.com/).

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [May 21, 2026, 9:31pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/2 "2026-05-21T21:31:57Z")

</div>

A few updates:

- First external contribution to the project from [PaioPaio](https://github.com/PaioPaio) : classic ResNet ([**Deep residual learning** for **image recognition**](http://openaccess.thecvf.com/content_cvpr_2016/html/He_Deep_Residual_Learning_CVPR_2016_paper.html))
- There is now a working CI system that tests every single backbone. I wasn’t able to use Github actions for this due to size of the weights / parity fixture cache, so I built a custom interface with Tachikoma.jl that runs everything inside of a dedicated virtual machine. CI is triggered manually to prevent bitcoin mining 😅
- Documenter.jl is now setup for the repo and I spent a decent amount of time trying to better refine and organize everything including the main readme.
- I had to update how we handle the tolerance check for features. The new formula is `max-abs-diff / max-abs(timm ref)`. Absolute tolerance on its own was proving rather difficult to work with across a wide range of models with different activation scales and depths. We should be able to use 1e-5 rtol as the new threshold using this formula and have it work across all backbones. Open to any suggestions here.
- The main thing left before registering the package is making sure that all of our interfaces are sound / we are not doing anything that causes type instability, etc.

---

<div class="post-metadata">

**Author:** ![madppiper](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/madppiper/32/220029_2.png) [@madppiper](https://discourse.julialang.org/u/madppiper)\
**Post date:** [May 22, 2026, 6:37pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/3 "2026-05-22T18:37:44Z")

</div>

So so cool!!!

---

<div class="post-metadata">

**Author:** ![madppiper](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/madppiper/32/220029_2.png) [@madppiper](https://discourse.julialang.org/u/madppiper)\
**Post date:** [May 23, 2026, 6:41pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/4 "2026-05-23T18:41:47Z")

</div>

Can’t believe nobody else posted a response than me. So let me repeat what I said:

This is actually really cool! Rest assured I’ll give it a go. Hi five!

---

<div class="post-metadata">

**Author:** ![TheLateKronos](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thelatekronos/32/12824_2.png) [@TheLateKronos](https://discourse.julialang.org/u/TheLateKronos)\
**Post date:** [May 24, 2026, 6:03am UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/5 "2026-05-24T06:03:17Z")

</div>

Congratz on the package 🚀

This sounds like a solid package to develop the ecosystem, very nice indeed.

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [May 26, 2026, 4:16am UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/6 "2026-05-26T04:16:02Z")

</div>

v0.1.0 has been published to the general registry! The package has been renamed Luximm.jl to satisfy naming guidelines.

This is a breaking release in terms of interface from the pre-release version, but it streamlines everything making it much easier to use. It’s much more `timm`-like even if its not fully possible to capture the exact same semantics with how Lux.jl handles parameters and state.

```julia-auto
using Luximm, Lux, Random

# ResNet50 with the trained 1000-class ImageNet head.
# `create_pretrained` is family-agnostic; the symbol selects the family.
# It returns the model and a closure that loads the released weights
# into `(ps, st)` once you've run `Lux.setup`.
model, load = create_pretrained(:resnet50_a1_in1k)
ps, st = Lux.setup(Xoshiro(0), model)
ps, st = load(ps, st)
st = Lux.testmode(st) # BatchNorm/Dropout in eval mode

x = randn(Float32, 224, 224, 3, 1)
logits, _ = model(x, ps, st) # (1000, 1)
top1 = argmax(vec(logits)) # ImageNet class index

```

## QoL Improvements

- You no longer need to remember the exact names for the model creation / weight loading functions for each model family. We now have a standard interface across all families: `create_pretrained` and `create_model`.
- No more passing the same arguments to both the model creation and weight loading functions. `create_pretrained` returns a function capturing the arguments needed to initialize the weights.

I won’t rule out changes to the interface for v1.0.0 someday, but the interface should remain stable for the forseeable future.

---

<div class="post-metadata">

**Author:** ![greatpet](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Post date:** [May 26, 2026, 9:58am UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/7 "2026-05-26T09:58:57Z")

</div>

What’s the main difference from Metalhead.jl? (1) Lux instead of Flux, obviously. (2) Ability to load pretrained weights from HF? (3) A non-overlapping coverage of CNN architectures - I guess Metalhead has a broader coverage, while Luximm has more up-to-date models?

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [May 26, 2026, 1:20pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/8 "2026-05-26T13:20:17Z")

</div>

> [@greatpet](#):
>
> What’s the main difference from [Metalhead.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/Metalhead)?

**TLDR:** The goal of Luximm is different: it aspires to be a Julia counterpart to timm. Luximm aims to eventually cover most of the numerous backbones timm offers, all with pretrained weights. The recency bias in v0.1.0 simply reflects the backbones I work with day-to-day. The focus for this release was building the foundation to add many more architectures easily. I have another ~5 planned for the next month, and I intend to keep pushing them out on a regular basis after that.

Pretrained weights are central to the value proposition for me. As a machine learning engineer who works with these backbones regularly, training from scratch isn’t usually practical: pretraining consumes a large amount of resources, and generalization tends to suffer significantly without it. So every architecture in Luximm ships with pretrained weights where they are available.

EDIT: I’m also open to the idea of trying to split off a common core of functionality that could be used by both Lux.jl and Flux.jl to load these models. I don’t know exactly what that looks like, perhaps a TimmCommon.jl package that handled most of the work of interacting with Hugging Face / providing a general way to easily map the loaded weights into your framework of choice. Feel free to reach out if you are interested in something like this!

---

<div class="post-metadata">

**Author:** ![yolhan\_mannes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yolhan_mannes/32/220485_2.png) [@yolhan\_mannes](https://discourse.julialang.org/u/yolhan_mannes)\
**Post date:** [May 26, 2026, 4:33pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/9 "2026-05-26T16:33:56Z")

</div>

Hi very nice package have you tested if it’s compatible with Reactant and the whole MLIR tools around.

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [May 26, 2026, 4:45pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/10 "2026-05-26T16:45:41Z")

</div>

It absolutely works with Reactant! At the end of the day it’s just standard Lux.jl and NNlib. It’s not something we actively test currently as part of CI, but my current project is using one of the backbones with Reactant.

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [June 7, 2026, 11:33pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/11 "2026-06-07T23:33:02Z")

</div>

v0.1.1 has been released and it builds the foundation to provide ViT based models. Future releases will target many of the common ones used by researcher and practitioners (ie CLIP and DINO). In addition to that, I included a few more classic architectures:

- VGG16/19
- SE-ResNet
- CoAtNet

I’m aiming on pushing out new sets of weights / architectures around twice a month until I feel like we have coverage of the most essential backbones.

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [August 12, 2026, 9:20pm UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/12 "2026-08-12T21:20:13Z")

</div>

v0.2.0 is out now, adding support for returning multi-scale features that are needed for decoders such as U-Net, FPN, etc. This release also adds support for the CLIP family of ViT models.

Also to answer your question @yolhan_mannes, I’ve been exclusively using these with Reactant for the past two months. For my workloads I’m getting better performance than `torch.compile` for the GPUs I have available, often by a factor of 2x or more. Reactant and Enzyme are truly incredible.

---

<div class="post-metadata">

**Author:** ![madppiper](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/madppiper/32/220029_2.png) [@madppiper](https://discourse.julialang.org/u/madppiper)\
**Post date:** [August 13, 2026, 9:51am UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/13 "2026-08-13T09:51:58Z")

</div>

This is great. Good job!

---

<div class="post-metadata">

**Author:** ![JackDevine](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jackdevine/32/1048_2.png) [@JackDevine](https://discourse.julialang.org/u/JackDevine)\
**Post date:** [August 14, 2026, 12:27am UTC](https://discourse.julialang.org/t/ann-luximm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/14 "2026-08-14T00:27:32Z")

</div>

Very cool
