# Significantly Higher VRAM Usage and Slower Training on Flux Compared to PyTorch

**URL:** <https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124>\
**Category:** Machine Learning\
**Created:** [May 14, 2026, 9:29pm UTC](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124 "2026-05-14T21:29:09Z")\
**Posts on this page:** 1\
**Showing post:** 16

<div class="post-metadata">

**Author:** ![JoshuaBillson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshuabillson/32/52706_2.png) [@JoshuaBillson](https://discourse.julialang.org/u/JoshuaBillson)\
**Post date:** [May 18, 2026, 5:01am UTC](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124/16 "2026-05-18T05:01:00Z")

</div>

> Also see here, I’m working on a timm port for [Lux.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/Lux): [[ANN] Jimm.jl: Lux ports of timm image backbones, with HuggingFace pretrained weights](https://discourse.julialang.org/t/ann-jimm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/1)

That’s almost exactly what I’m working on (my working name was even Jimm). I’ll take a look and see if I can contribute. So far, I’ve implemented all variants of Timm’s `VisionTransformer`, `ConvNeXt` (both v1 and v2), and `Eva` (basically ViT with rotary positional embeddings used by SAM3). I also have implementations for `Swin`, `PVT`, and `Twins`, but I didn’t get around to adding pre-trained weights yet. It should be relatively straightforward to convert from `Flux` to `Lux`.

---

_[View the full topic](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124)._
