# Significantly Higher VRAM Usage and Slower Training on Flux Compared to PyTorch

**URL:** <https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124>\
**Category:** Machine Learning\
**Created:** [May 14, 2026, 9:29pm UTC](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124 "2026-05-14T21:29:09Z")\
**Posts on this page:** 1\
**Showing post:** 15

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [May 18, 2026, 4:54am UTC](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124/15 "2026-05-18T04:54:45Z")

</div>

> [@JoshuaBillson](#):
>
> Does `StableHLO` produce native Julia models, where you can extract and modify layers, or is it essentially a black box like ONNX? The reason I ask is that my current project involves implementing `Timm` models as pure Julia equivalents in `Flux`, then I define a `load_params!` method that takes a PyTorch `state_dict` from the matching `Timm` model/layer and loads the parameters into the corresponding `Flux` layer. This has the advantage of producing a model that can be used like any other `Flux` layer, but obviously requires a fair amount of work to duplicate the original PyTorch code.

It doesn’t produce native Julia models, but you can use it inside the rest of your native Julia model. Well, I suppose as long as you compile it to Reactant.jl.

Also see here, I’m working on a timm port for Lux.jl: [[ANN] Jimm.jl: Lux ports of timm image backbones, with HuggingFace pretrained weights](https://discourse.julialang.org/t/ann-jimm-jl-lux-ports-of-timm-image-backbones-with-huggingface-pretrained-weights/137153/1)

It would likely be pretty easy to do the same sort of workflow I did here to port things over to Flux.jl, but I’m not personally familiar with it.

---

_[View the full topic](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124)._
