# Machine Learning

**URL:** https://discourse.julialang.org/c/domain/ml/24.md

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

---

## [About the Machine Learning category](https://discourse.julialang.org/t/about-the-machine-learning-category/315)

<div class="topic-metadata">

**Author:** [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Replies:** 4\
**Last updated:** [May 2, 2021, 4:44pm UTC](https://discourse.julialang.org/t/about-the-machine-learning-category/315 "2021-05-02T16:44:16Z")

</div>

Machine Learning in Julia, with a particular focus on Deep Learning Flux Knet Tensorflow.jl Differentiable programming

---

## [ReactantServer.jl: serve more models per GPU with Reactant.jl and XLA (gauging interest)](https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358)

<div class="topic-metadata">

**Author:** [@csvance](https://discourse.julialang.org/u/csvance)\
**Replies:** 18\
**Last updated:** [October 6, 2026, 12:37am UTC](https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358 "2026-10-06T00:37:25Z")

</div>

I have been building a Julia inference server and want to find out whether others would find it useful before investing more in polishing it for general use. Feedback, criticism, and “we already have this, it is called X…

---

## [Will Reactant.jl become a machine learning framework?](https://discourse.julialang.org/t/will-reactant-jl-become-a-machine-learning-framework/125478)

<div class="topic-metadata">

**Author:** [@Tarny\_GG\_Channie](https://discourse.julialang.org/u/Tarny_GG_Channie)\
**Replies:** 17\
**Last updated:** [September 24, 2026, 7:05pm UTC](https://discourse.julialang.org/t/will-reactant-jl-become-a-machine-learning-framework/125478 "2026-09-24T19:05:50Z")

</div>

I made an observation that MLIR + Enzyme could form quite a capable ML framework, and then, well… like the last time I theorized about how a good language would work, and Julians were one step ahead of me, creating Julia…

---

## [Flux train! no longer works](https://discourse.julialang.org/t/flux-train-no-longer-works/139543)

<div class="topic-metadata">

**Author:** [@MathJulia](https://discourse.julialang.org/u/MathJulia)\
**Replies:** 1\
**Last updated:** [September 18, 2026, 5:41pm UTC](https://discourse.julialang.org/t/flux-train-no-longer-works/139543 "2026-09-18T17:41:08Z")

</div>

I have a neural network which has two parts. one encoder and the other main network. Due to recent changes in Flux.jl API, my code no longer works. I keep some of weights of main network and train them beside encoder. ho…

---

## [\[ANN\] MachineLearningForecast.jl: forecasting with any MLJ regressor](https://discourse.julialang.org/t/ann-machinelearningforecast-jl-forecasting-with-any-mlj-regressor/139302)

<div class="topic-metadata">

**Author:** [@Resul.Akay](https://discourse.julialang.org/u/Resul.Akay)\
**Replies:** 3\
**Last updated:** [September 10, 2026, 10:04am UTC](https://discourse.julialang.org/t/ann-machinelearningforecast-jl-forecasting-with-any-mlj-regressor/139302 "2026-09-10T10:04:13Z")

</div>

MachineLearningForecast.jl: Forecasting with any MLJ regressor I’m happy to share MachineLearningForecast.jl, a package for time-series forecasting that turns any MLJ Deterministic regressor into a forecaster. EvoTrees, …

---

## [Training a (conditional) Universal Differential Equation using binary observations of a latent state](https://discourse.julialang.org/t/training-a-conditional-universal-differential-equation-using-binary-observations-of-a-latent-state/138431)

<div class="topic-metadata">

**Author:** [@NCCP](https://discourse.julialang.org/u/NCCP)\
**Replies:** 1\
**Last updated:** [September 5, 2026, 9:49am UTC](https://discourse.julialang.org/t/training-a-conditional-universal-differential-equation-using-binary-observations-of-a-latent-state/138431 "2026-09-05T09:49:40Z")

</div>

Hello everyone, I am training a conditional Universal Differential Equation in Julia. The model is a regular UDE with one additional trainable input parameter for each particle. The neural-network parameters are shared…

---

## [Exact Network Surgery and Reactive Computational Graphs in Julia with NeuroDSL](https://discourse.julialang.org/t/exact-network-surgery-and-reactive-computational-graphs-in-julia-with-neurodsl/138303)

<div class="topic-metadata">

**Author:** [@Khemais\_Abdallah](https://discourse.julialang.org/u/Khemais_Abdallah)\
**Replies:** 14\
**Last updated:** [July 22, 2026, 12:21pm UTC](https://discourse.julialang.org/t/exact-network-surgery-and-reactive-computational-graphs-in-julia-with-neurodsl/138303 "2026-07-22T12:21:01Z")

</div>

Exact Network Surgery and Reactive Computational Graphs in Julia with NeuroDSL Hi everyone, I’d like to share some recent theoretical and systems results from NeuroDSL, a persistent, reactive computational graph framewo…

---

## [Mutable Computational Graphs in Pure Julia: Performing "Hot Surgery" on a LLaMA Transformer](https://discourse.julialang.org/t/mutable-computational-graphs-in-pure-julia-performing-hot-surgery-on-a-llama-transformer/138018)

<div class="topic-metadata">

**Author:** [@Khemais\_Abdallah](https://discourse.julialang.org/u/Khemais_Abdallah)\
**Replies:** 0\
**Last updated:** [July 8, 2026, 11:53pm UTC](https://discourse.julialang.org/t/mutable-computational-graphs-in-pure-julia-performing-hot-surgery-on-a-llama-transformer/138018 "2026-07-08T23:53:25Z")

</div>

Hi everyone, I am Abdallah Khemais. As an Instructor of Applied Mathematics and AI at ISITCOM Sousse Tunisia, with over two decades of experience, my work focuses on the intersection of formal mathematical theory and hig…

---

## [Small step size in LBFGS for neural network with complex values](https://discourse.julialang.org/t/small-step-size-in-lbfgs-for-neural-network-with-complex-values/137946)

<div class="topic-metadata">

**Author:** [@Knud\_Sorensen](https://discourse.julialang.org/u/Knud_Sorensen)\
**Replies:** 3\
**Last updated:** [July 5, 2026, 6:46pm UTC](https://discourse.julialang.org/t/small-step-size-in-lbfgs-for-neural-network-with-complex-values/137946 "2026-07-05T18:46:42Z")

</div>

I wrote my own LBFGS code for a neural network with complex values. It worked until the step size became to small to change the network because of floating point position. Has anyone had similar problems and know how t…

---

## [Invertible Neural Networks](https://discourse.julialang.org/t/invertible-neural-networks/137909)

<div class="topic-metadata">

**Author:** [@langestefan](https://discourse.julialang.org/u/langestefan)\
**Replies:** 0\
**Last updated:** [July 2, 2026, 2:40pm UTC](https://discourse.julialang.org/t/invertible-neural-networks/137909 "2026-07-02T14:40:20Z")

</div>

I was looking at doing some modelling work using invertible neural-networks in Julia with Lux.jl. There aren’t many options, but I found GitHub - slimgroup/InvertibleNetworks.jl: A Julia framework for invertible neural …

---

## [SciML: Issues using \`GaussAdjoint\` on an \`EnsembleProblem\` with callbacks](https://discourse.julialang.org/t/sciml-issues-using-gaussadjoint-on-an-ensembleproblem-with-callbacks/137503)

<div class="topic-metadata">

**Author:** [@LucasGrasso](https://discourse.julialang.org/u/LucasGrasso)\
**Replies:** 11\
**Last updated:** [June 21, 2026, 8:30pm UTC](https://discourse.julialang.org/t/sciml-issues-using-gaussadjoint-on-an-ensembleproblem-with-callbacks/137503 "2026-06-21T20:30:04Z")

</div>

Hello there, I am rather new to Julia, and I am trying to implement a Neural ODE model that relies on event callbacks to terminate integration. I am using the GaussAdjoint, given that it supports callbacks and that it w…

---

## [Speeding up antialiased non-linearity](https://discourse.julialang.org/t/speeding-up-antialiased-non-linearity/137456)

<div class="topic-metadata">

**Author:** [@fps](https://discourse.julialang.org/u/fps)\
**Replies:** 2\
**Last updated:** [June 10, 2026, 7:30am UTC](https://discourse.julialang.org/t/speeding-up-antialiased-non-linearity/137456 "2026-06-10T07:30:54Z")

</div>

Hi, I’m playing around with some guitar amplifier modeling and in the course of that I implemented the anti-derivative anti-aliased non-linearity of “Note on Alias Suppression in Digital Distortion”, Martin Vicanek in a…

---

## [Trouble using @compact along reactant.jl](https://discourse.julialang.org/t/trouble-using-compact-along-reactant-jl/137497)

<div class="topic-metadata">

**Author:** [@yjr](https://discourse.julialang.org/u/yjr)\
**Replies:** 4\
**Last updated:** [June 8, 2026, 5:27pm UTC](https://discourse.julialang.org/t/trouble-using-compact-along-reactant-jl/137497 "2026-06-08T17:27:30Z")

</div>

Hello, I’m trying to set a simple network as described by the code below, but when I compiled it to GPU using reactant.jl it works if the sequence dimension (2nd dimension) is smaller than 100, but if I grow it bigger i…

---

## [Memory allocations in Flux evaluation/training (w/MVE)](https://discourse.julialang.org/t/memory-allocations-in-flux-evaluation-training-w-mve/137175)

<div class="topic-metadata">

**Author:** [@CrashBurnRepeat](https://discourse.julialang.org/u/CrashBurnRepeat)\
**Replies:** 39\
**Last updated:** [June 6, 2026, 10:00am UTC](https://discourse.julialang.org/t/memory-allocations-in-flux-evaluation-training-w-mve/137175 "2026-06-06T10:00:29Z")

</div>

I’m currently using Flux to do some GAN experiments and I’m finding that the training is painfully slow. On the CPU, a single epoch can take ~30min for a 286x32 sample dataloader. GPU evaluation is faster, but still feel…

---

## [Reactant.jl recompiles function each time it is called](https://discourse.julialang.org/t/reactant-jl-recompiles-function-each-time-it-is-called/137429)

<div class="topic-metadata">

**Author:** [@fastwave](https://discourse.julialang.org/u/fastwave)\
**Replies:** 7\
**Last updated:** [June 3, 2026, 5:10pm UTC](https://discourse.julialang.org/t/reactant-jl-recompiles-function-each-time-it-is-called/137429 "2026-06-03T17:10:36Z")

</div>

I am trying to use an optimiser from Manopt.jl to train a Lux.jl neural network using Reactant.jl. It is orders of magnitude slower than without Reactant.jl. Does anyone know how to fix it? The compiled functions, when …

---

## [Lux & Reactant on Colab TPUs](https://discourse.julialang.org/t/lux-reactant-on-colab-tpus/126926)

<div class="topic-metadata">

**Author:** [@avikpal](https://discourse.julialang.org/u/avikpal)\
**Replies:** 5\
**Last updated:** [May 21, 2026, 5:30pm UTC](https://discourse.julialang.org/t/lux-reactant-on-colab-tpus/126926 "2026-05-21T17:30:22Z")

</div>

Colab now has Lux and Reactant pre-installed. Here’s a starter script for TPUs (change the runtime, and it will automatically work on CPU or GPU)

---

## [Reactant.jl CUDA Version](https://discourse.julialang.org/t/reactant-jl-cuda-version/137211)

<div class="topic-metadata">

**Author:** [@csvance](https://discourse.julialang.org/u/csvance)\
**Replies:** 0\
**Last updated:** [May 20, 2026, 11:45am UTC](https://discourse.julialang.org/t/reactant-jl-cuda-version/137211 "2026-05-20T11:45:16Z")

</div>

Is there a recommended CUDA driver version + version of CUDA.jl I should be using with Reactant.jl? It looks like there are conflicting versions of CUDA somewhere. I0000 00:00:1779236186.222218 3515013 service.cc:178\] X…

---

## [Significantly Higher VRAM Usage and Slower Training on Flux Compared to PyTorch](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124)

<div class="topic-metadata">

**Author:** [@JoshuaBillson](https://discourse.julialang.org/u/JoshuaBillson)\
**Replies:** 25\
**Last updated:** [May 18, 2026, 2:59pm UTC](https://discourse.julialang.org/t/significantly-higher-vram-usage-and-slower-training-on-flux-compared-to-pytorch/137124 "2026-05-18T14:59:25Z")

</div>

I’ve been working on a project where I’m converting Timm models to Flux along with their pre-trained weights. Everything is going well so far, but I noticed that Flux has significantly higher VRAM consumption and takes a…

---

## [Minimal PINN example in Julia that works with current packages?](https://discourse.julialang.org/t/minimal-pinn-example-in-julia-that-works-with-current-packages/137147)

<div class="topic-metadata">

**Author:** [@photor](https://discourse.julialang.org/u/photor)\
**Replies:** 4\
**Last updated:** [May 18, 2026, 8:59am UTC](https://discourse.julialang.org/t/minimal-pinn-example-in-julia-that-works-with-current-packages/137147 "2026-05-18T08:59:03Z")

</div>

Hi, I’m trying to write a simple PINN (e.g., Newton cooling ODE) in Julia. I’ve tried Flux + Zygote but keep hitting issues from recent API changes and slow higher‑order AD. What’s the current recommended stack and a min…

---

## [Error and Segfault with DeviceIterator for parallel DataLoader on sharded ReactantDevice](https://discourse.julialang.org/t/error-and-segfault-with-deviceiterator-for-parallel-dataloader-on-sharded-reactantdevice/137074)

<div class="topic-metadata">

**Author:** [@benjwweber](https://discourse.julialang.org/u/benjwweber)\
**Replies:** 1\
**Last updated:** [May 14, 2026, 10:28pm UTC](https://discourse.julialang.org/t/error-and-segfault-with-deviceiterator-for-parallel-dataloader-on-sharded-reactantdevice/137074 "2026-05-14T22:28:16Z")

</div>

I am trying to load a data set via a DataLoader/DeviceIterator to a sharded ReactantDevice. I have replicated it here with a mock data set. However I run into a variety of errors or segfaults if I do this. Is there anyth…

---

## [Supressing "Info:" in XGBoost package?](https://discourse.julialang.org/t/supressing-info-in-xgboost-package/137121)

<div class="topic-metadata">

**Author:** [@Snowy](https://discourse.julialang.org/u/Snowy)\
**Replies:** 1\
**Last updated:** [May 14, 2026, 4:16pm UTC](https://discourse.julialang.org/t/supressing-info-in-xgboost-package/137121 "2026-05-14T16:16:51Z")

</div>

Hi, When training an XGBoost the xgboost.jl package provides information about the rmse value: This is helpful when training first but when running in production I might not want my terminal to be filled with this, …

---

## [How to train dense nets several times faster than with Adam](https://discourse.julialang.org/t/how-to-train-dense-nets-several-times-faster-than-with-adam/137009)

<div class="topic-metadata">

**Author:** [@uwestoehr](https://discourse.julialang.org/u/uwestoehr)\
**Replies:** 4\
**Last updated:** [May 8, 2026, 10:36am UTC](https://discourse.julialang.org/t/how-to-train-dense-nets-several-times-faster-than-with-adam/137009 "2026-05-08T10:36:59Z")

</div>

I had the pleasure to give a lecture about the basics of neural nets at a university and purposely used Julia also for the practical lessons. As a side effect I wanted to explain my students how an optimizer can be deri…

---

## [How to contribute an optimizer to the Optimisers.jl package](https://discourse.julialang.org/t/how-to-contribute-an-optimizer-to-the-optimisers-jl-package/136975)

<div class="topic-metadata">

**Author:** [@uwestoehr](https://discourse.julialang.org/u/uwestoehr)\
**Replies:** 0\
**Last updated:** [May 4, 2026, 4:07am UTC](https://discourse.julialang.org/t/how-to-contribute-an-optimizer-to-the-optimisers-jl-package/136975 "2026-05-04T04:07:17Z")

</div>

Since almost 2 years I use Flux and Lux and thus the Optimisers.jl package. Since I need for scientific reasons also tests with optimizers not supported by the Optimisers.jl package, I wrote them for my own. I want now …

---

## [AD over runtime-selected fields in nested structs — design patterns?](https://discourse.julialang.org/t/ad-over-runtime-selected-fields-in-nested-structs-design-patterns/136882)

<div class="topic-metadata">

**Author:** [@stevenhughes](https://discourse.julialang.org/u/stevenhughes)\
**Replies:** 4\
**Last updated:** [April 26, 2026, 12:10pm UTC](https://discourse.julialang.org/t/ad-over-runtime-selected-fields-in-nested-structs-design-patterns/136882 "2026-04-26T12:10:49Z")

</div>

We have a framework where users define models as nested structs and declare at runtime which fields are optimization variables. The framework then needs to differentiate constraints/objectives with respect to those selec…

---

## [Ignore masked tokens (Flux.jl)](https://discourse.julialang.org/t/ignore-masked-tokens-flux-jl/75591)

<div class="topic-metadata">

**Author:** [@spragud2](https://discourse.julialang.org/u/spragud2)\
**Replies:** 6\
**Last updated:** [April 24, 2026, 7:56pm UTC](https://discourse.julialang.org/t/ignore-masked-tokens-flux-jl/75591 "2026-04-24T19:56:14Z")

</div>

Hi, I’m working on implementing a seq2seq model in Flux. Is there a standard method, or if not a suitable approach, for ignoring masked tokens when computing the gradient? E.g. sentences of differing lengths, following …

---

## [New package for clustering - need for an arXiv endorser](https://discourse.julialang.org/t/new-package-for-clustering-need-for-an-arxiv-endorser/128983)

<div class="topic-metadata">

**Author:** [@uwestoehr](https://discourse.julialang.org/u/uwestoehr)\
**Replies:** 3\
**Last updated:** [April 13, 2026, 2:30pm UTC](https://discourse.julialang.org/t/new-package-for-clustering-need-for-an-arxiv-endorser/128983 "2026-04-13T14:30:42Z")

</div>

Hi I am Uwe, a MEMS engineer. During evaluation of some measurement data I had an idea for a clustering algorithm and started with Julia to be able to implement this algorithm as I had it in mind. That succeeded and the …

---

## [AMDGPU.jl issues](https://discourse.julialang.org/t/amdgpu-jl-issues/136171)

<div class="topic-metadata">

**Author:** [@jcbritobr](https://discourse.julialang.org/u/jcbritobr)\
**Replies:** 4\
**Last updated:** [March 22, 2026, 6:11pm UTC](https://discourse.julialang.org/t/amdgpu-jl-issues/136171 "2026-03-22T18:11:34Z")

</div>

Hello, good evening. Im having issue to work with amdgpu.jl and ROCM Hip. The issue happens when device = MLDataDevices.AMDGPUDevice() # get device. # Multilayer Perceptron: model = Chain( Dense(2 =\>…

---

## [Best Julia Package for Neural Networks](https://discourse.julialang.org/t/best-julia-package-for-neural-networks/45358)

<div class="topic-metadata">

**Author:** [@Jojo\_Dad](https://discourse.julialang.org/u/Jojo_Dad)\
**Replies:** 13\
**Last updated:** [March 12, 2026, 10:46am UTC](https://discourse.julialang.org/t/best-julia-package-for-neural-networks/45358 "2026-03-12T10:46:44Z")

</div>

What is currently the best machine learning package in Julia? First, in terms of ease of use for students. Second, in terms of capability & performance for experts?

---

## [Why is Flux/NNlib falling back to im2col instead of MIOpen on AMD GPU?](https://discourse.julialang.org/t/why-is-flux-nnlib-falling-back-to-im2col-instead-of-miopen-on-amd-gpu/135867)

<div class="topic-metadata">

**Author:** [@ajrohr2](https://discourse.julialang.org/u/ajrohr2)\
**Replies:** 3\
**Last updated:** [March 10, 2026, 1:30am UTC](https://discourse.julialang.org/t/why-is-flux-nnlib-falling-back-to-im2col-instead-of-miopen-on-amd-gpu/135867 "2026-03-10T01:30:35Z")

</div>

Hello! I’ve been trying to get Flux’s convolutional layers, specifically the Flux.NNlib.conv! function, to use the MIOpen version rather than im2col or direct. I’ve verified that I have MIOpen available with AMDGPU.funct…

---

## [Optimizing Lux.jl and Reactant.jl performance](https://discourse.julialang.org/t/optimizing-lux-jl-and-reactant-jl-performance/135942)

<div class="topic-metadata">

**Author:** [@KSepetanc](https://discourse.julialang.org/u/KSepetanc)\
**Replies:** 3\
**Last updated:** [March 6, 2026, 8:58pm UTC](https://discourse.julialang.org/t/optimizing-lux-jl-and-reactant-jl-performance/135942 "2026-03-06T20:58:07Z")

</div>

I am facing two separate issues. One is related to setting Reactant initialization options to improve performance and the other is about avoiding unnecessary latencies in neural network training loop via compilation. No…

[Next page](https://discourse.julialang.org/c/domain/ml/24.md?page=1)
