# Machine Learning

**URL:** https://discourse.julialang.org/c/domain/ml/24.md?page=3

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 4

---

## [Build a Build deep neural network classifier using a long short-term memory (LSTM) architecture](https://discourse.julialang.org/t/build-a-build-deep-neural-network-classifier-using-a-long-short-term-memory-lstm-architecture/129606)

<div class="topic-metadata">

**Author:** [@Ltukondjere](https://discourse.julialang.org/u/Ltukondjere)\
**Replies:** 3\
**Last updated:** [June 4, 2025, 4:30am UTC](https://discourse.julialang.org/t/build-a-build-deep-neural-network-classifier-using-a-long-short-term-memory-lstm-architecture/129606 "2025-06-04T04:30:03Z")

</div>

Good day, Please help me. I recently started with Julia and i am trying to use it to build a deep neural network classifier using a long short-term memory (LSTM) architecture. I manage to load my data, did text preproce…

---

## [\[Resolved\] ERROR: SciMLSensitivity](https://discourse.julialang.org/t/resolved-error-scimlsensitivity/129531)

<div class="topic-metadata">

**Author:** [@aquarelleX332](https://discourse.julialang.org/u/aquarelleX332)\
**Replies:** 7\
**Last updated:** [June 3, 2025, 9:00am UTC](https://discourse.julialang.org/t/resolved-error-scimlsensitivity/129531 "2025-06-03T09:00:43Z")

</div>

Hello all, I am dealing with a code including CRNN in the case of 2 consecutives chemical reactions. The code can be found here. Amongst other packages I am using: SciMLSensitivity v7.84.0 Zygote v0.7.7 The error is…

---

## [Questions about the DiffEqFlux multiple shooting example](https://discourse.julialang.org/t/questions-about-the-diffeqflux-multiple-shooting-example/129418)

<div class="topic-metadata">

**Author:** [@julianewbsourcream](https://discourse.julialang.org/u/julianewbsourcream)\
**Replies:** 4\
**Last updated:** [June 2, 2025, 5:17pm UTC](https://discourse.julialang.org/t/questions-about-the-diffeqflux-multiple-shooting-example/129418 "2025-06-02T17:17:28Z")

</div>

Hi, I am training a neural ODE on a periodic function and wanted to try out DiffEqFlux.multiple\_shoot(). After reviewing the online example @ docs.sciml.ai/DiffEqFlux/stable/examples/multiple\_shooting/, I was hoping so…

---

## [Facing high loss values when training the SciML model](https://discourse.julialang.org/t/facing-high-loss-values-when-training-the-sciml-model/129273)

<div class="topic-metadata">

**Author:** [@Rohan\_Kumar\_123](https://discourse.julialang.org/u/Rohan_Kumar_123)\
**Replies:** 1\
**Last updated:** [May 29, 2025, 10:59am UTC](https://discourse.julialang.org/t/facing-high-loss-values-when-training-the-sciml-model/129273 "2025-05-29T10:59:47Z")

</div>

Greetings, i am currently working on a problem to predict the state vectors of a satellite using a dataset which is generated by a propagator. I have made a Universal Differential Equation(UDE) to predict the state vecto…

---

## [Training a Neural ODE system using the output of a Method-Of-Lines computation](https://discourse.julialang.org/t/training-a-neural-ode-system-using-the-output-of-a-method-of-lines-computation/129313)

<div class="topic-metadata">

**Author:** [@Rajat\_Shirali](https://discourse.julialang.org/u/Rajat_Shirali)\
**Replies:** 2\
**Last updated:** [May 26, 2025, 9:51am UTC](https://discourse.julialang.org/t/training-a-neural-ode-system-using-the-output-of-a-method-of-lines-computation/129313 "2025-05-26T09:51:53Z")

</div>

Hello everyone! I’m currently looking to develop my own customised Julia code for training Neural ODEs using a Physics-Informed approach. I want to create an interactive PDE-based model; however, there are two constraint…

---

## [Lack of conexionist NN package?](https://discourse.julialang.org/t/lack-of-conexionist-nn-package/129304)

<div class="topic-metadata">

**Author:** [@PepeRoConde](https://discourse.julialang.org/u/PepeRoConde)\
**Replies:** 1\
**Last updated:** [May 25, 2025, 9:48am UTC](https://discourse.julialang.org/t/lack-of-conexionist-nn-package/129304 "2025-05-25T09:48:58Z")

</div>

I found SOM.jl, NetworkDynamics.jl and Boltzmann.jl to be alone of other packages but not together between themselves. Wouldn’t it be nice a Conexionist Neural Networks pacake? or it be more appropiate to add those into …

---

## [Flux PINN 1D Burgers](https://discourse.julialang.org/t/flux-pinn-1d-burgers/93262)

<div class="topic-metadata">

**Author:** [@parf](https://discourse.julialang.org/u/parf)\
**Replies:** 11\
**Last updated:** [May 23, 2025, 4:54pm UTC](https://discourse.julialang.org/t/flux-pinn-1d-burgers/93262 "2025-05-23T16:54:32Z")

</div>

Hi! I am trying to construct PINN to solve 1D Burgers equation in Flux.jl without using NeuralPDE.jl. The velocity field u(t,x) is defined by the net\_u(\[t,x\]) and the batch for training consists of N\_u=32 points correspo…

---

## [Help training a NN that's used as a parameter in an ODE](https://discourse.julialang.org/t/help-training-a-nn-thats-used-as-a-parameter-in-an-ode/129231)

<div class="topic-metadata">

**Author:** [@alexl123](https://discourse.julialang.org/u/alexl123)\
**Replies:** 1\
**Last updated:** [May 22, 2025, 2:01am UTC](https://discourse.julialang.org/t/help-training-a-nn-thats-used-as-a-parameter-in-an-ode/129231 "2025-05-22T02:01:42Z")

</div>

Hello! I’ve been pulling my hair out trying to implement an ODE where one of the parameters is the output of a NN net. ChatGPT has me going in circles. Depending on what fixes I try I either get a: Error in epoch 1: Me…

---

## [Using a trained MLJ model for prediction on non-Table objects](https://discourse.julialang.org/t/using-a-trained-mlj-model-for-prediction-on-non-table-objects/129071)

<div class="topic-metadata">

**Author:** [@elipivo](https://discourse.julialang.org/u/elipivo)\
**Replies:** 1\
**Last updated:** [May 18, 2025, 9:42pm UTC](https://discourse.julialang.org/t/using-a-trained-mlj-model-for-prediction-on-non-table-objects/129071 "2025-05-18T21:42:09Z")

</div>

After training an MLJ model on data that meets the Tables.jl interface requirements, I’m interested in then deploying these models into an environment where predict is called on objects that don’t necessarily meet the Ta…

---

## [How to embed expert knowledge /constraints in ML training?](https://discourse.julialang.org/t/how-to-embed-expert-knowledge-constraints-in-ml-training/129025)

<div class="topic-metadata">

**Author:** [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Replies:** 11\
**Last updated:** [May 16, 2025, 2:34pm UTC](https://discourse.julialang.org/t/how-to-embed-expert-knowledge-constraints-in-ml-training/129025 "2025-05-16T14:34:00Z")

</div>

I want to train a model to predict the timber volume growth in a forest as a function of pedological/site variables and climatic data. So, I have something like this: ip\_point X Y vol vol\_growth sand\_pc soil\_depth p…

---

## [Are forward-mode derivatives calculated in parallel?](https://discourse.julialang.org/t/are-forward-mode-derivatives-calculated-in-parallel/128870)

<div class="topic-metadata">

**Author:** [@Bizzi](https://discourse.julialang.org/u/Bizzi)\
**Replies:** 1\
**Last updated:** [May 9, 2025, 10:57am UTC](https://discourse.julialang.org/t/are-forward-mode-derivatives-calculated-in-parallel/128870 "2025-05-09T10:57:47Z")

</div>

When I calculate the output of a function and a directional derivative using forward-mode autodiff (ForwardDiff, TaylorDiff etc), is the calculation of the derivatives carried out in parallel? By parallel I mean: assumi…

---

## [Error when running Code on GPU](https://discourse.julialang.org/t/error-when-running-code-on-gpu/128711)

<div class="topic-metadata">

**Author:** [@Efstein](https://discourse.julialang.org/u/Efstein)\
**Replies:** 0\
**Last updated:** [May 5, 2025, 8:47am UTC](https://discourse.julialang.org/t/error-when-running-code-on-gpu/128711 "2025-05-05T08:47:17Z")

</div>

Hi everybody! I am still quite new to Julia, so I hope I’m not missing something obvious. My goal is to solve a system of chemical equations with NeuralPDE. This runs fine, although very slow, on the CPU, but when I try…

---

## [Type-stable function with Flux Chain and Dense](https://discourse.julialang.org/t/type-stable-function-with-flux-chain-and-dense/128689)

<div class="topic-metadata">

**Author:** [@bienpierre](https://discourse.julialang.org/u/bienpierre)\
**Replies:** 3\
**Last updated:** [May 4, 2025, 9:00pm UTC](https://discourse.julialang.org/t/type-stable-function-with-flux-chain-and-dense/128689 "2025-05-04T21:00:22Z")

</div>

Dear All, I’m trying to make the following function containing Flux Chain and Dense type stable. using Flux @kwdef mutable struct Fnn{T \<: Function} neuron::Int = 5 layer::Int = 3 σ::T = Flux.relu end fu…

---

## [Error when a neural ode is implemented](https://discourse.julialang.org/t/error-when-a-neural-ode-is-implemented/119963)

<div class="topic-metadata">

**Author:** [@Ashima\_Kalathingal](https://discourse.julialang.org/u/Ashima_Kalathingal)\
**Replies:** 3\
**Last updated:** [May 2, 2025, 8:09pm UTC](https://discourse.julialang.org/t/error-when-a-neural-ode-is-implemented/119963 "2025-05-02T20:09:03Z")

</div>

I am trying to implement a neural ode in Julia. The loss function reduces when training. But after the training the parameters are not updated. It goes back to initial value. When I run the algorithm the following error…

---

## [Best libraries for experimenting with small scale models?](https://discourse.julialang.org/t/best-libraries-for-experimenting-with-small-scale-models/128106)

<div class="topic-metadata">

**Author:** [@stur86](https://discourse.julialang.org/u/stur86)\
**Replies:** 7\
**Last updated:** [April 25, 2025, 2:38am UTC](https://discourse.julialang.org/t/best-libraries-for-experimenting-with-small-scale-models/128106 "2025-04-25T02:38:27Z")

</div>

I’m looking for tools to do ML experiments in Julia. Specifically, I’d like trying to design some small-ish models (hundreds to thousands of parameters) and experiment with different approaches that aren’t just different…

---

## [Memory usage increasing with each epoch](https://discourse.julialang.org/t/memory-usage-increasing-with-each-epoch/121798)

<div class="topic-metadata">

**Author:** [@JoshuaBillson](https://discourse.julialang.org/u/JoshuaBillson)\
**Replies:** 18\
**Last updated:** [April 14, 2025, 12:30pm UTC](https://discourse.julialang.org/t/memory-usage-increasing-with-each-epoch/121798 "2025-04-14T12:30:43Z")

</div>

I’m having a problem where memory usage is gradually increasing with each epoch when training large neural networks with Flux (v0.14.22) and CUDA (v5.5.2). At the same time, training appears to get progressively slower a…

---

## [What is the status of ReinforcementLearning.jl](https://discourse.julialang.org/t/what-is-the-status-of-reinforcementlearning-jl/126326)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 24\
**Last updated:** [April 8, 2025, 2:25pm UTC](https://discourse.julialang.org/t/what-is-the-status-of-reinforcementlearning-jl/126326 "2025-04-08T14:25:04Z")

</div>

Hi All, I would like to ask, what is the status of ReinforcementLearning.jl? Specifically, if there is any active community around and it is therefore worth to learn it, use it, and fix the bugs. To an outside observer…

---

## [Custom Optimisers and Projections in Flux/Zygote: Is There a Canonical Way?](https://discourse.julialang.org/t/custom-optimisers-and-projections-in-flux-zygote-is-there-a-canonical-way/127749)

<div class="topic-metadata">

**Author:** [@anscoil](https://discourse.julialang.org/u/anscoil)\
**Replies:** 1\
**Last updated:** [April 5, 2025, 6:42pm UTC](https://discourse.julialang.org/t/custom-optimisers-and-projections-in-flux-zygote-is-there-a-canonical-way/127749 "2025-04-05T18:42:59Z")

</div>

Hi everyone, After spending quite some time solving inverse problems in optics with custom tools and manual algorithmic differentiation, I’ve recently been exploring more canonical approaches using modern ML frameworks. …

---

## [Automatic Differentiation for existing AMG libraries](https://discourse.julialang.org/t/automatic-differentiation-for-existing-amg-libraries/127722)

<div class="topic-metadata">

**Author:** [@CiREus](https://discourse.julialang.org/u/CiREus)\
**Replies:** 3\
**Last updated:** [April 4, 2025, 7:23pm UTC](https://discourse.julialang.org/t/automatic-differentiation-for-existing-amg-libraries/127722 "2025-04-04T19:23:26Z")

</div>

Hi all, I recently joined Julia for my MSc thesis and was having some doubts about the Automatic Differentiation tools (Enzyme, Zygote, JuliaDiff, etc.). In short, I wanted to use existing linear algebra libraries cont…

---

## [Run the Gemma 2 9B LLM in Julia and without a GPU?](https://discourse.julialang.org/t/run-the-gemma-2-9b-llm-in-julia-and-without-a-gpu/124155)

<div class="topic-metadata">

**Author:** [@nsajko](https://discourse.julialang.org/u/nsajko)\
**Replies:** 5\
**Last updated:** [April 2, 2025, 11:01am UTC](https://discourse.julialang.org/t/run-the-gemma-2-9b-llm-in-julia-and-without-a-gpu/124155 "2025-04-02T11:01:37Z")

</div>

I’m ignorant about LLMs, machine learning, etc.; but I want to run the Gemma 2 9B large language model. I’m not interested in using it as a chat bot, rather the reason I want to use it is the fact that it defines the eva…

---

## [Reactant cannot find GPU device](https://discourse.julialang.org/t/reactant-cannot-find-gpu-device/127124)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 2\
**Last updated:** [March 19, 2025, 8:09pm UTC](https://discourse.julialang.org/t/reactant-cannot-find-gpu-device/127124 "2025-03-19T20:09:02Z")

</div>

I have an old server with two GPUs (TITAN Xp, and RTX 2080). I am using that mostly for developing an playing. Now, I wanted to compare the speed of Lux with CUDA and with Reactant, but the reactant library cannot find …

---

## [How can large language models work with julia?](https://discourse.julialang.org/t/how-can-large-language-models-work-with-julia/126976)

<div class="topic-metadata">

**Author:** [@WuSiren](https://discourse.julialang.org/u/WuSiren)\
**Replies:** 4\
**Last updated:** [March 18, 2025, 7:19am UTC](https://discourse.julialang.org/t/how-can-large-language-models-work-with-julia/126976 "2025-03-18T07:19:56Z")

</div>

Large language models such as ChatGPT, DeepSeek, etc. are all the rage these days. How can we use them in the world of Julia? What kind of fun things can we do? What is the state of the art relevant to the usage of large…

---

## [Best way to prevent AD to differentiate through useless or zero gradient function](https://discourse.julialang.org/t/best-way-to-prevent-ad-to-differentiate-through-useless-or-zero-gradient-function/126927)

<div class="topic-metadata">

**Author:** [@FerreolS](https://discourse.julialang.org/u/FerreolS)\
**Replies:** 2\
**Last updated:** [March 13, 2025, 9:05pm UTC](https://discourse.julialang.org/t/best-way-to-prevent-ad-to-differentiate-through-useless-or-zero-gradient-function/126927 "2025-03-13T21:05:03Z")

</div>

I want to minimize a function of the type: x = \\textrm{argmin}\_x \\min\_a || f(x) a - y ||^2 as it is quadratic in a we know the optimum in a is a(x) = (f(x)^T f(x))^{-1} f(x)^T y so we minimize x = \\text{argmin}\_x …

---

## [Lux + Enzyme and Zygote + NeuralODE, segmentation fault](https://discourse.julialang.org/t/lux-enzyme-and-zygote-neuralode-segmentation-fault/123645)

<div class="topic-metadata">

**Author:** [@vleon1234](https://discourse.julialang.org/u/vleon1234)\
**Replies:** 6\
**Last updated:** [March 3, 2025, 10:26pm UTC](https://discourse.julialang.org/t/lux-enzyme-and-zygote-neuralode-segmentation-fault/123645 "2025-03-03T22:26:40Z")

</div>

Hi all, I’m using Julia v1.11.2, and installed the used packages on 12/6 (last Friday), so I assume I have the most recent package versions. I’ve been following along the SciML tutorials for Neural ordinary differentia…

---

## [GPU backend-agnostic way to create efficiently random number on the GPU](https://discourse.julialang.org/t/gpu-backend-agnostic-way-to-create-efficiently-random-number-on-the-gpu/126110)

<div class="topic-metadata">

**Author:** [@Alexander-Barth](https://discourse.julialang.org/u/Alexander-Barth)\
**Replies:** 3\
**Last updated:** [March 3, 2025, 7:44am UTC](https://discourse.julialang.org/t/gpu-backend-agnostic-way-to-create-efficiently-random-number-on-the-gpu/126110 "2025-03-03T07:44:34Z")

</div>

I am trying to make my code Flux code portable across CUDA and AMDGPU. I would like to know if there is a GPU backend-agnostic way to create efficiently random number on the GPU? AMDGPU.randn is by far the fastest, but I…

---

## [Resetting some states in a batch in Lux.jl recurrent layers](https://discourse.julialang.org/t/resetting-some-states-in-a-batch-in-lux-jl-recurrent-layers/126402)

<div class="topic-metadata">

**Author:** [@emiwar](https://discourse.julialang.org/u/emiwar)\
**Replies:** 0\
**Last updated:** [February 27, 2025, 11:47pm UTC](https://discourse.julialang.org/t/resetting-some-states-in-a-batch-in-lux-jl-recurrent-layers/126402 "2025-02-27T23:47:52Z")

</div>

Hi, I’m writing a Deep RL (PPO) implementation using Lux.jl and want to add support for recurrent networks. My issue is that I have a batch of environments and at every time step some of them may reset due to the agent …

---

## [Implementing a ConvNet that doesn't allocate during inference (SimpleChains.jl?)](https://discourse.julialang.org/t/implementing-a-convnet-that-doesnt-allocate-during-inference-simplechains-jl/126365)

<div class="topic-metadata">

**Author:** [@sefffal](https://discourse.julialang.org/u/sefffal)\
**Replies:** 3\
**Last updated:** [February 27, 2025, 5:24pm UTC](https://discourse.julialang.org/t/implementing-a-convnet-that-doesnt-allocate-during-inference-simplechains-jl/126365 "2025-02-27T17:24:57Z")

</div>

Hi all, Myself and a student are trying to implement a simple, relatively small convolutional neural network. The trick is, we need it to run as part of our existing real time controller (implemented in Julia, see Case …

---

## [Help using cuDNN in Julia](https://discourse.julialang.org/t/help-using-cudnn-in-julia/126232)

<div class="topic-metadata">

**Author:** [@notGiGi](https://discourse.julialang.org/u/notGiGi)\
**Replies:** 1\
**Last updated:** [February 25, 2025, 5:02pm UTC](https://discourse.julialang.org/t/help-using-cudnn-in-julia/126232 "2025-02-25T17:02:52Z")

</div>

Hello, everyone. I’m trying to use cuDNN in Julia, but so far it’s not working as expected. I’m very confused. It is supposed to be included in CUDA.jl so that I can use using CUDA.CUDNN, but when I do that, I get the er…

---

## [Training Universal PINNs using NeuralPDE?](https://discourse.julialang.org/t/training-universal-pinns-using-neuralpde/125948)

<div class="topic-metadata">

**Author:** [@John\_Nardini](https://discourse.julialang.org/u/John_Nardini)\
**Replies:** 4\
**Last updated:** [February 18, 2025, 2:34pm UTC](https://discourse.julialang.org/t/training-universal-pinns-using-neuralpde/125948 "2025-02-18T14:34:44Z")

</div>

I am interested Universal PINNs (UPINNs) to describe some data. UPINNs are similar to a PINN, but we assume that the underlying DE model has known and unknown dynamics: \\dfrac{d\\vec{u}}{dt} = F\_{known}(\\vec{u}) + F\_{unk…

---

## [Distributed Data Parallel training with 2 GPUs fails with Flux.jl on AMD GPUs](https://discourse.julialang.org/t/distributed-data-parallel-training-with-2-gpus-fails-with-flux-jl-on-amd-gpus/125993)

<div class="topic-metadata">

**Author:** [@Alexander-Barth](https://discourse.julialang.org/u/Alexander-Barth)\
**Replies:** 7\
**Last updated:** [February 18, 2025, 8:41am UTC](https://discourse.julialang.org/t/distributed-data-parallel-training-with-2-gpus-fails-with-flux-jl-on-amd-gpus/125993 "2025-02-18T08:41:29Z")

</div>

I am using Flux (with this PR Fix missing imports in FluxMPIExt by Alexander-Barth · Pull Request #2589 · FluxML/Flux.jl · GitHub to fix some MPI related imports) on two AMD GPUs with Distributed Data Parallel (DDP). Ho…

[Previous page](https://discourse.julialang.org/c/domain/ml/24.md?page=2)

[Next page](https://discourse.julialang.org/c/domain/ml/24.md?page=4)
