# Machine Learning

**URL:** https://discourse.julialang.org/c/domain/ml/24.md?page=12

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 13

---

## [Flux new explicit API not work but old implicit API works for a simple RNN](https://discourse.julialang.org/t/flux-new-explicit-api-not-work-but-old-implicit-api-works-for-a-simple-rnn/104262)

<div class="topic-metadata">

**Author:** [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Replies:** 0\
**Last updated:** [September 26, 2023, 1:36pm UTC](https://discourse.julialang.org/t/flux-new-explicit-api-not-work-but-old-implicit-api-works-for-a-simple-rnn/104262 "2023-09-26T13:36:37Z")

</div>

I am trying to reproduce the tutorial A Basic RNN using Flux.jl v0.14.6. Using the old Flux API as in the tutorial, the model can be successfully trained. The code is using Flux num\_samples = 1000 num\_epochs = 50 func…

---

## [Looking to prevent overfitting while using DataDrivenDiffEq.jl to estimate the equation for a time dependent diff eq](https://discourse.julialang.org/t/looking-to-prevent-overfitting-while-using-datadrivendiffeq-jl-to-estimate-the-equation-for-a-time-dependent-diff-eq/103931)

<div class="topic-metadata">

**Author:** [@mjc715](https://discourse.julialang.org/u/mjc715)\
**Replies:** 7\
**Last updated:** [September 25, 2023, 12:07pm UTC](https://discourse.julialang.org/t/looking-to-prevent-overfitting-while-using-datadrivendiffeq-jl-to-estimate-the-equation-for-a-time-dependent-diff-eq/103931 "2023-09-25T12:07:18Z")

</div>

In this code, we are looking to see if we can use DataDrivenDiffEq.jl to accurate predict an equation for a time dependent differential equation. To test this, we have been creating an equation, adding noise to the solut…

---

## [Inconsistent matrix multiply output from Flux.Dense depending on shape of input](https://discourse.julialang.org/t/inconsistent-matrix-multiply-output-from-flux-dense-depending-on-shape-of-input/104116)

<div class="topic-metadata">

**Author:** [@JuliaMLTools](https://discourse.julialang.org/u/JuliaMLTools)\
**Replies:** 11\
**Last updated:** [September 22, 2023, 3:05pm UTC](https://discourse.julialang.org/t/inconsistent-matrix-multiply-output-from-flux-dense-depending-on-shape-of-input/104116 "2023-09-22T15:05:15Z")

</div>

Anyone know a workaround for this inconsistency in matrix multiply?

---

## [Parameter optimization using Multiple shooting](https://discourse.julialang.org/t/parameter-optimization-using-multiple-shooting/104104)

<div class="topic-metadata">

**Author:** [@Marco\_Nesta](https://discourse.julialang.org/u/Marco_Nesta)\
**Replies:** 1\
**Last updated:** [September 21, 2023, 12:18pm UTC](https://discourse.julialang.org/t/parameter-optimization-using-multiple-shooting/104104 "2023-09-21T12:18:19Z")

</div>

'm using the same structure as in the link, but how can I view the updated parameters? this is my code: using DifferentialEquations, Plots, Flux,Optim, DiffEqFlux, DataInterpolations,Random, ComponentArrays, Lux …

---

## [Struggling to load a pre-trained PyTorch model into Julia](https://discourse.julialang.org/t/struggling-to-load-a-pre-trained-pytorch-model-into-julia/104068)

<div class="topic-metadata">

**Author:** [@Jian\_ZUO](https://discourse.julialang.org/u/Jian_ZUO)\
**Replies:** 2\
**Last updated:** [September 20, 2023, 7:34pm UTC](https://discourse.julialang.org/t/struggling-to-load-a-pre-trained-pytorch-model-into-julia/104068 "2023-09-20T19:34:24Z")

</div>

Hi all, I am coming back to this unsolved problem. The background is that I have trained a reinforcement learning agent and I want to load the pre-trained model (pytorch) into Julia for inference (it is much faster to …

---

## [Saving and loading architectures with multiple blocks](https://discourse.julialang.org/t/saving-and-loading-architectures-with-multiple-blocks/104030)

<div class="topic-metadata">

**Author:** [@tom-plaa](https://discourse.julialang.org/u/tom-plaa)\
**Replies:** 3\
**Last updated:** [September 20, 2023, 3:49pm UTC](https://discourse.julialang.org/t/saving-and-loading-architectures-with-multiple-blocks/104030 "2023-09-20T15:49:25Z")

</div>

Hi! I’m working on a project where another person has been trying out different architectures in Flux.jl for a given problem, and the one that turned out to be the best one is giving us problems when saving and loading …

---

## [Support for Intel Mac AMD GPUs](https://discourse.julialang.org/t/support-for-intel-mac-amd-gpus/103933)

<div class="topic-metadata">

**Author:** [@keynescoefen](https://discourse.julialang.org/u/keynescoefen)\
**Replies:** 4\
**Last updated:** [September 18, 2023, 12:09pm UTC](https://discourse.julialang.org/t/support-for-intel-mac-amd-gpus/103933 "2023-09-18T12:09:00Z")

</div>

Hi, I saw this topic here and wondered whether there were any updates: ? It says that only AMD architectures on Linux are supported. Is there any hope of changing that in the future or is it a lost cause? The thing i…

---

## [How to solve 2D PDE for heat conduction with x and y having different length (A rectangle)](https://discourse.julialang.org/t/how-to-solve-2d-pde-for-heat-conduction-with-x-and-y-having-different-length-a-rectangle/103907)

<div class="topic-metadata">

**Author:** [@ShadmanSakief](https://discourse.julialang.org/u/ShadmanSakief)\
**Replies:** 1\
**Last updated:** [September 15, 2023, 5:17pm UTC](https://discourse.julialang.org/t/how-to-solve-2d-pde-for-heat-conduction-with-x-and-y-having-different-length-a-rectangle/103907 "2023-09-15T17:17:05Z")

</div>

using NeuralPDE, Lux, Optimization, OptimizationOptimJL import ModelingToolkit: Interval @parameters x y @variables u(..) Dxx = Differential(x)^2 Dyy = Differential(y)^2 Dx = Differential(x) Dy = Differential(y) 2…

---

## [LLama2-7b difference in inference when between Float16 and Float32](https://discourse.julialang.org/t/llama2-7b-difference-in-inference-when-between-float16-and-float32/103826)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 1\
**Last updated:** [September 13, 2023, 10:58pm UTC](https://discourse.julialang.org/t/llama2-7b-difference-in-inference-when-between-float16-and-float32/103826 "2023-09-13T22:58:11Z")

</div>

Hello, I want to present the results of one more tests I wanted to do. I have noticed that llama2-7b gives weird (read incorrect) results when used with Float16. I have set-up a small test demostrating the behavior to n…

---

## [Gradient of llama2 computed by Zygote seems to be incorrect](https://discourse.julialang.org/t/gradient-of-llama2-computed-by-zygote-seems-to-be-incorrect/103810)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 6\
**Last updated:** [September 13, 2023, 7:35pm UTC](https://discourse.julialang.org/t/gradient-of-llama2-computed-by-zygote-seems-to-be-incorrect/103810 "2023-09-13T19:35:49Z")

</div>

Hi, I have tested gradient of llama2-7b model with respect its input and it gives me wrong results (tested only on GPU). I have created a MWE as follows using ProfileSummarizer using Transformers using Flux using TextE…

---

## [Do we have differentiable sort in Julia?](https://discourse.julialang.org/t/do-we-have-differentiable-sort-in-julia/102713)

<div class="topic-metadata">

**Author:** [@pat-alt](https://discourse.julialang.org/u/pat-alt)\
**Replies:** 14\
**Last updated:** [September 12, 2023, 11:47am UTC](https://discourse.julialang.org/t/do-we-have-differentiable-sort-in-julia/102713 "2023-09-12T11:47:04Z")

</div>

Python has support for fast differentiable sorting which is an implementation of Blondel et al. (2020). I don’t think we have this in Julia at the moment? I have specific use case for ConformalPrediction.jl, but beyond …

---

## [Custom train function in SimpleChains.jl](https://discourse.julialang.org/t/custom-train-function-in-simplechains-jl/103663)

<div class="topic-metadata">

**Author:** [@Luca\_Pepe\_Sciarria](https://discourse.julialang.org/u/Luca_Pepe_Sciarria)\
**Replies:** 0\
**Last updated:** [September 8, 2023, 12:30pm UTC](https://discourse.julialang.org/t/custom-train-function-in-simplechains-jl/103663 "2023-09-08T12:30:31Z")

</div>

Hi SimpleChains team! I’m currently using SimpleChains and I want to implement my own train function based on the available template in Flux: for d in training\_set # Our super logic gs = gradient(params(m)) do …

---

## [Predicting missing physics term for discrete time dynamical system](https://discourse.julialang.org/t/predicting-missing-physics-term-for-discrete-time-dynamical-system/103595)

<div class="topic-metadata">

**Author:** [@kunal](https://discourse.julialang.org/u/kunal)\
**Replies:** 3\
**Last updated:** [September 8, 2023, 4:18am UTC](https://discourse.julialang.org/t/predicting-missing-physics-term-for-discrete-time-dynamical-system/103595 "2023-09-08T04:18:24Z")

</div>

My use case is very much similar to Using DiffEqFlux.jl for system identification with partially known system dynamics . However my objective is to utilize this in case of a discrete time space system. I was following t…

---

## [Second order derivatives with ChainRules](https://discourse.julialang.org/t/second-order-derivatives-with-chainrules/103606)

<div class="topic-metadata">

**Author:** [@mariusd](https://discourse.julialang.org/u/mariusd)\
**Replies:** 1\
**Last updated:** [September 7, 2023, 1:48pm UTC](https://discourse.julialang.org/t/second-order-derivatives-with-chainrules/103606 "2023-09-07T13:48:54Z")

</div>

How can I get the second order derivarive of a function using ChainRules? Bellow is my approach for the sin function. See the following code: using ChainRules using ChainRulesCore # from the docs, first derivative of …

---

## [Custom Flux Layer Connect](https://discourse.julialang.org/t/custom-flux-layer-connect/36262)

<div class="topic-metadata">

**Author:** [@TobiS](https://discourse.julialang.org/u/TobiS)\
**Replies:** 3\
**Last updated:** [September 6, 2023, 7:56am UTC](https://discourse.julialang.org/t/custom-flux-layer-connect/36262 "2023-09-06T07:56:19Z")

</div>

Hello everyone, I would like to implement the architecture of N-BEATS network (https://arxiv.org/pdf/1905.10437.pdf) in julia Flux. A simple example of a part of the architecture is shown in the following figure. Is…

---

## [Any examples applying Physics Informed Neural Networks to ODEs?](https://discourse.julialang.org/t/any-examples-applying-physics-informed-neural-networks-to-odes/103512)

<div class="topic-metadata">

**Author:** [@affans](https://discourse.julialang.org/u/affans)\
**Replies:** 1\
**Last updated:** [September 5, 2023, 2:04pm UTC](https://discourse.julialang.org/t/any-examples-applying-physics-informed-neural-networks-to-odes/103512 "2023-09-05T14:04:25Z")

</div>

I recently have been involved in PINNs and their applications. While I understand the mathematics of NNs, ODE solvers, PDEs, etc… I have very little experience in implementing them and in Julia especially. Does anyone ha…

---

## [Modeling multimodal data in Flux](https://discourse.julialang.org/t/modeling-multimodal-data-in-flux/103257)

<div class="topic-metadata">

**Author:** [@usiam](https://discourse.julialang.org/u/usiam)\
**Replies:** 3\
**Last updated:** [August 31, 2023, 12:34am UTC](https://discourse.julialang.org/t/modeling-multimodal-data-in-flux/103257 "2023-08-31T00:34:14Z")

</div>

Hey. I am trying to model some multimodal time-series data. The data has 3 “components” - a time-series component (shape = \[timesteps, channels, batchsize\]), and 2 separate flattened components (shape = \[numfeatures, bat…

---

## [DomainError with Loss is Inf on data item 1, stopping training:](https://discourse.julialang.org/t/domainerror-with-loss-is-inf-on-data-item-1-stopping-training/103343)

<div class="topic-metadata">

**Author:** [@chadagreene](https://discourse.julialang.org/u/chadagreene)\
**Replies:** 1\
**Last updated:** [August 30, 2023, 11:22am UTC](https://discourse.julialang.org/t/domainerror-with-loss-is-inf-on-data-item-1-stopping-training/103343 "2023-08-30T11:22:19Z")

</div>

I’m trying to learn Julia and ML principles simultaneously. I’ve successfully trained a model to fit a simple straight line, following the procedure described in the intro docs, but when I try to go a step further and ad…

---

## [Not using constants in trees in SymbolicRegression.jl](https://discourse.julialang.org/t/not-using-constants-in-trees-in-symbolicregression-jl/103231)

<div class="topic-metadata">

**Author:** [@VPBML](https://discourse.julialang.org/u/VPBML)\
**Replies:** 0\
**Last updated:** [August 26, 2023, 7:55pm UTC](https://discourse.julialang.org/t/not-using-constants-in-trees-in-symbolicregression-jl/103231 "2023-08-26T19:55:25Z")

</div>

Hi, I wonder how can we tell SymbolicRegression.jl to not use any constants (e.g., 2.0, 15.6) when creating its trees, only the variables. I am using SRRegressor via MLJ interface. Thank you. Regards

---

## [Running Machine Learning Experiments with Flux](https://discourse.julialang.org/t/running-machine-learning-experiments-with-flux/102661)

<div class="topic-metadata">

**Author:** [@josemanuel22](https://discourse.julialang.org/u/josemanuel22)\
**Replies:** 5\
**Last updated:** [August 24, 2023, 9:55am UTC](https://discourse.julialang.org/t/running-machine-learning-experiments-with-flux/102661 "2023-08-24T09:55:11Z")

</div>

I’m running some ML experiments to draw certain conclusions about whether my new method works as expected. I would like, so to speak, to run multiple experiments in the same file, as it would otherwise be cumbersome. The…

---

## [Compute gradients in neuralODE with Zygote](https://discourse.julialang.org/t/compute-gradients-in-neuralode-with-zygote/103106)

<div class="topic-metadata">

**Author:** [@tommy\_J](https://discourse.julialang.org/u/tommy_J)\
**Replies:** 3\
**Last updated:** [August 24, 2023, 7:50am UTC](https://discourse.julialang.org/t/compute-gradients-in-neuralode-with-zygote/103106 "2023-08-24T07:50:41Z")

</div>

Hello everyone, I would like to know how to check the gradients for a single batch in this example. After the computation of l1 I did: batch1 = train\_loader.data\[1\] time1 = train\_loader.data\[2\] loss, back = Zygote.…

---

## [Taking the gradient of an einsum](https://discourse.julialang.org/t/taking-the-gradient-of-an-einsum/103003)

<div class="topic-metadata">

**Author:** [@yuvalw](https://discourse.julialang.org/u/yuvalw)\
**Replies:** 4\
**Last updated:** [August 21, 2023, 8:55am UTC](https://discourse.julialang.org/t/taking-the-gradient-of-an-einsum/103003 "2023-08-21T08:55:21Z")

</div>

Hello, I would really appreciate some help optimizing my optimization. What I calculate makes the most sense (to me) as an einsum but I cannot seem to take its gradient properly. What I managed to do instead is iterate o…

---

## [Generic way to change float precision in FluxML](https://discourse.julialang.org/t/generic-way-to-change-float-precision-in-fluxml/102893)

<div class="topic-metadata">

**Author:** [@pat-alt](https://discourse.julialang.org/u/pat-alt)\
**Replies:** 3\
**Last updated:** [August 18, 2023, 6:26am UTC](https://discourse.julialang.org/t/generic-way-to-change-float-precision-in-fluxml/102893 "2023-08-18T06:26:44Z")

</div>

I’ve recently started using FastAI.jl and firstly want to say how much I appreciate the efforts of anyone involved here - this is really cool! When going through chapter 5 of the fast.ai book, I just came across the poi…

---

## [Mix-Precision training using Models of Metalhead (Flux.jl) or Boltz (Lux.jl)](https://discourse.julialang.org/t/mix-precision-training-using-models-of-metalhead-flux-jl-or-boltz-lux-jl/102857)

<div class="topic-metadata">

**Author:** [@lgmendesDynamics](https://discourse.julialang.org/u/lgmendesDynamics)\
**Replies:** 3\
**Last updated:** [August 16, 2023, 11:30pm UTC](https://discourse.julialang.org/t/mix-precision-training-using-models-of-metalhead-flux-jl-or-boltz-lux-jl/102857 "2023-08-16T23:30:20Z")

</div>

It is possible to train Deep Learning models for computer vision available in packages like Metalhead (Flux) or Boltz (Lux) with mix-precision?

---

## [ConfusionMatrix display in VSCode when there are ~ 16 categories](https://discourse.julialang.org/t/confusionmatrix-display-in-vscode-when-there-are-16-categories/102763)

<div class="topic-metadata">

**Author:** [@roeya](https://discourse.julialang.org/u/roeya)\
**Replies:** 3\
**Last updated:** [August 16, 2023, 9:34pm UTC](https://discourse.julialang.org/t/confusionmatrix-display-in-vscode-when-there-are-16-categories/102763 "2023-08-16T21:34:16Z")

</div>

When trying to display small confusion matrix for example for 4 categories classifier I get: ┌───────────────────────────────────────────────────────┐ │ Ground Truth …

---

## [Load a pre-trained Pytorch model in Julia](https://discourse.julialang.org/t/load-a-pre-trained-pytorch-model-in-julia/96932)

<div class="topic-metadata">

**Author:** [@Jian\_ZUO](https://discourse.julialang.org/u/Jian_ZUO)\
**Replies:** 5\
**Last updated:** [August 16, 2023, 3:30pm UTC](https://discourse.julialang.org/t/load-a-pre-trained-pytorch-model-in-julia/96932 "2023-08-16T15:30:15Z")

</div>

Hi all, What are the best practices to load a pre-trained pytorch model in Julia (e.g. Flux.jl)? Thank you!

---

## [How efficient is the automatic differentation and optimization in comparison to PyTorch](https://discourse.julialang.org/t/how-efficient-is-the-automatic-differentation-and-optimization-in-comparison-to-pytorch/102824)

<div class="topic-metadata">

**Author:** [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Replies:** 6\
**Last updated:** [August 16, 2023, 3:24pm UTC](https://discourse.julialang.org/t/how-efficient-is-the-automatic-differentation-and-optimization-in-comparison-to-pytorch/102824 "2023-08-16T15:24:16Z")

</div>

Hi all, how efficient is the current Julia AD and optimization landscape in comparsion to let’s say PyTorch? Python For example in PyTorch I can do arr1 = arr1.requires\_grad\_(True) arr2 = arr2.requires\_grad\_(True) opt…

---

## [Individual ctivation function for each network output](https://discourse.julialang.org/t/individual-ctivation-function-for-each-network-output/102676)

<div class="topic-metadata">

**Author:** [@tommy\_J](https://discourse.julialang.org/u/tommy_J)\
**Replies:** 1\
**Last updated:** [August 11, 2023, 8:38pm UTC](https://discourse.julialang.org/t/individual-ctivation-function-for-each-network-output/102676 "2023-08-11T20:38:19Z")

</div>

Hi, Can I use a different activation function for each output of my network? using Lux net = Chain(Dense(3, 10, tanh), Dense(10, 3, tanh)) I know that output 2 can only be positive, but 1 and 3 can be negative. Is i…

---

## [Flux model on GPU?](https://discourse.julialang.org/t/flux-model-on-gpu/102595)

<div class="topic-metadata">

**Author:** [@johnbb](https://discourse.julialang.org/u/johnbb)\
**Replies:** 4\
**Last updated:** [August 9, 2023, 12:29pm UTC](https://discourse.julialang.org/t/flux-model-on-gpu/102595 "2023-08-09T12:29:49Z")

</div>

How can I determine whether a Flux model is on a GPU (of any kind)? For example, if m = Dense(1, 1) |\> gpu/cpu should it somehow be based on typeof(m) or are there other alternatives? My use case is that I have a (traini…

---

## [Is there a C/C++/Rust/Go library that can do inference on Flux models?](https://discourse.julialang.org/t/is-there-a-c-c-rust-go-library-that-can-do-inference-on-flux-models/102581)

<div class="topic-metadata">

**Author:** [@tchebycheff](https://discourse.julialang.org/u/tchebycheff)\
**Replies:** 11\
**Last updated:** [August 9, 2023, 12:08pm UTC](https://discourse.julialang.org/t/is-there-a-c-c-rust-go-library-that-can-do-inference-on-flux-models/102581 "2023-08-09T12:08:39Z")

</div>

If not, is it useful to have one at least for simple layers like Dense, Conv, Scale, etc.?

[Previous page](https://discourse.julialang.org/c/domain/ml/24.md?page=11)

[Next page](https://discourse.julialang.org/c/domain/ml/24.md?page=13)
