# Machine Learning

**URL:** https://discourse.julialang.org/c/domain/ml/24.md?page=17

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 18

---

## [Saving a model built with Transformers.jl](https://discourse.julialang.org/t/saving-a-model-built-with-transformers-jl/96132)

<div class="topic-metadata">

**Author:** [@gemuesi](https://discourse.julialang.org/u/gemuesi)\
**Replies:** 6\
**Last updated:** [March 15, 2023, 1:48pm UTC](https://discourse.julialang.org/t/saving-a-model-built-with-transformers-jl/96132 "2023-03-15T13:48:28Z")

</div>

Hi there, I am pretty new to anything related to this topic in general, but I found Flux to be quite accessible and wanted to have a look at the Transformers.jl package as well. I’ve been playing around with the copy ta…

---

## [Lux initialization](https://discourse.julialang.org/t/lux-initialization/96049)

<div class="topic-metadata">

**Author:** [@Stefano\_Giampiccolo](https://discourse.julialang.org/u/Stefano_Giampiccolo)\
**Replies:** 2\
**Last updated:** [March 14, 2023, 12:22pm UTC](https://discourse.julialang.org/t/lux-initialization/96049 "2023-03-14T12:22:59Z")

</div>

Hello everyone, I apologize for the probably naive question, but if I wanted to initialize the layer parameters with glorot\_uniform in Lux with a gain smaller than the default, how do I then obtain the parameters and sta…

---

## [Flux Transformer Out of Memory](https://discourse.julialang.org/t/flux-transformer-out-of-memory/95922)

<div class="topic-metadata">

**Author:** [@Gadersd](https://discourse.julialang.org/u/Gadersd)\
**Replies:** 25\
**Last updated:** [March 13, 2023, 2:36pm UTC](https://discourse.julialang.org/t/flux-transformer-out-of-memory/95922 "2023-03-13T14:36:13Z")

</div>

I started off with PyTorch, but after hearing about Julia Flux I decided to try it out to reap the promised elegance and ease of use. I ported my transformer implementation to Julia and verified that the outputs match th…

---

## [Training a neural ode with unknow time span](https://discourse.julialang.org/t/training-a-neural-ode-with-unknow-time-span/95992)

<div class="topic-metadata">

**Author:** [@kaido975](https://discourse.julialang.org/u/kaido975)\
**Replies:** 4\
**Last updated:** [March 13, 2023, 2:48pm UTC](https://discourse.julialang.org/t/training-a-neural-ode-with-unknow-time-span/95992 "2023-03-13T14:48:03Z")

</div>

I am trying to train a neural ode with a start time = 0 and an unknown end time. Here is the working code. using Lux, DiffEqFlux, DifferentialEquations, Optimization, OptimizationOptimJL, Random, Plots, ComponentArrays…

---

## [Flux with Matrix input dimensions](https://discourse.julialang.org/t/flux-with-matrix-input-dimensions/43610)

<div class="topic-metadata">

**Author:** [@Ruan\_Stander](https://discourse.julialang.org/u/Ruan_Stander)\
**Replies:** 2\
**Last updated:** [March 12, 2023, 6:30pm UTC](https://discourse.julialang.org/t/flux-with-matrix-input-dimensions/43610 "2023-03-12T18:30:15Z")

</div>

Hi Just started using flux, cant get on top of input/layer dimensions where input is a matrix Adapting https://github.com/FluxML/model-zoo/blob/master/vision/mnist/mlp.jl to my input space but its not working My Input…

---

## [What happened to XLA.jl](https://discourse.julialang.org/t/what-happened-to-xla-jl/88088)

<div class="topic-metadata">

**Author:** [@Honza9723](https://discourse.julialang.org/u/Honza9723)\
**Replies:** 16\
**Last updated:** [March 12, 2023, 3:24pm UTC](https://discourse.julialang.org/t/what-happened-to-xla-jl/88088 "2023-03-12T15:24:22Z")

</div>

Dear All, I would like to ask, what happened to XLA.jl. When I checked the repository, it says it was archived. Is the project of building the Julia-to-XLA interface dead? Best, Jan

---

## [Holomorphic function not being recognized as such by ChainRules.jl](https://discourse.julialang.org/t/holomorphic-function-not-being-recognized-as-such-by-chainrules-jl/95806)

<div class="topic-metadata">

**Author:** [@Hareruya](https://discourse.julialang.org/u/Hareruya)\
**Replies:** 2\
**Last updated:** [March 9, 2023, 5:11pm UTC](https://discourse.julialang.org/t/holomorphic-function-not-being-recognized-as-such-by-chainrules-jl/95806 "2023-03-09T17:11:56Z")

</div>

Greetings. I am currently trying to implement some code using @scalar\_rule for a custom gradient for a very important group of functions in Physics and Engineering, the spherical bessel functions, in this case particula…

---

## [How do I a logistic regression without the intercept in GLM.jl?](https://discourse.julialang.org/t/how-do-i-a-logistic-regression-without-the-intercept-in-glm-jl/95720)

<div class="topic-metadata">

**Author:** [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Replies:** 4\
**Last updated:** [March 8, 2023, 9:36am UTC](https://discourse.julialang.org/t/how-do-i-a-logistic-regression-without-the-intercept-in-glm-jl/95720 "2023-03-08T09:36:27Z")

</div>

I can’t figure out how to fit a model without the intercept The documentatioin just says the intercept is fitted by default but doesn’t show how too do it without the intercept? is that even possible?

---

## [\[ForwardDiff\] how to consider expression as constant](https://discourse.julialang.org/t/forwarddiff-how-to-consider-expression-as-constant/95651)

<div class="topic-metadata">

**Author:** [@mariusd](https://discourse.julialang.org/u/mariusd)\
**Replies:** 5\
**Last updated:** [March 7, 2023, 12:41pm UTC](https://discourse.julialang.org/t/forwarddiff-how-to-consider-expression-as-constant/95651 "2023-03-07T12:41:49Z")

</div>

See the following code. Inside the function f, g(x) is called and I need the result to be used as both a variable and as constant. Now of course I can calculate x20=g(x) outside and provide this as an extra argument to …

---

## [Problem with training a spiking neural network with Flux](https://discourse.julialang.org/t/problem-with-training-a-spiking-neural-network-with-flux/95149)

<div class="topic-metadata">

**Author:** [@Daniele\_Linaro](https://discourse.julialang.org/u/Daniele_Linaro)\
**Replies:** 7\
**Last updated:** [March 6, 2023, 4:30pm UTC](https://discourse.julialang.org/t/problem-with-training-a-spiking-neural-network-with-flux/95149 "2023-03-06T16:30:53Z")

</div>

Hello, I am trying to use Flux to train a simple spiking neural network to perform input classification. The network has 100 inputs that represent incoming spike streams: half of the inputs are “active” (i.e., they rece…

---

## [\[Zygote\] Derivative of derivative compilation time](https://discourse.julialang.org/t/zygote-derivative-of-derivative-compilation-time/95537)

<div class="topic-metadata">

**Author:** [@mariusd](https://discourse.julialang.org/u/mariusd)\
**Replies:** 4\
**Last updated:** [March 4, 2023, 9:57pm UTC](https://discourse.julialang.org/t/zygote-derivative-of-derivative-compilation-time/95537 "2023-03-04T21:57:53Z")

</div>

In the following code I calculate derivative fx, then derivative of derivative fxy, and then fxyz. The code works and gives the correct result, however the compilation time is huge for the fxyz derivative. Is there a wa…

---

## [\[ReverseDiff.jl\] How to do nested differentiation?](https://discourse.julialang.org/t/reversediff-jl-how-to-do-nested-differentiation/95542)

<div class="topic-metadata">

**Author:** [@mariusd](https://discourse.julialang.org/u/mariusd)\
**Replies:** 0\
**Last updated:** [March 4, 2023, 2:00pm UTC](https://discourse.julialang.org/t/reversediff-jl-how-to-do-nested-differentiation/95542 "2023-03-04T14:00:12Z")

</div>

See the following example. I do calculate the derivative of f w.r.t. to x, called fx. Then I calculate the derivative of fx w.r.t. to y, called fxy. The function fxyz is computed similarly. The output of fx is correct, …

---

## [\[FluxML + Functors\] How to walk the model and ∇model simultaneously?](https://discourse.julialang.org/t/fluxml-functors-how-to-walk-the-model-and-model-simultaneously/95170)

<div class="topic-metadata">

**Author:** [@mariusd](https://discourse.julialang.org/u/mariusd)\
**Replies:** 7\
**Last updated:** [February 27, 2023, 12:51pm UTC](https://discourse.julialang.org/t/fluxml-functors-how-to-walk-the-model-and-model-simultaneously/95170 "2023-02-27T12:51:53Z")

</div>

The following two lines of code are commonly used when training a model in Julia. ∇model = gradient(m → loss(m, x, y), model)\[1\] opt\_state, model = Optimisers.update!(opt\_state, model, ∇model) For debugging purspose…

---

## [CliqueNet and backward connection neural network](https://discourse.julialang.org/t/cliquenet-and-backward-connection-neural-network/95186)

<div class="topic-metadata">

**Author:** [@bienpierre](https://discourse.julialang.org/u/bienpierre)\
**Replies:** 1\
**Last updated:** [February 25, 2023, 4:15pm UTC](https://discourse.julialang.org/t/cliquenet-and-backward-connection-neural-network/95186 "2023-02-25T16:15:05Z")

</div>

Hello, I would like to implement the CliquetNet neural networks with fully connected dense with Flux: The network has some backward connection, I do not know the best way to implement the connection from the Flux.Ski…

---

## [Forecasting with LSTM + MLJ How to ? (A detailed guide or tutorial is needed)](https://discourse.julialang.org/t/forecasting-with-lstm-mlj-how-to-a-detailed-guide-or-tutorial-is-needed/94835)

<div class="topic-metadata">

**Author:** [@huqedato](https://discourse.julialang.org/u/huqedato)\
**Replies:** 1\
**Last updated:** [February 25, 2023, 2:35pm UTC](https://discourse.julialang.org/t/forecasting-with-lstm-mlj-how-to-a-detailed-guide-or-tutorial-is-needed/94835 "2023-02-25T14:35:08Z")

</div>

For a future task/project (starting this late spring) we will be required to create and deploy a network for predictive analysis on a large time-series dataset with multiple complex features. The tools of work (set by m…

---

## [Implementation of self-attention in Transformers.jl?](https://discourse.julialang.org/t/implementation-of-self-attention-in-transformers-jl/94732)

<div class="topic-metadata">

**Author:** [@rkube](https://discourse.julialang.org/u/rkube)\
**Replies:** 9\
**Last updated:** [February 21, 2023, 4:51am UTC](https://discourse.julialang.org/t/implementation-of-self-attention-in-transformers-jl/94732 "2023-02-21T04:51:03Z")

</div>

Hi, I’m following this tutorial on implementing transformer architectures and have a question on how to best implement a self-attention block. My current code for a single attention head looks like this: # Self-attent…

---

## [How to apply an activation function to a subset of output units?](https://discourse.julialang.org/t/how-to-apply-an-activation-function-to-a-subset-of-output-units/33202)

<div class="topic-metadata">

**Author:** [@johnbb](https://discourse.julialang.org/u/johnbb)\
**Replies:** 4\
**Last updated:** [February 21, 2023, 12:34am UTC](https://discourse.julialang.org/t/how-to-apply-an-activation-function-to-a-subset-of-output-units/33202 "2023-02-21T00:34:19Z")

</div>

Using Flux I would like to apply an activation function to all but the first output unit. More precisely, I would like to constrain all but the first unit to be non-negative. How can I achieve that? Thanks.

---

## [Not fast enough backward pass on transformer model](https://discourse.julialang.org/t/not-fast-enough-backward-pass-on-transformer-model/94735)

<div class="topic-metadata">

**Author:** [@reachtarunhere](https://discourse.julialang.org/u/reachtarunhere)\
**Replies:** 5\
**Last updated:** [February 21, 2023, 12:25am UTC](https://discourse.julialang.org/t/not-fast-enough-backward-pass-on-transformer-model/94735 "2023-02-21T00:25:07Z")

</div>

I have implemented a basic transformer like model and have a reference version in PyTorch 2.0 (which does a few extra ops like masking but not less work basically) I am getting comparable performance in the forward pass …

---

## [Bug when training a custom model using Flux](https://discourse.julialang.org/t/bug-when-training-a-custom-model-using-flux/94440)

<div class="topic-metadata">

**Author:** [@PokeLu](https://discourse.julialang.org/u/PokeLu)\
**Replies:** 2\
**Last updated:** [February 18, 2023, 8:41pm UTC](https://discourse.julialang.org/t/bug-when-training-a-custom-model-using-flux/94440 "2023-02-18T20:41:57Z")

</div>

Hello, I am pretty new to Julia and Flux. I am trying to build a simple neural network but using an attention layer. I wrote the code as follows, which works fine in the inference(feed-forward) mode: using Flux struct…

---

## [Flux benchmark being too slow vs Jax](https://discourse.julialang.org/t/flux-benchmark-being-too-slow-vs-jax/91608)

<div class="topic-metadata">

**Author:** [@krestomantsi](https://discourse.julialang.org/u/krestomantsi)\
**Replies:** 11\
**Last updated:** [February 15, 2023, 4:18pm UTC](https://discourse.julialang.org/t/flux-benchmark-being-too-slow-vs-jax/91608 "2023-02-15T16:18:14Z")

</div>

Hello everyone. I wanted to benchmark Flux vs Jax so I created this basic example. The tldr is that the Flux version (posted below) takes about 40-50 seconds on an nvidia A6000 (compilation excluded). The exact same arch…

---

## [Flux with AMD GPU(s)?](https://discourse.julialang.org/t/flux-with-amd-gpu-s/85218)

<div class="topic-metadata">

**Author:** [@johnbb](https://discourse.julialang.org/u/johnbb)\
**Replies:** 34\
**Last updated:** [February 15, 2023, 1:09pm UTC](https://discourse.julialang.org/t/flux-with-amd-gpu-s/85218 "2023-02-15T13:09:03Z")

</div>

Has anyone used Flux with an AMD GPU? I will be involved in a project that likely will be using the LUMI supercomputer in Finland which is based on AMD.

---

## [LIBSVM.jl with imbalanced dataset](https://discourse.julialang.org/t/libsvm-jl-with-imbalanced-dataset/94635)

<div class="topic-metadata">

**Author:** [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Replies:** 0\
**Last updated:** [February 14, 2023, 4:36pm UTC](https://discourse.julialang.org/t/libsvm-jl-with-imbalanced-dataset/94635 "2023-02-14T16:36:47Z")

</div>

I am trying to use LIBSVM.jl with an imbalanced dataset. Providing weights for each class doesn’t seem to affect the result: using CSV using DataFrames using LIBSVM df = CSV.read("svm.csv", DataFrame) X = \[df.x1 df.x…

---

## [Uplift model similar to PyLift or R's Lift Package?](https://discourse.julialang.org/t/uplift-model-similar-to-pylift-or-rs-lift-package/94268)

<div class="topic-metadata">

**Author:** [@Billpete002](https://discourse.julialang.org/u/Billpete002)\
**Replies:** 1\
**Last updated:** [February 14, 2023, 4:36am UTC](https://discourse.julialang.org/t/uplift-model-similar-to-pylift-or-rs-lift-package/94268 "2023-02-14T04:36:53Z")

</div>

Hi, I am looking for a package that works similarly to PyLift or R’s lift package. The idea is that I have two models: Self-service and the other is Sales-Assisted. I used Catboost.jl and examined manually that leads go…

---

## [Native eigenvals for differentiable programming](https://discourse.julialang.org/t/native-eigenvals-for-differentiable-programming/27126)

<div class="topic-metadata">

**Author:** [@GJBoth](https://discourse.julialang.org/u/GJBoth)\
**Replies:** 23\
**Last updated:** [February 13, 2023, 3:32pm UTC](https://discourse.julialang.org/t/native-eigenvals-for-differentiable-programming/27126 "2023-02-13T15:32:20Z")

</div>

Hi all, I’m working on a problem where I require an autodifferentiable eigenvalue solver for real, square matrices (but non-hermitian). Since Linear algebra’s version is based on LAPACK, I can’t use autodiff and I haven…

---

## [DiffEqFlux.jl: Questions about Neural ODEs](https://discourse.julialang.org/t/diffeqflux-jl-questions-about-neural-odes/94409)

<div class="topic-metadata">

**Author:** [@scj](https://discourse.julialang.org/u/scj)\
**Replies:** 1\
**Last updated:** [February 10, 2023, 1:46pm UTC](https://discourse.julialang.org/t/diffeqflux-jl-questions-about-neural-odes/94409 "2023-02-10T13:46:05Z")

</div>

I’m relatively new to both Julia and deep learning. I want to learn the dynamical model of a system using a set of observed trajectories. I’ve heard neural ODEs may be of promise, but they seem quite new, and I haven’t b…

---

## [ReverseDiff for loss function with Zygote derivatives](https://discourse.julialang.org/t/reversediff-for-loss-function-with-zygote-derivatives/94382)

<div class="topic-metadata">

**Author:** [@PatrickMcFarlane](https://discourse.julialang.org/u/PatrickMcFarlane)\
**Replies:** 1\
**Last updated:** [February 10, 2023, 3:30am UTC](https://discourse.julialang.org/t/reversediff-for-loss-function-with-zygote-derivatives/94382 "2023-02-10T03:30:24Z")

</div>

I’m trying to use Flux.jl to find the solution to a function equation with a form a bit like g(f(x)\[1\]) = h(f’(x)), where f(x) is an unkown function to be approximated by a neural network. The loss function for this set…

---

## [Type stability with Flux gradient of loss function requiring parameters](https://discourse.julialang.org/t/type-stability-with-flux-gradient-of-loss-function-requiring-parameters/94156)

<div class="topic-metadata">

**Author:** [@PatrickMcFarlane](https://discourse.julialang.org/u/PatrickMcFarlane)\
**Replies:** 4\
**Last updated:** [February 8, 2023, 1:37pm UTC](https://discourse.julialang.org/t/type-stability-with-flux-gradient-of-loss-function-requiring-parameters/94156 "2023-02-08T13:37:07Z")

</div>

Hello, I’m using Flux to replicate an economics application of machine learning (originally done in Python/Tensor flow here). In this problem, neural networks are used to solve a functional equation of the form g(f(x),…

---

## [How to use FluxMPI on slurm?](https://discourse.julialang.org/t/how-to-use-fluxmpi-on-slurm/94251)

<div class="topic-metadata">

**Author:** [@Emrick\_Liu](https://discourse.julialang.org/u/Emrick_Liu)\
**Replies:** 0\
**Last updated:** [February 8, 2023, 4:16am UTC](https://discourse.julialang.org/t/how-to-use-fluxmpi-on-slurm/94251 "2023-02-08T04:16:08Z")

</div>

Hi, I tried to use FluxMPI on Slurm, couldn’t figure out the configuration. Can anyone provide a minimum example? Currently, I’m basically doing something like this: #SBATCH --ntasks=3 export JULIA\_CUDA\_MEMORY\_POOL=no…

---

## [Julia/Flux creating a model correctly - using Chain Embedding layer reshaping & Dense layers](https://discourse.julialang.org/t/julia-flux-creating-a-model-correctly-using-chain-embedding-layer-reshaping-dense-layers/94089)

<div class="topic-metadata">

**Author:** [@roeya](https://discourse.julialang.org/u/roeya)\
**Replies:** 3\
**Last updated:** [February 7, 2023, 1:05am UTC](https://discourse.julialang.org/t/julia-flux-creating-a-model-correctly-using-chain-embedding-layer-reshaping-dense-layers/94089 "2023-02-07T01:05:52Z")

</div>

I am trying to follow an old article about word embedding in Julia language with Flux: Julia Word Embedding Layer in Flux - Self Trained - Getting to Know Julia As it is an old outdated article I managed to handle the e…

---

## [3 Layered Neural Network for a simple curve fit not working](https://discourse.julialang.org/t/3-layered-neural-network-for-a-simple-curve-fit-not-working/94078)

<div class="topic-metadata">

**Author:** [@RSH](https://discourse.julialang.org/u/RSH)\
**Replies:** 2\
**Last updated:** [February 6, 2023, 12:30am UTC](https://discourse.julialang.org/t/3-layered-neural-network-for-a-simple-curve-fit-not-working/94078 "2023-02-06T00:30:29Z")

</div>

I am currently attempting to utilize a three-layered neural network to model the function cos(x). Despite extensive efforts to debug my code, I have been unable to identify and rectify the source of error. It is my belie…

[Previous page](https://discourse.julialang.org/c/domain/ml/24.md?page=16)

[Next page](https://discourse.julialang.org/c/domain/ml/24.md?page=18)
