# Machine Learning

**URL:** https://discourse.julialang.org/c/domain/ml/24.md?page=9

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 10

---

## [How do I save a TunedModel with MLJ? DeterministicTunedModel has no field artifact\_location](https://discourse.julialang.org/t/how-do-i-save-a-tunedmodel-with-mlj-deterministictunedmodel-has-no-field-artifact-location/111307)

<div class="topic-metadata">

**Author:** [@niltsz](https://discourse.julialang.org/u/niltsz)\
**Replies:** 3\
**Last updated:** [March 8, 2024, 8:03am UTC](https://discourse.julialang.org/t/how-do-i-save-a-tunedmodel-with-mlj-deterministictunedmodel-has-no-field-artifact-location/111307 "2024-03-08T08:03:44Z")

</div>

I want to train a self tuning model with MLJ. I set up my model, it trains, but when I want to save the machine, I get the error message that DeterministicTunedModel has no field artifact\_location. Any idea how to solv…

---

## [Changing the size of the \`target\` using MLJ's Learning Networks](https://discourse.julialang.org/t/changing-the-size-of-the-target-using-mljs-learning-networks/111002)

<div class="topic-metadata">

**Author:** [@barakber](https://discourse.julialang.org/u/barakber)\
**Replies:** 6\
**Last updated:** [March 7, 2024, 8:19pm UTC](https://discourse.julialang.org/t/changing-the-size-of-the-target-using-mljs-learning-networks/111002 "2024-03-07T20:19:51Z")

</div>

Hi. I have a Static transformer ( \<: Static ) that changes the size of the input (X). It’s using sliding windows, so that size(X\_tag)\[1\] = size(X)\[1\] - window\_size. I’m trying to figure out what’s the right way to chang…

---

## [MLJ.save() and restoring with machine() don't work](https://discourse.julialang.org/t/mlj-save-and-restoring-with-machine-dont-work/110535)

<div class="topic-metadata">

**Author:** [@Paulo\_Refosco](https://discourse.julialang.org/u/Paulo_Refosco)\
**Replies:** 11\
**Last updated:** [March 7, 2024, 8:06pm UTC](https://discourse.julialang.org/t/mlj-save-and-restoring-with-machine-dont-work/110535 "2024-03-07T20:06:47Z")

</div>

Hello! I made a simple XGBoost model with the pipeline below, fitted it with data and tried to save it and then restore it, but when restoring it (even if in the same notebook) I get error on trying to apply predict. An…

---

## [How do I tune a pipeline in MLJ?](https://discourse.julialang.org/t/how-do-i-tune-a-pipeline-in-mlj/111155)

<div class="topic-metadata">

**Author:** [@ablaom](https://discourse.julialang.org/u/ablaom)\
**Replies:** 1\
**Last updated:** [March 4, 2024, 8:08pm UTC](https://discourse.julialang.org/t/how-do-i-tune-a-pipeline-in-mlj/111155 "2024-03-04T20:08:00Z")

</div>

A slack user has asked the question in the title. More generally, how does one tune hyperparameters that are nested in composed model? One obtains composed models, for example, when applying Stack, EnsembleModel, Iterat…

---

## [A Julia DSL for language models](https://discourse.julialang.org/t/a-julia-dsl-for-language-models/110244)

<div class="topic-metadata">

**Author:** [@cpfiffer](https://discourse.julialang.org/u/cpfiffer)\
**Replies:** 14\
**Last updated:** [February 28, 2024, 2:51pm UTC](https://discourse.julialang.org/t/a-julia-dsl-for-language-models/110244 "2024-02-28T14:51:57Z")

</div>

@svilupp and I were talking on the Slack about what a DSL might look like for generative models. I wanted to stick this on the forum for a more sticky conversational place, so I hope this can serve as kind of an evolving…

---

## [Zygote with Tullio gives wrong gradients/pullbacks using CUDA](https://discourse.julialang.org/t/zygote-with-tullio-gives-wrong-gradients-pullbacks-using-cuda/110767)

<div class="topic-metadata">

**Author:** [@kcqpo](https://discourse.julialang.org/u/kcqpo)\
**Replies:** 1\
**Last updated:** [February 26, 2024, 2:01pm UTC](https://discourse.julialang.org/t/zygote-with-tullio-gives-wrong-gradients-pullbacks-using-cuda/110767 "2024-02-26T14:01:55Z")

</div>

using Tullio, Zygote, CUDA, KernelAbstractions, OMEinsum # Show outer product of strings A = \["x", "y", "z", "w"\] res = Array{String}(undef, length(A), length(A)) for (i, r) in enumerate(A) for (j, c) in enumerate(A…

---

## [Deep learning with Flux.jl and gumbel-softmax trick](https://discourse.julialang.org/t/deep-learning-with-flux-jl-and-gumbel-softmax-trick/110810)

<div class="topic-metadata">

**Author:** [@Flavio\_Barros](https://discourse.julialang.org/u/Flavio_Barros)\
**Replies:** 2\
**Last updated:** [February 27, 2024, 12:48pm UTC](https://discourse.julialang.org/t/deep-learning-with-flux-jl-and-gumbel-softmax-trick/110810 "2024-02-27T12:48:05Z")

</div>

I am new to Julia, but I’m evaluating Flux.jl as an option to implement a Variational Autoencoder (VAE). But, unfortunately, I’m not finding an example using the gumbel-softmax trick. I found on Pytorch, but I wonder how…

---

## [Fitted Flux model is "off" by one](https://discourse.julialang.org/t/fitted-flux-model-is-off-by-one/110709)

<div class="topic-metadata">

**Author:** [@affans](https://discourse.julialang.org/u/affans)\
**Replies:** 2\
**Last updated:** [February 24, 2024, 9:12pm UTC](https://discourse.julialang.org/t/fitted-flux-model-is-off-by-one/110709 "2024-02-24T21:12:09Z")

</div>

I am trying to reproduce a simple NN model in Julia using Lux.jl, originally using the R package forecast which trains an AR(2) model using nnet internally. I am almost able to reproduce it, but my fitted model is off by…

---

## [Dumb question: why autodifferentiation is needed in Neural Networks](https://discourse.julialang.org/t/dumb-question-why-autodifferentiation-is-needed-in-neural-networks/110673)

<div class="topic-metadata">

**Author:** [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Replies:** 9\
**Last updated:** [February 24, 2024, 6:11pm UTC](https://discourse.julialang.org/t/dumb-question-why-autodifferentiation-is-needed-in-neural-networks/110673 "2024-02-24T18:11:21Z")

</div>

I know little other than the basic NN, but why autodiff is needed for DeepLearning? Using backchain it is fairly easy to get the gradient without autodiff, or at least if standard activation functions are provided. For …

---

## [Help with Flux.jl, Metal.jl (Apple Silicon) and Conv layers](https://discourse.julialang.org/t/help-with-flux-jl-metal-jl-apple-silicon-and-conv-layers/105632)

<div class="topic-metadata">

**Author:** [@JakobPeterlin](https://discourse.julialang.org/u/JakobPeterlin)\
**Replies:** 2\
**Last updated:** [February 23, 2024, 4:27pm UTC](https://discourse.julialang.org/t/help-with-flux-jl-metal-jl-apple-silicon-and-conv-layers/105632 "2024-02-23T16:27:15Z")

</div>

Hello, after trying to use Flux.jl to reimplement some PyTorch code (SpinalVGG), I encountered some difficulties when trying to run my code on an M2 Max’s GPU with Metal.jl. The error boils down to the following code, w…

---

## [How can I solve complex-valued ordinary differential equations (ODEs) using neural networks, given limitations with complex data types in libraries like Lux?](https://discourse.julialang.org/t/how-can-i-solve-complex-valued-ordinary-differential-equations-odes-using-neural-networks-given-limitations-with-complex-data-types-in-libraries-like-lux/110581)

<div class="topic-metadata">

**Author:** [@sahakyan2003](https://discourse.julialang.org/u/sahakyan2003)\
**Replies:** 2\
**Last updated:** [February 22, 2024, 5:56pm UTC](https://discourse.julialang.org/t/how-can-i-solve-complex-valued-ordinary-differential-equations-odes-using-neural-networks-given-limitations-with-complex-data-types-in-libraries-like-lux/110581 "2024-02-22T17:56:50Z")

</div>

Here I describe the code where I want to train the neural network but face the problem Here are the libraries I use. using NeuralPDE using DifferentialEquations using Plots using Lux, Random The function system\_of\_de! …

---

## [Can't run a Resnet code](https://discourse.julialang.org/t/cant-run-a-resnet-code/110303)

<div class="topic-metadata">

**Author:** [@delcencen](https://discourse.julialang.org/u/delcencen)\
**Replies:** 3\
**Last updated:** [February 19, 2024, 1:12pm UTC](https://discourse.julialang.org/t/cant-run-a-resnet-code/110303 "2024-02-19T13:12:10Z")

</div>

Hello, in this code i ave the 2 following errors : -ERROR: LoadError: UndefVarError: \`Resnet\` not defined -ERROR: LoadError: UndefVarError: \`get\_classes\` not defined the code of my book : using Flux using M…

---

## [Do not update neural network weights with a value of 0](https://discourse.julialang.org/t/do-not-update-neural-network-weights-with-a-value-of-0/110335)

<div class="topic-metadata">

**Author:** [@kuanchiun](https://discourse.julialang.org/u/kuanchiun)\
**Replies:** 5\
**Last updated:** [February 19, 2024, 2:42am UTC](https://discourse.julialang.org/t/do-not-update-neural-network-weights-with-a-value-of-0/110335 "2024-02-19T02:42:38Z")

</div>

Hello, I have just been exposed to the field of deep learning for a short time. Recently I was trying to create the 3-dimensional Neural Network, by giving the 3D Array as a pretrained weights. Each 2D matrix in the wei…

---

## [Using LSTM cell in Lux with explicit parameters](https://discourse.julialang.org/t/using-lstm-cell-in-lux-with-explicit-parameters/110199)

<div class="topic-metadata">

**Author:** [@kunal](https://discourse.julialang.org/u/kunal)\
**Replies:** 0\
**Last updated:** [February 14, 2024, 4:06am UTC](https://discourse.julialang.org/t/using-lstm-cell-in-lux-with-explicit-parameters/110199 "2024-02-14T04:06:26Z")

</div>

Hi, I was referring to the Simple LSTM tutorial, which combines classifier(Dense layer) with LSTMCell using AbstractExplicitContainerLayer. Also, in my case I am passing explicit parameters to the neural net. NN = L…

---

## [How hard is it to rebuild a tensorflow model in julia?](https://discourse.julialang.org/t/how-hard-is-it-to-rebuild-a-tensorflow-model-in-julia/110349)

<div class="topic-metadata">

**Author:** [@cdawg](https://discourse.julialang.org/u/cdawg)\
**Replies:** 2\
**Last updated:** [February 17, 2024, 9:05pm UTC](https://discourse.julialang.org/t/how-hard-is-it-to-rebuild-a-tensorflow-model-in-julia/110349 "2024-02-17T21:05:37Z")

</div>

It has been a pain for me to get tensorflow working on my arm MacBook to run a recently published denoising model. Id love to port the model Julia because it would open up some options and make it easier to share with o…

---

## [Moving a custom loss function to GPU](https://discourse.julialang.org/t/moving-a-custom-loss-function-to-gpu/109637)

<div class="topic-metadata">

**Author:** [@josemanuel22](https://discourse.julialang.org/u/josemanuel22)\
**Replies:** 5\
**Last updated:** [February 15, 2024, 9:19pm UTC](https://discourse.julialang.org/t/moving-a-custom-loss-function-to-gpu/109637 "2024-02-15T21:19:49Z")

</div>

I have developed the following custom loss function, I am trying to run it on the GPU efficiently but I am unable to do so; I get the following error in the line Matrix(ω' \* yₖ\_batch) ERROR: MethodError: no method match…

---

## [Slow LSTM on GPU in Flux](https://discourse.julialang.org/t/slow-lstm-on-gpu-in-flux/84228)

<div class="topic-metadata">

**Author:** [@JLDC](https://discourse.julialang.org/u/JLDC)\
**Replies:** 21\
**Last updated:** [February 15, 2024, 2:14am UTC](https://discourse.julialang.org/t/slow-lstm-on-gpu-in-flux/84228 "2024-02-15T02:14:58Z")

</div>

Hi, I’ve been running into an issue where I try to train a neural network using Flux and the training times are much slower on GPU than on CPU. So I tried comparing an MWE using Flux and an MWE using PyCall and running …

---

## [Speeding up gradient of logpdf](https://discourse.julialang.org/t/speeding-up-gradient-of-logpdf/109708)

<div class="topic-metadata">

**Author:** [@bertschi](https://discourse.julialang.org/u/bertschi)\
**Replies:** 19\
**Last updated:** [February 12, 2024, 3:08pm UTC](https://discourse.julialang.org/t/speeding-up-gradient-of-logpdf/109708 "2024-02-12T15:08:06Z")

</div>

Somehow, I seem to be unable to speed up gradient calculations of a simple logpdf. As a minimal working example, consider the following computation: using Distributions import AbstractDifferentiation as AD import Forwar…

---

## [Crystal Graphs w/ GeometricFlux.jl](https://discourse.julialang.org/t/crystal-graphs-w-geometricflux-jl/109773)

<div class="topic-metadata">

**Author:** [@Stefan\_Bringuier](https://discourse.julialang.org/u/Stefan_Bringuier)\
**Replies:** 0\
**Last updated:** [February 6, 2024, 12:25am UTC](https://discourse.julialang.org/t/crystal-graphs-w-geometricflux-jl/109773 "2024-02-06T00:25:40Z")

</div>

I’m fairly new to GeometricFlux.jl and working on trying to implement the Crystal Graph CNN approach by Xie and Grossman. I have the following FeaturedGraph, target pair dataset. I guess what I’m wondering is 1.) How sho…

---

## [Lux (And Flux), "parallel" Network Input. When Input is flat, Zygote gradient works, when input is not flat it doesn't](https://discourse.julialang.org/t/lux-and-flux-parallel-network-input-when-input-is-flat-zygote-gradient-works-when-input-is-not-flat-it-doesnt/107763)

<div class="topic-metadata">

**Author:** [@Qfl3x](https://discourse.julialang.org/u/Qfl3x)\
**Replies:** 10\
**Last updated:** [February 5, 2024, 6:10pm UTC](https://discourse.julialang.org/t/lux-and-flux-parallel-network-input-when-input-is-flat-zygote-gradient-works-when-input-is-not-flat-it-doesnt/107763 "2024-02-05T18:10:42Z")

</div>

It’s easier to see the code: using Lux using Zygote using ComponentArrays using Random x\_test\_flat = rand(16) x\_test\_p = \[rand(4) for \_ in 1:4\] NNmodel\_1 = Lux.Chain(Lux.Dense(16 =\> 18, sigmoid), …

---

## [Parser for safetensors](https://discourse.julialang.org/t/parser-for-safetensors/109366)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 6\
**Last updated:** [February 5, 2024, 12:35pm UTC](https://discourse.julialang.org/t/parser-for-safetensors/109366 "2024-02-05T12:35:15Z")

</div>

There is a new format to store weights of neural networks advocated by huggingface, called safetensors. I have encountered when I wanted to download phi-2 model to Transformers.jl (see this pull request Add Phi model by…

---

## [Memoization in mcmc](https://discourse.julialang.org/t/memoization-in-mcmc/109694)

<div class="topic-metadata">

**Author:** [@swsw](https://discourse.julialang.org/u/swsw)\
**Replies:** 0\
**Last updated:** [February 4, 2024, 9:13am UTC](https://discourse.julialang.org/t/memoization-in-mcmc/109694 "2024-02-04T09:13:04Z")

</div>

Hi there, I wondered if how Turing or RxInfer deals w/ the following situation. For example, I have 10 parameters x1,…x10, the likelihood function is as follows. x1^2 + … + x10^2 Apparently, when updating x1, the o…

---

## [Physics-informed training of a surrogate model involving finite element analysis](https://discourse.julialang.org/t/physics-informed-training-of-a-surrogate-model-involving-finite-element-analysis/103735)

<div class="topic-metadata">

**Author:** [@touste](https://discourse.julialang.org/u/touste)\
**Replies:** 4\
**Last updated:** [February 2, 2024, 10:29pm UTC](https://discourse.julialang.org/t/physics-informed-training-of-a-surrogate-model-involving-finite-element-analysis/103735 "2024-02-02T22:29:53Z")

</div>

Hi all, I’m looking at ways to accelerate mechanical simulations for endovascular applications. We published a paper recently where we present EndoBeams.jl, a finite element framework for beam elements in contact with a…

---

## [Understanding \`Flux.Data.DataLoader\` when training an LSTM model](https://discourse.julialang.org/t/understanding-flux-data-dataloader-when-training-an-lstm-model/109581)

<div class="topic-metadata">

**Author:** [@Hugo](https://discourse.julialang.org/u/Hugo)\
**Replies:** 1\
**Last updated:** [February 2, 2024, 3:34am UTC](https://discourse.julialang.org/t/understanding-flux-data-dataloader-when-training-an-lstm-model/109581 "2024-02-02T03:34:36Z")

</div>

Hello, folks For a while I have been playing with Flux.jl to learn how to put together ML models, train them, and use them for predictive tasks. At the moment, I am stuck with the training stage of an LSTM model. Probl…

---

## [Enzyme autodiff: Why am I getting allocations?](https://discourse.julialang.org/t/enzyme-autodiff-why-am-i-getting-allocations/109307)

<div class="topic-metadata">

**Author:** [@Larbino1](https://discourse.julialang.org/u/Larbino1)\
**Replies:** 20\
**Last updated:** [January 29, 2024, 1:19pm UTC](https://discourse.julialang.org/t/enzyme-autodiff-why-am-i-getting-allocations/109307 "2024-01-29T13:19:54Z")

</div>

I’m using Enzyme to perform automatic differentiation. I thought that as the primal code is type stable and non allocating, the enzyme autodiff function would also be. However for me the minimized example below gives th…

---

## [cuDNN, julia-1.10 and linux](https://discourse.julialang.org/t/cudnn-julia-1-10-and-linux/106658)

<div class="topic-metadata">

**Author:** [@quietlight](https://discourse.julialang.org/u/quietlight)\
**Replies:** 10\
**Last updated:** [January 28, 2024, 9:34pm UTC](https://discourse.julialang.org/t/cudnn-julia-1-10-and-linux/106658 "2024-01-28T21:34:28Z")

</div>

I have trained a Resnet18 using Flux on julia-1.10 rc1, actually I have trained multiple models with no trouble at all, but this week when I try to predict using my model I get a compilation error for my private package: …

---

## [Why doesn't the loss calculated by Flux \`withgradient\` match what I have calculated?](https://discourse.julialang.org/t/why-doesnt-the-loss-calculated-by-flux-withgradient-match-what-i-have-calculated/109228)

<div class="topic-metadata">

**Author:** [@kirklong](https://discourse.julialang.org/u/kirklong)\
**Replies:** 2\
**Last updated:** [January 26, 2024, 9:05pm UTC](https://discourse.julialang.org/t/why-doesnt-the-loss-calculated-by-flux-withgradient-match-what-i-have-calculated/109228 "2024-01-26T21:05:57Z")

</div>

Just posted this to StackOverflow, but crossposting here for Julia reach. TL;DR – when making a simple CNN I noticed weird behavior between the reported loss during training and the loss I actually calculated from the “…

---

## [A discrepancy in self-attention between python and Julia (Transformers)](https://discourse.julialang.org/t/a-discrepancy-in-self-attention-between-python-and-julia-transformers/109195)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 7\
**Last updated:** [January 26, 2024, 7:27am UTC](https://discourse.julialang.org/t/a-discrepancy-in-self-attention-between-python-and-julia-transformers/109195 "2024-01-26T07:27:00Z")

</div>

Hi, I am continuing the porting of microsoft Phi model to the Transformers.jl. The most complicated is the self-attention, where I am lost, but gradually removing abstraction layers of Neuralattentionlib.jl to get to th…

---

## [GPU memory in Julia](https://discourse.julialang.org/t/gpu-memory-in-julia/109216)

<div class="topic-metadata">

**Author:** [@Wenbo\_Li](https://discourse.julialang.org/u/Wenbo_Li)\
**Replies:** 5\
**Last updated:** [January 25, 2024, 6:02pm UTC](https://discourse.julialang.org/t/gpu-memory-in-julia/109216 "2024-01-25T18:02:03Z")

</div>

I am training unet with 3d input size 512512128. However, with the same model, same loss function, same size of input and label in float32, python indeed can fit and train while julia not. Also, julia need much longer ti…

---

## [Forcing Lux to use Enzyme instead of Zygote?](https://discourse.julialang.org/t/forcing-lux-to-use-enzyme-instead-of-zygote/109193)

<div class="topic-metadata">

**Author:** [@Naceur\_Bouziani](https://discourse.julialang.org/u/Naceur_Bouziani)\
**Replies:** 10\
**Last updated:** [January 25, 2024, 3:35pm UTC](https://discourse.julialang.org/t/forcing-lux-to-use-enzyme-instead-of-zygote/109193 "2024-01-25T15:35:10Z")

</div>

Hi everyone. I’m confused. The example that used to work @GitHub - SciML/NeuralPDE.jl: Physics-Informed Neural Networks (PINN) and Deep BSDE Solvers of Differential Equations for Scientific Machine Learning (SciML) acc…

[Previous page](https://discourse.julialang.org/c/domain/ml/24.md?page=8)

[Next page](https://discourse.julialang.org/c/domain/ml/24.md?page=10)
