# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=4

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 5

---

## [Stochastic Gradient Descent (SGD) in julia?](https://discourse.julialang.org/t/stochastic-gradient-descent-sgd-in-julia/133846)

<div class="topic-metadata">

**Author:** [@Sigmund](https://discourse.julialang.org/u/Sigmund)\
**Replies:** 3\
**Last updated:** [November 13, 2025, 12:18pm UTC](https://discourse.julialang.org/t/stochastic-gradient-descent-sgd-in-julia/133846 "2025-11-13T12:18:38Z")

</div>

Im trying to run the newest version of OptimizationManopt.jl together with NeuralPDE.jl. When trying to add NeuralPDE.jl to the environment i get a long stacktrace with unsatisfiable requirements. When adding both to an …

---

## [Understanding "random" serialization when using @spawn in a "parallel region"](https://discourse.julialang.org/t/understanding-random-serialization-when-using-spawn-in-a-parallel-region/133824)

<div class="topic-metadata">

**Author:** [@mmesiti](https://discourse.julialang.org/u/mmesiti)\
**Replies:** 11\
**Last updated:** [November 13, 2025, 11:59am UTC](https://discourse.julialang.org/t/understanding-random-serialization-when-using-spawn-in-a-parallel-region/133824 "2025-11-13T11:59:50Z")

</div>

I have some Julia (1.10.10) code that I am benchmarking on a HPC cluster node with 76 physical cores, running with 76 threads, and I am seeing HUGE variability between runs with everything equal (same hardware, same code…

---

## [\`rand(::MyType, N)\` allocates a lot, how to define return type properly?](https://discourse.julialang.org/t/rand-mytype-n-allocates-a-lot-how-to-define-return-type-properly/133807)

<div class="topic-metadata">

**Author:** [@tamasgal](https://discourse.julialang.org/u/tamasgal)\
**Replies:** 4\
**Last updated:** [November 12, 2025, 1:19pm UTC](https://discourse.julialang.org/t/rand-mytype-n-allocates-a-lot-how-to-define-return-type-properly/133807 "2025-11-12T13:19:48Z")

</div>

Apologies if this has been discussed before but I am a bit lost in the docs and examples out there regarding a type-safe implementation of a proper rand interface. I guess it’s OK to have yet another another topic on thi…

---

## [Backup 1 line when using readline()](https://discourse.julialang.org/t/backup-1-line-when-using-readline/133826)

<div class="topic-metadata">

**Author:** [@Jake](https://discourse.julialang.org/u/Jake)\
**Replies:** 1\
**Last updated:** [November 12, 2025, 12:38pm UTC](https://discourse.julialang.org/t/backup-1-line-when-using-readline/133826 "2025-11-12T12:38:16Z")

</div>

I am going through a mixed binary/ASCII file implementing UFF58b file format in UFFFiles.jl. It will make my life easier if I can read a line using readline(), call a subroutine and then read the same line again. So I …

---

## [Array addition of oneAPI.jl slower](https://discourse.julialang.org/t/array-addition-of-oneapi-jl-slower/133663)

<div class="topic-metadata">

**Author:** [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Replies:** 10\
**Last updated:** [November 10, 2025, 8:30pm UTC](https://discourse.julialang.org/t/array-addition-of-oneapi-jl-slower/133663 "2025-11-10T20:30:05Z")

</div>

Why oneAPI.jl array addition is slow? julia\> using BenchmarkTools, oneAPI julia\> c = rand(100,100); julia\> @btime $c.+1 4.950 μs (3 allocations: 78.21 KiB) julia\> a = oneArray(rand(100,100)); julia\> @btime $a.+1 4…

---

## [Topological sort (performance)](https://discourse.julialang.org/t/topological-sort-performance/17153)

<div class="topic-metadata">

**Author:** [@govorunov](https://discourse.julialang.org/u/govorunov)\
**Replies:** 43\
**Last updated:** [November 10, 2025, 5:36pm UTC](https://discourse.julialang.org/t/topological-sort-performance/17153 "2025-11-10T17:36:27Z")

</div>

Hi. I needed to implement a topological sort, so I found some implementations on the Internet here: Topological sort - Rosetta Code I initially needed it in Python and experimented with Julia just out of curiosity. Her…

---

## [Threads.@threads on views into a vector doesn't give much of a speed-up](https://discourse.julialang.org/t/threads-threads-on-views-into-a-vector-doesnt-give-much-of-a-speed-up/133713)

<div class="topic-metadata">

**Author:** [@jlbosse](https://discourse.julialang.org/u/jlbosse)\
**Replies:** 3\
**Last updated:** [November 6, 2025, 9:16pm UTC](https://discourse.julialang.org/t/threads-threads-on-views-into-a-vector-doesnt-give-much-of-a-speed-up/133713 "2025-11-06T21:16:40Z")

</div>

I am trying to speed my code up using the Threads.@threads macro and splitting an array into multiple views so that each thread works on a separate view, but am not getting the performance gains I had hoped for. Conside…

---

## [Large allocations with MTK event](https://discourse.julialang.org/t/large-allocations-with-mtk-event/133680)

<div class="topic-metadata">

**Author:** [@klinders](https://discourse.julialang.org/u/klinders)\
**Replies:** 1\
**Last updated:** [November 5, 2025, 11:35pm UTC](https://discourse.julialang.org/t/large-allocations-with-mtk-event/133680 "2025-11-05T23:35:22Z")

</div>

Dear all, Recently, you helped me get events working in MTK. Thank you for all the help! Continuing on the topic, I ran into a performance issue with events. I have the following code to measure the time and allocation…

---

## [Dagger.jl mutable data on different workers](https://discourse.julialang.org/t/dagger-jl-mutable-data-on-different-workers/133676)

<div class="topic-metadata">

**Author:** [@Salmon](https://discourse.julialang.org/u/Salmon)\
**Replies:** 0\
**Last updated:** [November 5, 2025, 11:31am UTC](https://discourse.julialang.org/t/dagger-jl-mutable-data-on-different-workers/133676 "2025-11-05T11:31:54Z")

</div>

Hello people, I am new to Dagger and I have a beginner question. I am planning to create a bunch (lets say 10) of instances of a problem. Each instance should be allocated on a separate worker - Each instance should f…

---

## [Memory kills this code](https://discourse.julialang.org/t/memory-kills-this-code/133583)

<div class="topic-metadata">

**Author:** [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Replies:** 5\
**Last updated:** [November 3, 2025, 9:33pm UTC](https://discourse.julialang.org/t/memory-kills-this-code/133583 "2025-11-03T21:33:45Z")

</div>

This code is giving accurate plot but needs too much memory and is slow. I have been working since few months to get 3D volume plot of this Astrophysical data linked here. You can download this data by clicking on Downlo…

---

## [Is this conversion+store safe?](https://discourse.julialang.org/t/is-this-conversion-store-safe/133533)

<div class="topic-metadata">

**Author:** [@Tortar](https://discourse.julialang.org/u/Tortar)\
**Replies:** 4\
**Last updated:** [October 30, 2025, 12:52pm UTC](https://discourse.julialang.org/t/is-this-conversion-store-safe/133533 "2025-10-30T12:52:11Z")

</div>

I’m wondering if this is a correct usage of these unsafe functions: julia\> using StaticArrays julia\> a = MVector{4,UInt}(1,2,3,4); julia\> b = UInt.((5,6,7,8)); julia\> function f(a, b) dst = Base.unsafe\_con…

---

## [How to avoid recompilation when a large Julia function is called with many different argument combinations](https://discourse.julialang.org/t/how-to-avoid-recompilation-when-a-large-julia-function-is-called-with-many-different-argument-combinations/133516)

<div class="topic-metadata">

**Author:** [@r2ached](https://discourse.julialang.org/u/r2ached)\
**Replies:** 3\
**Last updated:** [October 29, 2025, 8:25pm UTC](https://discourse.julialang.org/t/how-to-avoid-recompilation-when-a-large-julia-function-is-called-with-many-different-argument-combinations/133516 "2025-10-29T20:25:37Z")

</div>

Hi everyone, I’m working on a Julia application where one top-level function has an input called configuration. This function gets called many times with different configurations. The problem I’m having is that Julia r…

---

## [Downloads.download or run(\`wget\`) much slower than direct \`wget\`](https://discourse.julialang.org/t/downloads-download-or-run-wget-much-slower-than-direct-wget/133501)

<div class="topic-metadata">

**Author:** [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Replies:** 2\
**Last updated:** [October 29, 2025, 4:31pm UTC](https://discourse.julialang.org/t/downloads-download-or-run-wget-much-slower-than-direct-wget/133501 "2025-10-29T16:31:47Z")

</div>

I am trying to download a relativelly large dataset from Zenodo (6GB), and if I use Downloads.download or run(wget) in a Julia v1.12 session, it takes 2h, while if I directly use wget in a console it takes 16 minutes. I…

---

## [Improving performance of CUDA GPU kernel: LU factorization](https://discourse.julialang.org/t/improving-performance-of-cuda-gpu-kernel-lu-factorization/132971)

<div class="topic-metadata">

**Author:** [@Leo\_I](https://discourse.julialang.org/u/Leo_I)\
**Replies:** 17\
**Last updated:** [October 28, 2025, 10:37am UTC](https://discourse.julialang.org/t/improving-performance-of-cuda-gpu-kernel-lu-factorization/132971 "2025-10-28T10:37:38Z")

</div>

I am learning GPU programming via CUDA.jl. I wish to implement an efficient LU factorization for matrices over a finite field (I implemented it as UInt8/UInt16/... with operations + mod and \* mod). For the sake of simpl…

---

## [Make mutating function more AD-friendly](https://discourse.julialang.org/t/make-mutating-function-more-ad-friendly/133356)

<div class="topic-metadata">

**Author:** [@MrMR](https://discourse.julialang.org/u/MrMR)\
**Replies:** 10\
**Last updated:** [October 23, 2025, 9:52am UTC](https://discourse.julialang.org/t/make-mutating-function-more-ad-friendly/133356 "2025-10-23T09:52:03Z")

</div>

Dear all, since this is my first post, I want to thank you for the help I got in many years I used Julia (I always found the answer to my problems in previous posts!). Diving into the problem: what I need to do is to …

---

## [The same async program run on (Linux+AMD) vs (Win+Intel), over 6x performance difference](https://discourse.julialang.org/t/the-same-async-program-run-on-linux-amd-vs-win-intel-over-6x-performance-difference/133276)

<div class="topic-metadata">

**Author:** [@WalterMadelim](https://discourse.julialang.org/u/WalterMadelim)\
**Replies:** 5\
**Last updated:** [October 23, 2025, 9:38am UTC](https://discourse.julialang.org/t/the-same-async-program-run-on-linux-amd-vs-win-intel-over-6x-performance-difference/133276 "2025-10-23T09:38:33Z")

</div>

Long story short: I wrote a async programming code which uses multithreads. The same code is run on two servers. The Linux server is 6.5 times faster than the other (both attaining the same correct results). I’m not hur…

---

## [Reading from socket causes massive amounts of memory allocation](https://discourse.julialang.org/t/reading-from-socket-causes-massive-amounts-of-memory-allocation/133206)

<div class="topic-metadata">

**Author:** [@world-peace](https://discourse.julialang.org/u/world-peace)\
**Replies:** 12\
**Last updated:** [October 23, 2025, 9:23am UTC](https://discourse.julialang.org/t/reading-from-socket-causes-massive-amounts-of-memory-allocation/133206 "2025-10-23T09:23:21Z")

</div>

Reading from network sockets appears to cause massive amounts of memory allocation. Does anyone know why? Is this a bug? raw\_data = Vector{UInt8}(undef, message\_length) read!(socket, raw\_data) What I find strange is t…

---

## [Worse runtimes in Julia v1.12](https://discourse.julialang.org/t/worse-runtimes-in-julia-v1-12/133231)

<div class="topic-metadata">

**Author:** [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Replies:** 20\
**Last updated:** [October 22, 2025, 6:13am UTC](https://discourse.julialang.org/t/worse-runtimes-in-julia-v1-12/133231 "2025-10-22T06:13:43Z")

</div>

The runtimes are getting worse in Julia v1.12. I have a couple of tests in GeoStatsFunctions.jl that are now failing: CompositeVariogram: Test Failed at /home/runner/work/GeoStatsFunctions.jl/GeoStatsFunctions.jl/test/…

---

## [Strange speed gain from eachindex and writing in a useless array?](https://discourse.julialang.org/t/strange-speed-gain-from-eachindex-and-writing-in-a-useless-array/132724)

<div class="topic-metadata">

**Author:** [@nicolas](https://discourse.julialang.org/u/nicolas)\
**Replies:** 13\
**Last updated:** [October 21, 2025, 7:53am UTC](https://discourse.julialang.org/t/strange-speed-gain-from-eachindex-and-writing-in-a-useless-array/132724 "2025-10-21T07:53:33Z")

</div>

This code doesn’t do anything useful, it is just intended as a MWE. my\_func! outputs the sum s in ≈ 120μs on my computer. Strangely, when I uncomment the lines adding to the array useless, it seems to do the same computa…

---

## [Type instability with (fairly) simple Tuples](https://discourse.julialang.org/t/type-instability-with-fairly-simple-tuples/133304)

<div class="topic-metadata">

**Author:** [@johnomotani](https://discourse.julialang.org/u/johnomotani)\
**Replies:** 7\
**Last updated:** [October 20, 2025, 12:25pm UTC](https://discourse.julialang.org/t/type-instability-with-fairly-simple-tuples/133304 "2025-10-20T12:25:03Z")

</div>

Should I expect the following to be type-stable? I think the types are inferrable in principle, but maybe the Tuple-of-Pairs structure is just too ‘complicated’? julia\> using Cthulhu julia\> function foo(t::Pair{Symbol,…

---

## [Performance drop x10 if \`missing\` used or two values returned from function](https://discourse.julialang.org/t/performance-drop-x10-if-missing-used-or-two-values-returned-from-function/133200)

<div class="topic-metadata">

**Author:** [@Alex1](https://discourse.julialang.org/u/Alex1)\
**Replies:** 13\
**Last updated:** [October 20, 2025, 5:08am UTC](https://discourse.julialang.org/t/performance-drop-x10-if-missing-used-or-two-values-returned-from-function/133200 "2025-10-20T05:08:49Z")

</div>

Hi, while doning GARCH fitting I discovered strange performance drop: Case 1 Returning two values instead of one is x10 times slower 1.5ms vs 23μs. See place marked with CHANGE1. Is that a known issue and should not be …

---

## [@snoop\_inference taking a very long time?](https://discourse.julialang.org/t/snoop-inference-taking-a-very-long-time/133233)

<div class="topic-metadata">

**Author:** [@johnomotani](https://discourse.julialang.org/u/johnomotani)\
**Replies:** 2\
**Last updated:** [October 17, 2025, 2:03pm UTC](https://discourse.julialang.org/t/snoop-inference-taking-a-very-long-time/133233 "2025-10-17T14:03:51Z")

</div>

Running @snoop\_inference on a function from my package is taking a very long time - it’s been running for an hour since the function finished executing and hasn’t returned yet. Is this a known issue in any cases? Any sug…

---

## [Inconsistent allocation with threads and function arguments in 1.12](https://discourse.julialang.org/t/inconsistent-allocation-with-threads-and-function-arguments-in-1-12/133190)

<div class="topic-metadata">

**Author:** [@bclyons12](https://discourse.julialang.org/u/bclyons12)\
**Replies:** 7\
**Last updated:** [October 16, 2025, 6:27pm UTC](https://discourse.julialang.org/t/inconsistent-allocation-with-threads-and-function-arguments-in-1-12/133190 "2025-10-16T18:27:45Z")

</div>

In a large code suite, I’ve noticed substantial performance degradation using v1.12. I’m seeing large numbers of allocations in parts of the code that are threaded and involve passing functions as arguments to other func…

---

## [Performace problems with Windows scheduler and multithreading on mixed core CPUs](https://discourse.julialang.org/t/performace-problems-with-windows-scheduler-and-multithreading-on-mixed-core-cpus/132984)

<div class="topic-metadata">

**Author:** [@mwlidar](https://discourse.julialang.org/u/mwlidar)\
**Replies:** 11\
**Last updated:** [October 14, 2025, 11:10am UTC](https://discourse.julialang.org/t/performace-problems-with-windows-scheduler-and-multithreading-on-mixed-core-cpus/132984 "2025-10-14T11:10:35Z")

</div>

I have a function with a long running simulation (\> 20 min) and used Threads@spawn to start four of them in parallel. Disappointingly the runtime more than doubled compared to a single instance run. After some digging I …

---

## [Benchmarking ADTs behavior](https://discourse.julialang.org/t/benchmarking-adts-behavior/133115)

<div class="topic-metadata">

**Author:** [@nandoconde](https://discourse.julialang.org/u/nandoconde)\
**Replies:** 2\
**Last updated:** [October 14, 2025, 5:43am UTC](https://discourse.julialang.org/t/benchmarking-adts-behavior/133115 "2025-10-14T05:43:47Z")

</div>

I am trying to get the most performant way to retrieve some properties associated to an ADT (I am using Moshi.jl, but I suspect the behavior with LightSumTypes et al. would be similar). Below I paste the code I am testin…

---

## [Slowdown when writing to too many arrays within a loop](https://discourse.julialang.org/t/slowdown-when-writing-to-too-many-arrays-within-a-loop/133041)

<div class="topic-metadata">

**Author:** [@ekboehm](https://discourse.julialang.org/u/ekboehm)\
**Replies:** 15\
**Last updated:** [October 13, 2025, 5:45pm UTC](https://discourse.julialang.org/t/slowdown-when-writing-to-too-many-arrays-within-a-loop/133041 "2025-10-13T17:45:08Z")

</div>

I came across a significant slowdown in my code when attempting to write to too many matrices within one loop. Below is a silly (and slightly long, my apologies) MWE that just writes zeroes. In my actual code, I perform …

---

## [Preventing Allocations when using AlgebraicMultigrid.jl](https://discourse.julialang.org/t/preventing-allocations-when-using-algebraicmultigrid-jl/132983)

<div class="topic-metadata">

**Author:** [@hssn15](https://discourse.julialang.org/u/hssn15)\
**Replies:** 1\
**Last updated:** [October 9, 2025, 3:10am UTC](https://discourse.julialang.org/t/preventing-allocations-when-using-algebraicmultigrid-jl/132983 "2025-10-09T03:10:14Z")

</div>

I am trying to apply AMG solver in my numerical simulator to solve big sparse linear systems (Ax=b). in order to apply it multilevel object should be created each time values of matrix A (sparsity structure is always sam…

---

## [Help improving the performance of my implementation of a lower diagonal storage](https://discourse.julialang.org/t/help-improving-the-performance-of-my-implementation-of-a-lower-diagonal-storage/132867)

<div class="topic-metadata">

**Author:** [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Replies:** 15\
**Last updated:** [October 6, 2025, 3:29pm UTC](https://discourse.julialang.org/t/help-improving-the-performance-of-my-implementation-of-a-lower-diagonal-storage/132867 "2025-10-06T15:29:55Z")

</div>

Hi! I am trying to build a lower diagonal storage to reduce the memory footprint of some applications. In my case, when computing spherical harmonics, the coefficients are usual placed in a lower triangular matrix. This…

---

## [@debug allocate even when not in DEBUG mode](https://discourse.julialang.org/t/debug-allocate-even-when-not-in-debug-mode/132854)

<div class="topic-metadata">

**Author:** [@zmoitier](https://discourse.julialang.org/u/zmoitier)\
**Replies:** 6\
**Last updated:** [October 4, 2025, 5:56pm UTC](https://discourse.julialang.org/t/debug-allocate-even-when-not-in-debug-mode/132854 "2025-10-04T17:56:54Z")

</div>

I would like to add debug statement in a function that is meant to be call in loops, so I would like the performance to be independent of the @debug statement when not in DEBUG mode. As an MRE, I have the function foo, a…

---

## [How to get vectorized code for single-field structs?](https://discourse.julialang.org/t/how-to-get-vectorized-code-for-single-field-structs/132875)

<div class="topic-metadata">

**Author:** [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Replies:** 1\
**Last updated:** [October 4, 2025, 3:36pm UTC](https://discourse.julialang.org/t/how-to-get-vectorized-code-for-single-field-structs/132875 "2025-10-04T15:36:25Z")

</div>

I have noticed that structs with a single field do not vectorize the same way as the underlying type. Here is an example: struct A x::Int end function f(s::NTuple{N,T}, t::NTuple{N,T}, i) where {N,T} ntuple(j -\>…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=3)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=5)
