# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=5

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 6

---

## [Efficient automatic differentation for Julia version \`jax.scan\`?](https://discourse.julialang.org/t/efficient-automatic-differentation-for-julia-version-jax-scan/132853)

<div class="topic-metadata">

**Author:** [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Replies:** 4\
**Last updated:** [October 3, 2025, 4:14pm UTC](https://discourse.julialang.org/t/efficient-automatic-differentation-for-julia-version-jax-scan/132853 "2025-10-03T16:14:03Z")

</div>

Hi, I am using in some JAX array code jax.lax.scan which implements something like this in my example: function jax\_lax\_scan(f, x; accumulator\_init) acc = accumulator\_init for i in axes(x, 1) acc = …

---

## [Nerd-sniping: can you make this faster?](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793)

<div class="topic-metadata">

**Author:** [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Replies:** 38\
**Last updated:** [October 2, 2025, 9:16pm UTC](https://discourse.julialang.org/t/nerd-sniping-can-you-make-this-faster/132793 "2025-10-02T21:16:14Z")

</div>

This is the core function of a module to compute solvent accessible surface areas of molecules. It works fine, but it would benefit of being faster. Without a full reworking of the data structures, etc, performance boil…

---

## [\[ANN\] \`Undigits.jl\` undoes \`Base.digits\` (not registered)](https://discourse.julialang.org/t/ann-undigits-jl-undoes-base-digits-not-registered/132791)

<div class="topic-metadata">

**Author:** [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Replies:** 2\
**Last updated:** [September 30, 2025, 6:21pm UTC](https://discourse.julialang.org/t/ann-undigits-jl-undoes-base-digits-not-registered/132791 "2025-09-30T18:21:34Z")

</div>

Undigits.jl This package provides undigits, the reverse of digits. julia\> undigits(\[2, 4\]) 42 julia\> undigits(\[2, 4\]; rev=true) 24 julia\> undigits(UInt8, \[1, 0, 1, 1\]; base=2) 0x0d julia\> n = UInt16(17); undigits(ty…

---

## [\[ANN\] InverseGammaFunction.jl (not registered)](https://discourse.julialang.org/t/ann-inversegammafunction-jl-not-registered/132661)

<div class="topic-metadata">

**Author:** [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Replies:** 9\
**Last updated:** [September 30, 2025, 5:40pm UTC](https://discourse.julialang.org/t/ann-inversegammafunction-jl-not-registered/132661 "2025-09-30T17:40:34Z")

</div>

This package provides a function to compute the inverse of the gamma function for positive (real) arguments. It supports both branches. It is a very quick implementation using root finding. I don’t plan to register it. …

---

## [Allocation Behavior Depending on Code Size](https://discourse.julialang.org/t/allocation-behavior-depending-on-code-size/132576)

<div class="topic-metadata">

**Author:** [@mraic1](https://discourse.julialang.org/u/mraic1)\
**Replies:** 16\
**Last updated:** [September 25, 2025, 3:36pm UTC](https://discourse.julialang.org/t/allocation-behavior-depending-on-code-size/132576 "2025-09-25T15:36:24Z")

</div>

Dear all, I am experimenting a bit on stream processing. To cope with this, Julia has a beautifully designed system of channels and tasks. However, the latter appear rather slow, even very slow when the channels are not…

---

## [Can I avoid dynamic dispatch here ? (recursive parameter in struct )](https://discourse.julialang.org/t/can-i-avoid-dynamic-dispatch-here-recursive-parameter-in-struct/132609)

<div class="topic-metadata">

**Author:** [@filchristou](https://discourse.julialang.org/u/filchristou)\
**Replies:** 6\
**Last updated:** [September 24, 2025, 4:01pm UTC](https://discourse.julialang.org/t/can-i-avoid-dynamic-dispatch-here-recursive-parameter-in-struct/132609 "2025-09-24T16:01:05Z")

</div>

There has been a while that I sense I might have some misconceptions about dynamic dispatch, and maybe this use case represents my doubts well.. So.. find myself in the unpleasant situation of having to use abstract con…

---

## [How to implement performance regression tests?](https://discourse.julialang.org/t/how-to-implement-performance-regression-tests/132578)

<div class="topic-metadata">

**Author:** [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Replies:** 5\
**Last updated:** [September 23, 2025, 9:54pm UTC](https://discourse.julialang.org/t/how-to-implement-performance-regression-tests/132578 "2025-09-23T21:54:26Z")

</div>

I added a performance test to my package KiteModels.jl: Benchmark simplify operation by ufechner7 · Pull Request #250 · OpenSourceAWE/KiteModels.jl · GitHub. If you run the script, you get an output like: julia\> includ…

---

## [Apple M4 Max AMX Linear Algebra performance versus CPU and GPU](https://discourse.julialang.org/t/apple-m4-max-amx-linear-algebra-performance-versus-cpu-and-gpu/132537)

<div class="topic-metadata">

**Author:** [@PetarM](https://discourse.julialang.org/u/PetarM)\
**Replies:** 2\
**Last updated:** [September 22, 2025, 12:17am UTC](https://discourse.julialang.org/t/apple-m4-max-amx-linear-algebra-performance-versus-cpu-and-gpu/132537 "2025-09-22T00:17:46Z")

</div>

Hello everybody! I made a video investigation of Apple’s AMX accelerator Linear Algebra performance in Julia: https://www.youtube.com/watch?v=TjfA9LVgHXk According to my findings, the 2 AMX cores achieve almost 3 times…

---

## [Multi-threaded Array building](https://discourse.julialang.org/t/multi-threaded-array-building/132463)

<div class="topic-metadata">

**Author:** [@FerreolS](https://discourse.julialang.org/u/FerreolS)\
**Replies:** 22\
**Last updated:** [September 20, 2025, 6:03pm UTC](https://discourse.julialang.org/t/multi-threaded-array-building/132463 "2025-09-20T18:03:52Z")

</div>

Dear all, I have a code of this kind to build an array: N= 2000 A = zeros(Float64,N,N) B # some long (\>\>10\_000) vector of custom struct for idx in arrayofindex #\> lots (\>\>10\_000) of elements i,j,v = processing(B\[…

---

## [OhMyThreads, ChunkSplitters, and Cost Estimates](https://discourse.julialang.org/t/ohmythreads-chunksplitters-and-cost-estimates/132449)

<div class="topic-metadata">

**Author:** [@afleming](https://discourse.julialang.org/u/afleming)\
**Replies:** 2\
**Last updated:** [September 16, 2025, 9:27pm UTC](https://discourse.julialang.org/t/ohmythreads-chunksplitters-and-cost-estimates/132449 "2025-09-16T21:27:04Z")

</div>

I’ve run into a situation that I believe ought to be a “solved problem”, or at least have a canonical solution somewhere in the collective knowledge about OhMyThreads.jl and ChunkSplitters.jl. I have a large finite-volu…

---

## [Preventing Enzyme from differentiating through constant computations](https://discourse.julialang.org/t/preventing-enzyme-from-differentiating-through-constant-computations/132445)

<div class="topic-metadata">

**Author:** [@dameka](https://discourse.julialang.org/u/dameka)\
**Replies:** 1\
**Last updated:** [September 16, 2025, 5:58pm UTC](https://discourse.julialang.org/t/preventing-enzyme-from-differentiating-through-constant-computations/132445 "2025-09-16T17:58:27Z")

</div>

I am training a neural network by computing the gradient of the loss with Enzyme. Part of my computations involve associated Legendre polynomials. I did not believe these to be a crucial part of the differentiation becau…

---

## [Find the max element of a numeric vector iteratively with mask and early break](https://discourse.julialang.org/t/find-the-max-element-of-a-numeric-vector-iteratively-with-mask-and-early-break/130233)

<div class="topic-metadata">

**Author:** [@WalterMadelim](https://discourse.julialang.org/u/WalterMadelim)\
**Replies:** 10\
**Last updated:** [September 16, 2025, 9:01am UTC](https://discourse.julialang.org/t/find-the-max-element-of-a-numeric-vector-iteratively-with-mask-and-early-break/130233 "2025-09-16T09:01:36Z")

</div>

This is an elementary question about algorithm. My algorithm is probably already usable but I wonder if there is a faster algorithm that can be implemented with Julia (e.g. even not using findmax) import Random; Rando…

---

## [Identical method redefinition suspiciously optimizes runtime and allocations](https://discourse.julialang.org/t/identical-method-redefinition-suspiciously-optimizes-runtime-and-allocations/132156)

<div class="topic-metadata">

**Author:** [@afleming](https://discourse.julialang.org/u/afleming)\
**Replies:** 25\
**Last updated:** [September 15, 2025, 12:39am UTC](https://discourse.julialang.org/t/identical-method-redefinition-suspiciously-optimizes-runtime-and-allocations/132156 "2025-09-15T00:39:19Z")

</div>

I am using Revise.jl to work on a package for my research, and am noticing that “revising” a file to identical code with different variable names causes a function to stop allocating. I am not sure how many details othe…

---

## [Praise: CUDA.allowscalar(false) is great](https://discourse.julialang.org/t/praise-cuda-allowscalar-false-is-great/132355)

<div class="topic-metadata">

**Author:** [@xzackli](https://discourse.julialang.org/u/xzackli)\
**Replies:** 0\
**Last updated:** [September 13, 2025, 6:22pm UTC](https://discourse.julialang.org/t/praise-cuda-allowscalar-false-is-great/132355 "2025-09-13T18:22:54Z")

</div>

One single line makes me substantially more productive (at writing GPU code), and I’d like to highlight how much I love this pattern. I’ve had vague ideas on how to express this kind of concept before, but CUDA.allowscal…

---

## [Taking power with float exponent x^y is slower than exp(y\*log(x))?](https://discourse.julialang.org/t/taking-power-with-float-exponent-x-y-is-slower-than-exp-y-log-x/132270)

<div class="topic-metadata">

**Author:** [@Norman](https://discourse.julialang.org/u/Norman)\
**Replies:** 8\
**Last updated:** [September 11, 2025, 4:27pm UTC](https://discourse.julialang.org/t/taking-power-with-float-exponent-x-y-is-slower-than-exp-y-log-x/132270 "2025-09-11T16:27:28Z")

</div>

Updated results: There was some sloppiness on my benchmark approach. The discrepancy isn’t that large. julia\> @benchmark x^1.5 setup=(x=rand()) samples=100000 BenchmarkTools.Trial: 100000 samples with 998 evaluations pe…

---

## [Strange slowdown with @threads :greedy with BigInts](https://discourse.julialang.org/t/strange-slowdown-with-threads-greedy-with-bigints/132211)

<div class="topic-metadata">

**Author:** [@sumiya11](https://discourse.julialang.org/u/sumiya11)\
**Replies:** 7\
**Last updated:** [September 9, 2025, 5:06pm UTC](https://discourse.julialang.org/t/strange-slowdown-with-threads-greedy-with-bigints/132211 "2025-09-09T17:06:18Z")

</div>

Hi, I have a function that does inplace operations with BigInts (so almost no Julia-side allocations in @time report, see below). With --threads=8 it is twice slower than with --threads=1. Note: seems fine without :gree…

---

## [Error in program for coupled PDE](https://discourse.julialang.org/t/error-in-program-for-coupled-pde/132200)

<div class="topic-metadata">

**Author:** [@Reinoud](https://discourse.julialang.org/u/Reinoud)\
**Replies:** 0\
**Last updated:** [September 8, 2025, 9:00pm UTC](https://discourse.julialang.org/t/error-in-program-for-coupled-pde/132200 "2025-09-08T21:00:50Z")

</div>

Hello Could somebody tell me what went wrong? Here the program using OrdinaryDiffEq, ModelingToolkit, MethodOfLines, DomainSets, Plots, Optimization, OptimizationPolyalgorithms, SciMLSensitivity ForwardDiff,JLD2 # Par…

---

## [\`log\` calling \`fma\_emulated\` on hardware that doesn't need it](https://discourse.julialang.org/t/log-calling-fma-emulated-on-hardware-that-doesnt-need-it/132149)

<div class="topic-metadata">

**Author:** [@cscherrer](https://discourse.julialang.org/u/cscherrer)\
**Replies:** 2\
**Last updated:** [September 7, 2025, 1:43am UTC](https://discourse.julialang.org/t/log-calling-fma-emulated-on-hardware-that-doesnt-need-it/132149 "2025-09-07T01:43:47Z")

</div>

Hi, I was doing some profiling for an application and came across this for log It should be able to use the real FMA, not emulated. It’s a newish machine: julia\> versioninfo() Julia Version 1.11.6 Commit 9615af0f26…

---

## [Why does \`Union{Int,...}\` in a \`Vector\` not cause allocations but \`Union{Float64,...}\` does?](https://discourse.julialang.org/t/why-does-union-int-in-a-vector-not-cause-allocations-but-union-float64-does/132067)

<div class="topic-metadata">

**Author:** [@jules](https://discourse.julialang.org/u/jules)\
**Replies:** 8\
**Last updated:** [September 3, 2025, 11:44am UTC](https://discourse.julialang.org/t/why-does-union-int-in-a-vector-not-cause-allocations-but-union-float64-does/132067 "2025-09-03T11:44:22Z")

</div>

I have a custom missing type which stores a symbol and a reference to an object whose type is not predetermined, so it’s an Any field. I need to store vectors of these, some values will be missing but most won’t be. Here…

---

## [Multi-threaded processing of a Dict](https://discourse.julialang.org/t/multi-threaded-processing-of-a-dict/131996)

<div class="topic-metadata">

**Author:** [@jlbosse](https://discourse.julialang.org/u/jlbosse)\
**Replies:** 8\
**Last updated:** [September 1, 2025, 2:09pm UTC](https://discourse.julialang.org/t/multi-threaded-processing-of-a-dict/131996 "2025-09-01T14:09:08Z")

</div>

I have a function that iterates over all key-value pairs, does a computation based on the key and the value and then updates the value at that key. Due to the fact that in each iteration I only update the value and don’t…

---

## [SIMD.jl - vload() continuous blocks from higher-dimensional arrays?](https://discourse.julialang.org/t/simd-jl-vload-continuous-blocks-from-higher-dimensional-arrays/131981)

<div class="topic-metadata">

**Author:** [@jl\_enthusiast](https://discourse.julialang.org/u/jl_enthusiast)\
**Replies:** 4\
**Last updated:** [August 31, 2025, 9:45pm UTC](https://discourse.julialang.org/t/simd-jl-vload-continuous-blocks-from-higher-dimensional-arrays/131981 "2025-08-31T21:45:29Z")

</div>

Hi, I am currently doing some practice projects in Julia, which involves using SIMD.jl for writing/generating microkernels. My issue is trying to determine what is the “best” way of loading Vec types from a higher-dimen…

---

## [Using control flow in Reactant](https://discourse.julialang.org/t/using-control-flow-in-reactant/131946)

<div class="topic-metadata">

**Author:** [@dameka](https://discourse.julialang.org/u/dameka)\
**Replies:** 4\
**Last updated:** [August 29, 2025, 6:20pm UTC](https://discourse.julialang.org/t/using-control-flow-in-reactant/131946 "2025-08-29T18:20:58Z")

</div>

I’ve got some expensive code I’m running that I’d like to compile with Reactant. There is some control flow that’s not playing nice with Reactant. I’ve got an MWE producing a different error than my main code, but I’m ho…

---

## [Large execution time jumps](https://discourse.julialang.org/t/large-execution-time-jumps/131924)

<div class="topic-metadata">

**Author:** [@harri37](https://discourse.julialang.org/u/harri37)\
**Replies:** 2\
**Last updated:** [August 29, 2025, 9:26am UTC](https://discourse.julialang.org/t/large-execution-time-jumps/131924 "2025-08-29T09:26:36Z")

</div>

Hello everyone, I am working on building a computer algebra library in Julia and have encountered some strange behavior in how the execution time increases when I multiply polynomials. Essentially the code will execute…

---

## [Using Reactant with Lux and Enzyme to speed up training in physics context](https://discourse.julialang.org/t/using-reactant-with-lux-and-enzyme-to-speed-up-training-in-physics-context/131898)

<div class="topic-metadata">

**Author:** [@dameka](https://discourse.julialang.org/u/dameka)\
**Replies:** 16\
**Last updated:** [August 28, 2025, 8:56pm UTC](https://discourse.julialang.org/t/using-reactant-with-lux-and-enzyme-to-speed-up-training-in-physics-context/131898 "2025-08-28T20:56:28Z")

</div>

I am training a neural network in a physical context on observables; crucially, the network is producing values which are passed through some physics calculations before being compared to the training data. My code works…

---

## [Performance of \`exp(A)\` for 9x9 anti-Hermitian matrix: Julia vs. PyTorch vs. MATLAB (CPU & GPU)](https://discourse.julialang.org/t/performance-of-exp-a-for-9x9-anti-hermitian-matrix-julia-vs-pytorch-vs-matlab-cpu-gpu/131696)

<div class="topic-metadata">

**Author:** [@draftman9](https://discourse.julialang.org/u/draftman9)\
**Replies:** 29\
**Last updated:** [August 28, 2025, 2:25am UTC](https://discourse.julialang.org/t/performance-of-exp-a-for-9x9-anti-hermitian-matrix-julia-vs-pytorch-vs-matlab-cpu-gpu/131696 "2025-08-28T02:25:07Z")

</div>

Hi everyone, I’m a Ph.D. student in physics. A core part of my research involves simulating the time evolution of quantum systems, which boils down to frequently calculating the matrix exponential U = exp(A). For these…

---

## [Latex fonts on axis and labels in Makie](https://discourse.julialang.org/t/latex-fonts-on-axis-and-labels-in-makie/131887)

<div class="topic-metadata">

**Author:** [@tduretz](https://discourse.julialang.org/u/tduretz)\
**Replies:** 7\
**Last updated:** [August 27, 2025, 3:33pm UTC](https://discourse.julialang.org/t/latex-fonts-on-axis-and-labels-in-makie/131887 "2025-08-27T15:33:09Z")

</div>

I would like to use LaTeX fonts for every text of a Makie figure. It’s easy to apply them to labels as LaTeXStrings can be used. For other texts, in particular axis tick labels, it a bit less trivial. I was once told …

---

## [Dagger not fully utilizing CPU cores](https://discourse.julialang.org/t/dagger-not-fully-utilizing-cpu-cores/131733)

<div class="topic-metadata">

**Author:** [@Miroboru](https://discourse.julialang.org/u/Miroboru)\
**Replies:** 5\
**Last updated:** [August 21, 2025, 1:05pm UTC](https://discourse.julialang.org/t/dagger-not-fully-utilizing-cpu-cores/131733 "2025-08-21T13:05:43Z")

</div>

I have a time consuming simulation that is run over a set of signal to noise ratios, and thus is trivial to parallelize. I am currently using Dagger.jl to start 24 processes (I am on a 24 core threadripper system with 25…

---

## [When does it make sense to use \`Base.MultiplicativeInverses\`](https://discourse.julialang.org/t/when-does-it-make-sense-to-use-base-multiplicativeinverses/131711)

<div class="topic-metadata">

**Author:** [@kunzaatko](https://discourse.julialang.org/u/kunzaatko)\
**Replies:** 6\
**Last updated:** [August 19, 2025, 4:18pm UTC](https://discourse.julialang.org/t/when-does-it-make-sense-to-use-base-multiplicativeinverses/131711 "2025-08-19T16:18:05Z")

</div>

When does it make sense to use the Base.MultiplicativeInverses? I am considering storing the MultiplicativeInverse in my AbstractArray subtype which uses some index calculations based on some fields and the parent array.…

---

## [Need perf help on paralel var](https://discourse.julialang.org/t/need-perf-help-on-paralel-var/131716)

<div class="topic-metadata">

**Author:** [@yolhan\_mannes](https://discourse.julialang.org/u/yolhan_mannes)\
**Replies:** 4\
**Last updated:** [August 19, 2025, 9:36pm UTC](https://discourse.julialang.org/t/need-perf-help-on-paralel-var/131716 "2025-08-19T21:36:39Z")

</div>

Hello, I’m trying to add Statistics to AcceleratedKernels.jl in Add Statistics in KA, (only mean and var implemented) by yolhan83 · Pull Request #64 · JuliaGPU/AcceleratedKernels.jl · GitHub, GPU results seems nice but b…

---

## [Abusing \`convert\` as an alternative to \`Union{Nothing, Int64}\`](https://discourse.julialang.org/t/abusing-convert-as-an-alternative-to-union-nothing-int64/131534)

<div class="topic-metadata">

**Author:** [@nhz2](https://discourse.julialang.org/u/nhz2)\
**Replies:** 15\
**Last updated:** [August 19, 2025, 8:14pm UTC](https://discourse.julialang.org/t/abusing-convert-as-an-alternative-to-union-nothing-int64/131534 "2025-08-19T20:14:52Z")

</div>

In C it is common to have functions that return a positive integer result on success, or a negative integer if an error happens. When I use these functions, I tend to forget the possible error values and do integer arith…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=4)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=6)
