# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=14

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 15

---

## [Avoiding allocations and run-time dispatch with ForwardDiff](https://discourse.julialang.org/t/avoiding-allocations-and-run-time-dispatch-with-forwarddiff/123699)

<div class="topic-metadata">

**Author:** [@ohmsweetohm1](https://discourse.julialang.org/u/ohmsweetohm1)\
**Replies:** 2\
**Last updated:** [December 11, 2024, 1:29pm UTC](https://discourse.julialang.org/t/avoiding-allocations-and-run-time-dispatch-with-forwarddiff/123699 "2024-12-11T13:29:24Z")

</div>

I have an optimization problem which involves solving an ODE and finding the input to the “system” such that an objective is minimized. A MWE of a toy problem is below: import Optimization import ForwardDiff import Opti…

---

## [Accelerate column-wise subtraction of a vector to a matrix](https://discourse.julialang.org/t/accelerate-column-wise-subtraction-of-a-vector-to-a-matrix/123619)

<div class="topic-metadata">

**Author:** [@matnbo](https://discourse.julialang.org/u/matnbo)\
**Replies:** 10\
**Last updated:** [December 10, 2024, 8:06am UTC](https://discourse.julialang.org/t/accelerate-column-wise-subtraction-of-a-vector-to-a-matrix/123619 "2024-12-10T08:06:27Z")

</div>

I try to accelerate the computation of the following operation: I have a matrix X (n, p) and a vector v (p). For each j {= 1,…,p}, I want to compute the differences X\[i, j\] - v\[j\] {i = 1,…,n}. and, this, on CPU (not G…

---

## [Mutable structs with all constant fields outperform immutable structs for equality comparison](https://discourse.julialang.org/t/mutable-structs-with-all-constant-fields-outperform-immutable-structs-for-equality-comparison/123430)

<div class="topic-metadata">

**Author:** [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Replies:** 1\
**Last updated:** [December 4, 2024, 12:13am UTC](https://discourse.julialang.org/t/mutable-structs-with-all-constant-fields-outperform-immutable-structs-for-equality-comparison/123430 "2024-12-04T00:13:04Z")

</div>

MWE: julia\> x1 = \[rand(Int) for \_=1:1000\]; julia\> x2 = \[rand(Int) for \_=1:1000\]; julia\> v1 = \[rand(3,3) for \_=1:1000\]; julia\> v2 = \[rand(3,3) for \_=1:1000\]; julia\> struct myT a::Int b::Any end …

---

## [Simple recursive Fibonaci example. How to make it faster?](https://discourse.julialang.org/t/simple-recursive-fibonaci-example-how-to-make-it-faster/32369)

<div class="topic-metadata">

**Author:** [@Juan](https://discourse.julialang.org/u/Juan)\
**Replies:** 40\
**Last updated:** [December 5, 2024, 12:26pm UTC](https://discourse.julialang.org/t/simple-recursive-fibonaci-example-how-to-make-it-faster/32369 "2024-12-05T12:26:10Z")

</div>

I’ve found this benchmark comparing several languages with a simple Fibonaci code. https://github.com/drujensen/fib function fib(n) if n \<= 1 return 1 end return fib(n - 1) + fib(n - 2) end @btime(fib(46)) Julia …

---

## [OOM and saving solutions of \`SDEProblem\` only at specific timepoints](https://discourse.julialang.org/t/oom-and-saving-solutions-of-sdeproblem-only-at-specific-timepoints/123450)

<div class="topic-metadata">

**Author:** [@johannesnauta](https://discourse.julialang.org/u/johannesnauta)\
**Replies:** 2\
**Last updated:** [December 5, 2024, 9:33am UTC](https://discourse.julialang.org/t/oom-and-saving-solutions-of-sdeproblem-only-at-specific-timepoints/123450 "2024-12-05T09:33:33Z")

</div>

For large systems, I am running into memory issues (which lead to crashes) when solving an SDEProblem with an adaptive scheme. I have found that these are related to the dense output of the solver, yet I just cannot seem…

---

## [Numpy.sort vs Julia sort](https://discourse.julialang.org/t/numpy-sort-vs-julia-sort/123421)

<div class="topic-metadata">

**Author:** [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Replies:** 10\
**Last updated:** [December 4, 2024, 10:17pm UTC](https://discourse.julialang.org/t/numpy-sort-vs-julia-sort/123421 "2024-12-04T22:17:08Z")

</div>

I recently tried numpy.sort and it is really fast for a vector of 100 millions floats. It seems to be parallelized by default and/or using AVX-512. It is like 5 times faster on my system. If someone has a benchmark vs J…

---

## [Experiments with LoopVectorization and convolutions](https://discourse.julialang.org/t/experiments-with-loopvectorization-and-convolutions/123188)

<div class="topic-metadata">

**Author:** [@Marco\_Lombardi](https://discourse.julialang.org/u/Marco_Lombardi)\
**Replies:** 24\
**Last updated:** [December 3, 2024, 8:35pm UTC](https://discourse.julialang.org/t/experiments-with-loopvectorization-and-convolutions/123188 "2024-12-03T20:35:27Z")

</div>

I have been doing some experiments with the great LoopVectorization for convolution, and I must admit that in spite of my efforts to manually refactor my code, I have been unable to even approach the speed that this pack…

---

## [We love sorting -- latest news in sorting](https://discourse.julialang.org/t/we-love-sorting-latest-news-in-sorting/100790)

<div class="topic-metadata">

**Author:** [@blackeneth](https://discourse.julialang.org/u/blackeneth)\
**Replies:** 0\
**Last updated:** [June 24, 2023, 6:44pm UTC](https://discourse.julialang.org/t/we-love-sorting-latest-news-in-sorting/100790 "2023-06-24T18:44:37Z")

</div>

This post was inspired by this news article on Intel releasing version 2.0 of their SIMD sort code: Julia has something similar to this: which was presented at the 2019 JuliaCon. ChipSort, however, isn’t limited t…

---

## [Mesh 3D Creation](https://discourse.julialang.org/t/mesh-3d-creation/123414)

<div class="topic-metadata">

**Author:** [@Kamran\_Ali](https://discourse.julialang.org/u/Kamran_Ali)\
**Replies:** 4\
**Last updated:** [December 3, 2024, 4:33pm UTC](https://discourse.julialang.org/t/mesh-3d-creation/123414 "2024-12-03T16:33:41Z")

</div>

I have created a 3D Quad mesh for my model, created vertices and faces for bottom layer, top layer vertices, now i have to make it 3D, so if i apply the loft command, corner faces are created in the diagonal, issue is t…

---

## [Vector of Vectors / Vector of Tuples](https://discourse.julialang.org/t/vector-of-vectors-vector-of-tuples/123368)

<div class="topic-metadata">

**Author:** [@hssn15](https://discourse.julialang.org/u/hssn15)\
**Replies:** 6\
**Last updated:** [December 3, 2024, 4:27pm UTC](https://discourse.julialang.org/t/vector-of-vectors-vector-of-tuples/123368 "2024-12-03T16:27:49Z")

</div>

I am making a simulator and need to store data in pre-allocated storages. But, I am a bit confused and worried about the performance. I have two ways to create storage element. First way: storage = \[ (x1, y…

---

## [Benchmark is moving target?](https://discourse.julialang.org/t/benchmark-is-moving-target/123376)

<div class="topic-metadata">

**Author:** [@tk3369](https://discourse.julialang.org/u/tk3369)\
**Replies:** 2\
**Last updated:** [December 3, 2024, 10:06am UTC](https://discourse.julialang.org/t/benchmark-is-moving-target/123376 "2024-12-03T10:06:07Z")

</div>

I am doing Advent of Code again this year (not sure how long I’d stay haha) and tried to optimize a very simple function. Surprisingly, @btime is showing different number of allocations on 2nd run. Also, when I redefined…

---

## [Static array multiplied by its adjoint seems to allocate](https://discourse.julialang.org/t/static-array-multiplied-by-its-adjoint-seems-to-allocate/123232)

<div class="topic-metadata">

**Author:** [@miguelborrero](https://discourse.julialang.org/u/miguelborrero)\
**Replies:** 12\
**Last updated:** [December 2, 2024, 3:54pm UTC](https://discourse.julialang.org/t/static-array-multiplied-by-its-adjoint-seems-to-allocate/123232 "2024-12-02T15:54:39Z")

</div>

Hi there, I am writing out a likelihood estimation and I have run into some allocation issues when coding up the Hessian. Even though the Hessian function HLL seems to mirror what the gradient function ∇LL is doing. One…

---

## [Getting to zero allocations in NonlinearSolve.jl](https://discourse.julialang.org/t/getting-to-zero-allocations-in-nonlinearsolve-jl/123123)

<div class="topic-metadata">

**Author:** [@rafaelbailo](https://discourse.julialang.org/u/rafaelbailo)\
**Replies:** 9\
**Last updated:** [November 29, 2024, 6:05am UTC](https://discourse.julialang.org/t/getting-to-zero-allocations-in-nonlinearsolve-jl/123123 "2024-11-29T06:05:30Z")

</div>

I am in the process of migrating some code from NLsolve.jl to NonlinearSolve.jl. I have written a simple Newton-Raphson test, but I cannot get the allocations down to zero; both solve! and reinit! allocate. Can this be i…

---

## [Type stability in closures](https://discourse.julialang.org/t/type-stability-in-closures/123227)

<div class="topic-metadata">

**Author:** [@AntonReinhard](https://discourse.julialang.org/u/AntonReinhard)\
**Replies:** 13\
**Last updated:** [November 28, 2024, 11:00pm UTC](https://discourse.julialang.org/t/type-stability-in-closures/123227 "2024-11-28T23:00:13Z")

</div>

I’m doing a bunch of code generation and came across a strange behaviour involving type stability, closures and variable captures. I reduced the problem to the following example: function f(a, b) \_f = (x, y) -\> begi…

---

## [\`IdDict{UInt64, Float64}\` is much slower than \`Dict{UInt64, Float64}\` for retrieving values](https://discourse.julialang.org/t/iddict-uint64-float64-is-much-slower-than-dict-uint64-float64-for-retrieving-values/123144)

<div class="topic-metadata">

**Author:** [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Replies:** 2\
**Last updated:** [November 27, 2024, 8:36pm UTC](https://discourse.julialang.org/t/iddict-uint64-float64-is-much-slower-than-dict-uint64-float64-for-retrieving-values/123144 "2024-11-27T20:36:16Z")

</div>

MWE: julia\> k = rand(UInt, 5000); julia\> v = rand(Float64, 5000); julia\> using Random julia\> ks = shuffle(k); julia\> d1 = Dict(k .=\> v); julia\> d2 = IdDict(k .=\> v); julia\> using BenchmarkTools julia\> @benchmark …

---

## [Why JLD2.jl is 40x slower than Arrow.jl](https://discourse.julialang.org/t/why-jld2-jl-is-40x-slower-than-arrow-jl/122217)

<div class="topic-metadata">

**Author:** [@Sixzero](https://discourse.julialang.org/u/Sixzero)\
**Replies:** 24\
**Last updated:** [November 25, 2024, 9:59am UTC](https://discourse.julialang.org/t/why-jld2-jl-is-40x-slower-than-arrow-jl/122217 "2024-11-25T09:59:38Z")

</div>

I have this gist: JLD2 is actually simply for reading in a 30 MB file with keys =\> randn(500) is 40x slower than Arrow.jl. How? I mean how come JLD2 generality sacrificed 40x speed difference?

---

## [Using a runtime-length \`SVector\` within a closure allocates memory](https://discourse.julialang.org/t/using-a-runtime-length-svector-within-a-closure-allocates-memory/123020)

<div class="topic-metadata">

**Author:** [@ferrolho](https://discourse.julialang.org/u/ferrolho)\
**Replies:** 6\
**Last updated:** [November 24, 2024, 5:40pm UTC](https://discourse.julialang.org/t/using-a-runtime-length-svector-within-a-closure-allocates-memory/123020 "2024-11-24T17:40:15Z")

</div>

Consider the Test struct below. julia\> using StaticArrays julia\> struct Test{T,V} x::Int v::V Test(x::Int, v::V) where {T\<:Real,V\<:AbstractVector{T}} = new{T,V}(x, v) end Let’s …

---

## [Capturing "sub" dependencies during precompilation](https://discourse.julialang.org/t/capturing-sub-dependencies-during-precompilation/123013)

<div class="topic-metadata">

**Author:** [@sstroemer](https://discourse.julialang.org/u/sstroemer)\
**Replies:** 0\
**Last updated:** [November 24, 2024, 10:46am UTC](https://discourse.julialang.org/t/capturing-sub-dependencies-during-precompilation/123013 "2024-11-24T10:46:50Z")

</div>

I have a package X which depends on Y which depends on Z. Z is not a direct dependency of X, and Y is done by a “third-party”, so I can’t influence it. I’m using some example code @compile\_workload, but still see a lot …

---

## [Help improving particle simulation on GPU](https://discourse.julialang.org/t/help-improving-particle-simulation-on-gpu/122733)

<div class="topic-metadata">

**Author:** [@rveltz](https://discourse.julialang.org/u/rveltz)\
**Replies:** 0\
**Last updated:** [November 17, 2024, 8:23am UTC](https://discourse.julialang.org/t/help-improving-particle-simulation-on-gpu/122733 "2024-11-17T08:23:00Z")

</div>

Hi, I am trying to simulate particles on CUDA. I have a bottleneck and I am wondering if some of you have any advice. This is a MWE of a much more sophisticated example. Basically, I have a field of probabiblities dist…

---

## [Checked vs unchecked, e.g. (floored) divison](https://discourse.julialang.org/t/checked-vs-unchecked-e-g-floored-divison/122899)

<div class="topic-metadata">

**Author:** [@Palli](https://discourse.julialang.org/u/Palli)\
**Replies:** 1\
**Last updated:** [November 21, 2024, 4:07pm UTC](https://discourse.julialang.org/t/checked-vs-unchecked-e-g-floored-divison/122899 "2024-11-21T16:07:37Z")

</div>

Continuing the discussion from Why are there all these strange stumbling blocks in Julia?: I first continue, answer here by forking old discussion (by clicking corner in top left corner, for the first time, good to know…

---

## [Allocations when using getfield with a tuple/vector of symbols](https://discourse.julialang.org/t/allocations-when-using-getfield-with-a-tuple-vector-of-symbols/122756)

<div class="topic-metadata">

**Author:** [@Tetrakai](https://discourse.julialang.org/u/Tetrakai)\
**Replies:** 12\
**Last updated:** [November 20, 2024, 2:45am UTC](https://discourse.julialang.org/t/allocations-when-using-getfield-with-a-tuple-vector-of-symbols/122756 "2024-11-20T02:45:24Z")

</div>

MWE: using Parameters, BenchmarkTools @with\_kw struct Test A :: String = "A" B :: Int64 = 0 end test = Test() tup = (:A, :B) @btime getfield($test, :A) @btime getfield($test, $tup\[1\]) Is there a way to loop…

---

## [Most efficient way of adding elements within matrices in loops](https://discourse.julialang.org/t/most-efficient-way-of-adding-elements-within-matrices-in-loops/122789)

<div class="topic-metadata">

**Author:** [@BMI\_OR](https://discourse.julialang.org/u/BMI_OR)\
**Replies:** 4\
**Last updated:** [November 18, 2024, 11:31pm UTC](https://discourse.julialang.org/t/most-efficient-way-of-adding-elements-within-matrices-in-loops/122789 "2024-11-18T23:31:41Z")

</div>

I have a vector A and I need to update its values with the ones of matrix B. Each element of B needs to be added to some values in A. The mapping of indeces from matrix to vector is given by a matrix of vectors C. I have…

---

## [Understanding DataFrame allocations](https://discourse.julialang.org/t/understanding-dataframe-allocations/122792)

<div class="topic-metadata">

**Author:** [@miguelborrero](https://discourse.julialang.org/u/miguelborrero)\
**Replies:** 1\
**Last updated:** [November 18, 2024, 10:16pm UTC](https://discourse.julialang.org/t/understanding-dataframe-allocations/122792 "2024-11-18T22:16:22Z")

</div>

Hi there, The split-apply-combine strategy is something that I use a lot so Im trying to understand a bit more what goes under the hood so that I can write better code. I was just looking at the following example: df =…

---

## [Julia access to Apple GPU with MLX, and or Metal Performance Shaders (MPS)?](https://discourse.julialang.org/t/julia-access-to-apple-gpu-with-mlx-and-or-metal-performance-shaders-mps/117647)

<div class="topic-metadata">

**Author:** [@pitsianis](https://discourse.julialang.org/u/pitsianis)\
**Replies:** 5\
**Last updated:** [November 18, 2024, 8:13pm UTC](https://discourse.julialang.org/t/julia-access-to-apple-gpu-with-mlx-and-or-metal-performance-shaders-mps/117647 "2024-11-18T20:13:22Z")

</div>

Is there any on-going effort to provide a Julia interface to the following? Metal Performance Shaders (MPS) is a PyTorch framework backend for GPU training acceleration, providing scripts and capabilities to set up an…

---

## [Why is my one loop faster than the other?](https://discourse.julialang.org/t/why-is-my-one-loop-faster-than-the-other/122771)

<div class="topic-metadata">

**Author:** [@J-J](https://discourse.julialang.org/u/J-J)\
**Replies:** 6\
**Last updated:** [November 18, 2024, 2:35pm UTC](https://discourse.julialang.org/t/why-is-my-one-loop-faster-than-the-other/122771 "2024-11-18T14:35:57Z")

</div>

As a MWE, I wrote the following two functions in Julia, both of which depend on a global constant integer N and a global constant vector global\_vector: using BenchmarkTools const N = 1024 const global\_vector = rand(N) …

---

## [Check if is folder empty](https://discourse.julialang.org/t/check-if-is-folder-empty/122761)

<div class="topic-metadata">

**Author:** [@hhaensel](https://discourse.julialang.org/u/hhaensel)\
**Replies:** 3\
**Last updated:** [November 18, 2024, 2:35pm UTC](https://discourse.julialang.org/t/check-if-is-folder-empty/122761 "2024-11-18T14:35:02Z")

</div>

What’s the best way to check if a folder is empty if folders possibly contains 10\_000 files? isempty(readdir(mydir)) The point is that I want to list all empty subfolders of a directory, but mapping readdir on all non-…

---

## [Allocations even when using StaticArrays.jl](https://discourse.julialang.org/t/allocations-even-when-using-staticarrays-jl/122775)

<div class="topic-metadata">

**Author:** [@johnmyslinski](https://discourse.julialang.org/u/johnmyslinski)\
**Replies:** 2\
**Last updated:** [November 18, 2024, 2:09pm UTC](https://discourse.julialang.org/t/allocations-even-when-using-staticarrays-jl/122775 "2024-11-18T14:09:43Z")

</div>

Hi All, In this below MRE I am still seeing some allocations when using @btime. Based on my understanding, I would expect to see 0 allocations. Am I misunderstanding something? Are the 8 allocations I see a byproduct …

---

## [Can the overhead of \`myT{\<:T}\` compared to \`myT{T}\`, where \`T\` is a concrete type, be avoided?](https://discourse.julialang.org/t/can-the-overhead-of-myt-t-compared-to-myt-t-where-t-is-a-concrete-type-be-avoided/122721)

<div class="topic-metadata">

**Author:** [@frankwswang](https://discourse.julialang.org/u/frankwswang)\
**Replies:** 3\
**Last updated:** [November 17, 2024, 11:05am UTC](https://discourse.julialang.org/t/can-the-overhead-of-myt-t-compared-to-myt-t-where-t-is-a-concrete-type-be-avoided/122721 "2024-11-17T11:05:17Z")

</div>

I understand that Vector{\<:Float64} is not the same as Vector{Float64} as it allows the bottom type Union{} to be its element type and is (therefore) not a concrete type: julia\> Vector{Union{}} \<: Vector{\<:Float64} true…

---

## [Creating an Iterable of Dicts](https://discourse.julialang.org/t/creating-an-iterable-of-dicts/122527)

<div class="topic-metadata">

**Author:** [@Niall](https://discourse.julialang.org/u/Niall)\
**Replies:** 15\
**Last updated:** [November 16, 2024, 7:44pm UTC](https://discourse.julialang.org/t/creating-an-iterable-of-dicts/122527 "2024-11-16T19:44:55Z")

</div>

Hi. I’m interested in learning more about creating Iterables to save space. In particular, right now I want to create an Iterable that runs through the rows of a truth table. The following code generates an array that (a…

---

## [Dimension mismatch error](https://discourse.julialang.org/t/dimension-mismatch-error/122590)

<div class="topic-metadata">

**Author:** [@Vegetasan](https://discourse.julialang.org/u/Vegetasan)\
**Replies:** 10\
**Last updated:** [November 15, 2024, 8:41am UTC](https://discourse.julialang.org/t/dimension-mismatch-error/122590 "2024-11-15T08:41:50Z")

</div>

Hello everyone, I’m working on a project in Julia to predict reaction mechanisms for a given set of inputs using kinetic modeling. I’m fairly new to Julia, and everything seems to be running well except for a dimension …

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=13)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=15)
