# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=8

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 9

---

## [State of closures, Fix1/Fix2](https://discourse.julialang.org/t/state-of-closures-fix1-fix2/118997)

<div class="topic-metadata">

**Author:** [@Deduction42](https://discourse.julialang.org/u/Deduction42)\
**Replies:** 16\
**Last updated:** [June 6, 2025, 10:31pm UTC](https://discourse.julialang.org/t/state-of-closures-fix1-fix2/118997 "2025-06-06T22:31:51Z")

</div>

I’ve been seeing more Base.Fix1 and Base.Fix2 inside code-bases these days, even though I thought most of these performance issues were fixed with v1.0. Apparently they aren’t? In Performance Tips it looks like we still …

---

## [Surprising slowdown from using array comprehension](https://discourse.julialang.org/t/surprising-slowdown-from-using-array-comprehension/129583)

<div class="topic-metadata">

**Author:** [@jwortmann](https://discourse.julialang.org/u/jwortmann)\
**Replies:** 8\
**Last updated:** [June 4, 2025, 10:01pm UTC](https://discourse.julialang.org/t/surprising-slowdown-from-using-array-comprehension/129583 "2025-06-04T22:01:56Z")

</div>

Hello! I recently came across a (for me) quite difficult to debug performance issue in my code, and it still leaves me somewhat puzzled. An algorithm that I wrote seems to experience severe slowdowns when using an array …

---

## [How to better initialize a dictionary containing multidimensional array?](https://discourse.julialang.org/t/how-to-better-initialize-a-dictionary-containing-multidimensional-array/129632)

<div class="topic-metadata">

**Author:** [@Harrykjg-physics](https://discourse.julialang.org/u/Harrykjg-physics)\
**Replies:** 10\
**Last updated:** [June 4, 2025, 11:52am UTC](https://discourse.julialang.org/t/how-to-better-initialize-a-dictionary-containing-multidimensional-array/129632 "2025-06-04T11:52:38Z")

</div>

Hi everyone, I want to create a Dict() object where the keys are NTuple{6, Int} and the corresponding values are rank-6 arrays. I found the following function to create such Dict object with 64 items are much slower than…

---

## [Sum over view of BitArray](https://discourse.julialang.org/t/sum-over-view-of-bitarray/129585)

<div class="topic-metadata">

**Author:** [@jerry\_ji](https://discourse.julialang.org/u/jerry_ji)\
**Replies:** 3\
**Last updated:** [June 3, 2025, 8:39am UTC](https://discourse.julialang.org/t/sum-over-view-of-bitarray/129585 "2025-06-03T08:39:39Z")

</div>

q=trues(10000) d=view(q,1:10000) @btime sum($q) 17.034 ns (0 allocations: 0 bytes) 10000 @btime sum($d) 2.678 μs (0 allocations: 0 bytes) 10000 how can i improve performance of sum over continuous subset of a Bit…

---

## [Optimized Python is as good as Julia](https://discourse.julialang.org/t/optimized-python-is-as-good-as-julia/104432)

<div class="topic-metadata">

**Author:** [@mrkorch](https://discourse.julialang.org/u/mrkorch)\
**Replies:** 28\
**Last updated:** [June 2, 2025, 11:55am UTC](https://discourse.julialang.org/t/optimized-python-is-as-good-as-julia/104432 "2025-06-02T11:55:55Z")

</div>

Comparison of Python / Julia benchmarks for a maths problem. I am considering switching to Julia, but would like to know if it is worth it for such tasks. Here is the Python code I used: import random import numpy as n…

---

## [How to target Intel performance cores only in multithreading](https://discourse.julialang.org/t/how-to-target-intel-performance-cores-only-in-multithreading/91415)

<div class="topic-metadata">

**Author:** [@babakfifoo](https://discourse.julialang.org/u/babakfifoo)\
**Replies:** 5\
**Last updated:** [June 1, 2025, 10:42am UTC](https://discourse.julialang.org/t/how-to-target-intel-performance-cores-only-in-multithreading/91415 "2025-06-01T10:42:07Z")

</div>

Hi all, I have a code that uses multithreading. My system has 12700K with 6/2 performance cores and 8 efficiency cores The issue I have is that I want to use only performance cores since efficiency cores are not strong…

---

## [Very long time for addition of complex identity matrix to transposed sparse matrix](https://discourse.julialang.org/t/very-long-time-for-addition-of-complex-identity-matrix-to-transposed-sparse-matrix/129465)

<div class="topic-metadata">

**Author:** [@rveltz](https://discourse.julialang.org/u/rveltz)\
**Replies:** 10\
**Last updated:** [May 31, 2025, 5:20pm UTC](https://discourse.julialang.org/t/very-long-time-for-addition-of-complex-identity-matrix-to-transposed-sparse-matrix/129465 "2025-05-31T17:20:12Z")

</div>

Hi, I am surprised by these timings on 1.10. Can it be improved? using SparseArrays, LinearAlgebra A = sprandn(90000,90000, 1e-5) @time (A + complex(0,1.) \* I); # 0.001344 seconds (13 allocations: 7.342 MiB) @time (tr…

---

## [Batched matrix-multiplication optimization](https://discourse.julialang.org/t/batched-matrix-multiplication-optimization/129454)

<div class="topic-metadata">

**Author:** [@quantumdreamer](https://discourse.julialang.org/u/quantumdreamer)\
**Replies:** 15\
**Last updated:** [May 30, 2025, 10:52pm UTC](https://discourse.julialang.org/t/batched-matrix-multiplication-optimization/129454 "2025-05-30T22:52:55Z")

</div>

How can matrix multiplication in a for loop be optimized? Here is an example of mine: original case: Z = Array{ComplexF64,3}(undef,512,8,24192); X = Array{ComplexF64,3}(undef,512,16,24192); Y = Array{ComplexF64,3}(unde…

---

## [Unstable execution time with high standard error in Julia](https://discourse.julialang.org/t/unstable-execution-time-with-high-standard-error-in-julia/129295)

<div class="topic-metadata">

**Author:** [@PanT12](https://discourse.julialang.org/u/PanT12)\
**Replies:** 27\
**Last updated:** [May 28, 2025, 8:25am UTC](https://discourse.julialang.org/t/unstable-execution-time-with-high-standard-error-in-julia/129295 "2025-05-28T08:25:31Z")

</div>

Has anyone encountered unstable results from Julia, specifically issues with large variance? For example, when I repeat the following experiment 100 times and record the time taken for each run, I find that there are alw…

---

## [Slow deserialization of booleans](https://discourse.julialang.org/t/slow-deserialization-of-booleans/129260)

<div class="topic-metadata">

**Author:** [@Kirby\_Zhang](https://discourse.julialang.org/u/Kirby_Zhang)\
**Replies:** 8\
**Last updated:** [May 25, 2025, 12:58am UTC](https://discourse.julialang.org/t/slow-deserialization-of-booleans/129260 "2025-05-25T00:58:38Z")

</div>

When a Vector{Bool} is deserialized, stdlib/Serialization/src/Serialization.jl:1338 allocates for each element deserialized. I changed the A = Array{Bool, length(dims)}(undef, dims) to A = Vector{Bool}(undef, n) and I wa…

---

## [Sphere surface and axes artefact in Makie.jl](https://discourse.julialang.org/t/sphere-surface-and-axes-artefact-in-makie-jl/129261)

<div class="topic-metadata">

**Author:** [@leespen1](https://discourse.julialang.org/u/leespen1)\
**Replies:** 4\
**Last updated:** [May 23, 2025, 6:12pm UTC](https://discourse.julialang.org/t/sphere-surface-and-axes-artefact-in-makie-jl/129261 "2025-05-23T18:12:39Z")

</div>

I am trying to plot points on the Bloch sphere using GLMakie, but I get white lines around the lines I use to draw the sphere: Is there anything I can do to fix this? The script I used to generate the plot is below. …

---

## [Any memory locations known to be ok to read across platforms? Doing a VM](https://discourse.julialang.org/t/any-memory-locations-known-to-be-ok-to-read-across-platforms-doing-a-vm/129249)

<div class="topic-metadata">

**Author:** [@Palli](https://discourse.julialang.org/u/Palli)\
**Replies:** 2\
**Last updated:** [May 22, 2025, 9:21pm UTC](https://discourse.julialang.org/t/any-memory-locations-known-to-be-ok-to-read-across-platforms-doing-a-vm/129249 "2025-05-22T21:21:01Z")

</div>

These work on my 64-bit Linux, but some are likely not guaranteed to work: julia\> unsafe\_load(Ptr{UInt8}(Int(0x00400000))) # reading julia itself 0x7f Is it up to ELF to use that location, and thus guaranteed to work?…

---

## [Trouble Precompiling GLMakie.jl using ShareAdd.jl - Using Shared Environments](https://discourse.julialang.org/t/trouble-precompiling-glmakie-jl-using-shareadd-jl-using-shared-environments/126999)

<div class="topic-metadata">

**Author:** [@Maddy](https://discourse.julialang.org/u/Maddy)\
**Replies:** 2\
**Last updated:** [May 22, 2025, 12:57pm UTC](https://discourse.julialang.org/t/trouble-precompiling-glmakie-jl-using-shareadd-jl-using-shared-environments/126999 "2025-05-22T12:57:40Z")

</div>

I tried using ShareAdd.jl to use packages from a shared environment. All other packages load and precompile successfully except GLMakie. First, I see if GLMakie.jl has been added to the environment successfully: …

---

## [How to make the most of SIMD.jl when number of data elements is not divisible by SIMD width](https://discourse.julialang.org/t/how-to-make-the-most-of-simd-jl-when-number-of-data-elements-is-not-divisible-by-simd-width/129226)

<div class="topic-metadata">

**Author:** [@davidbp](https://discourse.julialang.org/u/davidbp)\
**Replies:** 5\
**Last updated:** [May 22, 2025, 10:51am UTC](https://discourse.julialang.org/t/how-to-make-the-most-of-simd-jl-when-number-of-data-elements-is-not-divisible-by-simd-width/129226 "2025-05-22T10:51:36Z")

</div>

Consider the following MWE that simply stores in c an elementwise vector multiplication (so c=a.\*b). Here I want to take SIMD chunks and do the operation in SIMD vectors. But the “naive solution” that I get is around 2x…

---

## [ERROR: OutOfMemoryError()](https://discourse.julialang.org/t/error-outofmemoryerror/129229)

<div class="topic-metadata">

**Author:** [@raman\_kumar](https://discourse.julialang.org/u/raman_kumar)\
**Replies:** 10\
**Last updated:** [May 21, 2025, 9:44pm UTC](https://discourse.julialang.org/t/error-outofmemoryerror/129229 "2025-05-21T21:44:37Z")

</div>

I got ERROR: OutOfMemoryError() while running this code. How can i mitigate it? Here 872 is number of mesh blocks of whose coordinates are x, y and z. x = rand(16,872); y = rand(4,872); z = rand(18,872); r = reshape(x,…

---

## [Kwargs... and allocations](https://discourse.julialang.org/t/kwargs-and-allocations/129183)

<div class="topic-metadata">

**Author:** [@Allan\_Baker](https://discourse.julialang.org/u/Allan_Baker)\
**Replies:** 26\
**Last updated:** [May 21, 2025, 3:05pm UTC](https://discourse.julialang.org/t/kwargs-and-allocations/129183 "2025-05-21T15:05:49Z")

</div>

I recently was trying to overload some real-time code (meaning I don’t want any GC or allocations). I needed to add in some kwargs to the help with overloading some intermediate functions. This causes allocations becau…

---

## [Compiler specialisation in automatic differentiation](https://discourse.julialang.org/t/compiler-specialisation-in-automatic-differentiation/129020)

<div class="topic-metadata">

**Author:** [@Philippe\_Maincon1](https://discourse.julialang.org/u/Philippe_Maincon1)\
**Replies:** 3\
**Last updated:** [May 19, 2025, 8:16am UTC](https://discourse.julialang.org/t/compiler-specialisation-in-automatic-differentiation/129020 "2025-05-19T08:16:46Z")

</div>

I am working with finite element method code. In this context, it is (very!!!) useful to use forward automatic differentiation to differentiate the function that, from degrees of freedom, compute the element’s contribut…

---

## [How to prevent unwanted "optimization" in SIMD code?](https://discourse.julialang.org/t/how-to-prevent-unwanted-optimization-in-simd-code/129081)

<div class="topic-metadata">

**Author:** [@lntricate](https://discourse.julialang.org/u/lntricate)\
**Replies:** 8\
**Last updated:** [May 18, 2025, 2:59am UTC](https://discourse.julialang.org/t/how-to-prevent-unwanted-optimization-in-simd-code/129081 "2025-05-18T02:59:48Z")

</div>

I’m working on writing a SIMD base64 encoding algorithm (from this paper), and I came across an annoying thing. Consider the following example method: function test\_mul(a::NTuple{16, VecElement{UInt16}}, b::NTuple{16, V…

---

## [Faster way to peform rank-1 update to matrix inverse](https://discourse.julialang.org/t/faster-way-to-peform-rank-1-update-to-matrix-inverse/129001)

<div class="topic-metadata">

**Author:** [@Kieran\_Marray](https://discourse.julialang.org/u/Kieran_Marray)\
**Replies:** 5\
**Last updated:** [May 15, 2025, 9:10am UTC](https://discourse.julialang.org/t/faster-way-to-peform-rank-1-update-to-matrix-inverse/129001 "2025-05-15T09:10:45Z")

</div>

I want to update a matrix inverse using the Sherman-Morrison formula (A + uv’)^-1 = A^-1 - (A^-1 uv’ A^-1)/(1 + v’A^-1u) for a large non-sparse matrix inside a loop (a coordinate descent algorithm). Here is a minimum …

---

## [Custom serializer for arrays of Union types](https://discourse.julialang.org/t/custom-serializer-for-arrays-of-union-types/129016)

<div class="topic-metadata">

**Author:** [@Kirby\_Zhang](https://discourse.julialang.org/u/Kirby_Zhang)\
**Replies:** 2\
**Last updated:** [May 15, 2025, 4:08am UTC](https://discourse.julialang.org/t/custom-serializer-for-arrays-of-union-types/129016 "2025-05-15T04:08:59Z")

</div>

I’m looking to speed up serialization of arrays that allow for missing values. Performance falls by an order of magnitude according to the following test: N = Int(1e8); buffer = Vector{UInt8}(undef, 8 \* N); io = IOBu…

---

## [Faster \`findmin\` without LoopVectorization.jl](https://discourse.julialang.org/t/faster-findmin-without-loopvectorization-jl/121742)

<div class="topic-metadata">

**Author:** [@jling](https://discourse.julialang.org/u/jling)\
**Replies:** 35\
**Last updated:** [May 14, 2025, 4:06pm UTC](https://discourse.julialang.org/t/faster-findmin-without-loopvectorization-jl/121742 "2025-05-14T16:06:48Z")

</div>

@graeme-a-stewart has this faster alternative of Base.findmin: julia\> using LoopVectorization, Chairmarks julia\> function fast\_findmin(dij, n) best = 1 @inbounds dij\_min = dij\[1\] @turbo…

---

## [Raspberry PI 5 GPIO](https://discourse.julialang.org/t/raspberry-pi-5-gpio/126652)

<div class="topic-metadata">

**Author:** [@Jake](https://discourse.julialang.org/u/Jake)\
**Replies:** 6\
**Last updated:** [May 13, 2025, 6:21pm UTC](https://discourse.julialang.org/t/raspberry-pi-5-gpio/126652 "2025-05-13T18:21:41Z")

</div>

I have an application that requires a trigger signal (Just one pin going from low to high and back to low). There are several options for GPIO control in Julia but the PI 5 adds a new twist as outlined here and here. T…

---

## [Sparse arrays allocation versus speed](https://discourse.julialang.org/t/sparse-arrays-allocation-versus-speed/128726)

<div class="topic-metadata">

**Author:** [@Boris](https://discourse.julialang.org/u/Boris)\
**Replies:** 9\
**Last updated:** [May 13, 2025, 5:18am UTC](https://discourse.julialang.org/t/sparse-arrays-allocation-versus-speed/128726 "2025-05-13T05:18:56Z")

</div>

Hi! I have a question about allocations and speed trade-off when using Sparse Arrays. TLDR question: Sparse matrix multiplication for my case is fast, but allocates a lot of memory. Using dense matrices my code is slowe…

---

## [GLMakie Camera Controls vs Scene Controls](https://discourse.julialang.org/t/glmakie-camera-controls-vs-scene-controls/127754)

<div class="topic-metadata">

**Author:** [@Maddy](https://discourse.julialang.org/u/Maddy)\
**Replies:** 1\
**Last updated:** [May 11, 2025, 6:44am UTC](https://discourse.julialang.org/t/glmakie-camera-controls-vs-scene-controls/127754 "2025-05-11T06:44:18Z")

</div>

I am trying to do a very simple thing - Create a Scene in GLMakie and then control it’s camera. using GLMakie: Scene, scatter!, Circle, display, text! s1 = Scene() scatter!(s1, (0, 0), marker = Circle, markersize = 200…

---

## [Possible speedup in matrix-diagonal products?](https://discourse.julialang.org/t/possible-speedup-in-matrix-diagonal-products/68474)

<div class="topic-metadata">

**Author:** [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Replies:** 6\
**Last updated:** [May 7, 2025, 9:11am UTC](https://discourse.julialang.org/t/possible-speedup-in-matrix-diagonal-products/68474 "2025-05-07T09:11:24Z")

</div>

Right multiplication by a Diagonal matrix effectively scales the columns of a matrix. Currently the default Matrix-Diagonal product appears to use copy\_similar: (\*)(A::AbstractMatrix, D::Diagonal) = rmul!(copy\_simil…

---

## [\`\`mod1\`\` based Periodic Indexing on GPUs](https://discourse.julialang.org/t/mod1-based-periodic-indexing-on-gpus/128657)

<div class="topic-metadata">

**Author:** [@bpsomu](https://discourse.julialang.org/u/bpsomu)\
**Replies:** 1\
**Last updated:** [May 3, 2025, 5:27pm UTC](https://discourse.julialang.org/t/mod1-based-periodic-indexing-on-gpus/128657 "2025-05-03T17:27:42Z")

</div>

I am using the following approach to apply boundary conditions using mod1 for my Simulation code, that I want to be performant on GPUs: struct DirichletArray{arr\<:AbstractArray, T} data::arr core\_boundary\_value:…

---

## [Contiguous Read Non-Contiguous Write vs Non-Contiguous Read Contigous Write Performance](https://discourse.julialang.org/t/contiguous-read-non-contiguous-write-vs-non-contiguous-read-contigous-write-performance/128604)

<div class="topic-metadata">

**Author:** [@mj2984](https://discourse.julialang.org/u/mj2984)\
**Replies:** 3\
**Last updated:** [May 3, 2025, 12:50am UTC](https://discourse.julialang.org/t/contiguous-read-non-contiguous-write-vs-non-contiguous-read-contigous-write-performance/128604 "2025-05-03T00:50:39Z")

</div>

function permutedims\_custom(Y,X) for idx in range(1,size(Y,2)) @simd for idy in range(1,size(Y,1)) @inbounds Y\[idy,idx\] = X\[idx,idy\] end end end function permutedims\_custom\_2(Y,X) …

---

## [Why mem allocation may surge when function reads single scalar global variable?](https://discourse.julialang.org/t/why-mem-allocation-may-surge-when-function-reads-single-scalar-global-variable/128394)

<div class="topic-metadata">

**Author:** [@brownlight](https://discourse.julialang.org/u/brownlight)\
**Replies:** 4\
**Last updated:** [May 2, 2025, 7:39pm UTC](https://discourse.julialang.org/t/why-mem-allocation-may-surge-when-function-reads-single-scalar-global-variable/128394 "2025-05-02T19:39:36Z")

</div>

Two functions: pi\_1( ) sums N terms in for loop; N is defined in global scope. pi\_2(nterms) same as pi\_1( ) except that N is passed as a parameter. Output (copied from notebook cells): N = 10\_000\_000 # First run @ti…

---

## [Performance optimization：Frequently use permutedims function](https://discourse.julialang.org/t/performance-optimization-frequently-use-permutedims-function/128474)

<div class="topic-metadata">

**Author:** [@quantumdreamer](https://discourse.julialang.org/u/quantumdreamer)\
**Replies:** 22\
**Last updated:** [May 1, 2025, 1:33am UTC](https://discourse.julialang.org/t/performance-optimization-frequently-use-permutedims-function/128474 "2025-05-01T01:33:13Z")

</div>

Here is my example： G = rand(8,256,14,288,6)+rand(8,256,14,288,6)\*1im; @elapsed begin G = permutedims(G,\[1,2,4,3,5\]); k = sqrt.(sum(abs2.(reshape(G,589824,1,1,14,6)),dims = 1)); G .\*= k; G = permutedims(…

---

## [Incrementally pretty print a table, one row at a time](https://discourse.julialang.org/t/incrementally-pretty-print-a-table-one-row-at-a-time/128544)

<div class="topic-metadata">

**Author:** [@luke-kiernan](https://discourse.julialang.org/u/luke-kiernan)\
**Replies:** 1\
**Last updated:** [April 30, 2025, 1:55am UTC](https://discourse.julialang.org/t/incrementally-pretty-print-a-table-one-row-at-a-time/128544 "2025-04-30T01:55:38Z")

</div>

I’m comparing some Newton’s method-type iterative solvers. I’d like to pretty print the progress of the iterative method, so I can tell if the search is making good progress. If I didn’t care about performance, I’d just …

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=7)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=9)
