# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=40

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 41

---

## [Fast recursion with big rationals](https://discourse.julialang.org/t/fast-recursion-with-big-rationals/101023)

<div class="topic-metadata">

**Author:** [@Denis\_Ivanov](https://discourse.julialang.org/u/Denis_Ivanov)\
**Replies:** 4\
**Last updated:** [July 1, 2023, 3:56pm UTC](https://discourse.julialang.org/t/fast-recursion-with-big-rationals/101023 "2023-07-01T15:56:24Z")

</div>

Hi! Now I’m doing one task for OEIS. In particular, I need to implement a fast algorithm for finding egyptian fractions of shortest length. There is the next algorithm for to find just the shortest length: If there…

---

## [Changing array inside mutable struct passed to function, how to improve performance?](https://discourse.julialang.org/t/changing-array-inside-mutable-struct-passed-to-function-how-to-improve-performance/100862)

<div class="topic-metadata">

**Author:** [@ordep](https://discourse.julialang.org/u/ordep)\
**Replies:** 15\
**Last updated:** [June 27, 2023, 7:26pm UTC](https://discourse.julialang.org/t/changing-array-inside-mutable-struct-passed-to-function-how-to-improve-performance/100862 "2023-06-27T19:26:24Z")

</div>

I am creating a simulator, and to avoid passing on several different arrays (with different datatypes) to other functions that will apend values to them, i was thinking about grouping the arrays inside a mutable struct, …

---

## [Why is the matrix multiplication with integer matrices much slower than with float ones?](https://discourse.julialang.org/t/why-is-the-matrix-multiplication-with-integer-matrices-much-slower-than-with-float-ones/100898)

<div class="topic-metadata">

**Author:** [@DY\_K](https://discourse.julialang.org/u/DY_K)\
**Replies:** 6\
**Last updated:** [June 27, 2023, 6:02pm UTC](https://discourse.julialang.org/t/why-is-the-matrix-multiplication-with-integer-matrices-much-slower-than-with-float-ones/100898 "2023-06-27T18:02:25Z")

</div>

The following is the simple function for the test. function vec\_prod(n) output = 0; vec = collect(1:n) mat = Array{Float64}(undef, n, n) mat .= vec .\* vec' matt1 = mat .\* mat matt2 = mat \* mat output = dot(matt1, …

---

## [Is 1.9 great or what?](https://discourse.julialang.org/t/is-1-9-great-or-what/100664)

<div class="topic-metadata">

**Author:** [@lewis](https://discourse.julialang.org/u/lewis)\
**Replies:** 16\
**Last updated:** [June 27, 2023, 5:51pm UTC](https://discourse.julialang.org/t/is-1-9-great-or-what/100664 "2023-06-27T17:51:03Z")

</div>

Great! Not or what. TTFP is essentially instant. And many other great improvements. So much work has been done over so many releases, so with no disrespect to those release I will say that 1.9 is the best since 1.0. A…

---

## [Advanced tricks to reduce memory allocations of ODE](https://discourse.julialang.org/t/advanced-tricks-to-reduce-memory-allocations-of-ode/100841)

<div class="topic-metadata">

**Author:** [@JordiBolibar](https://discourse.julialang.org/u/JordiBolibar)\
**Replies:** 7\
**Last updated:** [June 27, 2023, 11:19am UTC](https://discourse.julialang.org/t/advanced-tricks-to-reduce-memory-allocations-of-ode/100841 "2023-06-27T11:19:47Z")

</div>

Hello, I’m trying to further reduce memory allocations when solving an ODE using DifferentialEquations.jl. I have already implemented most of the tricks I’m aware of, but I still get a ton of memory allocation when call…

---

## [Allocations when creating array](https://discourse.julialang.org/t/allocations-when-creating-array/100811)

<div class="topic-metadata">

**Author:** [@CarlosContrerasQ12](https://discourse.julialang.org/u/CarlosContrerasQ12)\
**Replies:** 16\
**Last updated:** [June 27, 2023, 1:46am UTC](https://discourse.julialang.org/t/allocations-when-creating-array/100811 "2023-06-27T01:46:17Z")

</div>

Hi! I have the following code trying to simulate 1000 dimensional brownian paths const type=Float32 function simulate\_path(path\_length) xis=randn(type,(100,path\_length-1)) X=zeros(type,(100,path\_length)) sqd…

---

## [Performance & Profiling Tips for Beginner Code](https://discourse.julialang.org/t/performance-profiling-tips-for-beginner-code/100585)

<div class="topic-metadata">

**Author:** [@physh](https://discourse.julialang.org/u/physh)\
**Replies:** 14\
**Last updated:** [June 26, 2023, 9:18am UTC](https://discourse.julialang.org/t/performance-profiling-tips-for-beginner-code/100585 "2023-06-26T09:18:04Z")

</div>

Dear all, I have been using Julia for a few months now, writing code mostly without worrying too much about performance (profiling very briefly with @time/@btime to make sure nothing too horribly is happening though). I…

---

## [Unexpected performance mismatch in gradients for "compiled-tape-in-tape" experiment](https://discourse.julialang.org/t/unexpected-performance-mismatch-in-gradients-for-compiled-tape-in-tape-experiment/100810)

<div class="topic-metadata">

**Author:** [@jacobusmmsmit](https://discourse.julialang.org/u/jacobusmmsmit)\
**Replies:** 0\
**Last updated:** [June 25, 2023, 12:49pm UTC](https://discourse.julialang.org/t/unexpected-performance-mismatch-in-gradients-for-compiled-tape-in-tape-experiment/100810 "2023-06-25T12:49:34Z")

</div>

Hi all, I’m working on a way to make selectively compiling parts of a ReverseDiff tape easier and more accessible. The goal is to enable fearless compilation of functions with branches. One problem I’ve run into is rel…

---

## [How to achieve multi-threaded vectorized FMA operations in the for-loop for SAXPY?](https://discourse.julialang.org/t/how-to-achieve-multi-threaded-vectorized-fma-operations-in-the-for-loop-for-saxpy/100805)

<div class="topic-metadata">

**Author:** [@xinwu](https://discourse.julialang.org/u/xinwu)\
**Replies:** 2\
**Last updated:** [June 25, 2023, 10:42am UTC](https://discourse.julialang.org/t/how-to-achieve-multi-threaded-vectorized-fma-operations-in-the-for-loop-for-saxpy/100805 "2023-06-25T10:42:03Z")

</div>

Hi everyone, For serial vectorized FMA operations in the for-loop for SAXPY, one can use @fastmath @inbounds @simd for i in 1:n See also this discussion (How to enable vectorized fma instruction for multiply-add vecto…

---

## [Why using a mutable struct type argument to create instances creates a 50x slowdown?](https://discourse.julialang.org/t/why-using-a-mutable-struct-type-argument-to-create-instances-creates-a-50x-slowdown/100764)

<div class="topic-metadata">

**Author:** [@Tortar](https://discourse.julialang.org/u/Tortar)\
**Replies:** 7\
**Last updated:** [June 25, 2023, 12:54am UTC](https://discourse.julialang.org/t/why-using-a-mutable-struct-type-argument-to-create-instances-creates-a-50x-slowdown/100764 "2023-06-25T00:54:20Z")

</div>

I created this MWE based on some package code, and I’m not getting why the two different functions have not similar performance: using BenchmarkTools mutable struct A q\_0::Int q\_1::Int q\_2::Int end @noinli…

---

## [\[YouTube/GitHub\] What is the FASTEST Computer Language? 45 Languages Tested](https://discourse.julialang.org/t/youtube-github-what-is-the-fastest-computer-language-45-languages-tested/64506)

<div class="topic-metadata">

**Author:** [@essenciary](https://discourse.julialang.org/u/essenciary)\
**Replies:** 18\
**Last updated:** [June 24, 2023, 5:35am UTC](https://discourse.julialang.org/t/youtube-github-what-is-the-fastest-computer-language-45-languages-tested/64506 "2023-06-24T05:35:58Z")

</div>

A pretty cool project which includes Julia:

---

## [Optimise calculation with NxNxMxM matrix](https://discourse.julialang.org/t/optimise-calculation-with-nxnxmxm-matrix/100565)

<div class="topic-metadata">

**Author:** [@Stalz](https://discourse.julialang.org/u/Stalz)\
**Replies:** 17\
**Last updated:** [June 22, 2023, 2:33pm UTC](https://discourse.julialang.org/t/optimise-calculation-with-nxnxmxm-matrix/100565 "2023-06-22T14:33:13Z")

</div>

I have been working on a simulation in python where I try to solve the following integral: f(x,y) = \\iint A(x-\\xi', y-\\eta') g(\\xi',\\eta') h(x,\\xi',y,\\eta') \\,d\\xi' \\,d\\eta' With multiprocessing on a server I could g…

---

## [Extra allocations performing dot product in ODE callback](https://discourse.julialang.org/t/extra-allocations-performing-dot-product-in-ode-callback/100683)

<div class="topic-metadata">

**Author:** [@albertomercurio](https://discourse.julialang.org/u/albertomercurio)\
**Replies:** 3\
**Last updated:** [June 23, 2023, 3:10pm UTC](https://discourse.julialang.org/t/extra-allocations-performing-dot-product-in-ode-callback/100683 "2023-06-23T15:10:35Z")

</div>

Hello, I’m experiencing some extra memory allocations when calculating the dot product inside a ODE callback. Here I write a minimal example A = sprand(ComplexF64, 100, 100, 0.05) u0 = rand(ComplexF64, 100) y = rand(Co…

---

## [Understanding the performance and overhead of a vector of SOA vs a vector of AOS for SIMD and the effect of push!](https://discourse.julialang.org/t/understanding-the-performance-and-overhead-of-a-vector-of-soa-vs-a-vector-of-aos-for-simd-and-the-effect-of-push/100560)

<div class="topic-metadata">

**Author:** [@f.ij](https://discourse.julialang.org/u/f.ij)\
**Replies:** 1\
**Last updated:** [June 23, 2023, 11:13am UTC](https://discourse.julialang.org/t/understanding-the-performance-and-overhead-of-a-vector-of-soa-vs-a-vector-of-aos-for-simd-and-the-effect-of-push/100560 "2023-06-23T11:13:23Z")

</div>

I’ve been working on an interactive simulation tool for Monte Carlo simulations of Ising(-like) Models. Performance is key, and something I’ve been a bit confused by. I made a very simplified version of the program here,…

---

## [Efficient Hessian assembling within an interior point method](https://discourse.julialang.org/t/efficient-hessian-assembling-within-an-interior-point-method/100731)

<div class="topic-metadata">

**Author:** [@twopii](https://discourse.julialang.org/u/twopii)\
**Replies:** 0\
**Last updated:** [June 23, 2023, 8:33am UTC](https://discourse.julialang.org/t/efficient-hessian-assembling-within-an-interior-point-method/100731 "2023-06-23T08:33:07Z")

</div>

Hi all! I am currently considering an assembling problem. Given is a dense matrix A and a sparse matrix B and the goal is to efficiently compute the matrix H = B^\\top (A \\otimes A) B. If one expresses the matrix B as B…

---

## [Comprehensions versus pre-allocation](https://discourse.julialang.org/t/comprehensions-versus-pre-allocation/7352)

<div class="topic-metadata">

**Author:** [@yakir12](https://discourse.julialang.org/u/yakir12)\
**Replies:** 16\
**Last updated:** [June 22, 2023, 4:20pm UTC](https://discourse.julialang.org/t/comprehensions-versus-pre-allocation/7352 "2023-06-22T16:20:54Z")

</div>

When it’s possible, I rewrite the calculation of some array from the “first pre-allocation then for-loop” to just a comprehension (albeit fairly complicated sometimes). In my mind comprehensions seem faster because the p…

---

## [Question: derefing named tuples](https://discourse.julialang.org/t/question-derefing-named-tuples/100671)

<div class="topic-metadata">

**Author:** [@lewis](https://discourse.julialang.org/u/lewis)\
**Replies:** 1\
**Last updated:** [June 21, 2023, 9:01pm UTC](https://discourse.julialang.org/t/question-derefing-named-tuples/100671 "2023-06-21T21:01:00Z")

</div>

I need to pass a bunch of vectors (from a typed table) into a function that is called in a hot loop. I had previously avoided the time cost of derefing the columns from the table in the loop. Now, I am putting the neede…

---

## [Are there any pentadiagonal system solvers?](https://discourse.julialang.org/t/are-there-any-pentadiagonal-system-solvers/100665)

<div class="topic-metadata">

**Author:** [@Veenty](https://discourse.julialang.org/u/Veenty)\
**Replies:** 0\
**Last updated:** [June 21, 2023, 4:12pm UTC](https://discourse.julialang.org/t/are-there-any-pentadiagonal-system-solvers/100665 "2023-06-21T16:12:58Z")

</div>

I have a pentadiagonal system and I was hoping that there are fast implementations of solvers related to this. I was trying to use BandedMatrices.jl and then do LU decomposition. But in this case it returns only 1 bande…

---

## [Can constant propagation transform integer powers of -1 to an if/else?](https://discourse.julialang.org/t/can-constant-propagation-transform-integer-powers-of-1-to-an-if-else/100462)

<div class="topic-metadata">

**Author:** [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Replies:** 5\
**Last updated:** [June 16, 2023, 10:01pm UTC](https://discourse.julialang.org/t/can-constant-propagation-transform-integer-powers-of-1-to-an-if-else/100462 "2023-06-16T22:01:11Z")

</div>

julia\> @btime (n -\> \[iseven(k) ? 1 : (-1) for k in 1:n\])(500); 890.571 ns (1 allocation: 4.06 KiB) julia\> @btime (n -\> \[(-1)^k for k in 1:n\])(500); 6.170 μs (1 allocation: 4.06 KiB) It would be great if the latter …

---

## [Can Julia achieve fine grained control of performance without sacrificing ease of use?](https://discourse.julialang.org/t/can-julia-achieve-fine-grained-control-of-performance-without-sacrificing-ease-of-use/100437)

<div class="topic-metadata">

**Author:** [@Tarny\_GG\_Channie](https://discourse.julialang.org/u/Tarny_GG_Channie)\
**Replies:** 14\
**Last updated:** [June 16, 2023, 1:31pm UTC](https://discourse.julialang.org/t/can-julia-achieve-fine-grained-control-of-performance-without-sacrificing-ease-of-use/100437 "2023-06-16T13:31:53Z")

</div>

Julia was designed to have a good overall performance, being llvm-compiled with a good type inference. However, Julia has some subtle issues when fine-grained control over the inner working is needed, for example… -Wha…

---

## [FlameGraph ProfileView not showing matrix operations](https://discourse.julialang.org/t/flamegraph-profileview-not-showing-matrix-operations/100359)

<div class="topic-metadata">

**Author:** [@hsgg](https://discourse.julialang.org/u/hsgg)\
**Replies:** 2\
**Last updated:** [June 15, 2023, 9:54pm UTC](https://discourse.julialang.org/t/flamegraph-profileview-not-showing-matrix-operations/100359 "2023-06-15T21:54:04Z")

</div>

Hi, I am using FlameGraphs.jl and ProfileView.jl to profile my code with matrix operations. I am fairly new to these profiling tools. On MacOS M1 the matrix operations are not showing up in the flame graph. Here is a MW…

---

## [Compiled program from PackageCompiler is much slower than REPL?](https://discourse.julialang.org/t/compiled-program-from-packagecompiler-is-much-slower-than-repl/42640)

<div class="topic-metadata">

**Author:** [@linlinlin](https://discourse.julialang.org/u/linlinlin)\
**Replies:** 3\
**Last updated:** [June 15, 2023, 8:31am UTC](https://discourse.julialang.org/t/compiled-program-from-packagecompiler-is-much-slower-than-repl/42640 "2023-06-15T08:31:43Z")

</div>

Hi, Community members. Recently, I want to compile my code to an executable program, but I have the speed problem. I develop my program based on a package JuliaGrid. In this package, it calls many other packages, like…

---

## [Optimizing code with array assignments](https://discourse.julialang.org/t/optimizing-code-with-array-assignments/99857)

<div class="topic-metadata">

**Author:** [@fergu](https://discourse.julialang.org/u/fergu)\
**Replies:** 12\
**Last updated:** [June 15, 2023, 7:48am UTC](https://discourse.julialang.org/t/optimizing-code-with-array-assignments/99857 "2023-06-15T07:48:49Z")

</div>

TL;DR I have some performance critical code that I am trying to optimize as it will otherwise take unacceptably long to run. Testing seems to suggest that the main culprit is assignments to arrays rather than the computa…

---

## [Parallelizing file processing with the least amount of overhead](https://discourse.julialang.org/t/parallelizing-file-processing-with-the-least-amount-of-overhead/100169)

<div class="topic-metadata">

**Author:** [@CodeGodz](https://discourse.julialang.org/u/CodeGodz)\
**Replies:** 2\
**Last updated:** [June 14, 2023, 9:58am UTC](https://discourse.julialang.org/t/parallelizing-file-processing-with-the-least-amount-of-overhead/100169 "2023-06-14T09:58:28Z")

</div>

Hey all! I wrote ViewReader a while ago and now thought adding some parallel file processing might be fun. ViewReader reads UInt8 (chars) to a buffer array so each thread could have its own buffer array. This might be ne…

---

## [How to enable vectorized fma instruction for multiply-add vectors?](https://discourse.julialang.org/t/how-to-enable-vectorized-fma-instruction-for-multiply-add-vectors/100219)

<div class="topic-metadata">

**Author:** [@xinwu](https://discourse.julialang.org/u/xinwu)\
**Replies:** 10\
**Last updated:** [June 14, 2023, 8:31am UTC](https://discourse.julialang.org/t/how-to-enable-vectorized-fma-instruction-for-multiply-add-vectors/100219 "2023-06-14T08:31:02Z")

</div>

Hi, the vectorized fma instructions (unrolled by 4x) can easily be generated for the C code below via clang -Ofast -S -Wall -std=c11 -march=skylake vecfma.c. void vecfma(float \* restrict c, float \* restrict a, float \* …

---

## [Optimzing many linear solves](https://discourse.julialang.org/t/optimzing-many-linear-solves/100304)

<div class="topic-metadata">

**Author:** [@RobertGregg](https://discourse.julialang.org/u/RobertGregg)\
**Replies:** 2\
**Last updated:** [June 14, 2023, 1:59am UTC](https://discourse.julialang.org/t/optimzing-many-linear-solves/100304 "2023-06-14T01:59:59Z")

</div>

I have a function that runs many linear solves on subsets of columns from a matrix I give it. Something like: function score(data, parents, child) #Subset columns from the data matrix @views begin X = d…

---

## [Multithreading with shared memory caches](https://discourse.julialang.org/t/multithreading-with-shared-memory-caches/100194)

<div class="topic-metadata">

**Author:** [@joaquinpelle](https://discourse.julialang.org/u/joaquinpelle)\
**Replies:** 27\
**Last updated:** [June 13, 2023, 6:46pm UTC](https://discourse.julialang.org/t/multithreading-with-shared-memory-caches/100194 "2023-06-13T18:46:25Z")

</div>

Hi there, I have a loop that uses a cache array to avoid excessive memory allocations. The structure is captured in the following example: A = rand(10000) B = similar(A) cache = zeros(3) for i in eachindex(A) f!(ca…

---

## [Why does the following setindex call allocate?](https://discourse.julialang.org/t/why-does-the-following-setindex-call-allocate/100216)

<div class="topic-metadata">

**Author:** [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Replies:** 6\
**Last updated:** [June 12, 2023, 10:15am UTC](https://discourse.julialang.org/t/why-does-the-following-setindex-call-allocate/100216 "2023-06-12T10:15:04Z")

</div>

julia\> function f!(v, A, r) v\[eachindex(r)\] = @view A\[r\] v end f! (generic function with 1 method) julia\> v = zeros(10); A = ones(4,10); julia\> @btime f!($v, $A, 4:3:13); 56.050 ns (2 all…

---

## [PyCall minimal overhead](https://discourse.julialang.org/t/pycall-minimal-overhead/41127)

<div class="topic-metadata">

**Author:** [@LaurentPlagne](https://discourse.julialang.org/u/LaurentPlagne)\
**Replies:** 15\
**Last updated:** [June 12, 2023, 8:01am UTC](https://discourse.julialang.org/t/pycall-minimal-overhead/41127 "2023-06-12T08:01:00Z")

</div>

Hi, I work on a Julia benchmark Data Base project aiming to compare different implementations of basic computational kernels implemented in various languages. I wonder about the most efficient way to call a Python snip…

---

## [Performance of hasmethod vs try-catch on MethodError](https://discourse.julialang.org/t/performance-of-hasmethod-vs-try-catch-on-methoderror/99827)

<div class="topic-metadata">

**Author:** [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Replies:** 22\
**Last updated:** [June 11, 2023, 4:31am UTC](https://discourse.julialang.org/t/performance-of-hasmethod-vs-try-catch-on-methoderror/99827 "2023-06-11T04:31:24Z")

</div>

To help me better understand Julia’s performance behavior, I am trying to understand the large difference between two strategies for checking for undefined methods: hasmethod vs try-catch. I’ve read through Are exceptio…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=39)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=41)
