# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=56

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 57

---

## [Is it possible to prefetch memory in Julia?](https://discourse.julialang.org/t/is-it-possible-to-prefetch-memory-in-julia/65566)

<div class="topic-metadata">

**Author:** [@cjdoris](https://discourse.julialang.org/u/cjdoris)\
**Replies:** 13\
**Last updated:** [October 15, 2022, 9:54pm UTC](https://discourse.julialang.org/t/is-it-possible-to-prefetch-memory-in-julia/65566 "2022-10-15T21:54:14Z")

</div>

The title says it all! The only references I can find online are 5 years old and don’t work on Julia v1

---

## [@spawn large memory allocation reduced when some code abstracted out in a function](https://discourse.julialang.org/t/spawn-large-memory-allocation-reduced-when-some-code-abstracted-out-in-a-function/88691)

<div class="topic-metadata">

**Author:** [@pitsianis](https://discourse.julialang.org/u/pitsianis)\
**Replies:** 6\
**Last updated:** [October 15, 2022, 6:24am UTC](https://discourse.julialang.org/t/spawn-large-memory-allocation-reduced-when-some-code-abstracted-out-in-a-function/88691 "2022-10-15T06:24:38Z")

</div>

In another discussion, bizarre behavior that requires further attention and explanation came up. A recursive textbook implementation of quicksort makes no memory allocations as it modifies its argument input vector. H…

---

## [Vector Multiplication on Submatrices](https://discourse.julialang.org/t/vector-multiplication-on-submatrices/88686)

<div class="topic-metadata">

**Author:** [@willsharpless](https://discourse.julialang.org/u/willsharpless)\
**Replies:** 9\
**Last updated:** [October 14, 2022, 10:29pm UTC](https://discourse.julialang.org/t/vector-multiplication-on-submatrices/88686 "2022-10-14T22:29:19Z")

</div>

If I have a matrix which is a (horizontal) collection of matrices, A := \\begin{bmatrix} A\_0 & A\_1 & A\_2 & \\dots & A\_t \\end{bmatrix} , \\quad A\_i \\in \\mathbb{R}^{n \\times n}, A \\in \\mathbb{R}^{n \\times tn} and I want …

---

## [Julia distributed and multithreaded](https://discourse.julialang.org/t/julia-distributed-and-multithreaded/88649)

<div class="topic-metadata">

**Author:** [@this\_josh](https://discourse.julialang.org/u/this_josh)\
**Replies:** 14\
**Last updated:** [October 13, 2022, 6:57pm UTC](https://discourse.julialang.org/t/julia-distributed-and-multithreaded/88649 "2022-10-13T18:57:53Z")

</div>

In my code I leverage multithreading through the use of things like Threads.@threads and by starting julia with Julia --threads=auto. Alongside my code I now wish to utilise Juniper.jl which utilises distributed computin…

---

## [How to speed up rowsum function?](https://discourse.julialang.org/t/how-to-speed-up-rowsum-function/88664)

<div class="topic-metadata">

**Author:** [@Strange\_Xue](https://discourse.julialang.org/u/Strange_Xue)\
**Replies:** 6\
**Last updated:** [October 13, 2022, 4:57pm UTC](https://discourse.julialang.org/t/how-to-speed-up-rowsum-function/88664 "2022-10-13T16:57:40Z")

</div>

There is a rowsum function in R, it’s very helpful and fast when constructing some likelihood function, rowsum can apply a function to a group subsetted from a matrix then concatenate these resulted vectors to a new mat…

---

## [Allocation when sorting a view](https://discourse.julialang.org/t/allocation-when-sorting-a-view/86558)

<div class="topic-metadata">

**Author:** [@CodeGodz](https://discourse.julialang.org/u/CodeGodz)\
**Replies:** 16\
**Last updated:** [October 13, 2022, 1:32pm UTC](https://discourse.julialang.org/t/allocation-when-sorting-a-view/86558 "2022-10-13T13:32:17Z")

</div>

Just reading through “In place sorting for views” and cannot wrap my head around the following allocations in my case: x = \[2,1,10,15,20\] @btime sort!(view($x, 2:4)) 53.308 ns (2 allocations: 96 bytes) To check I also…

---

## [Speeding up repetitive calls to ODEProblem](https://discourse.julialang.org/t/speeding-up-repetitive-calls-to-odeproblem/87644)

<div class="topic-metadata">

**Author:** [@cdawg](https://discourse.julialang.org/u/cdawg)\
**Replies:** 11\
**Last updated:** [October 11, 2022, 3:41pm UTC](https://discourse.julialang.org/t/speeding-up-repetitive-calls-to-odeproblem/87644 "2022-10-11T15:41:49Z")

</div>

Hey ODE people. Thanks for all the mind-bending great work on top-of-the-line packages. I have stupid code solving the same ODE with different boundary conditions thousands of times. Unfortunately for this use case, thi…

---

## [Moving from \`Float64\` to \`Float32\` not improving performance](https://discourse.julialang.org/t/moving-from-float64-to-float32-not-improving-performance/88516)

<div class="topic-metadata">

**Author:** [@JordiBolibar](https://discourse.julialang.org/u/JordiBolibar)\
**Replies:** 25\
**Last updated:** [October 10, 2022, 5:20pm UTC](https://discourse.julialang.org/t/moving-from-float64-to-float32-not-improving-performance/88516 "2022-10-10T17:20:35Z")

</div>

Hi, I’m trying to optimize a script. I first did the usual stuff to avoid memory allocation, but when moving all data structures from Float64 to Float32, which should result in a reduced memory usage, the code has become…

---

## [Want output of dijikstra\_algorithm like A star Algorithm in edges Field instead of int64... or How we can convert integer into edges between nodes?](https://discourse.julialang.org/t/want-output-of-dijikstra-algorithm-like-a-star-algorithm-in-edges-field-instead-of-int64-or-how-we-can-convert-integer-into-edges-between-nodes/88531)

<div class="topic-metadata">

**Author:** [@Kumail\_Haider](https://discourse.julialang.org/u/Kumail_Haider)\
**Replies:** 0\
**Last updated:** [October 10, 2022, 5:15pm UTC](https://discourse.julialang.org/t/want-output-of-dijikstra-algorithm-like-a-star-algorithm-in-edges-field-instead-of-int64-or-how-we-can-convert-integer-into-edges-between-nodes/88531 "2022-10-10T17:15:17Z")

</div>

julia\> using Graphs, GraphPlot julia\> using Symbolics, Statistics, LinearAlgebra julia\> using Plots, LaTeXStrings julia\> g = wheel\_graph(10) {10, 18} undirected simple Int64 graph julia\> ds = dijkstra\_shortest\_paths…

---

## [C++ code much faster than Julia how can I optimize it?](https://discourse.julialang.org/t/c-code-much-faster-than-julia-how-can-i-optimize-it/87868)

<div class="topic-metadata">

**Author:** [@Adegasel](https://discourse.julialang.org/u/Adegasel)\
**Replies:** 40\
**Last updated:** [October 10, 2022, 6:09am UTC](https://discourse.julialang.org/t/c-code-much-faster-than-julia-how-can-i-optimize-it/87868 "2022-10-10T06:09:39Z")

</div>

I’m trying to adapt this C++ code to Julia (here my translation). However C++ original code is much faster than Julia. How can I optimize it?

---

## [Avoiding allocations in small subvector ops](https://discourse.julialang.org/t/avoiding-allocations-in-small-subvector-ops/88500)

<div class="topic-metadata">

**Author:** [@Stephen\_Vavasis](https://discourse.julialang.org/u/Stephen_Vavasis)\
**Replies:** 2\
**Last updated:** [October 10, 2022, 1:03am UTC](https://discourse.julialang.org/t/avoiding-allocations-in-small-subvector-ops/88500 "2022-10-10T01:03:07Z")

</div>

In the code below, there is an allocation each time the statement u\[i1:i1+1\].=... is executed. A few years ago, I wrote a quick-and-dirty macro to implement statements like this with a for-loop like the commented-out lo…

---

## [Shift-Inverse diagonalization in Julia](https://discourse.julialang.org/t/shift-inverse-diagonalization-in-julia/87994)

<div class="topic-metadata">

**Author:** [@devanshu](https://discourse.julialang.org/u/devanshu)\
**Replies:** 7\
**Last updated:** [October 7, 2022, 9:57am UTC](https://discourse.julialang.org/t/shift-inverse-diagonalization-in-julia/87994 "2022-10-07T09:57:29Z")

</div>

What is the best technique to use the shift-inverse diagonalisation technique in Julia? I have used Arpack.jl but it doesn’t seem to be much faster; I am not sure how it handles the inverse part. So I would like to under…

---

## [Package is not updating, but why?](https://discourse.julialang.org/t/package-is-not-updating-but-why/88338)

<div class="topic-metadata">

**Author:** [@bernhard](https://discourse.julialang.org/u/bernhard)\
**Replies:** 5\
**Last updated:** [October 6, 2022, 3:09pm UTC](https://discourse.julialang.org/t/package-is-not-updating-but-why/88338 "2022-10-06T15:09:52Z")

</div>

I have a private package (which I cannot share) where DataFrames is not updating. I cannot find any reason in the Project.toml file why it should be fixed to v1.3.6. Should Pkg.status(; outdated=true) provide the reas…

---

## [Does passing a dataframe declared outside a function as an argument improves performance?](https://discourse.julialang.org/t/does-passing-a-dataframe-declared-outside-a-function-as-an-argument-improves-performance/88318)

<div class="topic-metadata">

**Author:** [@mb96](https://discourse.julialang.org/u/mb96)\
**Replies:** 4\
**Last updated:** [October 6, 2022, 12:15pm UTC](https://discourse.julialang.org/t/does-passing-a-dataframe-declared-outside-a-function-as-an-argument-improves-performance/88318 "2022-10-06T12:15:23Z")

</div>

Hi there, I had a question regarding performance of functions in general which act or use DataFrames. Let say I read in a csv file into a DataFrame in a my main program (code). Then I declare a function which uses such …

---

## [Arithmetic performance of expression](https://discourse.julialang.org/t/arithmetic-performance-of-expression/88196)

<div class="topic-metadata">

**Author:** [@Jojo](https://discourse.julialang.org/u/Jojo)\
**Replies:** 11\
**Last updated:** [October 4, 2022, 8:01pm UTC](https://discourse.julialang.org/t/arithmetic-performance-of-expression/88196 "2022-10-04T20:01:56Z")

</div>

Hello!, Recently I’ve noticed that Julia performs better when I don’t associate the terms in a expression im using in a for loop. Why is (1), faster than (2)? According to @btime (2) is 0.3 seconds slower than (1) aft…

---

## [Potential performance regressions in Julia 1.8 for special un-precompiled type dispatches and how to fix them](https://discourse.julialang.org/t/potential-performance-regressions-in-julia-1-8-for-special-un-precompiled-type-dispatches-and-how-to-fix-them/86359)

<div class="topic-metadata">

**Author:** [@sloede](https://discourse.julialang.org/u/sloede)\
**Replies:** 25\
**Last updated:** [October 4, 2022, 4:02pm UTC](https://discourse.julialang.org/t/potential-performance-regressions-in-julia-1-8-for-special-un-precompiled-type-dispatches-and-how-to-fix-them/86359 "2022-10-04T16:02:53Z")

</div>

TL;DR With Julia 1.8.0 (as opposed to Julia 1.7.3), for OrdinaryDiffEq v6.24.0 and Trixi.jl v0.4.44, we observe that package installation time increased by 20-50% package loading time increased by 30-50% compilation ti…

---

## [Too many allocations for this couple of small functions](https://discourse.julialang.org/t/too-many-allocations-for-this-couple-of-small-functions/88170)

<div class="topic-metadata">

**Author:** [@daviddoij](https://discourse.julialang.org/u/daviddoij)\
**Replies:** 4\
**Last updated:** [October 4, 2022, 9:00am UTC](https://discourse.julialang.org/t/too-many-allocations-for-this-couple-of-small-functions/88170 "2022-10-04T09:00:49Z")

</div>

Hi there, newbie here solved Project Euler #4 problem with Julia and, althogh the timing is good, I got too many allocations (around 100k) for such a small program. I’ve tried with explicit types in the variables and d…

---

## [Vectors with elements of same type but different parametric type](https://discourse.julialang.org/t/vectors-with-elements-of-same-type-but-different-parametric-type/88190)

<div class="topic-metadata">

**Author:** [@DanielVandH](https://discourse.julialang.org/u/DanielVandH)\
**Replies:** 5\
**Last updated:** [October 4, 2022, 7:19am UTC](https://discourse.julialang.org/t/vectors-with-elements-of-same-type-but-different-parametric-type/88190 "2022-10-04T07:19:20Z")

</div>

I need to work with a set of points that might each have a different type. Consider as a rough example: abstract type AbstractLabel end struct LabelA \<: AbstractLabel end struct LabelB \<: AbstractLabel end struct LabelC…

---

## [Optimising function for broadcast](https://discourse.julialang.org/t/optimising-function-for-broadcast/87918)

<div class="topic-metadata">

**Author:** [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Replies:** 9\
**Last updated:** [October 4, 2022, 4:25am UTC](https://discourse.julialang.org/t/optimising-function-for-broadcast/87918 "2022-10-04T04:25:29Z")

</div>

function f( a, b ) c = 2a b + c end f.( 1, 1:100 ) I believe c = 2a will be calculated 100 times (every iteration of the broadcast). Is there a way to have c evaluated once and cached ? without r…

---

## [Mysterious allocation in function returning Union type](https://discourse.julialang.org/t/mysterious-allocation-in-function-returning-union-type/88189)

<div class="topic-metadata">

**Author:** [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Replies:** 3\
**Last updated:** [October 4, 2022, 1:48am UTC](https://discourse.julialang.org/t/mysterious-allocation-in-function-returning-union-type/88189 "2022-10-04T01:48:40Z")

</div>

I’m looking at functions with return type Union{T,Nothing} for some concrete type T. Most of the time they don’t allocate. But where does the allocation in f(a1, nothing) in the code below come from? Is there a way to av…

---

## [Problem converting serial code to parallel code with FLoops](https://discourse.julialang.org/t/problem-converting-serial-code-to-parallel-code-with-floops/87610)

<div class="topic-metadata">

**Author:** [@DanielVandH](https://discourse.julialang.org/u/DanielVandH)\
**Replies:** 9\
**Last updated:** [October 3, 2022, 9:52pm UTC](https://discourse.julialang.org/t/problem-converting-serial-code-to-parallel-code-with-floops/87610 "2022-10-03T21:52:55Z")

</div>

I have the following code which loops over some indices and adds values to two separate parts of an array. using FLoops, Random function update\_u!(du, u, k, q, j) v = \[k, (k + 1) % length(u) + 1, (k + 7) % length(u…

---

## [A problem about performance](https://discourse.julialang.org/t/a-problem-about-performance/88182)

<div class="topic-metadata">

**Author:** [@XuJingye2022](https://discourse.julialang.org/u/XuJingye2022)\
**Replies:** 30\
**Last updated:** [October 3, 2022, 8:40pm UTC](https://discourse.julialang.org/t/a-problem-about-performance/88182 "2022-10-03T20:40:15Z")

</div>

My version is 1.6.7, # global variable v = rand(10000) function fun1() s = 0.0 for i in v::Vector{Float64} s += i end end function fun2(x::Vector{Float64}) s = 0.0 for i in x s += i…

---

## [Fastest way to calculate a rasterised Voronoi diagram without a GPU](https://discourse.julialang.org/t/fastest-way-to-calculate-a-rasterised-voronoi-diagram-without-a-gpu/77016)

<div class="topic-metadata">

**Author:** [@jacobusmmsmit](https://discourse.julialang.org/u/jacobusmmsmit)\
**Replies:** 17\
**Last updated:** [October 3, 2022, 8:05pm UTC](https://discourse.julialang.org/t/fastest-way-to-calculate-a-rasterised-voronoi-diagram-without-a-gpu/77016 "2022-10-03T20:05:39Z")

</div>

Context: Hi everyone, I’m using Turing.jl with NUTS, to infer the parameters a chaotic ODE system. Due to the chaos, doing inference on the positions of the particles as the time window of the inference gets longer beco…

---

## [Can you call Julia methods with LLVM call?](https://discourse.julialang.org/t/can-you-call-julia-methods-with-llvm-call/31082)

<div class="topic-metadata">

**Author:** [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Replies:** 15\
**Last updated:** [October 1, 2022, 9:13pm UTC](https://discourse.julialang.org/t/can-you-call-julia-methods-with-llvm-call/31082 "2022-10-01T21:13:26Z")

</div>

Is there a way to call Julia methods with llvmcall in a way transparent to the optimizer, so that these methods can be inlined, etc? Defining julia\> foo(a,b,c) = sin(a \* b + c) foo (generic function with 1 method) Each…

---

## [Julia startup speed cut in half. Was: (Unofficial) Julia 1.9 for lower latency (startup)](https://discourse.julialang.org/t/julia-startup-speed-cut-in-half-was-unofficial-julia-1-9-for-lower-latency-startup/83338)

<div class="topic-metadata">

**Author:** [@Palli](https://discourse.julialang.org/u/Palli)\
**Replies:** 17\
**Last updated:** [September 30, 2022, 3:15pm UTC](https://discourse.julialang.org/t/julia-startup-speed-cut-in-half-was-unofficial-julia-1-9-for-lower-latency-startup/83338 "2022-09-30T15:15:40Z")

</div>

The timing for (not unofficial) 1.9 master shows 7.6% faster startup than for Julia 1.7.0: $ hyperfine '~/Downloads/julia-a60c76ea57/bin/julia -e ""' Benchmark 1: ~/Downloads/julia-a60c76ea57/bin/julia -e "" Time (mea…

---

## [About the additional memory allocation that appears in the OMEinsum package](https://discourse.julialang.org/t/about-the-additional-memory-allocation-that-appears-in-the-omeinsum-package/87938)

<div class="topic-metadata">

**Author:** [@F-YF](https://discourse.julialang.org/u/F-YF)\
**Replies:** 2\
**Last updated:** [September 29, 2022, 2:40am UTC](https://discourse.julialang.org/t/about-the-additional-memory-allocation-that-appears-in-the-omeinsum-package/87938 "2022-09-29T02:40:20Z")

</div>

julia\> using OMEinsum julia\> a=rand(1024,1024); julia\> b=rand(1024,1024,2); julia\> using BenchmarkTools julia\> @btime ein"xy,yzw-\>xzw"(a,b); 23.790 ms (77 allocations: 16.00 MiB) julia\> @btime ein"yx,yzw-\>xzw"(a,b…

---

## [Using LLVM intrinsic functions inside \`llvmcall\`](https://discourse.julialang.org/t/using-llvm-intrinsic-functions-inside-llvmcall/87955)

<div class="topic-metadata">

**Author:** [@matthias314](https://discourse.julialang.org/u/matthias314)\
**Replies:** 6\
**Last updated:** [September 28, 2022, 11:16pm UTC](https://discourse.julialang.org/t/using-llvm-intrinsic-functions-inside-llvmcall/87955 "2022-09-28T23:16:19Z")

</div>

I’m playing around with llvmcall. I can use LLVM IR instructions (like add), but I cannot get intrinsic functions (like @llvm.abs.i32) to work. From the post Debugging float SIMD Intrinsics via llvmcall I get the impress…

---

## [Matrix-vector product faster than matrix addition?](https://discourse.julialang.org/t/matrix-vector-product-faster-than-matrix-addition/87954)

<div class="topic-metadata">

**Author:** [@goerz](https://discourse.julialang.org/u/goerz)\
**Replies:** 5\
**Last updated:** [September 28, 2022, 11:04pm UTC](https://discourse.julialang.org/t/matrix-vector-product-faster-than-matrix-addition/87954 "2022-09-28T23:04:51Z")

</div>

I’m running some benchmarks, and to my surprize I’m finding adding two matrices much slower than a matrix-vector multiplication. The relevant benchmark code is this (see link for context): function benchmark\_mv\_vs\_mpm()…

---

## [Multi-threading monotonic dynamic](https://discourse.julialang.org/t/multi-threading-monotonic-dynamic/87642)

<div class="topic-metadata">

**Author:** [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Replies:** 6\
**Last updated:** [September 28, 2022, 12:38pm UTC](https://discourse.julialang.org/t/multi-threading-monotonic-dynamic/87642 "2022-09-28T12:38:57Z")

</div>

Hi, I am trying to use multi-threading in a monotonic dynamic way (in the sense of openmp). By monotonic , I mean that if a thread exected iteration i then the thread must execute iterations larger than i subsequently. …

---

## [Multithreading a for loop](https://discourse.julialang.org/t/multithreading-a-for-loop/87920)

<div class="topic-metadata">

**Author:** [@Sagar\_123](https://discourse.julialang.org/u/Sagar_123)\
**Replies:** 6\
**Last updated:** [September 28, 2022, 12:27pm UTC](https://discourse.julialang.org/t/multithreading-a-for-loop/87920 "2022-09-28T12:27:59Z")

</div>

function collect\_nodes\_frac\_serial(nodes, weights, pos, grid\_spacing) @turbo for i = 1:length(pos) nodes\[i\] = div(pos\[i\],grid\_spacing) weights\[i\] = (pos\[i\] - nodes\[i\]\*grid\_spacing)/grid\_spacing en…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=55)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=57)
