# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=89

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 90

---

## [Could the Profile.print() tree view sort in the other order?](https://discourse.julialang.org/t/could-the-profile-print-tree-view-sort-in-the-other-order/63881)

<div class="topic-metadata">

**Author:** [@gleyland](https://discourse.julialang.org/u/gleyland)\
**Replies:** 1\
**Last updated:** [July 1, 2021, 10:59am UTC](https://discourse.julialang.org/t/could-the-profile-print-tree-view-sort-in-the-other-order/63881 "2021-07-01T10:59:43Z")

</div>

Hi, I’m aware there are other ways of viewing profile output, but Profile.print() is pretty handy in a pinch. It would be even handier if the tree was sorted by count. The docs say that sortedby “Controls the order in…

---

## [Speed of Julia vs. Fortran and other languages](https://discourse.julialang.org/t/speed-of-julia-vs-fortran-and-other-languages/63850)

<div class="topic-metadata">

**Author:** [@Beliavsky](https://discourse.julialang.org/u/Beliavsky)\
**Replies:** 6\
**Last updated:** [July 1, 2021, 12:54pm UTC](https://discourse.julialang.org/t/speed-of-julia-vs-fortran-and-other-languages/63850 "2021-07-01T12:54:37Z")

</div>

The Benchmarks category of Fortran code on GitHub lists many benchmarks where Julia is considered, listed below. I will not make any generalizations about the results. Assessment of Programming Languages for Computation…

---

## [Why is this raytracer slow?](https://discourse.julialang.org/t/why-is-this-raytracer-slow/59176)

<div class="topic-metadata">

**Author:** [@linwaytin](https://discourse.julialang.org/u/linwaytin)\
**Replies:** 29\
**Last updated:** [May 5, 2021, 3:11pm UTC](https://discourse.julialang.org/t/why-is-this-raytracer-slow/59176 "2021-05-05T15:11:17Z")

</div>

I bumped into this repo. https://github.com/edin/raytracer It compares implementations in different languages. I’m surprised that julia is more than 5 times slower than C. The code seems idiomatic. Does anyone have s…

---

## [Why is Matrix multiplication with an UpperTriangular slower than with normal Matrix?](https://discourse.julialang.org/t/why-is-matrix-multiplication-with-an-uppertriangular-slower-than-with-normal-matrix/63824)

<div class="topic-metadata">

**Author:** [@zsoerenm](https://discourse.julialang.org/u/zsoerenm)\
**Replies:** 2\
**Last updated:** [June 30, 2021, 1:14pm UTC](https://discourse.julialang.org/t/why-is-matrix-multiplication-with-an-uppertriangular-slower-than-with-normal-matrix/63824 "2021-06-30T13:14:55Z")

</div>

Why is Matrix multiplication with an UpperTriangular slower than with normal Matrix? I would expect it to be faster: using LinearAlgebra, BenchmarkTools A = randn(60,60); B = randn(60,60); @btime $A \* $B' # 19.693 μs (…

---

## [Slow \`convert\` for arrays. Use multiple dispatch instead?](https://discourse.julialang.org/t/slow-convert-for-arrays-use-multiple-dispatch-instead/63822)

<div class="topic-metadata">

**Author:** [@manuelbb-upb](https://discourse.julialang.org/u/manuelbb-upb)\
**Replies:** 3\
**Last updated:** [June 30, 2021, 1:09pm UTC](https://discourse.julialang.org/t/slow-convert-for-arrays-use-multiple-dispatch-instead/63822 "2021-06-30T13:09:27Z")

</div>

I was wondering why something like x = rand(10) convert(Vector{Float64}, x) is relatively slow, when x does not even need to be converted. I noticed that in Julia base the convert method is defined as convert(::Type{…

---

## [Any benchmark comparing Julia on Windows vs Linux vs OSX?](https://discourse.julialang.org/t/any-benchmark-comparing-julia-on-windows-vs-linux-vs-osx/18394)

<div class="topic-metadata">

**Author:** [@Juan](https://discourse.julialang.org/u/Juan)\
**Replies:** 7\
**Last updated:** [June 29, 2021, 8:55pm UTC](https://discourse.julialang.org/t/any-benchmark-comparing-julia-on-windows-vs-linux-vs-osx/18394 "2021-06-29T20:55:24Z")

</div>

Is there any benchmark comparing the speed and memory usage of Julia on different Operating Systems?

---

## [Performance issue with this bootstrapp](https://discourse.julialang.org/t/performance-issue-with-this-bootstrapp/63551)

<div class="topic-metadata">

**Author:** [@lrnv](https://discourse.julialang.org/u/lrnv)\
**Replies:** 11\
**Last updated:** [June 29, 2021, 1:34pm UTC](https://discourse.julialang.org/t/performance-issue-with-this-bootstrapp/63551 "2021-06-29T13:34:16Z")

</div>

Hi, I have the following piece of code that performs a simple bootstrap, but which is taking a very long time. I was wondering If I made obvious mistakes in the implmentation, although @code\_warntype seems to tell me ev…

---

## [Optimal approach to simulating large system of time-dependent ODEs?](https://discourse.julialang.org/t/optimal-approach-to-simulating-large-system-of-time-dependent-odes/63709)

<div class="topic-metadata">

**Author:** [@kharing](https://discourse.julialang.org/u/kharing)\
**Replies:** 1\
**Last updated:** [June 28, 2021, 10:39pm UTC](https://discourse.julialang.org/t/optimal-approach-to-simulating-large-system-of-time-dependent-odes/63709 "2021-06-28T22:39:00Z")

</div>

I’m looking to solve a large (N ~ 1000) system of 1st order, linear ODEs with time-dependent coefficients, \\frac{d u}{d t} = A(t) \\cdot u(t). Because both the numerical constants involved in the coefficients and the nu…

---

## [SuiteSparse.CHOLMOD.lowrankupdate! does not speed up calculations](https://discourse.julialang.org/t/suitesparse-cholmod-lowrankupdate-does-not-speed-up-calculations/63690)

<div class="topic-metadata">

**Author:** [@sidelkin](https://discourse.julialang.org/u/sidelkin)\
**Replies:** 0\
**Last updated:** [June 28, 2021, 11:17am UTC](https://discourse.julialang.org/t/suitesparse-cholmod-lowrankupdate-does-not-speed-up-calculations/63690 "2021-06-28T11:17:48Z")

</div>

Hello. Can someone explain why using SuiteSparse.CHOLMOD.lowrankupdate! does not give acceleration in this example? using BenchmarkTools, LinearAlgebra, SparseArrays, Test, SuiteSparse N = 100; M = 20 b = rand(N) spA =…

---

## [Basic stochastic model](https://discourse.julialang.org/t/basic-stochastic-model/63680)

<div class="topic-metadata">

**Author:** [@Cyber](https://discourse.julialang.org/u/Cyber)\
**Replies:** 0\
**Last updated:** [June 28, 2021, 8:19am UTC](https://discourse.julialang.org/t/basic-stochastic-model/63680 "2021-06-28T08:19:26Z")

</div>

Hello all. I am trying to simulate stochasitc motions for some particles in a specific domain. It is fine for just 2 particles. But if I would like to simulate more than 2 particles, my code becomes too slow.(I only know…

---

## [Why matrix multiplication is much slower than PyTorch](https://discourse.julialang.org/t/why-matrix-multiplication-is-much-slower-than-pytorch/63661)

<div class="topic-metadata">

**Author:** [@yingqiuz](https://discourse.julialang.org/u/yingqiuz)\
**Replies:** 4\
**Last updated:** [June 27, 2021, 8:49pm UTC](https://discourse.julialang.org/t/why-matrix-multiplication-is-much-slower-than-pytorch/63661 "2021-06-27T20:49:27Z")

</div>

For example, matrix multiplication of 10,000 x 10,100 matrices, single threaded In julia: BLAS.set\_num\_threads(1) A = randn(10000, 10000) B = randn(10000, 10000) C = Matrix{Float64}(undef, 10000, 10000) @benchmark mul…

---

## [A simple SIMD.jl loop that is slower than a vanilla \`@inbounds @simd\`](https://discourse.julialang.org/t/a-simple-simd-jl-loop-that-is-slower-than-a-vanilla-inbounds-simd/63655)

<div class="topic-metadata">

**Author:** [@Krastanov](https://discourse.julialang.org/u/Krastanov)\
**Replies:** 8\
**Last updated:** [June 27, 2021, 8:21pm UTC](https://discourse.julialang.org/t/a-simple-simd-jl-loop-that-is-slower-than-a-vanilla-inbounds-simd/63655 "2021-06-27T20:21:43Z")

</div>

I have a loop that performs some bit-fiddling operations. Because of the way some out-of-loop constants are used, @simd (and occasionally even LoopVectorization.jl) does not always give the fastest possible machine code.…

---

## [Why is this small \`@inline\` function much slower than an equivalent macro?](https://discourse.julialang.org/t/why-is-this-small-inline-function-much-slower-than-an-equivalent-macro/63620)

<div class="topic-metadata">

**Author:** [@Krastanov](https://discourse.julialang.org/u/Krastanov)\
**Replies:** 2\
**Last updated:** [June 26, 2021, 11:19pm UTC](https://discourse.julialang.org/t/why-is-this-small-inline-function-much-slower-than-an-equivalent-macro/63620 "2021-06-26T23:19:50Z")

</div>

My understanding is that in Julia macros are meant to be used for expressiveness and to make DSLs, not for performance hacks. But in this bit-fiddling SIMD loop, a macro is much faster than the corresponding inline funct…

---

## [In place assignment of scalars vs vectors](https://discourse.julialang.org/t/in-place-assignment-of-scalars-vs-vectors/63542)

<div class="topic-metadata">

**Author:** [@lorenzo.a.ricciardi](https://discourse.julialang.org/u/lorenzo.a.ricciardi)\
**Replies:** 20\
**Last updated:** [June 25, 2021, 2:40pm UTC](https://discourse.julialang.org/t/in-place-assignment-of-scalars-vs-vectors/63542 "2021-06-25T14:40:06Z")

</div>

Hi all, I’m writing a performance critical function that will be called many times within an optimization loop. I have several questions regarding how to make it as efficient as possible, but I’ll break those down so t…

---

## [Conditional array comprehension vs. simple for loop - 750k allocations vs. 1 allocation - why?](https://discourse.julialang.org/t/conditional-array-comprehension-vs-simple-for-loop-750k-allocations-vs-1-allocation-why/63249)

<div class="topic-metadata">

**Author:** [@jlawrie](https://discourse.julialang.org/u/jlawrie)\
**Replies:** 14\
**Last updated:** [June 25, 2021, 8:26am UTC](https://discourse.julialang.org/t/conditional-array-comprehension-vs-simple-for-loop-750k-allocations-vs-1-allocation-why/63249 "2021-06-25T08:26:54Z")

</div>

Hello! Thanks for taking a moment to have a look at my question :slight\_smile: Full disclosure: I’m only a few weeks into learning Julia, brand new to Discourse and an interdisciplinary engineer (not a computer scienti…

---

## [ILU preconditioned GMRES inside TRBDF2 for ODEs generated by NetworkDynamics.jl](https://discourse.julialang.org/t/ilu-preconditioned-gmres-inside-trbdf2-for-odes-generated-by-networkdynamics-jl/63385)

<div class="topic-metadata">

**Author:** [@ziolai](https://discourse.julialang.org/u/ziolai)\
**Replies:** 6\
**Last updated:** [June 24, 2021, 4:47pm UTC](https://discourse.julialang.org/t/ilu-preconditioned-gmres-inside-trbdf2-for-odes-generated-by-networkdynamics-jl/63385 "2021-06-24T16:47:29Z")

</div>

We are solving differential equations defined on a graph using NetworkDynamics.jl and DifferentialEquations.jl. In this context we have the two questions outlined below. Thx. Domenico Lahaye. 1/ Switching the automatic…

---

## [Better way of replacing Matrix values with another Array values](https://discourse.julialang.org/t/better-way-of-replacing-matrix-values-with-another-array-values/63440)

<div class="topic-metadata">

**Author:** [@eduardovrs](https://discourse.julialang.org/u/eduardovrs)\
**Replies:** 6\
**Last updated:** [June 24, 2021, 11:26am UTC](https://discourse.julialang.org/t/better-way-of-replacing-matrix-values-with-another-array-values/63440 "2021-06-24T11:26:45Z")

</div>

Dear Julia community, I want to change some values of an existing Matrix with some values that are on another array, but I find that this part of my code is the bottleneck of what I am doing. Is there a better way to do…

---

## [Garbage collector behaviour when memory is almost full](https://discourse.julialang.org/t/garbage-collector-behaviour-when-memory-is-almost-full/28169)

<div class="topic-metadata">

**Author:** [@rssdev10](https://discourse.julialang.org/u/rssdev10)\
**Replies:** 7\
**Last updated:** [June 24, 2021, 4:59am UTC](https://discourse.julialang.org/t/garbage-collector-behaviour-when-memory-is-almost-full/28169 "2021-06-24T04:59:15Z")

</div>

I have a script for calculation of millions of strings distances with StringDistances.jl. I’m using 64 CPUs virtual machine with 240GB RAM and found that my script was killed by out of memory. I cleared the script to min…

---

## [KeyError on additional processes while using Distributed](https://discourse.julialang.org/t/keyerror-on-additional-processes-while-using-distributed/63437)

<div class="topic-metadata">

**Author:** [@vedang](https://discourse.julialang.org/u/vedang)\
**Replies:** 0\
**Last updated:** [June 23, 2021, 12:29pm UTC](https://discourse.julialang.org/t/keyerror-on-additional-processes-while-using-distributed/63437 "2021-06-23T12:29:25Z")

</div>

I am new to using Julia. I am trying to solve an EnsembleProblem using EnsembleDistributed() option in DifferentialEquations.jl. I have created multiple processes using: using Distributed addprocs(1,exeflags="--project…

---

## [Huge performance difference between inbuilt gcdx() and my function](https://discourse.julialang.org/t/huge-performance-difference-between-inbuilt-gcdx-and-my-function/63350)

<div class="topic-metadata">

**Author:** [@msravi](https://discourse.julialang.org/u/msravi)\
**Replies:** 3\
**Last updated:** [June 22, 2021, 8:09am UTC](https://discourse.julialang.org/t/huge-performance-difference-between-inbuilt-gcdx-and-my-function/63350 "2021-06-22T08:09:12Z")

</div>

Hi All, I had written a Bezout’s identity function (before I knew Julia already had one in the form of gcdx), and although my implementation is very close to the one that comes with Julia, I see that the inbuilt functio…

---

## [How to do SIMD code with wide-register accumulators (@simd vs LoopVectorization.jl vs SIMD.jl)](https://discourse.julialang.org/t/how-to-do-simd-code-with-wide-register-accumulators-simd-vs-loopvectorization-jl-vs-simd-jl/63322)

<div class="topic-metadata">

**Author:** [@Krastanov](https://discourse.julialang.org/u/Krastanov)\
**Replies:** 11\
**Last updated:** [June 22, 2021, 12:33am UTC](https://discourse.julialang.org/t/how-to-do-simd-code-with-wide-register-accumulators-simd-vs-loopvectorization-jl-vs-simd-jl/63322 "2021-06-22T00:33:37Z")

</div>

I have a loop that would benefit from SIMD operations, but it has one “accumulator” variable that confuses most SIMD tools available. Here is the pseudo code, and below is the actual code with benchmarks. My questions is…

---

## [Threads parallelization on different nested loops](https://discourse.julialang.org/t/threads-parallelization-on-different-nested-loops/63295)

<div class="topic-metadata">

**Author:** [@calvar](https://discourse.julialang.org/u/calvar)\
**Replies:** 0\
**Last updated:** [June 21, 2021, 1:40pm UTC](https://discourse.julialang.org/t/threads-parallelization-on-different-nested-loops/63295 "2021-06-21T13:40:55Z")

</div>

Hi, I am learning Julia and I was testing how to parallelize a nested loop with the @threads macro, to compute the product between matrices. This is the code I am trying to parallelize: function dot\_prod!(A, B, C, i, j…

---

## [MLJFlux is a lot slower than the same algorithm written in Flux](https://discourse.julialang.org/t/mljflux-is-a-lot-slower-than-the-same-algorithm-written-in-flux/63253)

<div class="topic-metadata">

**Author:** [@ultrapoci](https://discourse.julialang.org/u/ultrapoci)\
**Replies:** 3\
**Last updated:** [June 21, 2021, 12:30pm UTC](https://discourse.julialang.org/t/mljflux-is-a-lot-slower-than-the-same-algorithm-written-in-flux/63253 "2021-06-21T12:30:36Z")

</div>

Disclaimer: I am by no means an expert in Flux, MLJFlux, Julia or even machine learning. I may have done something very stupid without noticing. I have a neural network classifier written in Flux. It works fine, it give…

---

## [How to choose vec size in SIMD.jl](https://discourse.julialang.org/t/how-to-choose-vec-size-in-simd-jl/63269)

<div class="topic-metadata">

**Author:** [@Krastanov](https://discourse.julialang.org/u/Krastanov)\
**Replies:** 5\
**Last updated:** [June 21, 2021, 4:32am UTC](https://discourse.julialang.org/t/how-to-choose-vec-size-in-simd-jl/63269 "2021-06-21T04:32:47Z")

</div>

Looking at the SIMD.jl example vadd!(xs, ys, Vec{8,Float64}) there is a specific size 8 being set for the vec size. How do I know whether to use 4 or 8 or more? I assume it is hardware dependent, but probably there is so…

---

## [Understanding major order performance when broadcasting in column vs row operations](https://discourse.julialang.org/t/understanding-major-order-performance-when-broadcasting-in-column-vs-row-operations/63139)

<div class="topic-metadata">

**Author:** [@gzagatti](https://discourse.julialang.org/u/gzagatti)\
**Replies:** 9\
**Last updated:** [June 21, 2021, 3:07am UTC](https://discourse.julialang.org/t/understanding-major-order-performance-when-broadcasting-in-column-vs-row-operations/63139 "2021-06-21T03:07:54Z")

</div>

I understand that Julia stores arrays in column-major order which means that columns are stacked onto one another. Thus, adjacent rows in the same column are adjacent in memory. If that is the case, I am failing to see …

---

## [Generators vs loops vs broadcasting: Calculate PI via Monte Carlo Sampling](https://discourse.julialang.org/t/generators-vs-loops-vs-broadcasting-calculate-pi-via-monte-carlo-sampling/63235)

<div class="topic-metadata">

**Author:** [@JanKap](https://discourse.julialang.org/u/JanKap)\
**Replies:** 7\
**Last updated:** [June 20, 2021, 10:15pm UTC](https://discourse.julialang.org/t/generators-vs-loops-vs-broadcasting-calculate-pi-via-monte-carlo-sampling/63235 "2021-06-20T22:15:41Z")

</div>

Dear all, I’m trying to solve some toy problems in Julia and benchmark them since there often are many different ways to go. I have a strong Matlab background therefore it feels a bit weird to go back to simple loops (an…

---

## [Improving Computational Performance](https://discourse.julialang.org/t/improving-computational-performance/63210)

<div class="topic-metadata">

**Author:** [@nullmap](https://discourse.julialang.org/u/nullmap)\
**Replies:** 6\
**Last updated:** [June 20, 2021, 10:13pm UTC](https://discourse.julialang.org/t/improving-computational-performance/63210 "2021-06-20T22:13:29Z")

</div>

I have three questions and I was curious if someone can provide some clarity on the results I’m seeing with a couple variations of a simple polynomial evaluation computation. I wrote a Horner’s method implementation whi…

---

## [More effective function parameter?](https://discourse.julialang.org/t/more-effective-function-parameter/63160)

<div class="topic-metadata">

**Author:** [@Jojo\_Dad](https://discourse.julialang.org/u/Jojo_Dad)\
**Replies:** 9\
**Last updated:** [June 19, 2021, 10:38pm UTC](https://discourse.julialang.org/t/more-effective-function-parameter/63160 "2021-06-19T22:38:31Z")

</div>

I have a function that I frequently use function expectation(f::Function; μ=0.0, σ=1.0) and it works great for expectation(x -\> (x-1)^2) as well as a many other examples. However, when I have a variable, e.g., avg=…

---

## [Efficient way to simulate large system of ODEs](https://discourse.julialang.org/t/efficient-way-to-simulate-large-system-of-odes/63057)

<div class="topic-metadata">

**Author:** [@colebrookson](https://discourse.julialang.org/u/colebrookson)\
**Replies:** 5\
**Last updated:** [June 18, 2021, 10:18pm UTC](https://discourse.julialang.org/t/efficient-way-to-simulate-large-system-of-odes/63057 "2021-06-18T22:18:49Z")

</div>

Edit: So the way I proposed doing it here does work, I just made a mistake that @Oscar\_Smith pointed out, so using the functions this way is fine. Original Q: Hi there! I’m still a bit new to Julia and trying to figure…

---

## [Pairwise computation slower than Python (Cython) code (BallTree very slow!)](https://discourse.julialang.org/t/pairwise-computation-slower-than-python-cython-code-balltree-very-slow/62273)

<div class="topic-metadata">

**Author:** [@florpi](https://discourse.julialang.org/u/florpi)\
**Replies:** 27\
**Last updated:** [June 18, 2021, 3:08pm UTC](https://discourse.julialang.org/t/pairwise-computation-slower-than-python-cython-code-balltree-very-slow/62273 "2021-06-18T15:08:07Z")

</div>

Hi everyone ! I’m new to Julia, and so far have found it very nice and neat. However, my Julia implementation is still slower than the python one I was trying to beat. The goal is to compute the mean radial pairwise vel…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=88)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=90)
