# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=107

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 108

---

## [Flux model on CPU runs slowly](https://discourse.julialang.org/t/flux-model-on-cpu-runs-slowly/47723)

<div class="topic-metadata">

**Author:** [@AquaIndigo](https://discourse.julialang.org/u/AquaIndigo)\
**Replies:** 3\
**Last updated:** [October 4, 2020, 2:54am UTC](https://discourse.julialang.org/t/flux-model-on-cpu-runs-slowly/47723 "2020-10-04T02:54:10Z")

</div>

I find the Flux model on CPU runs more slowly than that on GPU: julia\> m = Chain( Dense(2250, 500, σ), Dense(500, 50, tanh), Dense(50, 7, σ)); julia\> X = Float32.(X); julia\> size(X) (…

---

## [Executing array of functions much slower than executing functions in series](https://discourse.julialang.org/t/executing-array-of-functions-much-slower-than-executing-functions-in-series/47674)

<div class="topic-metadata">

**Author:** [@Yamen](https://discourse.julialang.org/u/Yamen)\
**Replies:** 10\
**Last updated:** [October 3, 2020, 8:32pm UTC](https://discourse.julialang.org/t/executing-array-of-functions-much-slower-than-executing-functions-in-series/47674 "2020-10-03T20:32:14Z")

</div>

I have a very performance sensitive application that needs to activate a neural network on a per row basis for a large number of rows. Consider a network with 4 input values and 1 output value, with a few layers in betw…

---

## [Iteration using array of arrays/matrices vs iteration over matrix subsets](https://discourse.julialang.org/t/iteration-using-array-of-arrays-matrices-vs-iteration-over-matrix-subsets/47661)

<div class="topic-metadata">

**Author:** [@jleman](https://discourse.julialang.org/u/jleman)\
**Replies:** 2\
**Last updated:** [October 3, 2020, 12:00am UTC](https://discourse.julialang.org/t/iteration-using-array-of-arrays-matrices-vs-iteration-over-matrix-subsets/47661 "2020-10-03T00:00:15Z")

</div>

It seems to me that in past versions of Julia the array of arrays or array of matrices approach for iteration usually outperformed iterating over matrix subsets. That gap seems to be smaller in the more recent version. …

---

## [Are Python packages called in julia faster than in python](https://discourse.julialang.org/t/are-python-packages-called-in-julia-faster-than-in-python/47644)

<div class="topic-metadata">

**Author:** [@MalteMederacke](https://discourse.julialang.org/u/MalteMederacke)\
**Replies:** 12\
**Last updated:** [October 2, 2020, 5:04pm UTC](https://discourse.julialang.org/t/are-python-packages-called-in-julia-faster-than-in-python/47644 "2020-10-02T17:04:10Z")

</div>

Hi guys, basicly this. I am not a computer scientist, but I am curious, if I call a python package with julia, does it run faster than in python? Or is it the same or even slower.

---

## [How to time julia's choleskey factorization?](https://discourse.julialang.org/t/how-to-time-julias-choleskey-factorization/47627)

<div class="topic-metadata">

**Author:** [@vkv](https://discourse.julialang.org/u/vkv)\
**Replies:** 8\
**Last updated:** [October 2, 2020, 2:44pm UTC](https://discourse.julialang.org/t/how-to-time-julias-choleskey-factorization/47627 "2020-10-02T14:44:03Z")

</div>

Hi, I’m trying to compare timings of julia’s cholesky factorization and LAPACK.potrf! function. The timings of julia’s cholesky function is varying from 50 \\mu s to 100 \\mu s and LAPACK.potrf! is blazingly fast varying…

---

## [Collecting homogenized vectors: modify iterator vs modify collector](https://discourse.julialang.org/t/collecting-homogenized-vectors-modify-iterator-vs-modify-collector/42957)

<div class="topic-metadata">

**Author:** [@BMO](https://discourse.julialang.org/u/BMO)\
**Replies:** 3\
**Last updated:** [September 30, 2020, 5:24pm UTC](https://discourse.julialang.org/t/collecting-homogenized-vectors-modify-iterator-vs-modify-collector/42957 "2020-09-30T17:24:38Z")

</div>

Nemo is a computer algebra package for the Julia programming language. For a multivariated polynomial p there is an iterator exponent\_vectors(p) whose values are vectors corresponding to the exponents of the terms of p (…

---

## [Why might I be seeing a large overhead for multiprocessing?](https://discourse.julialang.org/t/why-might-i-be-seeing-a-large-overhead-for-multiprocessing/47547)

<div class="topic-metadata">

**Author:** [@JamesK](https://discourse.julialang.org/u/JamesK)\
**Replies:** 8\
**Last updated:** [October 1, 2020, 5:07am UTC](https://discourse.julialang.org/t/why-might-i-be-seeing-a-large-overhead-for-multiprocessing/47547 "2020-10-01T05:07:07Z")

</div>

I have an 8 core machine. When I run a certain piece of code with Python’s multiprocessing.Pool library, with 1 core vs. 4 cores, I see an almost 4x speedup for the 4 core case, with very little overhead penalty, assumi…

---

## [Dataframes with groupby in parallel](https://discourse.julialang.org/t/dataframes-with-groupby-in-parallel/47186)

<div class="topic-metadata">

**Author:** [@bicepjai](https://discourse.julialang.org/u/bicepjai)\
**Replies:** 1\
**Last updated:** [October 1, 2020, 4:39am UTC](https://discourse.julialang.org/t/dataframes-with-groupby-in-parallel/47186 "2020-10-01T04:39:06Z")

</div>

I have a dataframe of size (2450225, 117) Just using pandas and multiprocessing module, gives me the results in 34 seconds. I have panda code below. I am trying to achieve the same with julia and hoping this would be a …

---

## [Correct way to dereference large memory?](https://discourse.julialang.org/t/correct-way-to-dereference-large-memory/47395)

<div class="topic-metadata">

**Author:** [@Ibb](https://discourse.julialang.org/u/Ibb)\
**Replies:** 6\
**Last updated:** [September 28, 2020, 1:34pm UTC](https://discourse.julialang.org/t/correct-way-to-dereference-large-memory/47395 "2020-09-28T13:34:02Z")

</div>

I want to read binary files into some large tuples by Ref, but my julia stuck at code below. julia\> a = Ref{NTuple{36000, UInt8}}() julia\> a\[\] #julia v1.5.2 will stuck here

---

## [Fast copy from stream to IOBuffer?](https://discourse.julialang.org/t/fast-copy-from-stream-to-iobuffer/42099)

<div class="topic-metadata">

**Author:** [@bobcassels](https://discourse.julialang.org/u/bobcassels)\
**Replies:** 7\
**Last updated:** [September 28, 2020, 12:49am UTC](https://discourse.julialang.org/t/fast-copy-from-stream-to-iobuffer/42099 "2020-09-28T00:49:54Z")

</div>

I need to read parts of a (binary) stream, and concatenate them, so I can read from that concatenated buffer. I’m reconstructing a thing that was packetized into segments. I see I can do write(iob, read(s, ...)), but I …

---

## [Can this random matrix simulation be sped up?](https://discourse.julialang.org/t/can-this-random-matrix-simulation-be-sped-up/47277)

<div class="topic-metadata">

**Author:** [@patrick](https://discourse.julialang.org/u/patrick)\
**Replies:** 16\
**Last updated:** [September 27, 2020, 11:09pm UTC](https://discourse.julialang.org/t/can-this-random-matrix-simulation-be-sped-up/47277 "2020-09-27T23:09:19Z")

</div>

I am a Julia novice and converted the following program from Python (for a simple practice exercise). It computes a N by N random matrix of quaternions, drawn from the Gaussian distribution, and then computes the smalles…

---

## [Matrix vector multiplication](https://discourse.julialang.org/t/matrix-vector-multiplication/47356)

<div class="topic-metadata">

**Author:** [@j\_tanu](https://discourse.julialang.org/u/j_tanu)\
**Replies:** 4\
**Last updated:** [September 27, 2020, 4:50pm UTC](https://discourse.julialang.org/t/matrix-vector-multiplication/47356 "2020-09-27T16:50:10Z")

</div>

I am doing a dense matrix-vector multiplication in a 64 core workstation. The order of the matrix is 11000. By default, BLAS is using only 8 threads. I have changed both the BLAS & JULIA(1.5.1) NUM\_THREADS and getting th…

---

## [DiffEqFlux.sciml\_train is slow](https://discourse.julialang.org/t/diffeqflux-sciml-train-is-slow/47155)

<div class="topic-metadata">

**Author:** [@checco.jike](https://discourse.julialang.org/u/checco.jike)\
**Replies:** 5\
**Last updated:** [September 25, 2020, 7:28pm UTC](https://discourse.julialang.org/t/diffeqflux-sciml-train-is-slow/47155 "2020-09-25T19:28:18Z")

</div>

Hi all! I’m trying to train the Universal Differential Equations learning with DiffEqFlux.sciml\_train Here is the code relative to the 1D problem: https://github.com/ChrisRackauckas/universal\_differential\_equations/blo…

---

## [Why does changing this indexing slows the code so much?](https://discourse.julialang.org/t/why-does-changing-this-indexing-slows-the-code-so-much/47246)

<div class="topic-metadata">

**Author:** [@disberd](https://discourse.julialang.org/u/disberd)\
**Replies:** 5\
**Last updated:** [September 25, 2020, 12:30pm UTC](https://discourse.julialang.org/t/why-does-changing-this-indexing-slows-the-code-so-much/47246 "2020-09-25T12:30:28Z")

</div>

Hi, I was trying to play around with the implementation of fsortperm in SortingLab.jl and wanted to try making an implementation to just compute directly the sorted indices rather than sorting both the values vector and…

---

## [Why is this code run-time dispatch/slow?](https://discourse.julialang.org/t/why-is-this-code-run-time-dispatch-slow/47178)

<div class="topic-metadata">

**Author:** [@JamesK](https://discourse.julialang.org/u/JamesK)\
**Replies:** 16\
**Last updated:** [September 25, 2020, 10:34am UTC](https://discourse.julialang.org/t/why-is-this-code-run-time-dispatch-slow/47178 "2020-09-25T10:34:29Z")

</div>

I made a minimum reproducible example for a performance issue I am having with my code: mutable struct Array\_container array\_field::Array{Float64} end function foo(a::Array\_container, b::Array\_container) a.arra…

---

## [Understanding and avoiding allocations with StructArrays](https://discourse.julialang.org/t/understanding-and-avoiding-allocations-with-structarrays/46880)

<div class="topic-metadata">

**Author:** [@touste](https://discourse.julialang.org/u/touste)\
**Replies:** 16\
**Last updated:** [September 25, 2020, 9:42am UTC](https://discourse.julialang.org/t/understanding-and-avoiding-allocations-with-structarrays/46880 "2020-09-25T09:42:15Z")

</div>

Hi all, I’m working on a piece of code where I have triangles which are made of points. Each point has a number of properties such as position, velocity, force… In some part of the code I need to calculate the force fo…

---

## [Performance Help on Large Matrix Manipulation](https://discourse.julialang.org/t/performance-help-on-large-matrix-manipulation/47029)

<div class="topic-metadata">

**Author:** [@jleman](https://discourse.julialang.org/u/jleman)\
**Replies:** 16\
**Last updated:** [September 24, 2020, 6:08am UTC](https://discourse.julialang.org/t/performance-help-on-large-matrix-manipulation/47029 "2020-09-24T06:08:35Z")

</div>

I have a large array of matricies and need to apply inverse FFT as shown in the attached diagram. I do this by restructuring the data using getindex . After completing the inverse FFT I then recreating the original matr…

---

## [How to speed up tasks?](https://discourse.julialang.org/t/how-to-speed-up-tasks/46511)

<div class="topic-metadata">

**Author:** [@bjarthur](https://discourse.julialang.org/u/bjarthur)\
**Replies:** 8\
**Last updated:** [September 23, 2020, 9:58pm UTC](https://discourse.julialang.org/t/how-to-speed-up-tasks/46511 "2020-09-23T21:58:59Z")

</div>

i have a function that i need to call 10s of thousands of times that only takes a few milliseconds to run. there is no I/O. and it doesn’t return anything, so no communication. it is strictly just compute. how do i b…

---

## [\[ANN\] CuCountMap.jl - CUDA.jl-enabled faster \`StatsBase.countmap\` for small types](https://discourse.julialang.org/t/ann-cucountmap-jl-cuda-jl-enabled-faster-statsbase-countmap-for-small-types/47125)

<div class="topic-metadata">

**Author:** [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Replies:** 0\
**Last updated:** [September 23, 2020, 9:45am UTC](https://discourse.julialang.org/t/ann-cucountmap-jl-cuda-jl-enabled-faster-statsbase-countmap-for-small-types/47125 "2020-09-23T09:45:30Z")

</div>

See GitHub - xiaodaigh/CuCountMap.jl: Fast \`StatsBase.countmap\` for small types on the GPU via CUDA.jl I can get about 3x the performance for small types on the GPU via CUDA.jl vs a purely CPU implementation. This is in…

---

## [Help eliminating allocations in a call to a Fortran subroutine](https://discourse.julialang.org/t/help-eliminating-allocations-in-a-call-to-a-fortran-subroutine/47078)

<div class="topic-metadata">

**Author:** [@zyth0s](https://discourse.julialang.org/u/zyth0s)\
**Replies:** 2\
**Last updated:** [September 22, 2020, 5:09pm UTC](https://discourse.julialang.org/t/help-eliminating-allocations-in-a-call-to-a-fortran-subroutine/47078 "2020-09-22T17:09:17Z")

</div>

I would like to ask you for help to improve a call to a critical routine written in Fortran. A minimal version of the library is the following Fortran library sourcemodule fortran\_lib contains subroutine density\_gra…

---

## [Improving the code speed by employing parallelism for asynchronous task](https://discourse.julialang.org/t/improving-the-code-speed-by-employing-parallelism-for-asynchronous-task/47041)

<div class="topic-metadata">

**Author:** [@Nova](https://discourse.julialang.org/u/Nova)\
**Replies:** 5\
**Last updated:** [September 22, 2020, 7:06am UTC](https://discourse.julialang.org/t/improving-the-code-speed-by-employing-parallelism-for-asynchronous-task/47041 "2020-09-22T07:06:01Z")

</div>

I was trying to employ parallelism to improve the speed of the code using multi-threading. However, I noticed that I get wrong answer using multi-threading. The reason for that is my code update a value in a for loop con…

---

## [Scattered Atomic Writes Into Array](https://discourse.julialang.org/t/scattered-atomic-writes-into-array/46990)

<div class="topic-metadata">

**Author:** [@cshenton](https://discourse.julialang.org/u/cshenton)\
**Replies:** 6\
**Last updated:** [September 21, 2020, 3:41pm UTC](https://discourse.julialang.org/t/scattered-atomic-writes-into-array/46990 "2020-09-21T15:41:06Z")

</div>

I’m porting some OpenCL code that does scattered atomic writes. These writes are quite sparse, meaning that threads rarely try to write cache lines at the same time. I’m trying to figure out how to do this in Julia but …

---

## [Load sysimage project dependent (in Atom/VSCode)](https://discourse.julialang.org/t/load-sysimage-project-dependent-in-atom-vscode/46995)

<div class="topic-metadata">

**Author:** [@SteffenPL](https://discourse.julialang.org/u/SteffenPL)\
**Replies:** 1\
**Last updated:** [September 21, 2020, 11:23am UTC](https://discourse.julialang.org/t/load-sysimage-project-dependent-in-atom-vscode/46995 "2020-09-21T11:23:14Z")

</div>

Is it possible to load sysimages project dependent, in the following sense: If the project folder contains a sysimage, then the sysimage is loaded, otherwise not. I would know how to do this via scripts, but I don’t kn…

---

## [Comparison of languages for parallel computing tasks](https://discourse.julialang.org/t/comparison-of-languages-for-parallel-computing-tasks/46806)

<div class="topic-metadata">

**Author:** [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Replies:** 11\
**Last updated:** [September 21, 2020, 9:11am UTC](https://discourse.julialang.org/t/comparison-of-languages-for-parallel-computing-tasks/46806 "2020-09-21T09:11:40Z")

</div>

The paper https://www.sciencedirect.com/science/article/abs/pii/S2210650220303734 https://lilloa.univ-lille.fr/handle/20.500.12210/29528 compares Chapel, Python, and Julia. For reasons that are at this point not crisp…

---

## [Slow Cholesky factorization for sparse matrices?](https://discourse.julialang.org/t/slow-cholesky-factorization-for-sparse-matrices/46909)

<div class="topic-metadata">

**Author:** [@Albert\_de\_montserrat](https://discourse.julialang.org/u/Albert_de_montserrat)\
**Replies:** 6\
**Last updated:** [September 19, 2020, 7:45pm UTC](https://discourse.julialang.org/t/slow-cholesky-factorization-for-sparse-matrices/46909 "2020-09-19T19:45:05Z")

</div>

I am trying to compare Julia’s Cholesky factorization against MATLAB for sparse matrices, and Julia seems almost 2 times slower than MATLAB. Julia: using SparseArrays,BenchmarkTools, LinearAlgebra n = Int(5e3) A = rand…

---

## [Improving Barnes-Hut n-body simulation performance](https://discourse.julialang.org/t/improving-barnes-hut-n-body-simulation-performance/20864)

<div class="topic-metadata">

**Author:** [@novoselrok](https://discourse.julialang.org/u/novoselrok)\
**Replies:** 14\
**Last updated:** [September 19, 2020, 5:03pm UTC](https://discourse.julialang.org/t/improving-barnes-hut-n-body-simulation-performance/20864 "2020-09-19T17:03:17Z")

</div>

Hi. I’m writing a Barnes-Hut simulator in Julia but I’m experiencing a slowdown compared to the C version. For 80k bodies: C - 25 seconds Julia - 35 seconds As expected it spends 99.9% time in the compute\_force func…

---

## [Its possible to use all cores or parallel cores to do some project/calculations?](https://discourse.julialang.org/t/its-possible-to-use-all-cores-or-parallel-cores-to-do-some-project-calculations/46639)

<div class="topic-metadata">

**Author:** [@Lucas\_Pollito](https://discourse.julialang.org/u/Lucas_Pollito)\
**Replies:** 41\
**Last updated:** [September 19, 2020, 4:05pm UTC](https://discourse.julialang.org/t/its-possible-to-use-all-cores-or-parallel-cores-to-do-some-project-calculations/46639 "2020-09-19T16:05:58Z")

</div>

I have three different and independent task/calculation to do. function interpQa() ag = 0.01:0.05:3 aσ = 0.03 : 0.025 : 0.63 if PrimeiraVez mQ = \[Q(g,sigma) for g in ag, sigma in aσ\] w…

---

## [Avoid allocations in the naive discrete Fourier transform](https://discourse.julialang.org/t/avoid-allocations-in-the-naive-discrete-fourier-transform/46847)

<div class="topic-metadata">

**Author:** [@stakaz](https://discourse.julialang.org/u/stakaz)\
**Replies:** 12\
**Last updated:** [September 19, 2020, 12:23am UTC](https://discourse.julialang.org/t/avoid-allocations-in-the-naive-discrete-fourier-transform/46847 "2020-09-19T00:23:38Z")

</div>

Hello, I try to implemente a straight forwared discrete Fourier transform abut I do not see what variables are allocated in the intermediate steps. For my undertanding it is a generator expression and should get the resu…

---

## [How to track dynamic dispatch](https://discourse.julialang.org/t/how-to-track-dynamic-dispatch/46826)

<div class="topic-metadata">

**Author:** [@tisztamo](https://discourse.julialang.org/u/tisztamo)\
**Replies:** 4\
**Last updated:** [September 18, 2020, 9:07am UTC](https://discourse.julialang.org/t/how-to-track-dynamic-dispatch/46826 "2020-09-18T09:07:29Z")

</div>

Similar to --track-allocation, I would like to see where in my code dynamic dispatch happened during a specific run. Is there a tool for that?

---

## [How to searchsortedfirst when each element requires expensive transformation](https://discourse.julialang.org/t/how-to-searchsortedfirst-when-each-element-requires-expensive-transformation/46720)

<div class="topic-metadata">

**Author:** [@MFairley](https://discourse.julialang.org/u/MFairley)\
**Replies:** 24\
**Last updated:** [September 17, 2020, 8:49pm UTC](https://discourse.julialang.org/t/how-to-searchsortedfirst-when-each-element-requires-expensive-transformation/46720 "2020-09-17T20:49:40Z")

</div>

I have a problem where I have an expensive function expensive() and I want to find the first integer i in 0:n such that expensive(i) \> a where a is some constant and we know that expensive(j) \> a for j \>= i. I want to m…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=106)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=108)
