# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=73

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 74

---

## [CUDA global synchronization HOWTO](https://discourse.julialang.org/t/cuda-global-synchronization-howto/74920)

<div class="topic-metadata">

**Author:** [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Replies:** 9\
**Last updated:** [January 20, 2022, 12:52pm UTC](https://discourse.julialang.org/t/cuda-global-synchronization-howto/74920 "2022-01-20T12:52:47Z")

</div>

Dear Julianners, I try to create an algorithm that runs an elementwise update operation and a reduction in 10k iteration and about 1\_000\_000 times, so the kernel restarts(2-8us) are really expensive in this scenario. T…

---

## [Arch D. Robison high performance Julia workshop from 2016 in 2022](https://discourse.julialang.org/t/arch-d-robison-high-performance-julia-workshop-from-2016-in-2022/74932)

<div class="topic-metadata">

**Author:** [@KZiemian](https://discourse.julialang.org/u/KZiemian)\
**Replies:** 0\
**Last updated:** [January 20, 2022, 12:13pm UTC](https://discourse.julialang.org/t/arch-d-robison-high-performance-julia-workshop-from-2016-in-2022/74932 "2022-01-20T12:13:07Z")

</div>

Arch D. Robison workshop Introduction to Writing High Performance Julia presented at JuliaCon 2016 thought me so many things and I believe that many Julia users still can learn valuable lessons from it. At the same time,…

---

## [Out of GPU memory with user defined function](https://discourse.julialang.org/t/out-of-gpu-memory-with-user-defined-function/74914)

<div class="topic-metadata">

**Author:** [@kadir-gunel](https://discourse.julialang.org/u/kadir-gunel)\
**Replies:** 1\
**Last updated:** [January 20, 2022, 8:40am UTC](https://discourse.julialang.org/t/out-of-gpu-memory-with-user-defined-function/74914 "2022-01-20T08:40:00Z")

</div>

Hello, I am getting out of memory error with CUDA while trying to multiply 2 matrices of size (300x40k) inside a user defined function. I understand that X and Y variables are copied hence the memory is insufficient …

---

## [What's a good and easy way to parallelize this?](https://discourse.julialang.org/t/whats-a-good-and-easy-way-to-parallelize-this/74893)

<div class="topic-metadata">

**Author:** [@amrods](https://discourse.julialang.org/u/amrods)\
**Replies:** 3\
**Last updated:** [January 20, 2022, 5:35am UTC](https://discourse.julialang.org/t/whats-a-good-and-easy-way-to-parallelize-this/74893 "2022-01-20T05:35:21Z")

</div>

Ultimately, I want to speed up fullsys!, and as you can see there are functions that fill arrays. I thought those filling operations could be parallelized. How can I get started with that? const szAge = 64 - 19 + 1 cons…

---

## [Where is this one allocation coming from?](https://discourse.julialang.org/t/where-is-this-one-allocation-coming-from/72693)

<div class="topic-metadata">

**Author:** [@amrods](https://discourse.julialang.org/u/amrods)\
**Replies:** 16\
**Last updated:** [January 20, 2022, 1:11am UTC](https://discourse.julialang.org/t/where-is-this-one-allocation-coming-from/72693 "2022-01-20T01:11:47Z")

</div>

I have to run functions like this many times. I’m interested in reducing allocations and don’t see where that one allocation is coming from. using BenchmarkTools using ComponentArrays using UnPack α0 = rand(2, 2, 5) α …

---

## [A macro to unroll by hand but not by hand?](https://discourse.julialang.org/t/a-macro-to-unroll-by-hand-but-not-by-hand/74633)

<div class="topic-metadata">

**Author:** [@lrnv](https://discourse.julialang.org/u/lrnv)\
**Replies:** 24\
**Last updated:** [January 19, 2022, 7:43pm UTC](https://discourse.julialang.org/t/a-macro-to-unroll-by-hand-but-not-by-hand/74633 "2022-01-19T19:43:34Z")

</div>

Hey, I have a computation that we already discussed there and which still takes up more than half of my total runtime (so days…). I did found a very clunky way to cut in half it’s runtime by manually unrolling a loop, …

---

## [Knet on Julia 1.7.0 (?)](https://discourse.julialang.org/t/knet-on-julia-1-7-0/74860)

<div class="topic-metadata">

**Author:** [@mazzanti](https://discourse.julialang.org/u/mazzanti)\
**Replies:** 2\
**Last updated:** [January 19, 2022, 12:07pm UTC](https://discourse.julialang.org/t/knet-on-julia-1-7-0/74860 "2022-01-19T12:07:12Z")

</div>

Hi, I just wanted to give Knet a go in my ubuntu machine running Julia 1.7.0, but surprisingly things fail… I just installed it as standard @(v1.7) pkg\> add Knet @(v1.7) pkg\> test Knet @(v1.7) pkg\> test Knet … Test…

---

## [JuliaOptics performance discussion](https://discourse.julialang.org/t/juliaoptics-performance-discussion/74835)

<div class="topic-metadata">

**Author:** [@martin.d.maas](https://discourse.julialang.org/u/martin.d.maas)\
**Replies:** 7\
**Last updated:** [January 19, 2022, 4:20am UTC](https://discourse.julialang.org/t/juliaoptics-performance-discussion/74835 "2022-01-19T04:20:24Z")

</div>

Hi! I just happened to remember that this abandoned package came out in this forum during a discussion of performance benchmarks, and users’ expectations of performance in Julia vs other languages. As I happened to be …

---

## [Is there a way to speed up this matrix (with only 1 & 0) multiplication & power operation?](https://discourse.julialang.org/t/is-there-a-way-to-speed-up-this-matrix-with-only-1-0-multiplication-power-operation/74747)

<div class="topic-metadata">

**Author:** [@Endeavour](https://discourse.julialang.org/u/Endeavour)\
**Replies:** 20\
**Last updated:** [January 18, 2022, 5:00pm UTC](https://discourse.julialang.org/t/is-there-a-way-to-speed-up-this-matrix-with-only-1-0-multiplication-power-operation/74747 "2022-01-18T17:00:56Z")

</div>

I have a matrix (cost\_mat\[~600,~600\]) with only 1 & 0 as its elements. You can think that each element of the matrix is individually put as 1 (with probability p) or 0. I need to find out (cost\_mat^N) where N(~600) is a …

---

## [Perform multiple replacements on a string in a single pass](https://discourse.julialang.org/t/perform-multiple-replacements-on-a-string-in-a-single-pass/43247)

<div class="topic-metadata">

**Author:** [@nstgc](https://discourse.julialang.org/u/nstgc)\
**Replies:** 19\
**Last updated:** [January 18, 2022, 2:13pm UTC](https://discourse.julialang.org/t/perform-multiple-replacements-on-a-string-in-a-single-pass/43247 "2022-01-18T14:13:17Z")

</div>

I have a couple of scripts which perform multiple replacements on the same string. It seems as though this can be sped up by searching for all potential matches. For example, something like julia\>replace("123",\[r"1" =\> …

---

## [macOS M1, Julia 1.6, and swtch\_pri](https://discourse.julialang.org/t/macos-m1-julia-1-6-and-swtch-pri/59826)

<div class="topic-metadata">

**Author:** [@GlenHenshaw](https://discourse.julialang.org/u/GlenHenshaw)\
**Replies:** 1\
**Last updated:** [January 18, 2022, 1:38am UTC](https://discourse.julialang.org/t/macos-m1-julia-1-6-and-swtch-pri/59826 "2022-01-18T01:38:29Z")

</div>

Hi all, Using a MacBook Air M1 and Julia 1.6. I’m running a fairly complex scientific code that does a ton of matrix vector calculations. When I profile it and display the results using PProf, it shows that approximate…

---

## [Runtime (memory) on M1 Macbooks: something is not right](https://discourse.julialang.org/t/runtime-memory-on-m1-macbooks-something-is-not-right/74721)

<div class="topic-metadata">

**Author:** [@Schneeschaufel](https://discourse.julialang.org/u/Schneeschaufel)\
**Replies:** 10\
**Last updated:** [January 17, 2022, 11:38am UTC](https://discourse.julialang.org/t/runtime-memory-on-m1-macbooks-something-is-not-right/74721 "2022-01-17T11:38:05Z")

</div>

The following link: Trying to understand memory usage If I do this (see link above) on my Macbook 13" with M1 processor and 8GB RAM and SSD (Julia Rosetta Version 1.7.1 (2021-12-22)): julia\> using LinearAlgebra julia\> …

---

## [Julia vs Zig surprise](https://discourse.julialang.org/t/julia-vs-zig-surprise/74540)

<div class="topic-metadata">

**Author:** [@Bardo](https://discourse.julialang.org/u/Bardo)\
**Replies:** 16\
**Last updated:** [January 16, 2022, 4:29pm UTC](https://discourse.julialang.org/t/julia-vs-zig-surprise/74540 "2022-01-16T16:29:38Z")

</div>

Recently ran into two performance comparisons including Julia. THE LINEAR ALGEBRA MAPPING PROBLEM. CURRENT STATE OF LINEAR ALGEBRA LANGUAGES AND LIBRARIES. shows that using BLAS for dense matrix operations, all librar…

---

## [Type stability with variable arguments](https://discourse.julialang.org/t/type-stability-with-variable-arguments/74640)

<div class="topic-metadata">

**Author:** [@atteson](https://discourse.julialang.org/u/atteson)\
**Replies:** 16\
**Last updated:** [January 16, 2022, 12:07am UTC](https://discourse.julialang.org/t/type-stability-with-variable-arguments/74640 "2022-01-16T00:07:54Z")

</div>

Is there a way to maintain type stability with variable arguments? Below is some simplified code illustrating the issue. I’d like f(x,y,z) to perform like g(x,y,z): function f( args... ) s = 0.0 for arg in arg…

---

## [Why is BitArray so slow?](https://discourse.julialang.org/t/why-is-bitarray-so-slow/14383)

<div class="topic-metadata">

**Author:** [@mbeach42](https://discourse.julialang.org/u/mbeach42)\
**Replies:** 28\
**Last updated:** [January 14, 2022, 10:43pm UTC](https://discourse.julialang.org/t/why-is-bitarray-so-slow/14383 "2022-01-14T22:43:46Z")

</div>

I was playing around with flipping numbers of an array, and I was surprised that doing so with BitArrays is much slower. I think the code below is pretty clear, using BenchmarkTools using Random a = rand(-1:2:1, 26, 11…

---

## [Julia Thread Affinity not persistent when calling MKL function](https://discourse.julialang.org/t/julia-thread-affinity-not-persistent-when-calling-mkl-function/74560)

<div class="topic-metadata">

**Author:** [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Replies:** 2\
**Last updated:** [January 14, 2022, 8:10pm UTC](https://discourse.julialang.org/t/julia-thread-affinity-not-persistent-when-calling-mkl-function/74560 "2022-01-14T20:10:45Z")

</div>

I observe a subtle issue where the pinning of Julia threads to specific cores is spoiled massively by running a seemingly harmless computation. By spoiled I mean that after running the computation all threads are pinned …

---

## [Efficient conditional in-place assignment for arrays?](https://discourse.julialang.org/t/efficient-conditional-in-place-assignment-for-arrays/74593)

<div class="topic-metadata">

**Author:** [@robsmith11](https://discourse.julialang.org/u/robsmith11)\
**Replies:** 7\
**Last updated:** [January 14, 2022, 12:26pm UTC](https://discourse.julialang.org/t/efficient-conditional-in-place-assignment-for-arrays/74593 "2022-01-14T12:26:29Z")

</div>

Is there any way to the get the efficient assignment of f1 with the nicer syntax of f2? Something like a view or conditional SubArray? julia\> function f1!(xs) for i in eachindex(xs) if xs\[i\] \< 0.5 …

---

## [Why is \`TaskLocalRNG\` faster than \`Xoshiro\` with multiple threads?](https://discourse.julialang.org/t/why-is-tasklocalrng-faster-than-xoshiro-with-multiple-threads/74577)

<div class="topic-metadata">

**Author:** [@robsmith11](https://discourse.julialang.org/u/robsmith11)\
**Replies:** 7\
**Last updated:** [January 14, 2022, 3:58am UTC](https://discourse.julialang.org/t/why-is-tasklocalrng-faster-than-xoshiro-with-multiple-threads/74577 "2022-01-14T03:58:52Z")

</div>

I was a bit confused by the docs: In a multi-threaded program, you should generally use different RNG objects from different threads or tasks in order to be thread-safe. However, the default RNG is thread-safe as of Ju…

---

## [Computational performance 3D finite difference stencil for different vectorization methods and precision](https://discourse.julialang.org/t/computational-performance-3d-finite-difference-stencil-for-different-vectorization-methods-and-precision/74382)

<div class="topic-metadata">

**Author:** [@Chiil](https://discourse.julialang.org/u/Chiil)\
**Replies:** 1\
**Last updated:** [January 13, 2022, 2:30pm UTC](https://discourse.julialang.org/t/computational-performance-3d-finite-difference-stencil-for-different-vectorization-methods-and-precision/74382 "2022-01-13T14:30:37Z")

</div>

I have a code that generates with a macro a finite difference stencil. I have disabled the macro in order to make a minimal working example, so please forgive me for this absurdly long line, it is not there in my normal …

---

## [Is there a better way to monitor a log file?](https://discourse.julialang.org/t/is-there-a-better-way-to-monitor-a-log-file/74538)

<div class="topic-metadata">

**Author:** [@anon69491625](https://discourse.julialang.org/u/anon69491625)\
**Replies:** 4\
**Last updated:** [January 13, 2022, 12:47pm UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-monitor-a-log-file/74538 "2022-01-13T12:47:28Z")

</div>

Hi all new to julia and so still learning the basics. I have a log file that another system updates about every 1 second. It’s just a simple text file and each “row” is about 1k bytes. I just want to get the “updates”…

---

## [Help diagnosing a slow iterator](https://discourse.julialang.org/t/help-diagnosing-a-slow-iterator/74413)

<div class="topic-metadata">

**Author:** [@tecosaur](https://discourse.julialang.org/u/tecosaur)\
**Replies:** 5\
**Last updated:** [January 13, 2022, 9:56am UTC](https://discourse.julialang.org/t/help-diagnosing-a-slow-iterator/74413 "2022-01-13T09:56:16Z")

</div>

Note, this was asked on Zulip a week ago, but received no responses. Reposting in the hope that someone else might see and comment I have a tree-like structure, composed of a mix of types, and I’m looking to iterate th…

---

## [Help optimizing IPFP algorithm](https://discourse.julialang.org/t/help-optimizing-ipfp-algorithm/74376)

<div class="topic-metadata">

**Author:** [@amrods](https://discourse.julialang.org/u/amrods)\
**Replies:** 6\
**Last updated:** [January 13, 2022, 3:56am UTC](https://discourse.julialang.org/t/help-optimizing-ipfp-algorithm/74376 "2022-01-13T03:56:38Z")

</div>

I’m trying to speed up this version of the IPFP algorithm, which alternates between solving 2 systems of nonlinear equations until convergence: using NLsolve using LinearAlgebra const szAge = 64 - 19 + 1 const szE = 2 …

---

## [Matrix multiplication of a view of QR.Q with Vector{VariableRef} is slow](https://discourse.julialang.org/t/matrix-multiplication-of-a-view-of-qr-q-with-vector-variableref-is-slow/74480)

<div class="topic-metadata">

**Author:** [@Thomas](https://discourse.julialang.org/u/Thomas)\
**Replies:** 6\
**Last updated:** [January 12, 2022, 3:17pm UTC](https://discourse.julialang.org/t/matrix-multiplication-of-a-view-of-qr-q-with-vector-variableref-is-slow/74480 "2022-01-12T15:17:55Z")

</div>

To assist with formulating a JuMP model for an SDP, I need to create affine expressions y (then the objectives and constraints can be cheaply constructed by getindex into y): F = qr(m::SparseMatrixCSC) # call to SuiteSp…

---

## [Confusing memory allocations when using the integrator of DifferentialEquations.jl](https://discourse.julialang.org/t/confusing-memory-allocations-when-using-the-integrator-of-differentialequations-jl/74436)

<div class="topic-metadata">

**Author:** [@duan](https://discourse.julialang.org/u/duan)\
**Replies:** 10\
**Last updated:** [January 12, 2022, 12:28am UTC](https://discourse.julialang.org/t/confusing-memory-allocations-when-using-the-integrator-of-differentialequations-jl/74436 "2022-01-12T00:28:46Z")

</div>

Here is a prototype of my problem: using OrdinaryDiffEq n = 2^18; u0 = rand(ComplexF64, n); function f(du, u, p, t) for i in 1:length(u) du\[i\] = -1.0im \* u\[i\] end return end mutable struct A u0…

---

## [Profiling Requires.jl](https://discourse.julialang.org/t/profiling-requires-jl/73570)

<div class="topic-metadata">

**Author:** [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Replies:** 1\
**Last updated:** [January 11, 2022, 7:37pm UTC](https://discourse.julialang.org/t/profiling-requires-jl/73570 "2022-01-11T19:37:58Z")

</div>

Is there any way to profile which Requires.jl usage are taking the longest time?

---

## [Precompilation No Effect ?!](https://discourse.julialang.org/t/precompilation-no-effect/74387)

<div class="topic-metadata">

**Author:** [@yoh-meyers](https://discourse.julialang.org/u/yoh-meyers)\
**Replies:** 1\
**Last updated:** [January 11, 2022, 1:05pm UTC](https://discourse.julialang.org/t/precompilation-no-effect/74387 "2022-01-11T13:05:49Z")

</div>

Hello, Am trying to wrap my head around pre-compilation of a package. Have read through SnoopCompile.jl · SnoopCompile and tried the steps myself. So basically in my package I have: if ccall(:jl\_generating\_output, Ci…

---

## [Improved installation for diffeqpy and similar packages](https://discourse.julialang.org/t/improved-installation-for-diffeqpy-and-similar-packages/74350)

<div class="topic-metadata">

**Author:** [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Replies:** 2\
**Last updated:** [January 10, 2022, 7:44pm UTC](https://discourse.julialang.org/t/improved-installation-for-diffeqpy-and-similar-packages/74350 "2022-01-10T19:44:24Z")

</div>

Here is an alternative, more user friendly, packaging of diffeqpy. (diffeqpy allows calling DifferentialEquations.jl from Python): https://github.com/jlapeyre/diffeq\_julia If you use diffeqpy, please give this a try. I…

---

## [How to wrap a vector so that it does simd?](https://discourse.julialang.org/t/how-to-wrap-a-vector-so-that-it-does-simd/74242)

<div class="topic-metadata">

**Author:** [@abulak](https://discourse.julialang.org/u/abulak)\
**Replies:** 0\
**Last updated:** [January 8, 2022, 2:40pm UTC](https://discourse.julialang.org/t/how-to-wrap-a-vector-so-that-it-does-simd/74242 "2022-01-08T14:40:14Z")

</div>

Hey, I’ve been recently hit with this issue: I have a very thin wrapper around Vector{T}, let’s say it’s struct MyVector{T} \<: AbstractVector{T} data::Vector{T} end # Array Interface Base.size(v::MyVector) = size(…

---

## [Faster alternate to @views for passing subarrays to functions](https://discourse.julialang.org/t/faster-alternate-to-views-for-passing-subarrays-to-functions/74195)

<div class="topic-metadata">

**Author:** [@Arpit\_Babbar](https://discourse.julialang.org/u/Arpit_Babbar)\
**Replies:** 3\
**Last updated:** [January 7, 2022, 4:39pm UTC](https://discourse.julialang.org/t/faster-alternate-to-views-for-passing-subarrays-to-functions/74195 "2022-01-07T16:39:33Z")

</div>

Hi, I have got 4 arrays a,b,c,d each of size nvar X nx where nvar is a small Int64, but nx is a pretty large Int64. I wish to iteratively act a function over the arguments (a\[:,i\],a\[:,i-1\],a\[:,i+1\],b\[:,i\],b\[:,i-1\],b\[:,i+…

---

## [ReadOnlyMemoryError() in an ODEProblem](https://discourse.julialang.org/t/readonlymemoryerror-in-an-odeproblem/74189)

<div class="topic-metadata">

**Author:** [@Frazze](https://discourse.julialang.org/u/Frazze)\
**Replies:** 6\
**Last updated:** [January 7, 2022, 2:58pm UTC](https://discourse.julialang.org/t/readonlymemoryerror-in-an-odeproblem/74189 "2022-01-07T14:58:23Z")

</div>

Hi everyone I’m studying a pde (swift-hohenberg model) converting it in the ODE form. I’ve discretized the laplacian: function Laplacian2D(Nx, lx) hx = lx/Nx D2x = CenteredDifference(2, 2, hx, Nx) Qx = Periodi…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=72)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=74)
