# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=123

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 124

---

## [Pmap and multi-threaded BLAS](https://discourse.julialang.org/t/pmap-and-multi-threaded-blas/31617)

<div class="topic-metadata">

**Author:** [@ahthomas](https://discourse.julialang.org/u/ahthomas)\
**Replies:** 2\
**Last updated:** [November 29, 2019, 9:29am UTC](https://discourse.julialang.org/t/pmap-and-multi-threaded-blas/31617 "2019-11-29T09:29:10Z")

</div>

When I run a call to one of the low level LinearAlgebra.BLAS functions inside a function executed on a worker process launched via pmap it appears that BLAS does not use multiple threads. Here is a simple example: using…

---

## [Why mul! is so fast with BitVector?](https://discourse.julialang.org/t/why-mul-is-so-fast-with-bitvector/31629)

<div class="topic-metadata">

**Author:** [@e3c6](https://discourse.julialang.org/u/e3c6)\
**Replies:** 1\
**Last updated:** [November 29, 2019, 12:45am UTC](https://discourse.julialang.org/t/why-mul-is-so-fast-with-bitvector/31629 "2019-11-29T00:45:49Z")

</div>

Consider the following example code: using LinearAlgebra, BenchmarkTools n = 200 xb = BitVector(rand(Bool, n)) xf = Float32.(xb) W = randn(Float32,n,n) y = zeros(Float32,n); function mymul!(y, A, x) y .= 0 @in…

---

## [Cosine seems slow](https://discourse.julialang.org/t/cosine-seems-slow/31507)

<div class="topic-metadata">

**Author:** [@sw72](https://discourse.julialang.org/u/sw72)\
**Replies:** 14\
**Last updated:** [November 27, 2019, 10:59am UTC](https://discourse.julialang.org/t/cosine-seems-slow/31507 "2019-11-27T10:59:28Z")

</div>

I’m finding the cos function slow in Julia compared with Mathematica. For example: a = RandomReal\[NormalDistribution\[\], {2000, 2000}\]; RepeatedTiming\[b = Cos\[a\];\] {0.018, Null} function test() a=randn(2000,2000…

---

## [Performance with small matrices: 5-argument mul! vs BLAS.gemm!](https://discourse.julialang.org/t/performance-with-small-matrices-5-argument-mul-vs-blas-gemm/31543)

<div class="topic-metadata">

**Author:** [@Egwene\_al\_Vere](https://discourse.julialang.org/u/Egwene_al_Vere)\
**Replies:** 1\
**Last updated:** [November 26, 2019, 7:26pm UTC](https://discourse.julialang.org/t/performance-with-small-matrices-5-argument-mul-vs-blas-gemm/31543 "2019-11-26T19:26:48Z")

</div>

Trying out the new 5-argument mul! introduced in Julia 1.3, (ABα+Cβ → C), for the calculation of matrix commutator comu(A,B)=AB-BA. From my testing, with small (2x2 or 3x3) matrices, mul! is actually faster than gemm!, a…

---

## [Repeated @benchmark causes RAM creep](https://discourse.julialang.org/t/repeated-benchmark-causes-ram-creep/31533)

<div class="topic-metadata">

**Author:** [@Crown421](https://discourse.julialang.org/u/Crown421)\
**Replies:** 4\
**Last updated:** [November 26, 2019, 5:19pm UTC](https://discourse.julialang.org/t/repeated-benchmark-causes-ram-creep/31533 "2019-11-26T17:19:08Z")

</div>

I am currently trying to quantify the performance/ scaling of the VML functions as added via VML.jl, and ran into a curious problem. Running this script, my RAM steadily increases until julia consumes 12-14Gb at the en…

---

## [\`using AMQPClient\` significantly slows down code](https://discourse.julialang.org/t/using-amqpclient-significantly-slows-down-code/31228)

<div class="topic-metadata">

**Author:** [@Marek\_Kukan](https://discourse.julialang.org/u/Marek_Kukan)\
**Replies:** 4\
**Last updated:** [November 25, 2019, 10:05am UTC](https://discourse.julialang.org/t/using-amqpclient-significantly-slows-down-code/31228 "2019-11-25T10:05:52Z")

</div>

Hello, I came across this weird performance issue, can someone please explain why using AMQPClient causes ~100x performance drop? julia\> m = Matrix{Any}(rand(1000,1000)); julia\> function f(m, n) for i in 1:…

---

## [Multivariate Normal Distribution](https://discourse.julialang.org/t/multivariate-normal-distribution/29973)

<div class="topic-metadata">

**Author:** [@Hugo\_RB](https://discourse.julialang.org/u/Hugo_RB)\
**Replies:** 17\
**Last updated:** [November 24, 2019, 8:27pm UTC](https://discourse.julialang.org/t/multivariate-normal-distribution/29973 "2019-11-24T20:27:02Z")

</div>

Hi there Could somebody give me a clear example about how to construct a MVND using the syntaxis given by the Distribution Pkg. I know that with Distributions.jl it is possible to generate a MVND just as it is possible t…

---

## [Functions inside functions using "global" variables](https://discourse.julialang.org/t/functions-inside-functions-using-global-variables/31434)

<div class="topic-metadata">

**Author:** [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Replies:** 2\
**Last updated:** [November 23, 2019, 11:08pm UTC](https://discourse.julialang.org/t/functions-inside-functions-using-global-variables/31434 "2019-11-23T23:08:23Z")

</div>

Hi! I need help about constructing functions inside functions that use variables that “seems” global: function a() aux = Float64\[\] b = ()-\>begin push!(aux, rand()) end b() b() return aux…

---

## [Regular Expression and Threads](https://discourse.julialang.org/t/regular-expression-and-threads/31415)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 3\
**Last updated:** [November 23, 2019, 7:07pm UTC](https://discourse.julialang.org/t/regular-expression-and-threads/31415 "2019-11-23T19:07:11Z")

</div>

Hi All, in our Flux ML project, we need to process large volume of data with a single models. We do it by dividing the minibatch into chunks, calculate gradient on each chunk in a separate thread, and reduce it (is it c…

---

## [For loop in function and multiplication of larger matrices, slow speed in parallel](https://discourse.julialang.org/t/for-loop-in-function-and-multiplication-of-larger-matrices-slow-speed-in-parallel/31258)

<div class="topic-metadata">

**Author:** [@luboshanus](https://discourse.julialang.org/u/luboshanus)\
**Replies:** 3\
**Last updated:** [November 22, 2019, 1:42pm UTC](https://discourse.julialang.org/t/for-loop-in-function-and-multiplication-of-larger-matrices-slow-speed-in-parallel/31258 "2019-11-22T13:42:58Z")

</div>

Hello, I’d like to understand how does the performance of large for loops in julia work? Running the same function on one core or on multiple cores does not give me similar time results. The computer has 48 physical co…

---

## [Improving Performance of this Code](https://discourse.julialang.org/t/improving-performance-of-this-code/31326)

<div class="topic-metadata">

**Author:** [@jmcastro2109](https://discourse.julialang.org/u/jmcastro2109)\
**Replies:** 12\
**Last updated:** [November 22, 2019, 11:59am UTC](https://discourse.julialang.org/t/improving-performance-of-this-code/31326 "2019-11-22T11:59:25Z")

</div>

Hello! I am trying to solve a big model and at the core I have the following loop which essentially solves a Backward Induction Problem. In what follows I provide a MWE. In my model, I am solving this problem many many…

---

## [Optimization of Code for NaN Operations](https://discourse.julialang.org/t/optimization-of-code-for-nan-operations/31339)

<div class="topic-metadata">

**Author:** [@natgeo-wong](https://discourse.julialang.org/u/natgeo-wong)\
**Replies:** 3\
**Last updated:** [November 21, 2019, 8:50pm UTC](https://discourse.julialang.org/t/optimization-of-code-for-nan-operations/31339 "2019-11-21T20:50:33Z")

</div>

Hi, I’ve been trying to perform operations on Arrays that may have missing values, using the below prototype function: function nanop(f::Function,data::Array;dim=1) ndim = ndims(data); dsize = size(data); if nd…

---

## [Performance of the iterator approach](https://discourse.julialang.org/t/performance-of-the-iterator-approach/31293)

<div class="topic-metadata">

**Author:** [@alanderos](https://discourse.julialang.org/u/alanderos)\
**Replies:** 3\
**Last updated:** [November 21, 2019, 7:27am UTC](https://discourse.julialang.org/t/performance-of-the-iterator-approach/31293 "2019-11-21T07:27:11Z")

</div>

I started using IterativeSolvers.jl in some of my work and found the iterator-based code to be easy on the eyes and rather slick. It seems to me that any Julia implementation of an iterative algorithm ought to adopt this…

---

## [Interpreting multiple lines on one level of ProfileView](https://discourse.julialang.org/t/interpreting-multiple-lines-on-one-level-of-profileview/31235)

<div class="topic-metadata">

**Author:** [@Ross\_Boylan](https://discourse.julialang.org/u/Ross_Boylan)\
**Replies:** 1\
**Last updated:** [November 19, 2019, 12:21am UTC](https://discourse.julialang.org/t/interpreting-multiple-lines-on-one-level-of-profileview/31235 "2019-11-19T00:21:15Z")

</div>

A handy feature of ProfileView is that hovering over the graph shows the line that goes with the bar one is over. However, for some of my results, I see multiple lines listed. What does this mean? @tim.holy I wasn’t…

---

## [Hack: AMD Ryzen/TR/Epyc + Intel Math Kernel Library (MKL)](https://discourse.julialang.org/t/hack-amd-ryzen-tr-epyc-intel-math-kernel-library-mkl/31226)

<div class="topic-metadata">

**Author:** [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Replies:** 0\
**Last updated:** [November 18, 2019, 2:50pm UTC](https://discourse.julialang.org/t/hack-amd-ryzen-tr-epyc-intel-math-kernel-library-mkl/31226 "2019-11-18T14:50:37Z")

</div>

Interesting performance hack for Ryzen/TR CPU + MKL → “up to 250% performance gains” (based on reddit post ) in theory it should work with Julia+MKL ; Anybody can test / validate ? “The method provided here does e…

---

## [Performance: replace conditional jump by some shift/and/or?](https://discourse.julialang.org/t/performance-replace-conditional-jump-by-some-shift-and-or/30966)

<div class="topic-metadata">

**Author:** [@rryi](https://discourse.julialang.org/u/rryi)\
**Replies:** 10\
**Last updated:** [November 14, 2019, 2:31pm UTC](https://discourse.julialang.org/t/performance-replace-conditional-jump-by-some-shift-and-or/30966 "2019-11-14T14:31:31Z")

</div>

I work on a data type which stores several fields bit-packed into a container c :: Int64. There are two different storage layouts, depending on the most significant bit. A simplified example: if bit63 is 0, my bit-packed…

---

## [Big endian conversion on custom datatypes](https://discourse.julialang.org/t/big-endian-conversion-on-custom-datatypes/31076)

<div class="topic-metadata">

**Author:** [@tamasgal](https://discourse.julialang.org/u/tamasgal)\
**Replies:** 3\
**Last updated:** [November 14, 2019, 10:01am UTC](https://discourse.julialang.org/t/big-endian-conversion-on-custom-datatypes/31076 "2019-11-14T10:01:00Z")

</div>

I have to deal with a couple of “big endian” structures in raw files and want to parse them into Julia structs. Here is an MWE, where I have some data (big endian, coming from the network or from a file), a struct and a…

---

## [Profile.init time units](https://discourse.julialang.org/t/profile-init-time-units/30888)

<div class="topic-metadata">

**Author:** [@Ross\_Boylan](https://discourse.julialang.org/u/Ross_Boylan)\
**Replies:** 14\
**Last updated:** [November 13, 2019, 5:05am UTC](https://discourse.julialang.org/t/profile-init-time-units/30888 "2019-11-13T05:05:09Z")

</div>

The documentation says the time units for the delay in Profile.init are in seconds, but the results I’m getting seem more consistent with 100th of a second. What’s going on? Julia 1.2. Profile.init(10000000, 0.3) Prof…

---

## [No constant expression elimination for e.g. '2^24-1'](https://discourse.julialang.org/t/no-constant-expression-elimination-for-e-g-2-24-1/30964)

<div class="topic-metadata">

**Author:** [@rryi](https://discourse.julialang.org/u/rryi)\
**Replies:** 7\
**Last updated:** [November 12, 2019, 11:26am UTC](https://discourse.julialang.org/t/no-constant-expression-elimination-for-e-g-2-24-1/30964 "2019-11-12T11:26:25Z")

</div>

I found differences in code generation for simple constant expressions: 1\<\<24-1 generates a constant, 2^24-1 generates a quite expensive C function call. Code example below. I did expect Julia to replace both expression…

---

## [Performance of Dict](https://discourse.julialang.org/t/performance-of-dict/30993)

<div class="topic-metadata">

**Author:** [@natgeo-wong](https://discourse.julialang.org/u/natgeo-wong)\
**Replies:** 1\
**Last updated:** [November 12, 2019, 9:17am UTC](https://discourse.julialang.org/t/performance-of-dict/30993 "2019-11-12T09:17:32Z")

</div>

Hi! I tend to use the Dict functionality of Julia pretty often to store information (e.g. attributes of various parameters that contain both strings and integers). But I’ve heard that Dict can slow down the performance…

---

## [In-place operations on matrices](https://discourse.julialang.org/t/in-place-operations-on-matrices/30908)

<div class="topic-metadata">

**Author:** [@iamsuddhasattwa](https://discourse.julialang.org/u/iamsuddhasattwa)\
**Replies:** 6\
**Last updated:** [November 10, 2019, 5:05pm UTC](https://discourse.julialang.org/t/in-place-operations-on-matrices/30908 "2019-11-10T17:05:11Z")

</div>

When working with huge matrices, the limitations of the RAM is a real issue, and one would like to perform matrix operations in place if possible. Given a matrix A which occupies a sizable chunk of RAM, what is the best …

---

## [SymTrididiagonal matrices performance and views allocations](https://discourse.julialang.org/t/symtrididiagonal-matrices-performance-and-views-allocations/24618)

<div class="topic-metadata">

**Author:** [@LaurentPlagne](https://discourse.julialang.org/u/LaurentPlagne)\
**Replies:** 12\
**Last updated:** [November 8, 2019, 4:54pm UTC](https://discourse.julialang.org/t/symtrididiagonal-matrices-performance-and-views-allocations/24618 "2019-11-08T16:54:10Z")

</div>

Hi, I work on a toy 2D CFD solver and I wonder about the idiomatic Julian way to compute efficiently the product for all j \\in \[1:ny\] Pxy\[:,j\]=Lx\*Sxy\[:j\] where Lx is a a SymTridiagonal matrix of rank nx and Pxy and S…

---

## [Switching from 1.1 to 1.2, more startup slowness after loading startup file](https://discourse.julialang.org/t/switching-from-1-1-to-1-2-more-startup-slowness-after-loading-startup-file/30832)

<div class="topic-metadata">

**Author:** [@Usor](https://discourse.julialang.org/u/Usor)\
**Replies:** 4\
**Last updated:** [November 7, 2019, 5:24pm UTC](https://discourse.julialang.org/t/switching-from-1-1-to-1-2-more-startup-slowness-after-loading-startup-file/30832 "2019-11-07T17:24:30Z")

</div>

I understand that there may be some slowness, and I don’t mind it so much, but when I switched from 1.1 to 1.2, Julia became much slower after manually loading a startup file. Though the cursor blinks after loading, pres…

---

## [Initializing an array of mutable objects is surprisingly slow](https://discourse.julialang.org/t/initializing-an-array-of-mutable-objects-is-surprisingly-slow/11780)

<div class="topic-metadata">

**Author:** [@cnliao](https://discourse.julialang.org/u/cnliao)\
**Replies:** 8\
**Last updated:** [November 7, 2019, 3:13pm UTC](https://discourse.julialang.org/t/initializing-an-array-of-mutable-objects-is-surprisingly-slow/11780 "2019-11-07T15:13:57Z")

</div>

How can one performantly preallocate a vector of mutable objects? Currently allocating and initializing a vector of mutables is ~60x slower than doing the same thing for mutables: # same result on v0.6.3 and latest mast…

---

## [Get intuition on how to improve matmul code](https://discourse.julialang.org/t/get-intuition-on-how-to-improve-matmul-code/30799)

<div class="topic-metadata">

**Author:** [@davidbp](https://discourse.julialang.org/u/davidbp)\
**Replies:** 14\
**Last updated:** [November 6, 2019, 5:31pm UTC](https://discourse.julialang.org/t/get-intuition-on-how-to-improve-matmul-code/30799 "2019-11-06T17:31:02Z")

</div>

Hello, After playing a little bit with the most naive matmul blocked version I have found on the internet I was wondering what could I do to improve the performance. function matmul\_blocked\_1!(A, B, C) bs = 10 …

---

## [Running a function in parallel](https://discourse.julialang.org/t/running-a-function-in-parallel/30756)

<div class="topic-metadata">

**Author:** [@Patrik\_Waldmann](https://discourse.julialang.org/u/Patrik_Waldmann)\
**Replies:** 2\
**Last updated:** [November 5, 2019, 6:14pm UTC](https://discourse.julialang.org/t/running-a-function-in-parallel/30756 "2019-11-05T18:14:59Z")

</div>

I need help getting a function to run in parallel. I want to distribute both data and packages to all workers, see the following example (which doesn’t work): #Simulated data (n observations, p variables, tr truevariabl…

---

## [Assignments, Copies and Deepcop](https://discourse.julialang.org/t/assignments-copies-and-deepcop/30723)

<div class="topic-metadata">

**Author:** [@natgeo-wong](https://discourse.julialang.org/u/natgeo-wong)\
**Replies:** 14\
**Last updated:** [November 5, 2019, 5:07pm UTC](https://discourse.julialang.org/t/assignments-copies-and-deepcop/30723 "2019-11-05T17:07:15Z")

</div>

This is a question regarding the use of “=” signs in Julia. I can do the following for-loop in MATLAB for ii = 1 : dt ηp1 = shallowwave1D(ηn,ηm1,r); ηm1 = ηn; ηn = ηp1; end Can the same be done in Julia? Or …

---

## [Question on semantics of loops, maps, and broadcast](https://discourse.julialang.org/t/question-on-semantics-of-loops-maps-and-broadcast/30706)

<div class="topic-metadata">

**Author:** [@klaff](https://discourse.julialang.org/u/klaff)\
**Replies:** 6\
**Last updated:** [November 5, 2019, 3:10am UTC](https://discourse.julialang.org/t/question-on-semantics-of-loops-maps-and-broadcast/30706 "2019-11-05T03:10:00Z")

</div>

I’ve been playing with writing things in different forms, for example the following: function f\_by\_loop!(data,x) @inbounds for i in eachindex(data) data\[i\] = min(x,data\[i\]) end end vs data .= min.(x,dat…

---

## [Type stability: ProfileView and @code\_warntype don't agree](https://discourse.julialang.org/t/type-stability-profileview-and-code-warntype-dont-agree/30619)

<div class="topic-metadata">

**Author:** [@jlbosse](https://discourse.julialang.org/u/jlbosse)\
**Replies:** 4\
**Last updated:** [November 2, 2019, 11:15pm UTC](https://discourse.julialang.org/t/type-stability-profileview-and-code-warntype-dont-agree/30619 "2019-11-02T23:15:56Z")

</div>

I have some code that I feel should run faster (see below). So I ran it with @profile and looked at the results with ProfileView.view() and as expected a lot of red showed up (i.e. runtime method lookup), the correspind…

---

## [Writable global const arrays](https://discourse.julialang.org/t/writable-global-const-arrays/30635)

<div class="topic-metadata">

**Author:** [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Replies:** 2\
**Last updated:** [November 2, 2019, 8:49pm UTC](https://discourse.julialang.org/t/writable-global-const-arrays/30635 "2019-11-02T20:49:32Z")

</div>

I tried to implement the xoshiro RNG for Base julia (cf https://github.com/JuliaLang/julia/issues/27614). So, for this to work properly, I thought about using something like const global\_xorostate = zeros(UInt64, Threa…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=122)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=124)
