# Julia slower than Matlab & Python? No

**URL:** <https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128>\
**Category:** Performance\
**Tags:** economics, tensorflow, matlab, pytorch\
**Created:** [January 8, 2020, 11:57pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128 "2020-01-08T23:57:37Z")\
**Posts on this page:** 20\
**Page:** 6

<div class="post-metadata">

**Author:** ![Albert\_Zevelev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/albert_zevelev/32/11844_2.png) [@Albert\_Zevelev](https://discourse.julialang.org/u/Albert_Zevelev)\
**Post date:** [January 19, 2020, 7:30pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/102 "2020-01-19T19:30:04Z")

</div>

It works now. Both “Simple” & “Fastest” are much faster, especially for bigger orders.

---

<div class="post-metadata">

**Author:** ![jlperla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlperla/32/34332_2.png) [@jlperla](https://discourse.julialang.org/u/jlperla)\
**Post date:** [January 19, 2020, 7:34pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/103 "2020-01-19T19:34:01Z")

</div>

> [@Mason](#):
>
> that goes through and detects cases where views should be used or intermediary arrays should be pre-allocated, see if it can prove that certain operations don’t need bounds checking, etc. could be quite interesting and useful.

This would be great, even if it started with baby steps. It would mean that we could respond to performance issues with a single macro to get started, which is especially useful for dabblers of the language. For example, adding`@inbounds` to every loop generally never hurts (it can make things crash, but so would manually adding it to ever loop) so why not start with that and build things up?

---

<div class="post-metadata">

**Author:** ![jlperla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlperla/32/34332_2.png) [@jlperla](https://discourse.julialang.org/u/jlperla)\
**Post date:** [January 19, 2020, 7:34pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/104 "2020-01-19T19:34:20Z")

</div>

> [@Albert\_Zevelev](#):
>
> It works now.

the OpenBlas fix or the MKL?

---

<div class="post-metadata">

**Author:** ![Albert\_Zevelev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/albert_zevelev/32/11844_2.png) [@Albert\_Zevelev](https://discourse.julialang.org/u/Albert_Zevelev)\
**Post date:** [January 19, 2020, 7:38pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/105 "2020-01-19T19:38:12Z")

</div>

Not sure which one.  
First I ran the Blas fix & it was faster.  
Then I checked & MKL now works & its faster.

---

<div class="post-metadata">

**Author:** ![jlperla](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlperla/32/34332_2.png) [@jlperla](https://discourse.julialang.org/u/jlperla)\
**Post date:** [January 19, 2020, 7:43pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/106 "2020-01-19T19:43:20Z")

</div>

Great! Is the speedup significant (e.g. \> 20%)?

Eventually it would be nice to see if the patched openblas is comparable to the MKL here.

---

<div class="post-metadata">

**Author:** ![Albert\_Zevelev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/albert_zevelev/32/11844_2.png) [@Albert\_Zevelev](https://discourse.julialang.org/u/Albert_Zevelev)\
**Post date:** [January 19, 2020, 7:48pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/107 "2020-01-19T19:48:00Z")

</div>

For option pricing, for big orders its 30-40% faster than w/o `LinearAlgebra.BLAS.set_num_threads(n)`

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [January 19, 2020, 9:06pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/108 "2020-01-19T21:06:08Z")

</div>

FWIW, on my machine, OpenBLAS:

```julia
julia> using LinearAlgebra

julia> BLAS.vendor()
:openblas64

julia> LinearAlgebra.peakflops(16_000)
4.478966395538375e11

julia> BLAS.set_num_threads(18)

julia> LinearAlgebra.peakflops(16_000)
8.023464806960204e11

```

vs MKL:

```julia
julia> using LinearAlgebra

julia> BLAS.vendor()
:mkl

julia> LinearAlgebra.peakflops(16_000)
2.1343568002173381e12

julia> BLAS.set_num_threads(18)

julia> LinearAlgebra.peakflops(16_000)
2.0939895828215435e12

julia> versioninfo()
Julia Version 1.5.0-DEV.89
Commit 5c75bec530* (2020-01-18 11:56 UTC)
Platform Info:
  OS: Linux (x86_64-generic-linux)
  CPU: Intel(R) Core(TM) i9-10980XE CPU @ 3.00GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-9.0.0 (ORCJIT, skylake)
Environment:
  JULIA_NUM_THREADS = 18

```

MKL seems around 2.6 times faster.  
FWIW, I set the avx512 all-core clock speed to 4.1 GHz in the bios.  
This means the CPU’s theoretical peak flops should be

```julia
julia> num_cores = 18;

julia> clock = 4.1e9;

julia> fma_per_clock = 2;

julia> flop_per_fma = 16; # avx512, double precision: 8 mul, 8 add

julia> num_cores * clock * fma_per_clock * flop_per_fma
2.3616e12

```

MKL gets fairly close. OpenBLAS does not.

EDIT:  
I’m going to have to take a good look at ModelingToolkit.  
It looks interesting. I wonder if I can take advantage of it in my own work.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [January 19, 2020, 11:41pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/109 "2020-01-19T23:41:14Z")

</div>

> [@ChrisRackauckas](#):
>
> It’s literally just a static computational graph language. It seems like you really really don’t understand it, so you might want to look into it more before commenting on what it can do…
> 
> Literally as simple as:
> 
> 1. Make variables
> 2. Put them into normal Julia code
> 3. Static graph comes out
> 4. Do things with the static graph

This is getting spicier than I like, but I’m going to continue because even after reading yet again more on the subject, I don’t understand how ModelingToolkit is relevant to the sort of problems that the original post in this gigantic thread is about. Consider this file: [https://github.com/vduarte/benchmarkingML/blob/master/Sovereign\_Default/julia.jl](https://github.com/vduarte/benchmarkingML/blob/master/Sovereign_Default/julia.jl) and pretend we don’t know the size of the input data [https://github.com/vduarte/benchmarkingML/blob/master/Sovereign\_Default/julia.jl#L4](https://github.com/vduarte/benchmarkingML/blob/master/Sovereign_Default/julia.jl#L4) and [https://github.com/vduarte/benchmarkingML/blob/master/Sovereign\_Default/julia.jl#L5](https://github.com/vduarte/benchmarkingML/blob/master/Sovereign_Default/julia.jl#L5) ahead of time.

How would one use ModelingToolkit to build up a graph and optimize the array operations in that code? Insofar as ModelingToolkit can understand array operations, it is concerned with the _elements_ of the arrays, not the arrays themselves as entities. Variables in ModelingToolkit are `<: Number`. We can do things like `@variables x[1:N, 1:M]` to make a `N x M` matrix of variables, but if you tried to push say a `256 x 256` matrix of Variables through the above function it’d take an inordinate amount of time, let alone `10_000 x 10_000`. With ModelingToolkit, the time it takes to trace matrix code would scale according to the size of the matrices. In TensorFlow or PyTorch, it wouldn’t (as far as I understand).

As a very very cut down example, consider:

```julia
f(A, X) = A .* X
using ModelingToolkit
@variables X[1:5000, 1:5000]
A = rand(size(X)...)
f(A, X)

```

It’d be great if ModelingToolkit could do things like `@variables A::Matrix X::Matrix` where the size of the matrix is undetermined and then `f(A, X)` would just give something loosely like `Expression(.*, A, X)`. Then I see how we could usefully trace through the above code and get optimizations, but it doesn’t seem like ModelingToolkit is in the business of doing those sorts of operations yet.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [January 20, 2020, 3:05am UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/110 "2020-01-20T03:05:47Z")

</div>

> [@Mason](#):
>
> It’d be great if ModelingToolkit could do things like `@variables A::Matrix X::Matrix` where the size of the matrix is undetermined and then `f(A, X)` would just give something loosely like `Expression(.*, A, X)` . Then I see how we could usefully trace through the above code and get optimizations, but it doesn’t seem like ModelingToolkit is in the business of doing those sorts of operations yet.

Yup, the current behavior is a bug that should get fixed.

---

<div class="post-metadata">

**Author:** ![KirillGerke](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kirillgerke/32/11383_2.png) [@KirillGerke](https://discourse.julialang.org/u/KirillGerke)\
**Post date:** [March 5, 2020, 8:51pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/111 "2020-03-05T20:51:58Z")

</div>

The paper is now published: [Benchmarking machine-learning software and hardware for quantitative economics - ScienceDirect](https://www.sciencedirect.com/science/article/pii/S0165188919301939)  
I think you did a great job in this thread and should consider writing a comment to that paper with all these points of criticism that evolved during rewriting the code.

---

<div class="post-metadata">

**Author:** ![jovansam](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jovansam/32/15023_2.png) [@jovansam](https://discourse.julialang.org/u/jovansam)\
**Post date:** [March 13, 2021, 8:03pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/112 "2021-03-13T20:03:04Z")

</div>

This is arguably the best thread I have come across on discourse! One year late, but never too late. Time to start optimizing Julia code. 😄 🛩

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 13, 2021, 8:37pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/113 "2021-03-13T20:37:24Z")

</div>

I wonder can this code be optimized further now. `LoopVectorization.jl` moved forward and there is `Tullio.jl` which probably can help with multithreading issues.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [March 15, 2021, 2:36pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/114 "2021-03-15T14:36:06Z")

</div>

11 posts were split to a new topic: [BLAS thread count vs Julia thread count](https://discourse.julialang.org/t/blas-thread-count-vs-julia-thread-count/57197)

---

<div class="post-metadata">

**Author:** ![Albert\_Zevelev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/albert_zevelev/32/11844_2.png) [@Albert\_Zevelev](https://discourse.julialang.org/u/Albert_Zevelev)\
**Post date:** [March 15, 2021, 3:57pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/115 "2021-03-15T15:57:47Z")

</div>

> [@Oscar\_Smith](#):
>
> I think a lot of it is that multithreading is only a few months old in Julia, so not much of Base takes advantage yet. If we could get broadcasting to use threads, that would be huge.

Has this changed much?  
Does more of Base take advantage of multithreading?  
Does broadcasting use threads?

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [March 15, 2021, 4:00pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/116 "2021-03-15T16:00:44Z")

</div>

Not yet. Broadcasting in particular will take quite a bit of work.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [March 15, 2021, 4:11pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/117 "2021-03-15T16:11:13Z")

</div>

I’d imagine the overhead of “check if this broadcasting is worth multi-threading or not” will slow down numerous small broadcasting everywhere by a (relatively) significant amount.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [March 15, 2021, 4:13pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/118 "2021-03-15T16:13:16Z")

</div>

> [@jling](#):
>
> I’d imagine the overhead of “check if this broadcasting is worth multi-threading or not” will slow down numerous small broadcasting everywhere by a (relatively) significant amount.

I think the immediate goal is for it to be explicit, i.e. `@threads y .= foo.(bar.(x))`. That way the caller can decide (by benchmarking) whether it is worth it or not, in the same way that you might decide whether to use `@threads for`.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [March 15, 2021, 4:14pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/119 "2021-03-15T16:14:39Z")

</div>

That’s really not the part that I’m worried about — that’s as simple as a few type/length checks, which are _really_ quick, especially if the kernels are outlined. The challenge is the “check if this broadcasting is _safe_ to multi-thread or not” accounting for side-effects, aliasing, crazy data structures, and more. Making it explicit is exactly a good first step.

---

<div class="post-metadata">

**Author:** ![antoine-levitt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antoine-levitt/32/4008_2.png) [@antoine-levitt](https://discourse.julialang.org/u/antoine-levitt)\
**Post date:** [March 15, 2021, 4:17pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/120 "2021-03-15T16:17:57Z")

</div>

Doesn’t [https://github.com/Jutho/Strided.jl](https://github.com/Jutho/Strided.jl) do this already?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [March 15, 2021, 4:18pm UTC](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128/121 "2021-03-15T16:18:06Z")

</div>

this looks pretty do-able, do we have a PR/WIP already somewhere?

[Previous page](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128.md?page=5)

[Next page](https://discourse.julialang.org/t/julia-slower-than-matlab-python-no/33128.md?page=7)
