# Comparing exp() performance on Julia versus numpy

**URL:** <https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325>\
**Category:** Performance\
**Created:** [August 5, 2022, 6:08am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325 "2022-08-05T06:08:46Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![torrance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torrance/32/38990_2.png) [@torrance](https://discourse.julialang.org/u/torrance)\
**Post date:** [August 5, 2022, 6:08am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/1 "2022-08-05T06:08:46Z")

</div>

So I was actually preparing a demo to show off Julia to colleagues when I encountered this issue.

I want to simply compute the exponential of an array. In Julia (v. 1.7.2):

```
arr = rand(100_000)
@benchmark exp.($arr)

# Time (median): 576.806 μs GC (median): 0.00%
# Memory estimate: 781.30 KiB, allocs estimate: 2.

```

Compared to Python:

```
arr = np.random.rand(100000)
timeit.timeit(lambda: np.exp(arr), number=100000) / 100000
# 0.0001027145623601973

```

So Python is about a factor of 6 faster.

Of course I’ve also tried wrapping the Julia dot call inside a function, as well as manually looping myself, e.g.:

```
@inbounds function myfunc(arr)
    result = similar(arr)
    @simd for i in eachindex(arr)
        result[i] = exp(arr[i]) 
    end
    return result
end

@benchmark myfunc($arr)
# Time (median): 569.613 μs

```

…which is identical timing to the dotted version.

Any idea what’s going on here? Some obvious sanity checks - both are Float64, and both are not using fastmath, both are allocating a new array and both are single threaded.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [August 5, 2022, 6:22am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/2 "2022-08-05T06:22:39Z")

</div>

Edit: saw now that you said no multithreading.

But are you actually certain that numpy is not implicitly doing multithreading under the hood?

---

<div class="post-metadata">

**Author:** ![torrance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torrance/32/38990_2.png) [@torrance](https://discourse.julialang.org/u/torrance)\
**Post date:** [August 5, 2022, 6:26am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/3 "2022-08-05T06:26:45Z")

</div>

@DNF I checked that: both are single threaded.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [August 5, 2022, 6:33am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/4 "2022-08-05T06:33:04Z")

</div>

I’m no good at reading llvm code, so I cannot confirm, but it might be that Julia fails to simd the computation. If you use LoopVectorization.jl, it works:

```julia
julia> @btime exp.($x);
  503.300 μs (2 allocations: 781.30 KiB)

julia> @btime @turbo exp.($x);
  91.100 μs (2 allocations: 781.30 KiB)

```

---

<div class="post-metadata">

**Author:** ![torrance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torrance/32/38990_2.png) [@torrance](https://discourse.julialang.org/u/torrance)\
**Post date:** [August 5, 2022, 6:53am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/5 "2022-08-05T06:53:11Z")

</div>

Yeah, I think that’s it! I see near identical timings between numpy and Julia with `@turbo`. `@turbo` also speeds up my my own explicit function, even though this already included a `@simd` macro.

**So guess the next question is: why doesn’t the dot call automatically apply simd instructions? And why does my explicit function with manual `@simd` not work?**

(I’ve checked whether this slowdown is attributed to the slow handling of tails, as it mentioned on the LoopVectorization page, but passing in arrays without tails does not make any difference in my own testing.)

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [August 5, 2022, 7:13am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/6 "2022-08-05T07:13:35Z")

</div>

It’s a good question. I am not able to figure it out, but I believe LoopVectorization.jl in some cases uses other implementations of some functions which are hard to simd. Probably, `exp` is one of them.

Julia is just using general mechanisms to calculate `exp` over your array, while numpy is free to do a special array version, maybe much like LoopVectorization does. Making simd happen in Julia in this case apparently takes more than a simple `@simd` annotation.

---

<div class="post-metadata">

**Author:** ![ikirill](https://avatars.discourse-cdn.com/v4/letter/i/43a26b/32.png) [@ikirill](https://discourse.julialang.org/u/ikirill)\
**Post date:** [August 5, 2022, 7:14am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/7 "2022-08-05T07:14:37Z")

</div>

I just tried this, and I’m getting 402us in Julia 1.7.3 and 435us with numpy. I would have expected them to be completely equal, assuming they are doing the same thing. I’m surprised there’s a difference one way _or_ the other.

---

<div class="post-metadata">

**Author:** ![suavesito](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/suavesito/32/34386_2.png) [@suavesito](https://discourse.julialang.org/u/suavesito)\
**Post date:** [August 5, 2022, 7:17am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/8 "2022-08-05T07:17:50Z")

</div>

Curiously, in my machine I got the numpy code to run `np.exp(arr)` on ~960\mu s, the no LV.jl code on ~712\mu s and the `@turbo`’ed code in ~142\mu s.

I got the same results using v1.8-rc3 or 1.7.3.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [August 5, 2022, 11:03am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/9 "2022-08-05T11:03:09Z")

</div>

Which CPU do you have?

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [August 5, 2022, 11:38am UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/10 "2022-08-05T11:38:17Z")

</div>

> [@ikirill](#):
>
> I would have expected them to be completely equal, assuming they are doing the same thin

That is probably not a good assumption.  
I suspect numpy uses your system libM,  
julia uses implementations of math functions written in julia – following the same techniques libM authors use in terms of balancing speed and accurasy but potentially ending up at a different point.  
Even if exp isn’t one of the functions we have tweaked and optimized it would still be based on the OpenLibM definition which is probably not the same libM that your system runs.

---

<div class="post-metadata">

**Author:** ![acxz](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/acxz/32/16759_2.png) [@acxz](https://discourse.julialang.org/u/acxz)\
**Post date:** [August 5, 2022, 12:55pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/11 "2022-08-05T12:55:09Z")

</div>

Just as a another data point:

vanilla julia:  
600 us

turbo julia:  
150 us

python numpy:  
650 us

CPU: Intel i7-8850H (6) @ 2.6GHz

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [August 5, 2022, 1:02pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/12 "2022-08-05T13:02:17Z")

</div>

Possibly related:

> <https://github.com/JuliaLang/julia/pull/46238>
>
> before it wasn't because the compiler can't prove that the getindex into the tab…le is inbounds.

---

<div class="post-metadata">

**Author:** ![torrance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/torrance/32/38990_2.png) [@torrance](https://discourse.julialang.org/u/torrance)\
**Post date:** [August 5, 2022, 1:14pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/13 "2022-08-05T13:14:00Z")

</div>

The speedup in Python is very dependent on how numpy has been compiled. I have tested on various system and see the numpy code variably matching the speed of the `@turbo` Julia or alternatively the plain julia. Obviously some are making use of simd instructions, and some are not.

**That’s fine, but doesn’t address the question I asked earlier:**

1. Why isn’t `exp.(arr)` automatically applying simd optimizations?
2. And why doesn’t my explicit `@simd` annotation in the manual loop function apply simd instructions, either?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 5, 2022, 1:48pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/14 "2022-08-05T13:48:43Z")

</div>

> [@torrance](#):
>
> Why isn’t `exp.(arr)` automatically applying simd optimizations?

we generally don’t do this.

> [@torrance](#):
>
> And why doesn’t my explicit `@simd` annotation in the manual loop function apply simd instructions, either?

see the linked PR, and try this on nightly to see if you observe any difference

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [August 5, 2022, 1:51pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/15 "2022-08-05T13:51:27Z")

</div>

Basically the reason is that Julia can do this if it is able to inline the code, but generically turning scalar code into vector code through the compiler only is a very complicated optimization.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 5, 2022, 1:53pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/16 "2022-08-05T13:53:05Z")

</div>

> [@Oscar\_Smith](#):
>
> but generically turning scalar code into vector code through the compiler only is a very complicated optimization.

maybe there is something to python vectorized style then /s

and, \*cough, `vmap`…

---

<div class="post-metadata">

**Author:** ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)\
**Post date:** [August 5, 2022, 1:54pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/17 "2022-08-05T13:54:51Z")

</div>

There is no reason to expect Julia to outperform Python on a single call like this. A significant difference in either direction is just low fruit for the other (in this case, the Julia version failing to SIMD). This function is already optimized in Python (or, rather, in C or whatever language was used to write it under the hood). Julia’s performance benefits only really factor in when you are trying to do more complicated sequences of operations that lack a bespoke implementation.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [August 5, 2022, 1:56pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/18 "2022-08-05T13:56:16Z")

</div>

One thing we could that would make this better is to have a better interface to allow the compiler to replace scalar code with custom provided vector code. I’m not fully sure what that interface looks like though.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 5, 2022, 1:59pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/19 "2022-08-05T13:59:28Z")

</div>

1.8 rc-3

```julia
julia> @btime exp.(x) setup=(x = rand(100_000));
  383.979 μs (2 allocations: 781.30 KiB)

```

current master:

```julia
Julia Version 1.9.0-DEV.1086
Commit 01c0778057* (2022-08-05 06:31 UTC)

julia> @btime exp.(x) setup=(x = rand(100_000));
  171.195 μs (2 allocations: 781.30 KiB)

```

(AMD CPU, so no avx512)

---

<div class="post-metadata">

**Author:** ![jishnub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jishnub/32/33620_2.png) [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Post date:** [August 5, 2022, 2:00pm UTC](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325/20 "2022-08-05T14:00:31Z")

</div>

> [@jling](#):
>
> > [@torrance](#):
> >
> > Why isn’t `exp.(arr)` automatically applying simd optimizations?
> 
> we generally don’t do this

Perhaps I’m misunderstanding, but doesn’t the `copyto!` invoked in broadcasting use the `@simd` annotation?

> <https://github.com/JuliaLang/julia/blob/01c07780573b49db95be2f384f31a869fb6ee4a7/base/broadcast.jl#L973-L975>

[Next page](https://discourse.julialang.org/t/comparing-exp-performance-on-julia-versus-numpy/85325.md?page=2)
