# Quite bad performance of Julia 0.6.4 vs Python+Numpy

**URL:** <https://discourse.julialang.org/t/quite-bad-performance-of-julia-0-6-4-vs-python-numpy/16978>\
**Category:** General Usage\
**Created:** [October 30, 2018, 5:44pm UTC](https://discourse.julialang.org/t/quite-bad-performance-of-julia-0-6-4-vs-python-numpy/16978 "2018-10-30T17:44:41Z")\
**Posts on this page:** 1\
**Showing post:** 26

<div class="post-metadata">

**Author:** ![davidbp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidbp/32/463_2.png) [@davidbp](https://discourse.julialang.org/u/davidbp)\
**Post date:** [November 9, 2018, 4:25pm UTC](https://discourse.julialang.org/t/quite-bad-performance-of-julia-0-6-4-vs-python-numpy/16978/26 "2018-11-09T16:25:17Z")

</div>

I think that if you call the BLAS then there is no way julia can avoid running the whole code (it cannot know if the code compiled in another language is afecting the program or not).

In this case all the performance problem comes form a tiny detail that is hard to get.

- `vv_k*W` will not call the BLAS in this case (because the types of vv\_k and W are different)

```julia
zer_0 = 0
pl_1 = 1
vv_k = rand([zer_0,pl_1],Nsamples,Nv+1)

```

- `vv_k*W` will call the blass in this case (because the types of vv\_k and W are the same)

```julia
zer_0 = 0.
pl_1 = 1.
vv_k = rand([zer_0,pl_1],Nsamples,Nv+1)

```

And the time it takes to compute that can be very different

```julia
# vv_k*W timing 
N = 3000 same type 0.431971 seconds
N = 3000 different type 36.492838 seconds

```

Here I did some tests

> [@Drastic performance hit matrix multiply different types. Internal cast julia vs numpy?](https://discourse.julialang.org/t/drastic-performance-hit-matrix-multiply-different-types-internal-cast-julia-vs-numpy/17120):
>
> Because of several performance problems I have seen in discourse I wanted to ask why this is happening. And how is it possible it is not happening in numpy? I know we can avoid the problem I present here beeing careful when defining arrays but the performance hit seems still too big. Here NUMPY N = 1000 time float64 and float64 0.05866699999999997 N = 1000 time float64 and float32 0.06433399999999995 N = 3000 time float64 and float64 1.2407199999999996 N = 3000 time float64 and float32 1…

The random number generation you mention and the rewritting of the code to avoid allocations are important performance tricks you can use but they are irrelevant compared to the BLAS/not BLAS call when matrices contain 1000x1000 elements or more.

---

_[View the full topic](https://discourse.julialang.org/t/quite-bad-performance-of-julia-0-6-4-vs-python-numpy/16978)._
