# Drastic performance hit matrix multiply different types. Internal cast julia vs numpy?

**URL:** <https://discourse.julialang.org/t/drastic-performance-hit-matrix-multiply-different-types-internal-cast-julia-vs-numpy/17120>\
**Category:** Numerics\
**Created:** [November 3, 2018, 7:52pm UTC](https://discourse.julialang.org/t/drastic-performance-hit-matrix-multiply-different-types-internal-cast-julia-vs-numpy/17120 "2018-11-03T19:52:38Z")\
**Posts on this page:** 1\
**Showing post:** 8

<div class="post-metadata">

**Author:** ![davidbp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidbp/32/463_2.png) [@davidbp](https://discourse.julialang.org/u/davidbp)\
**Post date:** [November 3, 2018, 10:42pm UTC](https://discourse.julialang.org/t/drastic-performance-hit-matrix-multiply-different-types-internal-cast-julia-vs-numpy/17120/8 "2018-11-03T22:42:02Z")

</div>

I understand how I can solve this, but was this the expected default behaviour? When I see the code\_lowered I agree the ! suffix suggest no copies. Neverhteless from my “high level code” X\*Y I already expect memory to be allocated. And I was expecting a Blas call !

```julia
N = 1000

X = rand(N,N)
Y = rand(N,N)
println("N = 1000 same type")
@btime X*Y

X = rand(N,N)
Y = rand(Float32, N ,N)
println("N = 1000 different type")
@btime X*Y

```

Gives me same allocations in MiB but a drastically worse performance

```julia
N = 1000 same type
  15.822 ms (2 allocations: 7.63 MiB)
N = 1000 different type
  1.326 s (8 allocations: 7.63 MiB)

```

I would rather have some extra copies and much better speed. I can see this as a personal preference but at the end of the day If I expected a call to a highly optimized Blas (and therefore speed).

This was the issue of the huge performance hit from here:

> [@Quite bad performance of Julia 0.6.4 vs Python+Numpy](https://discourse.julialang.org/t/quite-bad-performance-of-julia-0-6-4-vs-python-numpy/16978):
>
> Dear all, I guess I could use a piece of good help from the experts here. The fact is that I was about to translate a piece of Monte Carlo code from my old 0.6.4 program library, into Julia 1.0 (going through 0.7 first), until a friend of mine tried essentially the same code under Anaconda Python on the same computer. I was so disappointed by the results of test runs that I’m wondering what could be wrong on my code. I post here the Julia and the Python codes: Julia 0.6.4: function TEST\_ais(…

and I expect it to happen to a lot of people.

If that was the expected behaviour maybe there should be a line in performance tips stating “Matrix multiplication with different types will destroy your performance yet it won’t allocate memory. This was a design decision.”.

Arrays are a remarkable part of Julia, and I expect that for a lot of people LinearAlgebra will be also relevant. Since the language clearly shows speed as a big feature (the comparisson between languages here [Julia Micro-Benchmarks](https://julialang.org/benchmarks/) does not show memory) I felt the fastest solution should be the default one.

[https://docs.julialang.org/en/v1/manual/performance-tips/index.html](https://docs.julialang.org/en/v1/manual/performance-tips/index.html)

Certainly in other frameworks (pytorch/numpy/matlab) this is not the default behaviour.

---

_[View the full topic](https://discourse.julialang.org/t/drastic-performance-hit-matrix-multiply-different-types-internal-cast-julia-vs-numpy/17120)._
