# Int numerical calculation speed slower than Float?

**URL:** <https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705>\
**Category:** Performance\
**Created:** [February 16, 2020, 6:01am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705 "2020-02-16T06:01:54Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![nuclear718](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nuclear718/32/16163_2.png) [@nuclear718](https://discourse.julialang.org/u/nuclear718)\
**Post date:** [February 16, 2020, 6:01am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/1 "2020-02-16T06:01:54Z")

</div>

Why the calculation speed of “Int” matrix is much slower than “Float” matrix?

IDE: Jupyter  
Julia version: 1.1.0

Code for Float:  
s=1000  
d1=rand(s,s)  
d2=rand(s,s)  
@time (d1\*d2);  
Result: 0.032192 seconds (6 allocations: 7.630 MiB)

Code for Int:  
s=1000  
d3=rand(Int,s,s)  
d4=rand(Int, s,s)  
@time (d3\*d4);  
Result: 1.131875 seconds (12 allocations: 7.630 MiB)

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [February 16, 2020, 7:14am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/2 "2020-02-16T07:14:08Z")

</div>

Float matrices use blas. Int uses a generic fallback. Making the fallback method multithreaded for large matrices would fix much of the problem.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [February 16, 2020, 7:28am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/3 "2020-02-16T07:28:34Z")

</div>

I think that for floats, BLAS is used, while for integers it is native Julia code. The latter could probably be made faster, but it is not a common use case so it is waiting for someone to do it.

(also, please [quote code](https://discourse.julialang.org/t/psa-how-to-quote-code-with-backticks/7530))

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [February 16, 2020, 8:08am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/4 "2020-02-16T08:08:41Z")

</div>

Do you think it would be worth it for mixed type matmul to convert arguments before multiplying? I think that should speed things up a lot (with some memory downsides). The other big thing the fallback needs is better cache aware looping.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [February 16, 2020, 8:27am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/5 "2020-02-16T08:27:43Z")

</div>

No, I would not convert. First, integers have specific overflow semantics in Julia different from float, so I am not sure what is intended and what isn’t.

Second (and more importantly), you really have to go out of your way to get a matrix with a non-concrete element type when writing idiomatic code, so I am not sure it is a common use case. I would leave it up to the user to promote if that is needed.

---

<div class="post-metadata">

**Author:** ![mschauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mschauer/32/13946_2.png) [@mschauer](https://discourse.julialang.org/u/mschauer)\
**Post date:** [February 16, 2020, 8:59am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/6 "2020-02-16T08:59:22Z")

</div>

We should have a specialized very effective Int multiplication kernel though.

---

<div class="post-metadata">

**Author:** ![improbable22](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/improbable22/32/5464_2.png) [@improbable22](https://discourse.julialang.org/u/improbable22)\
**Post date:** [February 16, 2020, 9:14am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/7 "2020-02-16T09:14:21Z")

</div>

One way to do this is [LoopVectorization’s example](https://github.com/chriselrod/LoopVectorization.jl#matrix-multiply), which is faster but still not as fast as floats:

```julia
julia> C1 = Matrix{Int}(undef, M, N); A = rand(1:100, M, K); B = rand(1:100, K, N);

julia> C2 = similar(C1); C3 = similar(C1);

julia> @btime mygemmavx!($C1, $A, $B)
  77.412 μs (0 allocations: 0 bytes)

julia> @btime mygemm!($C2, $A, $B)
  245.869 μs (0 allocations: 0 bytes)

julia> @btime mul!($C3, $A, $B); # julia's generic_matmul
  164.278 μs (6 allocations: 336 bytes)

```

compared to Float64:

```julia
julia> @btime mygemmavx!($C1, $A, $B)
  14.599 μs (0 allocations: 0 bytes)

julia> @btime mygemm!($C2, $A, $B)
  290.296 μs (0 allocations: 0 bytes)

julia> @btime mul!($C3, $A, $B); # openblas, not MKL
  22.635 μs (0 allocations: 0 bytes)

```

Note BTW that in your example, `rand(Int, ...)` produces lots of large numbers, which will overflow:

```julia
julia> (float(d3) * float(d4)) .- (d3 * d4) |> extrema
(-4.2418818662328814e39, 4.4115065145744565e39)

julia> extrema(A)
(1, 100)

julia> (float(A) * float(B)) .- (A * B) |> extrema
(0.0, 0.0)

```

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [February 16, 2020, 4:33pm UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/8 "2020-02-16T16:33:24Z")

</div>

Using

```julia
julia> M, K, N = 72, 75, 71;

```

My results with `Float64` are:

```julia
julia> BLAS.set_num_threads(1)

julia> @btime mygemmavx!($C1, $A, $B)
  7.380 μs (0 allocations: 0 bytes)

julia> @btime mygemm!($C2, $A, $B)
  231.900 μs (0 allocations: 0 bytes)

julia> @btime mul!($C3, $A, $B); # julia's generic_matmul
  6.780 μs (0 allocations: 0 bytes)

```

And with `Int`:

```julia
julia> @btime mygemmavx!($C1, $A, $B)
  26.158 μs (0 allocations: 0 bytes)

julia> @btime mygemm!($C2, $A, $B)
  190.645 μs (0 allocations: 0 bytes)

julia> @btime mul!($C3, $A, $B); # julia's generic_matmul
  101.748 μs (6 allocations: 336 bytes)

```

So `Int` is about 3.5x slower than Float for me, while it is 5.3x slower for you.  
With integers, it uses the `vpmullq` instruction for integer multiplication. But this instruction appears to be slow, with a [reciprocal throughput of around 1.5-3](https://www.agner.org/optimize/instruction_tables.pdf), while the vpaddq instruction is around 0.33 or 0.5.  
The floating point versions use fused multiply-add instructions to combine both the multiplication and addition, and have a reciprical throughput of about 0.5.  
You can think of “reciprical throughput” as how many clock cycles it takes per completed instruction if a core is executing many simultaneously. It generally takes many more clock cycles to complete any given instruction (e.g., 4 for the fma instructions), but a core can work on many simultaneously, thus the rate at which they’re completed can be much faster.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [February 16, 2020, 6:55pm UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/9 "2020-02-16T18:55:34Z")

</div>

> [@mschauer](#):
>
> We should have a specialized very effective Int multiplication kernel though.

Realize that implementing a highly optimized matrix–matrix multiplication is nontrivial. Optimized BLAS libraries typically involve tens of thousands of lines of code and painstaking performance tuning. While there is no theoretical reason why this cannot be replicated in Julia, it is a huge undertaking.

---

<div class="post-metadata">

**Author:** ![improbable22](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/improbable22/32/5464_2.png) [@improbable22](https://discourse.julialang.org/u/improbable22)\
**Post date:** [February 16, 2020, 9:41pm UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/10 "2020-02-16T21:41:35Z")

</div>

> [@Elrod](#):
>
> `vpmullq` instruction for integer multiplication. But this instruction appears to be slow, with a [reciprocal throughput of around 1.5-3](https://www.agner.org/optimize/instruction_tables.pdf), while the vpaddq instruction is around 0.33 or 0.5.  
> The floating point versions use fused multiply-add instructions to combine both the multiplication and addition, and have a reciprical throughput of about 0.5.

Thanks, interesting. And the reason for this difference in the first place, perhaps that there’s just more demand for vectorised floating point stuff, which justifies spending a lot of silicon on it?

---

<div class="post-metadata">

**Author:** ![mschauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mschauer/32/13946_2.png) [@mschauer](https://discourse.julialang.org/u/mschauer)\
**Post date:** [February 17, 2020, 9:29am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/11 "2020-02-17T09:29:12Z")

</div>

I have no doubt. “We should” in open source can perhaps only mean “it would be appreciated”.

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [February 17, 2020, 9:45am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/12 "2020-02-17T09:45:37Z")

</div>

Not having BLAS implementation for a type really hurts.

This is why Arraymancer (a BLAS written in Nim that supports a bunch of types/is generic?),  
trashes everyone at integer matmul

> <https://github.com/mratsim/Arraymancer/blob/71ccad01953060fad121a03b49b97e81ed8ac12c/benchmarks/integer_matmul.jl#L40>

---

<div class="post-metadata">

**Author:** ![mschauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mschauer/32/13946_2.png) [@mschauer](https://discourse.julialang.org/u/mschauer)\
**Post date:** [February 17, 2020, 10:10am UTC](https://discourse.julialang.org/t/int-numerical-calculation-speed-slower-than-float/34705/13 "2020-02-17T10:10:24Z")

</div>

Out of curiosity, how comes that JuliaBLAS is closing in on BLAS for floats but is not fundamentally faster than generic matmul for ints?
