# abs(::Int32) slower than abs(::Int)

**URL:** https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011
**Category:** Performance
**Tags:** question
**Created:** [January 19, 2024, 2:56pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011 "2024-01-19T14:56:08Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![djholiver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djholiver/32/50470_2.png) [@djholiver](https://discourse.julialang.org/u/djholiver)
#### Post date: [January 19, 2024, 2:56pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/1 "2024-01-19T14:56:08Z")

</div>

Hi,

I have a need for a lot of large integer vectors, so have attempted to make use of Int32 rather than Int to minimise the memory footprint. As I need to ensure the values are positive, I call abs(i) on each element during iteration.

Before pursuing a change to my code, I did a quick (repeatable) timing check:

```julia
     i :: Int32 = -1
    @btime abs(i) # 2.300 ns (0 allocations: 0 bytes)
    @btime abs(-1) # 1.000 ns (0 allocations: 0 bytes)

```

I’m sure that there is a good reason for this, but is there any way to mitigate?

Regards,

---

<div class="post-metadata">

### Author: ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)
#### Post date: [January 19, 2024, 3:13pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/2 "2024-01-19T15:13:53Z")

</div>

Interpolate in your benchmarks as the BenchmarkTools docs suggest?

```julia
julia> i32 = Int32(-1)
-1

julia> i64 = -1
-1

julia> @btime abs($i64)
  1.500 ns (0 allocations: 0 bytes)
1

julia> @btime abs($i32)
  1.900 ns (0 allocations: 0 bytes)
1

julia> @btime abs($i32)
  1.500 ns (0 allocations: 0 bytes)
1

```

also benchmarks in the \<2ns category aren’t super reliable, try something closer to your actual use case maybe like

```julia
julia> x32 = rand(Int32, 100_000); x64 = Int.(rand(Int32, 100_000));

julia> @btime abs.($x32);
  23.200 μs (2 allocations: 390.67 KiB)

julia> @btime abs.($x64);
  53.000 μs (2 allocations: 781.30 KiB)

```

---

<div class="post-metadata">

### Author: ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)
#### Post date: [January 19, 2024, 3:17pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/3 "2024-01-19T15:17:22Z")

</div>

> [@nilshg](#):
>
> ```julia
> julia> @btime abs.($x32);
> 23.200 μs (2 allocations: 390.67 KiB)
> 
> julia> @btime abs.($x64);
> 53.000 μs (2 allocations: 781.30 KiB)
> 
> ```

There’s probably a difference in allocation time here. Perhaps pre-allocate the output vectors to better isolate the time of the `abs`?

~~There’s also a factor of 2 difference in simd effect, which is not relevant for scalar `abs`.~~ (simd may actually be relevant, per the OP).

---

<div class="post-metadata">

### Author: ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)
#### Post date: [January 19, 2024, 3:35pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/4 "2024-01-19T15:35:02Z")

</div>

Sorry I maybe I should have left it at “actual use case” and not made up one potential (and likely irrelevant) example… My point was mainly that fluctuations of \<1ns in benchmarks are unlikely to be informative of impacts on real world code with perceivable runtimes.

An allocation free benchmark:

```julia
julia> function f(x)
           for i ∈ eachindex(x)
               x[i] = abs(x[i])
           end
           return x
       end;

julia> @btime f($x32);
  4.543 μs (0 allocations: 0 bytes)

julia> @btime f($x64);
  9.000 μs (0 allocations: 0 bytes)

```

---

<div class="post-metadata">

### Author: ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)
#### Post date: [January 19, 2024, 7:31pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/5 "2024-01-19T19:31:56Z")

</div>

> [@nilshg](#):
>
> Interpolate in your benchmarks as the BenchmarkTools docs suggest?

You may need `Ref` interpolation:

```julia
julia> using BenchmarkTools

julia> @btime abs($(Ref(Int64(-1)))[]);
  1.500 ns (0 allocations: 0 bytes)

julia> @btime abs($(Ref(Int32(-1)))[]);
  1.500 ns (0 allocations: 0 bytes)

```

to prevent the compiler from evaluating the whole expression statically, as [explained in the BenchmarkTools manual](https://juliaci.github.io/BenchmarkTools.jl/stable/manual/#Understanding-compiler-optimizations).

---

<div class="post-metadata">

### Author: ![djholiver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djholiver/32/50470_2.png) [@djholiver](https://discourse.julialang.org/u/djholiver)
#### Post date: [January 19, 2024, 8:07pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/6 "2024-01-19T20:07:03Z")

</div>

Thanks so much for the extra detail - incredibly helpful.

So essentially, even though I can still repeat my original results (without interpolation), this is purely as a result of how I was benchmarking a value?

Why is it that the single evaluation is comparable (in your test) but the loops are different (in Nils?)

---

<div class="post-metadata">

### Author: ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)
#### Post date: [January 19, 2024, 8:11pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/7 "2024-01-19T20:11:58Z")

</div>

> [@djholiver](#):
>
> Why is it that the single evaluation is comparable (in your test) but the loops are different (in Nils?)

Loops over sufficiently large arrays will be limited by the speed of memory access, not by the `abs` function. With `Int64` there is twice as much memory to access compared to an `Int32` array of the same length.

---

<div class="post-metadata">

### Author: ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)
#### Post date: [January 19, 2024, 8:18pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011/8 "2024-01-19T20:18:30Z")

</div>

Also, AVX2 has instructions for computing the absolute value of 8x32-bit integers, but not 4x64-bit integers. So the latter is computed using a subtraction and a blend instruction, which is slower.  
(memory may still be the bottleneck in reality, not sure)
