# Potential solution to the overflow problem; 64-bit considered harmful, 32- or 21-bit better

**URL:** <https://discourse.julialang.org/t/potential-solution-to-the-overflow-problem-64-bit-considered-harmful-32-or-21-bit-better/69918>\
**Category:** General Usage\
**Tags:** integer-overflow\
**Created:** [October 17, 2021, 12:29am UTC](https://discourse.julialang.org/t/potential-solution-to-the-overflow-problem-64-bit-considered-harmful-32-or-21-bit-better/69918 "2021-10-17T00:29:46Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [October 17, 2021, 3:23pm UTC](https://discourse.julialang.org/t/potential-solution-to-the-overflow-problem-64-bit-considered-harmful-32-or-21-bit-better/69918/4 "2021-10-17T15:23:10Z")

</div>

> [@ericphanson](#):
>
> I think the basic issue there is that as soon as you do one `*` , the result is an `Int64`, and you’re back where we started.

No, because even if you get to `Int64`, and use that subsequently it will be checked (unlike for status quo), but yes, you want to avoid going that far for safety AND _performance_ reasons (with `Int21` you have a budget of three multiplies before slowdown, or more as explained with my polynomial example, or 43 additions). But **no matter the size of the resulting type you end up with, it will be checked**. I just explained, we could go to `Int128` next, but likely you want to go to `BigInt` directly (and possibly it will have a fast path for values that would have fit into `Int64` or `Int128`).

> [@Elrod](#):
>
> checking will not be faster.
> 
> You can confirm that Base.checked\_add’s generated code

**Yes, it will be faster** because there will be no check (as long as the result fits in one register), i.e. in the common case (for Int6 **3** inputs to addition, and seemingly `Int64` too; for multiply of `Int64` with `Base.checked_mul` it seems much worse than my idea), and you get a three-byte assembly code with my scheme, just an `add` (or `addq` or similar) which is as fast as non-checked 32-bit add (which seems to be shorter code than 64-bit add that gets me “`lea	rax, [rdi + rsi]`”.

The kicker is it will work with SIMD code too (and provide 2x memory bandwidth), except for all other overflow check schemes:

> <https://stackoverflow.com/questions/10511000/sse2-integer-overflow-checking>

> No flags are touched by the underlying PADDD instruction.
> 
> So to test this, you have to write additional code, depending on what you want to do.

> [@Elrod](#):
>
> `Base.checked_add`, `Base.checked_mul`, etc all already exist and are going to be about as efficient as it gets.

There’s some reason they’re not the default, and as you show the code expansion is huge. Non-taken branches are cheap, I can believe 15% overhead or less, that statistic seems rather high if the branches were the only added code, but the bad thing about those functions are the huge L1 Dcache pollution, and also most people don’t use or know of them. Also 1 letter operators much better… or 0 letters as in `2x`.

---

_[View the full topic](https://discourse.julialang.org/t/potential-solution-to-the-overflow-problem-64-bit-considered-harmful-32-or-21-bit-better/69918)._
