# Massive performance penalty for Float16 compared to Float32

**URL:** <https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864>\
**Category:** Performance\
**Tags:** performance\
**Created:** [November 3, 2017, 7:16pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864 "2017-11-03T19:16:50Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![zenon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zenon/32/6508_2.png) [@zenon](https://discourse.julialang.org/u/zenon)\
**Post date:** [November 3, 2017, 7:16pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/1 "2017-11-03T19:16:51Z")

</div>

Hello,

I just do the Coursera lectures about Deep Leanring by Andrew Ng, and want to implement what I learned in Julia (lecture is in Python), so I can check my understanding. So there is a small neural net, and it takes time, and I change the numerical type from Float64 to Float32 and get like 30% speed up.

Great, I think, and try Float16. That didn’t go so well. Time went up by a factor of 100 compared to Float64.  
Time. Not speed.

I tried to isolate the issue and found a penalty of factor 10 for point wise multiplication of matrices.

This is the code in my test file:

```julia
numType = Float16

A = rand(numType, 10000, 10000)
B = rand(numType, 10000, 10000)
C = Array{numType, 2}(10000,10000)

@time C .= A .* B

```

and the result is (without some ramp up, just start the file, so that’s just a very crude test:

First two starts with Float32

```julia
  0.235396 seconds (46.17 k allocations: 2.444 MiB)

  0.209013 seconds (46.17 k allocations: 2.444 MiB)

```

And the two with Float16.

```julia
  1.848572 seconds (46.60 k allocations: 2.468 MiB)

  1.847379 seconds (46.60 k allocations: 2.468 MiB)

```

I found in the docs that Float16 is “implemented in software”, but that’s less performant than I expected.

Am I doing something wrong?

Thank you, and kind regards, z.

---

<div class="post-metadata">

**Author:** ![zenon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zenon/32/6508_2.png) [@zenon](https://discourse.julialang.org/u/zenon)\
**Post date:** [November 3, 2017, 7:18pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/2 "2017-11-03T19:18:29Z")

</div>

Btw. Why is it allocating memory in the first place?

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [November 3, 2017, 7:23pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/3 "2017-11-03T19:23:12Z")

</div>

> [@zenon](#):
>
> Why is it allocating memory in the first place

Don’t benchmark in the global scope. Put that in a function and run twice.

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [November 3, 2017, 7:24pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/4 "2017-11-03T19:24:57Z")

</div>

`Float16` is in an odd place right now. For every calculation the `Float16` is first converted to a `Float32` in which the computation is performed and then converted back to `Float16`. This is necessary since most hardware has no native support for `Float16`. The only hardware where `Float16` is really relevant is GPUs otherwise it is primarily a storage type.

---

<div class="post-metadata">

**Author:** ![tim.holy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tim.holy/32/52_2.png) [@tim.holy](https://discourse.julialang.org/u/tim.holy)\
**Post date:** [November 3, 2017, 7:25pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/5 "2017-11-03T19:25:10Z")

</div>

Check out how `Float16` multiplication is defined:

```julia
julia> x, y = rand(Float16), rand(Float16)
(Float16(0.3818), Float16(0.825))

julia> @which x*y
*(a::Float16, b::Float16) in Base at float.jl:372

julia> @edit x*y

```

You’ll see it first converts to `Float32`, performs the multiplication, and then converts back to `Float16`. Hence it cannot be as fast as `Float32`. FPUs support `Float32` natively but not `Float16`.

Even more significantly, julia used parallelized BLAS routines (written in Fortran) for `Float32`. For `Float16` it is not much better than naive multiplication (it’s a little more cache friendly, but not parallelized nor as optimized as it is for `Float32`).

---

<div class="post-metadata">

**Author:** ![zenon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zenon/32/6508_2.png) [@zenon](https://discourse.julialang.org/u/zenon)\
**Post date:** [November 3, 2017, 8:26pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/6 "2017-11-03T20:26:16Z")

</div>

@ChrisRackauckas Ah, right. Should have known that.

```julia
> t32();
  0.158636 seconds

> t32();
  0.134481 seconds

> t16();
  1.785036 seconds

> t16();
  1.782366 seconds

```

for

```julia
function t16()

    numType = Float16

    A = rand(numType, 10000, 10000)
    B = rand(numType, 10000, 10000)
    C = Array{numType, 2}(10000,10000)

    @time C .= A .* B
end

function t32()

    numType = Float32

    A = rand(numType, 10000, 10000)
    B = rand(numType, 10000, 10000)
    C = Array{numType, 2}(10000,10000)

    @time C .= A .* B
end

```

---

<div class="post-metadata">

**Author:** ![zenon](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zenon/32/6508_2.png) [@zenon](https://discourse.julialang.org/u/zenon)\
**Post date:** [November 3, 2017, 8:27pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/7 "2017-11-03T20:27:14Z")

</div>

Thank you all for the explanations!

---

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [November 4, 2017, 12:16pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/8 "2017-11-04T12:16:58Z")

</div>

Intel added vector instructions to do conversions to/from 16-bit floats many years ago, and in fact, showed that (because of using half the memory, better cache utilization) that using 16-bit could be faster than 32-bit, for larger operations, and not that much slower for smaller vectors.

[https://software.intel.com/en-us/articles/performance-benefits-of-half-precision-floats](https://software.intel.com/en-us/articles/performance-benefits-of-half-precision-floats)

It seems that making sure that Julia can use the SIMD instructions when doing vector operations on 16-bit floats could acheive some nice performance benefits.

---

<div class="post-metadata">

**Author:** ![chobbes](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chobbes](https://discourse.julialang.org/u/chobbes)\
**Post date:** [June 19, 2022, 7:25am UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/9 "2022-06-19T07:25:32Z")

</div>

> [@vchuravy](#):
>
> For every calculation the `Float16` is first converted to a `Float32` in which the computation is performed and then converted back to `Float16`. This is necessary since most hardware has no native support for `Float16`. The only har

Is this still the case now in June 2022 that Float16 is first converted to Float32 before any calculation is invoked and converted back to Float16 after the calculation is done?

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [June 19, 2022, 8:18am UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/10 "2022-06-19T08:18:28Z")

</div>

There is not yet widespread support for Float16 in Floating Point hardware. So yes.  
However, see the next note by @giordano.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [June 19, 2022, 9:17am UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/11 "2022-06-19T09:17:34Z")

</div>

The Apple Silicon CPUs such as the M1 have hardware support for float16

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [June 19, 2022, 9:52am UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/12 "2022-06-19T09:52:39Z")

</div>

Is the float16 support fully interwoven into whatever SIMD they support?

also informative is [this from 2020](https://scicomp.stackexchange.com/questions/35187/is-half-precision-supported-by-modern-architecture)

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [June 19, 2022, 10:02am UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/13 "2022-06-19T10:02:33Z")

</div>

I ran [this benchmark](https://github.com/JuliaLang/julia/issues/40308#issuecomment-973127696) on a64fx, another CPU with hardware support for float16, and simd scaling was pretty good. I seem to recall I tried the same on M1, with comparable results.

---

<div class="post-metadata">

**Author:** ![chobbes](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chobbes](https://discourse.julialang.org/u/chobbes)\
**Post date:** [June 19, 2022, 12:04pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/14 "2022-06-19T12:04:31Z")

</div>

@JeffreySarnoff @giordano Thanks a lot! I don’t work on M1 or Fujitsu a64fx. Does this mean that I basically stand no chance at the moment?

A dumb question - who should I expect to have the Float16 problem solved? The CPU manufacturer or Julia community? Or both?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [June 19, 2022, 1:09pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/15 "2022-06-19T13:09:02Z")

</div>

On the Julia side, we could make float16 mammal pretty fast (@celrod), but for general use, we need cpu support

---

<div class="post-metadata">

**Author:** ![chobbes](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chobbes](https://discourse.julialang.org/u/chobbes)\
**Post date:** [June 19, 2022, 2:22pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/16 "2022-06-19T14:22:08Z")

</div>

Thanks, buddy. By ‘pretty fast’, do you mean that we can achieve something faster than fp32 so that there is an clear advantage of using fp16 if only half precision is needed?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [June 19, 2022, 5:08pm UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/17 "2022-06-19T17:08:21Z")

</div>

For small matrices float32 will be faster without hardware support, but for matrices bigger than cache, you might be able to go faster than float32 by but converting as you pack sub blocks

---

<div class="post-metadata">

**Author:** ![chobbes](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chobbes](https://discourse.julialang.org/u/chobbes)\
**Post date:** [June 20, 2022, 1:15am UTC](https://discourse.julialang.org/t/massive-performance-penalty-for-float16-compared-to-float32/6864/18 "2022-06-20T01:15:29Z")

</div>

Thanks for the hint! @Oscar_Smith
