# Lux.jl Chain of Dense layers don't see speed up going from 32 bits -\> 16 bits

**URL:** https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419
**Category:** Performance
**Tags:** machine-learning
**Created:** [August 6, 2022, 10:06pm UTC](https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419 "2022-08-06T22:06:30Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)
#### Post date: [August 6, 2022, 10:06pm UTC](https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419/1 "2022-08-06T22:06:30Z")

</div>

```julia
using Lux, BenchmarkTools, Random
const rng = Random.default_rng()
julia> let model = Chain(
                  Dense(23, 30, Lux.relu),
                  Dense(30, 25, Lux.relu),
                  Dense(25, 20, Lux.relu)
              )
              ps, st = Lux.setup(rng, model)
              @btime $model(inputs, $ps, $st)[1] setup=(inputs=rand(Float32, 23))
           end;
  364.164 ns (12 allocations: 1.17 KiB)

julia> let model = Chain(
                  Dense(23, 30, Lux.relu),
                  Dense(30, 25, Lux.relu),
                  Dense(25, 20, Lux.relu)
              )
              ps, st = Lux.setup(rng, model)
              @btime $model(inputs, $ps, $st)[1] setup=(inputs=rand(Float16, 23))
           end;
  400.045 ns (13 allocations: 1.33 KiB)

```

for reference, 64 bits is about 2x slower at `772.561 ns (12 allocations: 1.77 KiB)`

(1.8-rc3)

---

<div class="post-metadata">

### Author: ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)
#### Post date: [August 6, 2022, 10:14pm UTC](https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419/2 "2022-08-06T22:14:18Z")

</div>

Julia currently doesn’t use native float16 operations even if available and it adds some overhead due to the conversion overhead. If you were memory bound you might see some improvement but I doubt it. In addition I believe very few x86 chips have float16 operations so your mileage may vary. [Hardware Float16 on A64fx · Issue #40216 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/40216)

---

<div class="post-metadata">

### Author: ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)
#### Post date: [August 6, 2022, 10:15pm UTC](https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419/3 "2022-08-06T22:15:28Z")

</div>

that’s a bit unfortunate, if we can lift another ~2x from here we can compete with FPGA for some real time ML application…

it’s still possible we just need to find a CPU with better single core performance

---

<div class="post-metadata">

### Author: ![gbaraldi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbaraldi/32/22101_2.png) [@gbaraldi](https://discourse.julialang.org/u/gbaraldi)
#### Post date: [August 6, 2022, 10:17pm UTC](https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419/4 "2022-08-06T22:17:12Z")

</div>

If you have float16 operations you could try a build from source and comment out `PM->add(createDemoteFloat16Pass());` in aotcompile.cpp to check if there is a possible improvement there. If you do there might be a 2x improvement

---

<div class="post-metadata">

### Author: ![DrChainsaw](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/drchainsaw/32/8497_2.png) [@DrChainsaw](https://discourse.julialang.org/u/DrChainsaw)
#### Post date: [August 7, 2022, 8:05am UTC](https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419/5 "2022-08-07T08:05:56Z")

</div>

If those network sizes are representative of what you want to do then [GitHub - PumasAI/SimpleChains.jl: Simple chains](https://github.com/PumasAI/SimpleChains.jl) should give you a significant speedup.

---

<div class="post-metadata">

### Author: ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)
#### Post date: [August 7, 2022, 8:49am UTC](https://discourse.julialang.org/t/lux-jl-chain-of-dense-layers-dont-see-speed-up-going-from-32-bits-16-bits/85419/6 "2022-08-07T08:49:54Z")

</div>

jesus that’s fast

```julia
julia> let model = SimpleChain(
                static(23), # input dimension (optional)
                TurboDense{true}(relu, 30), # dense layer with bias that maps to 8 outputs and applies `tanh` activation
                TurboDense{true}(relu, 25),
                TurboDense{true}(relu, 20)
              )
              p = SimpleChains.init_params(a);
              @btime $model(inputs, $p) setup=(inputs=rand(Float32, 23))
           end;
  77.845 ns (0 allocations: 0 bytes)

```
