# Making Float16 LU better in generic\_lu.jl

**URL:** <https://discourse.julialang.org/t/making-float16-lu-better-in-generic-lu-jl/102612>\
**Category:** Numerics\
**Created:** [August 8, 2023, 7:29pm UTC](https://discourse.julialang.org/t/making-float16-lu-better-in-generic-lu-jl/102612 "2023-08-08T19:29:21Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ctkelley](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ctkelley/32/10684_2.png) [@ctkelley](https://discourse.julialang.org/u/ctkelley)\
**Post date:** [August 8, 2023, 7:29pm UTC](https://discourse.julialang.org/t/making-float16-lu-better-in-generic-lu-jl/102612/1 "2023-08-08T19:29:21Z")

</div>

I finally stumbled across `generic_lufact!` [here](https://github.com/JuliaLang/julia/blob/master/stdlib/LinearAlgebra/src/lu.jl). This seems to be where `Float16` goes to die.

I made a copy and threaded the critical loop and things got a lot better for half precision work on my Mac.

Threading this function gave me almost perfect parallelism in my half precision lu work. Adding an `@simd` before the inner loop made the improvement about 10x what I had a couple weeks ago. I made no attempt to connect the number of threads (8 in my case) to the size of the problem.

I used `Polyester.@batch` for this. That was faster for me than `FLoops.@floop` or `Threads.@threads`. But even they were 5-10x faster than doing nothing when coupled with `@simd` on the inner loop.

Is there a reason why things like `generic_lufact!` are not threaded?

All I did was change this

```julia
# Update the rest
            for j = k+1:n
                for i = k+1:m
                    A[i,j] -= A[i,k]*A[k,j]
                end
            end

```

to

```julia
# Update the rest
@batch for j = k+1:n
       @simd ivdep for i = k+1:m
            @inbounds A[i,j] -= A[i,k]*A[k,j]
       end
end

```

Even with this, my `Float16` LU takes 2.5–6x longer than LAPACK’s  
`Float64` LU. If anyone sees a way I can get some more speed out of this,  
I’d be glad to try it.

Short of coding the [Toledo algorithm](https://epubs.siam.org/doi/10.1137/S0895479896297744), like [LAPACK](https://www.netlib.org/lapack/explore-html/dd/d9a/group__double_g_ecomputational_ga0019443faea08275ca60a734d0593e60.html) and [RecursiveFactorization.jl](https://github.com/JuliaLinearAlgebra/RecursiveFactorization.jl) (neither of which support Float16) do, I don’t see any way to make much more progress. When LAPACK/BLAS realize the [mixed-precision dream](https://icl.utk.edu/bblas/sc18/files/NG_BLAS_SC18.pdf), this will not be necessary.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [August 8, 2023, 7:53pm UTC](https://discourse.julialang.org/t/making-float16-lu-better-in-generic-lu-jl/102612/2 "2023-08-08T19:53:25Z")

</div>

> [@ctkelley](#):
>
> Is there a reason why things like `generic_lufact!` are not threaded?

Some libraries use it as a base case for small problems. It is the fastest there, and these optimizations probably hurt that.

---

<div class="post-metadata">

**Author:** ![ctkelley](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ctkelley/32/10684_2.png) [@ctkelley](https://discourse.julialang.org/u/ctkelley)\
**Post date:** [August 8, 2023, 9:12pm UTC](https://discourse.julialang.org/t/making-float16-lu-better-in-generic-lu-jl/102612/3 "2023-08-08T21:12:32Z")

</div>

That makes sense. Optimization for me: pessimizaton for others.

---

<div class="post-metadata">

**Author:** ![ctkelley](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ctkelley/32/10684_2.png) [@ctkelley](https://discourse.julialang.org/u/ctkelley)\
**Post date:** [August 26, 2023, 3:11pm UTC](https://discourse.julialang.org/t/making-float16-lu-better-in-generic-lu-jl/102612/4 "2023-08-26T15:11:18Z")

</div>

I played with it and found that `minbatch = 16` was best for small problems and no worse for the larger ones.
