# Status of BFloat16

**URL:** <https://discourse.julialang.org/t/status-of-bfloat16/103109>\
**Category:** Performance\
**Created:** [August 23, 2023, 11:42am UTC](https://discourse.julialang.org/t/status-of-bfloat16/103109 "2023-08-23T11:42:19Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [August 23, 2023, 11:42am UTC](https://discourse.julialang.org/t/status-of-bfloat16/103109/1 "2023-08-23T11:42:19Z")

</div>

Hello,

I have tried Llama2 large language model in Julia following [https://github.com/chengchingwen/Transformers.jl/blob/master/example/Llama2\_example.ipynb](https://github.com/chengchingwen/Transformers.jl/blob/master/example/Llama2_example.ipynb). This works really nice and smoothly, but, the example uses `Float32`. To save memory, I wanted to use it with `Float16`, since the model card of llama2 says it that `torch_dtype` is `Float16` When I try that, the model starts to halucinate, so I guess that something overflow / underflows. I wanted to give a try to `BFloat16`, since they can better handle large differences in magnitude. Does anyone has an experience with BFloats and CUDA? Is there some `bf16` equivalent of `f16`?

I have tried this repository

> **[GitHub - JuliaMath/BFloat16s.jl: Nobody needed all those bits anyway](https://github.com/JuliaMath/BFloat16s.jl)**
>
> Nobody needed all those bits anyway. Contribute to JuliaMath/BFloat16s.jl development by creating an account on GitHub.

but I am not sure, how relevant it is.

Thanks for answers in advance.  
Tomas

---

<div class="post-metadata">

**Author:** ![ctkelley](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ctkelley/32/10684_2.png) [@ctkelley](https://discourse.julialang.org/u/ctkelley)\
**Post date:** [August 23, 2023, 11:52am UTC](https://discourse.julialang.org/t/status-of-bfloat16/103109/2 "2023-08-23T11:52:47Z")

</div>

That repo is a software implementation, so I suspect it will be slow. If it works as advertised, you should be able to get going for small problems to see if it solves your problems. There is Bfloat16 support in hardware out there (Apple M\* seems to have it somewhere, maybe in the neural engine) but software support could be hard to come by.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [August 23, 2023, 11:55am UTC](https://discourse.julialang.org/t/status-of-bfloat16/103109/3 "2023-08-23T11:55:43Z")

</div>

I would expect the software support to suck. But cuda has HW support, therefore I was hoping it would be possible to use it with `CUDA.jl`

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [August 23, 2023, 3:05pm UTC](https://discourse.julialang.org/t/status-of-bfloat16/103109/4 "2023-08-23T15:05:10Z")

</div>

CUDA.jl already support BFloat16s for some common API functions, like `gemm`, `gemv`, etc. Native kernel support for BFloat16 depends on Julia properly supporting the type, i.e., not through BFloat16s.jl’ emulation. Keep an eye on [Add support for BFloat16 · Issue #41075 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/41075) for the status of that.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [August 24, 2023, 6:57am UTC](https://discourse.julialang.org/t/status-of-bfloat16/103109/5 "2023-08-24T06:57:06Z")

</div>

Thanks, I will
