# 16 bit float on Transformers.jl

**URL:** <https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867>\
**Category:** General Usage\
**Tags:** transformers\
**Created:** [December 20, 2023, 5:01pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867 "2023-12-20T17:01:33Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![RainerHeintzmann](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rainerheintzmann/32/19726_2.png) [@RainerHeintzmann](https://discourse.julialang.org/u/RainerHeintzmann)\
**Post date:** [December 20, 2023, 5:01pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/1 "2023-12-20T17:01:33Z")

</div>

I am new to LLMs but thought to give it a shot using `Transformers.jl`. This worked rather well, after creating an account at HuggingFace and getting permission for Llama-7B from Meta. Yet, I was surprised how much GPU ram was needed even for this “small” model. Running the [notebook example](https://github.com/chengchingwen/Transformers.jl/blob/master/example/Llama2_example.ipynb), works, but it maxes out my GPU memory on a 47Gb RTX A6000.

Answering a question like the example “Can you explain to me briefly what is the Julia programming language?” takes about 1 minute on that GPU.

I wonder, whether one can speed up things by using 16 bit floats in the weigths? Is this supported by `Transformers.jl`? Can one use the same model and convert it or does one needs another one from `HuggingFace`?  
Or is the system always claiming all GPU memory anyway, and the speed is what is to be expected?

Thanks for any help!

---

<div class="post-metadata">

**Author:** ![mohamed.d180](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mohamed.d180/32/52028_2.png) [@mohamed.d180](https://discourse.julialang.org/u/mohamed.d180)\
**Post date:** [December 20, 2023, 8:06pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/2 "2023-12-20T20:06:12Z")

</div>

I don’t know if this will help :  
[State of quantization](https://github.com/chengchingwen/Transformers.jl/issues/154)

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [December 20, 2023, 8:43pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/3 "2023-12-20T20:43:40Z")

</div>

16bit with Transformers will work out of the box, but it will hallucinate. I believe the problem is that when we convert weights and embeddings to Float16, the entire computation is carried in Float16. What HF’s transformers do, I believe, is that they store weights in Float16, but perform the computation in Float32, which gives them the best of both words. I think we would need to i) either convert weights just before use to Float32, or define a matmul for multiplication of Float16 and Float32 matrices.

---

<div class="post-metadata">

**Author:** ![pitsianis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pitsianis/32/26588_2.png) [@pitsianis](https://discourse.julialang.org/u/pitsianis)\
**Post date:** [December 21, 2023, 2:56am UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/4 "2023-12-21T02:56:49Z")

</div>

You may want to consider [GitHub - cafaxo/Llama2.jl: llama2.c but in Julia](https://github.com/cafaxo/Llama2.jl)

Using `llama-2-7b-chat.ggmlv3.q4_K_S.bin`,  
for the same prompt, I got 10 tokens per second on a MacBook Pro.

---

<div class="post-metadata">

**Author:** ![RainerHeintzmann](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rainerheintzmann/32/19726_2.png) [@RainerHeintzmann](https://discourse.julialang.org/u/RainerHeintzmann)\
**Post date:** [December 21, 2023, 3:00pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/5 "2023-12-21T15:00:18Z")

</div>

Thanks for this hint.  
I was using the Tiny-Lllama example (the download link being in the code), which indeed gave me roughly 200 tokens per second. Nice.

But then I tried downloading `llama-2-7b-chat.ggmlv3.q4_K_S.bin` but I am not sure where to find it. It was listed on HuggingFace, but I could not find a way to download anything, trying for about an hour. In the description it says somehting about a download button, but I could not find one (being logged in). It also sais something about this being an old format.  
Is there any link for this?

Can these (Clearly smaller) models also be run using Transformers.jl?

---

<div class="post-metadata">

**Author:** ![pitsianis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pitsianis/32/26588_2.png) [@pitsianis](https://discourse.julialang.org/u/pitsianis)\
**Post date:** [December 21, 2023, 3:31pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/6 "2023-12-21T15:31:00Z")

</div>

> [@RainerHeintzmann](#):
>
> But then I tried downloading `llama-2-7b-chat.ggmlv3.q4_K_S.bin` but I am not sure where to find it.

> **[llama-2-7b-chat.ggmlv3.q4\_K\_S.bin · TheBloke/Llama-2-7B-Chat-GGML at main](https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGML/blob/main/llama-2-7b-chat.ggmlv3.q4_K_S.bin)**
>
> We’re on a journey to advance and democratize artificial intelligence through open source and open science.

---

<div class="post-metadata">

**Author:** ![RainerHeintzmann](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rainerheintzmann/32/19726_2.png) [@RainerHeintzmann](https://discourse.julialang.org/u/RainerHeintzmann)\
**Post date:** [December 21, 2023, 4:40pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/7 "2023-12-21T16:40:40Z")

</div>

Thanks a lot! This worked. Yet, when running it, I am now surprised that there seems to be little or no use to the GPU? This model not use any GPU ram, and the speed was also not that great (1,7 tokens per second).  
Is that `Llama2.jl` toolbox a CPU-only Llama implementation?

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [December 28, 2023, 7:40pm UTC](https://discourse.julialang.org/t/16-bit-float-on-transformers-jl/107867/8 "2023-12-28T19:40:13Z")

</div>

No,

you can call llama from Transformers.jl, this is what I did. See the [notebook](https://github.com/chengchingwen/Transformers.jl/blob/master/example/Llama2_example.ipynb)
