# LLaMA in Julia?

**URL:** <https://discourse.julialang.org/t/llama-in-julia/95979>\
**Category:** Offtopic\
**Created:** [March 13, 2023, 1:02am UTC](https://discourse.julialang.org/t/llama-in-julia/95979 "2023-03-13T01:02:28Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![robsmith11](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robsmith11/32/29641_2.png) [@robsmith11](https://discourse.julialang.org/u/robsmith11)\
**Post date:** [March 13, 2023, 1:02am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/1 "2023-03-13T01:02:28Z")

</div>

LLaMA.cpp [1] has been getting a lot of attention on Hacker News [2] for its ability to run a Large Language Model (LLM) on any recent CPU with modest memory requirements. I’ve been meaning to get a better understanding of LLMs, so porting LLaMA over to Julia and being able to run it on my laptop seems like a good way to do that.

Has anyone else already started a similar project or have any thoughts? I briefly went through the C++ code and it looks fairly straightforward and a good fit for Julia from what I can tell.

[1] [GitHub - ggerganov/llama.cpp: Port of Facebook's LLaMA model in C/C++](https://github.com/ggerganov/llama.cpp)  
[2] [Using LLaMA with M1 Mac and Python 3.11 | Hacker News](https://news.ycombinator.com/item?id=35122689)

---

<div class="post-metadata">

**Author:** ![woclass](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/woclass/32/212699_2.png) [@woclass](https://discourse.julialang.org/u/woclass)\
**Post date:** [March 13, 2023, 6:18am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/2 "2023-03-13T06:18:47Z")

</div>

> - Plain C/C++ implementation **without** dependencies

`ggml.c`: 10502 lines (8554 loc) · 316 KB  
GGML Tensor Library

> **[GitHub - ggerganov/ggml: Tensor library for machine learning](https://github.com/ggerganov/ggml)**
>
> Tensor library for machine learning. Contribute to ggerganov/ggml development by creating an account on GitHub.

* * *

Perhaps picoGPT would be simpler?

> `picoGPT` is an unnecessarily tiny and minimal implementation of [GPT-2](https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf) in plain [NumPy](https://numpy.org/). The entire forward pass code is [40 lines of code](https://github.com/jaymody/picoGPT/blob/main/gpt2_pico.py#L3-L41).

> **[GitHub - jaymody/picoGPT: An unnecessarily tiny implementation of GPT-2 in...](https://github.com/jaymody/picoGPT)**
>
> An unnecessarily tiny implementation of GPT-2 in NumPy. - GitHub - jaymody/picoGPT: An unnecessarily tiny implementation of GPT-2 in NumPy.

---

<div class="post-metadata">

**Author:** ![Acey](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/acey/32/8438_2.png) [@Acey](https://discourse.julialang.org/u/Acey)\
**Post date:** [July 22, 2023, 8:55pm UTC](https://discourse.julialang.org/t/llama-in-julia/95979/3 "2023-07-22T20:55:39Z")

</div>

any updates on this? I am also interested in using Julia for LLMs.

---

<div class="post-metadata">

**Author:** ![robsmith11](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robsmith11/32/29641_2.png) [@robsmith11](https://discourse.julialang.org/u/robsmith11)\
**Post date:** [July 26, 2023, 2:44am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/4 "2023-07-26T02:44:15Z")

</div>

I haven’t had time to do anything, but looks like someone is working on a project in julia:

> **[GitHub - jdiaz97/llama2.jl: llama2 in Julia](https://github.com/jdiaz97/llama2.jl)**
>
> llama2 in Julia. Contribute to jdiaz97/llama2.jl development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![Acey](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/acey/32/8438_2.png) [@Acey](https://discourse.julialang.org/u/Acey)\
**Post date:** [July 26, 2023, 12:21pm UTC](https://discourse.julialang.org/t/llama-in-julia/95979/5 "2023-07-26T12:21:21Z")

</div>

thank you - I have seen it!

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [July 26, 2023, 12:23pm UTC](https://discourse.julialang.org/t/llama-in-julia/95979/6 "2023-07-26T12:23:58Z")

</div>

I believe this implementation is currently in better shape: [https://gist.github.com/cafaxo/83b795f11c645e217688449dd91147a3](https://gist.github.com/cafaxo/83b795f11c645e217688449dd91147a3)

---

<div class="post-metadata">

**Author:** ![lostella](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lostella/32/356_2.png) [@lostella](https://discourse.julialang.org/u/lostella)\
**Post date:** [July 29, 2023, 5:42am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/7 "2023-07-29T05:42:27Z")

</div>

This has been moved to [GitHub - cafaxo/Llama2.jl](https://github.com/cafaxo/Llama2.jl)

---

<div class="post-metadata">

**Author:** ![sivark](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sivark/32/960_2.png) [@sivark](https://discourse.julialang.org/u/sivark)\
**Post date:** [August 5, 2023, 6:22am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/8 "2023-08-05T06:22:25Z")

</div>

Some recent stuff from Jiahao Chen: [GitHub - rai-llc/LanguageModels.jl: Load nanoGPT-style transformers in Julia. Code ported from @karpathy's llama2.c](https://github.com/rai-llc/LanguageModels.jl)

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [August 5, 2023, 10:54am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/9 "2023-08-05T10:54:24Z")

</div>

This is great endeavour. Please, do not take next comments as criticism, but as a suggestions of what to do next.

but on the end, the question will be speed. How does the implementation compares to that in `llama.c`?  
Also, I have recently the pleasure the use `Transformers.jl` and was impressed how the package is. It implements the python counterpart with Julia flexibility, which is just nice. Unfortunately, it is about half the performance (speed) of Julia. So the next question would be, how the tiny implementation of llama compares to the implementation with `Transformers.jl`, which might not exist yet. I think this is important questions, because these tiny libraries are incredible for showing versatility, Transformers should be go-to package for LLMs and we should learn tricks from these small packages to improve it.

Private confession: for some experiments, I had to use the python counterpart, because Falcon class of models use tokenizer which is not supported in Julia.

---

<div class="post-metadata">

**Author:** ![cafaxo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cafaxo/32/5234_2.png) [@cafaxo](https://discourse.julialang.org/u/cafaxo)\
**Post date:** [August 6, 2023, 11:43am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/10 "2023-08-06T11:43:20Z")

</div>

I am the author of Llama2.jl. It currently runs the Llama2 7B model (q4\_K\_S GGML quantization) at 9 tokens/second (it slows down to about 7 tokens/second as it approaches sequence length 512).  
llama.cpp runs the same model at about 14 tokens/second. (all on an M1 Air)  
As far as I know, llama2.c does not support loading GGML weights at all yet.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [August 7, 2023, 5:18am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/11 "2023-08-07T05:18:17Z")

</div>

Why there is about half speed comparing to the C version? What is their secret sauce they use?

I wanted to try to load LLama2 to transformers.jl, but I got the reply `GatedRepo`, because I have not accepted the license.

---

<div class="post-metadata">

**Author:** ![ImreSamu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/imresamu/32/20677_2.png) [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Post date:** [August 7, 2023, 7:32am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/12 "2023-08-07T07:32:45Z")

</div>

> [@Tomas\_Pevny](#):
>
> Why there is about half speed comparing to the C version? What is their secret sauce they use?

`-Ofast -march=native .... `

- [GitHub - karpathy/llama2.c: Inference Llama 2 in one file of pure C](https://github.com/karpathy/llama2.c#performance)

---

<div class="post-metadata">

**Author:** ![cafaxo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cafaxo/32/5234_2.png) [@cafaxo](https://discourse.julialang.org/u/cafaxo)\
**Post date:** [August 7, 2023, 8:05am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/13 "2023-08-07T08:05:38Z")

</div>

I was comparing speed against [llama.cpp](https://github.com/ggerganov/llama.cpp), not llama2.c.  
I do not yet know what causes the difference.

---

<div class="post-metadata">

**Author:** ![ImreSamu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/imresamu/32/20677_2.png) [@ImreSamu](https://discourse.julialang.org/u/ImreSamu)\
**Post date:** [August 7, 2023, 8:43am UTC](https://discourse.julialang.org/t/llama-in-julia/95979/14 "2023-08-07T08:43:02Z")

</div>

> [@cafaxo](#):
>
> I was comparing speed against [llama.cpp](https://github.com/ggerganov/llama.cpp), not llama2.c.  
> I do not yet know what causes the difference

cflags ?

```julia
    OPT = -Ofast 

```

```julia
	CFLAGS += -march=native -mtune=native
	CXXFLAGS += -march=native -mtune=native

```

> <https://github.com/ggerganov/llama.cpp/blob/3d9a55181603e85a26378a850a14068034e5002d/Makefile#L133-L134>
