# Community Interest Check: LLMs from Scratch in Pure Julia

**URL:** <https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796>\
**Category:** Offtopic\
**Tags:** package\
**Created:** [October 26, 2024, 7:10pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796 "2024-10-26T19:10:38Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![mantzaris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mantzaris/32/3852_2.png) [@mantzaris](https://discourse.julialang.org/u/mantzaris)\
**Post date:** [October 26, 2024, 7:10pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/1 "2024-10-26T19:10:38Z")

</div>

Is there active interest in developing LLMs from the ground up in pure Julia or is there ongoing work on this that I have not come across?

## Current

I’ve been reviewing existing Julia LLM projects:

- Transformers.jl (~1K commits, last updated 4 months ago) seems to focus on providing interfaces to pre-trained models like BERT/Llama
- TransformerBlocks.jl (\<100 commits, not recently maintained) offers some building blocks but isn’t actively developed
- There appears to be a gap in pure Julia implementations for training LLMs from scratch

## Project Interest

1. Developing pure Julia implementations for LLM training (not just inference)
2. Building the necessary distributed training infrastructure
3. Creating efficient Julia-native attention mechanisms and optimizers
4. Leveraging Julia’s strengths

## Looking For

- Others interested in building LLMs from scratch in Julia
- Insights from those who’ve attempted similar projects
- Discussion about technical challenges and Julia-specific questions

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [October 26, 2024, 7:31pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/2 "2024-10-26T19:31:33Z")

</div>

The existence of [GitHub - cafaxo/Llama2.jl: Julia package for inference and training of Llama-style language models](https://github.com/cafaxo/Llama2.jl) should answer some of your questions.

---

<div class="post-metadata">

**Author:** ![TheCedarPrince](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thecedarprince/32/17323_2.png) [@TheCedarPrince](https://discourse.julialang.org/u/TheCedarPrince)\
**Post date:** [October 26, 2024, 8:21pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/3 "2024-10-26T20:21:35Z")

</div>

I think you might be interested in this talk by @jpsamaroo and @dhairyagandhi96:

[![](https://global.discourse-cdn.com/julialang/original/3X/b/e/be7886a6326a2412b8825ff7bb52103899d94139.jpeg "Train a Llama(2) in Julia! | Gandhi, P Samaroo | JuliaCon 2024") ](https://www.youtube.com/watch?v=klH9mbNu6GQ)

If I am not crazy, I believe they explored just this question you are asking. Otherwise, CC @svilupp and @cpfiffer too!

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [October 26, 2024, 9:00pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/4 "2024-10-26T21:00:10Z")

</div>

It’s probably not better to go _pure_ Julia (rather than use state-of-the-art code and algorithms), unless as a learning exercise. But if you do, consider “1-bit networks” (from 2023 and from this week):

[https://arxiv.org/pdf/2410.16144](https://arxiv.org/pdf/2410.16144)

It’s very likely if you redo some software, you reimplement an outdated way. E.g. transformers are likely going away in current form.

> **[BitNet: Scaling 1-bit Transformers for Large Language Models - Microsoft...](https://www.microsoft.com/en-us/research/publication/bitnet-scaling-1-bit-transformers-for-large-language-models/)**
>
> The increasing size of large language models has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. In this work, we introduce BitNet, a scalable and stable 1-bit Transformer architecture...

We’ve likely gone to the end of the line with quantization with such 1-, 2-bit networks, and it helps keep the size down. To stay competitive with training you need thousands of GPUs, and software that can target so many, so it seems out of the question to use pure Julia. But maybe you can go half way there, leave out some parts like distributing to many GPUs, use DeepSpeed or something for that.

Training from scratch is still very costly, so no need to, since you can finetune a model for Julia use. But then you need to choose the best model to start from and formats/quantization as in llama.cpp or this new bitnet.cpp from Microsoft. See on the former (and relation to Llama2.jl):

> [@Is Rust (and/or Python) the new high-level API (here ANN/GGUF/LLama example)](https://discourse.julialang.org/t/is-rust-and-or-python-the-new-high-level-api-here-ann-gguf-llama-example/121721):
>
> This state-of-the-art library has Rust and Python API, but ironically no (C++) high-level one despite written in C++: I’m guessing C-like C++ is kinda necessary for bindings to other [languages] [TensorFlow, also written in C++, only had officially stable Python API (while Julia’s API better until unmaintained).] If we want to call this library then we can for sure, using PythonCall.jl, also an option to call Rust’s API. So which would you prefer? You might think why not use: GGUF mod…

KAN networks (they can be drop-in replacement into MLP part of transformers, if I recall) are worth-while to reimplement in Julia:

> [@Has Julia met your need for AI and ML?](https://discourse.julialang.org/t/has-julia-met-your-need-for-ai-and-ml/119818/5):
>
> The reason Python is used, is because it’s a good interface to high-speed code written in other fast languages. Julia can also be a good interface, and that faster language. Training massive AI/ML models relies on libraries like Microsoft’s DeepSpeed (“[Python 69.0%](https://github.com/microsoft/DeepSpeed/search?l=python) [C++ 20.3%](https://github.com/microsoft/DeepSpeed/search?l=c%2B%2B)[Cuda 9.8%](https://github.com/microsoft/DeepSpeed/search?l=cuda)”, Julia could theoretically have replaced all the languages, the numbers are for source code, for runtime about 0% is Python), for distributed training, this is only of interest to big companies that can afford to…

KAN networks are likely not compatible with 1-bit networks, I mean their weights larger, but might still be a good thing, if you get away with fewer. Also I think not intirely contradictory, since you can still have a transformer and other parts with 1-bit weights, where KAN is not replaceing the MLP part. But isn’t the MLP part the largest part of the total?

I think also worthwhile to help with this:

> [@Julia implementation of Liquid Neural Networks - anyone wanna help make one?](https://discourse.julialang.org/t/julia-implementation-of-liquid-neural-networks-anyone-wanna-help-make-one/100953):
>
> Following a thread on #machine-learning on Julia slack, i’m interested in putting together a Julia implementation of [“Liquid Neural Networks”](https://cap.csail.mit.edu/sites/default/files/research-pdfs/Robust%20flight%20navigation%20out%20of%20distribution%20with%20liquid%20neural%20networks.pdf). These are time-adaptive neural networks, inspired by neuroscience models, where each neuron’s activity is a differential equation. The synapses are also plastic and continue to adapt after training. I think this is very interesting, and Julia has a great ecosystem for automatic differentiation with differential equations. Anyone interested in teaming toge…

Best models will likely use new ways of multiplying not yet in software (but you could emulate slowly(?) for compatibility until hardware catches up, or maybe just use Float8, of bflot, I don’t recall, might be compatible with it):  
[https://arxiv.org/html/2410.00907v2#S2](https://arxiv.org/html/2410.00907v2#S2)

> [@In Binary Floating Point, you can multiply without multiplying](https://discourse.julialang.org/t/in-binary-floating-point-you-can-multiply-without-multiplying/121218/21):
>
> This is very intriguing (though I think the benefit of the methods doesn’t translate to 4-bit floating point, it may not matter as much, since likely useless for training neural networks; while useful for inference, though even then 2/3-bit integers taking over). The proposed ℒ-Mul method will lead to a significantly reduced energy consumption for both model training and inference. […] multiplying two 32-bit floating point numbers (fp32) costs four times energy as adding two fp32 numbers, and …

---

<div class="post-metadata">

**Author:** ![svilupp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/svilupp/32/34933_2.png) [@svilupp](https://discourse.julialang.org/u/svilupp)\
**Post date:** [October 26, 2024, 9:07pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/5 "2024-10-26T21:07:37Z")

</div>

fun topic! Is there interest? Absolutely!

Would there be demand for pure Julia implementations? Definitely!

My understanding was that people are stretched so thin on existing projects that we need more people interested and willing to hack!

Personally, I’m crazy about the applications of GenAI and building on top of it rather than training, but I bet that differs for everyone.

If you’re keen to hack deeper than just pure inference, but still easy to start maybe you wanna dip your feet with [Entropix](https://github.com/xjdr-alt/entropix)? Have you played with it? It does a lot of clever stuff with really small models - having that perfect balance of performance, practical, and runnable locally. That could be a fun starter!

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [October 26, 2024, 9:26pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/6 "2024-10-26T21:26:36Z")

</div>

[https://timkellogg.me/blog/2024/10/10/entropix](https://timkellogg.me/blog/2024/10/10/entropix)

> You might be surprised to learn that LLMs know when they’re confused, but [that’s been known for a little while](https://arxiv.org/abs/2410.02707).

Funny how the link there is to a paper from this month, so “known for a little while” in AI research that means what, about 3 weeks?! [Or does the new paper reference older paper/ideas?] Not really to surprising with the rapid changes even if meaning 3 weeks.

Thanks, I didn’t know of Entropix, seems interesting.

> Last I heard, entropix is splitting the repository, one effort going toward huge models and pushing the limits for where this can go. The other is focused on local LLMs, squeezing out every last drop of intelligence.

From same blog:

> **[GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in...](https://arxiv.org/abs/2410.05229)**
>
> Recent advancements in Large Language Models (LLMs) have sparked interest in their formal reasoning capabilities, particularly in mathematics. The GSM8K benchmark is widely used to assess the mathematical reasoning of models on grade-school-level...

---

<div class="post-metadata">

**Author:** ![greatpet](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Post date:** [October 26, 2024, 11:09pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/7 "2024-10-26T23:09:07Z")

</div>

How good / complete are the existing Julia packages for transformer networks in general, not necessarily for LLMs?

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [November 13, 2024, 2:57pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/8 "2024-11-13T14:57:58Z")

</div>

Consider implementing DIFF Transformer from October paper (it seems better across the board), I started translating to Julia:

```julia
function DiffAttn(X, W_q, W_k, W_v, λ)
  Q1, Q2 = split(X * W_q)
  K1, K2 = split(X * W_k)
  V = X * W_v
  # Qi, Ki: [b, n, d]; V: [b, n, 2d]
  s = 1 / sqrt(d) # torch.rsqrt
  A1 = Q1 * K1.transpose(−1, −2) * s
  A2 = Q2 * K2.transpose(−1, −2) * s
  return (softmax(A1) − λ * softmax(A2)) * V
end

I leave in the Python pseudocode as is:

def MultiHead(X, W_q, W_k, W_v, W_o, λ):
  O = GroupNorm([DiffAttn(X, W_qi, W_ki,
  W_vi, λ) for i in range(h)])
  O = O ∗ (1 − λinit)
  return Concat(O) @ W_o

```

Microsoft’s Python code here: [unilm/Diff-Transformer at master · microsoft/unilm · GitHub](https://github.com/microsoft/unilm/tree/master/Diff-Transformer)

It allows for more “post-training quantization”, with the current transformer suffering (way more, see figure 8) at 4-bit. Best quantization I’ve seen is even less than 4-bit, but I think never for post-training, and I think this will also work, no less, probably even better when combined with such methods:

> The results indicate that DIFF Transformer natively mitigates activation outliers in attention scores, providing new opportunities for low-bit FlashAttention [8] implementations.

[https://arxiv.org/pdf/2410.05258](https://arxiv.org/pdf/2410.05258)

> Figure 1: Transformer often over-attends to irrelevant context (i.e., attention noise). DIFF Transformer amplifies attention to answer spans and cancels noise, enhancing the capability of context modeling.

> In this [Microsoft] paper, we introduce Differential Transformer (a.k.a. DIFF Transformer), a foundation architecture for large language models. […] Specifically, we partition the query and key vectors into two groups and compute two separate softmax attention maps. […] The differential attention mechanism eliminates attention noise, encouraging models to focus on critical information. The approach is analogous to noise-canceling headphones and differential amplifiers [19] in electrical engineering, where the difference between two signals cancels out common-mode noise. In the middle of Figure 1, we also present the normalized distribution of attention scores for DIFF Transformer. We observe that DIFF Transformer assigns significantly higher scores to the correct answer and much lower scores to irrelevant context compared to Transformer. […] We conduct extensive experiments on language modeling. We scale up DIFF Transformer in terms of parameter count, training tokens, and context length. The scaling curves indicate that DIFF Transformer requires only about 65% of model size or training tokens needed by Transformer to achieve comparable language modeling performance. Moreover, **DIFF Transformer outperforms Transformer in various downstream tasks.** The long-sequence evaluation also shows that **DIFF Transformer is highly effective in utilizing the increasing context. In addition, the experimental results**  **demonstrate that DIFF Transformer has intriguing advantages for large language models.** For example, the proposed method substantially outperforms Transformer in key information retrieval, **hallucination mitigation** , and in-context learning. DIFF Transformer also reduces outliers in model activations, which provides new opportunities for quantization. **The findings establish DIFF Transformer as an effective and distinctive foundation architecture for large language models**.

> Figure 3: Language modeling loss of scaling up parameter count and training tokens. **DIFF Transformer requires only about 65% of model size or training tokens to match Transformer’s performance.**

> **3.5 In-Context Learning**  
> We evaluate in-context learning from two perspectives, including many-shot classification and robustness of in-context learning. **In-context learning is a fundamental capability of language models, which indicates how well a model can utilize input context.**  
> […]  
> The results show that **DIFF Transformer consistently outperforms Transformer across datasets and varying numbers of demonstration samples. Moreover, the improvement in average accuracy is substantial, ranging from 5.2% to 21.6%.**

> **Robustness of In-Context Learning** Figure 7 compares the robustness of in-context learning between Transformer and DIFF Transformer. […]  
> The results indicate that **our approach is more robust for in-context learning**. In contrast, **Transformer tends to be distracted by order permutations [25], resulting in a huge margin between the best and worst results.**  
> […]  
> Compared with Transformer, **our method mitigates contextual hallucination on summarization and question answering**. The performance improvement possibly stems from **DIFF Transformer’s better focus on essential information needed for the task, instead of irrelevant context**. This aligns with previous observation [16] that one primary reason for **contextual hallucination in Transformer is the misallocation of attention scores**.

> #### F Gradient Flow of DIFF Transformer
> 
> We show that the **gradient flow in differential attention is similar to that of conventional softmax attention. With this property, the same hyperparameters used in Transformer can be applied directly to the corresponding DIFF Transformer without concerns about training instability**

They use “internal version of [21]” meaning SuperScaler (but it seems optional, only for training speed):  
SuperScaler: Supporting Flexible DNN Parallelization via a Unified Abstraction  
[https://arxiv.org/pdf/2301.08984](https://arxiv.org/pdf/2301.08984)

> As a result, SuperScaler can not only generate empirical paral-  
> lelization plans, but also construct new plans that achieve up  
> to 3.5× speedup compared to state-of-the-art solutions like  
> DeepSpeed, Megatron and Alpa  
> [torch.rsqrt — PyTorch 2.5 documentation](https://pytorch.org/docs/stable/generated/torch.rsqrt.html)

Available [RMSNorm — PyTorch 2.5 documentation](https://pytorch.org/docs/stable/generated/torch.nn.modules.normalization.RMSNorm.html) so why do they implement their own RMSNorm? I suppose such is available in Julia, and most of (such foundation) code they rely on.

It references (and uses Group normalization [paper]) and:  
Magneto: A foundation Transformer. In International Conference on Machine Learning, pp. 36077–36092. PMLR, 2023.

> **[wang23u.pdf](https://proceedings.mlr.press/v202/wang23u/wang23u.pdf)**
>
> 497.41 KB

and:

> **[The Devil in Linear Transformer](https://arxiv.org/abs/2210.10340)**
>
> Linear transformers aim to reduce the quadratic space-time complexity of vanilla transformers. However, they usually suffer from degraded performances on various tasks and corpus. In this paper, we examine existing kernel-based linear transformers...

[https://arxiv.org/pdf/2210.06423](https://arxiv.org/pdf/2210.06423)

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [November 13, 2024, 3:29pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/9 "2024-11-13T15:29:08Z")

</div>

Also intriguing paper from days ago, and likely can be combined with the paper I just posted:

> **[BitNet a4.8: 4-bit Activations for 1-bit LLMs](https://arxiv.org/abs/2411.04965)**
>
> Recent research on the 1-bit Large Language Models (LLMs), such as BitNet b1.58, presents a promising direction for reducing the inference cost of LLMs while maintaining their performance. In this work, we introduce BitNet a4.8, enabling 4-bit...

> In this work, we introduce BitNet a4.8, a **hybrid quantization and sparsification strategy** that enables 4-bit activations for 1-bit LLMs. By carefully analyzing the activation distribution of 1-bit LLMs, we selectively apply 4-bit quantization or sparsification based on the distribution patterns of these activations. Specifically, as shown in Figure 1, BitNet a4.8 employs 4-bit activations for the inputs to attention and FFN, while utilizing sparsification with 8 bits for intermediate states. To improve the training efficiency, BitNet a4.8 is trained from 8-bit to 4-bit activations with a two-stage recipe, which requires only a few training tokens to adapt BitNet b1.58 to the low-bit activations at the end of training. **Extensive experiments demonstrate that BitNet a4.8 achieves competitive performance to BitNet b1.58 with the same training cost while being significantly more efficient at inference time**.

Version 4 in May:

> **[GitHub - AIoT-MLSys-Lab/Efficient-LLMs-Survey: \[TMLR 2024\] Efficient Large Language Models: A...](https://github.com/AIoT-MLSys-Lab/Efficient-LLMs-Survey)**
>
> \[TMLR 2024\] Efficient Large Language Models: A Survey

Most recent update there:  
[https://hqjiang.com/minference.html](https://hqjiang.com/minference.html)

> Our proposed technique can be directly applied to existing LLMs without any modifications to the pre-training setup or additional fine-tuning.

> Now, you can process 1M context 10x faster in a single A100 using Long-context LLMs like LLaMA-3-8B-1M, GLM-4-1M, with even better accuracy, try MInference 1.0 right now!

Rather cool:

> **[Pyramid Flow](https://pyramid-flow.github.io/)**
>
> Pyramidal Flow Matching for Efficient Video Generative Modeling

> ### Text-to-Video Generation (1280x768, 10s, 24fps)

Others do longer up to 2 min, but at 24fps (I think way lower)? What’s the max fps?

---

<div class="post-metadata">

**Author:** ![mantzaris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mantzaris/32/3852_2.png) [@mantzaris](https://discourse.julialang.org/u/mantzaris)\
**Post date:** [November 13, 2024, 4:16pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/10 "2024-11-13T16:16:44Z")

</div>

Have there been other key/substantial improvements recently in the other components across the llm stack? Are most of the key changes in the transformer blocks? Is the attention mechanism mostly the same across implementations?

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [November 15, 2024, 10:07am UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/11 "2024-11-15T10:07:55Z")

</div>

If you’re asking me, there so much I can write on this, but I want to be confident in my answers, so I let others answer.

About training, I thought, that training scratch, would be basically an impossible task, and also missing infrastructure code (in Julia if we can not use already available). This seems like a huge deal from August:

> **[GitHub - NousResearch/DisTrO: Distributed Training Over-The-Internet](https://github.com/NousResearch/DisTrO)**
>
> Distributed Training Over-The-Internet

> This is the repository for DisTrO (Distributed Training Over-The-Internet), a family of low latency distributed optimizers that reduce inter-GPU communication requirements by three to four orders of magnitude.

This means people could help out together, I think with their home GPUs, but it’s unclear it lowers GPU requirements though, so we can likely not get 300,000 GPUs or something (or that many people) to help.

Either you fine-tune (doable) or from scratch, but I’m thinking do only those extreme exist? I suppose you do not start totally from scratch all the time, would be best to start from some early checkpoint.

I’ve not looked into DisTrO closely, is it a replacement for Adam, Lion etc. (likely not) or builds on such? Probably at least replaces Deepspeed.

DoReMi is also a very intriguing development, training a smaller model, then larger based on it:  
[https://neurips.cc/virtual/2023/poster/70588](https://neurips.cc/virtual/2023/poster/70588)

> DoReMi improves perplexity across _all_ domains

Not to be confused with DoReMi there (looking it up again, got me confused…), this paper is on robotics:

> **[DoReMi](https://sites.google.com/view/doremi-paper)**
>
> Abstract: In this paper, we propose DoReMi, a novel language model grounding framework that enables immediate Detection and Recovery from Misalignments between plan and execution. Specifically, we leverage LLMs to play a dual role, aiding not only in...

On this:

> [@mantzaris](#):
>
> 1. Creating efficient Julia-native attention mechanisms and optimizers

What do you mean? I’m not sure I know enough about “attention”, it basically means transformers, right and we have them already?

[https://h2o.ai/wiki/self-attention/](https://h2o.ai/wiki/self-attention/)

[https://h2o.ai/wiki/attention-mechanism/](https://h2o.ai/wiki/attention-mechanism/)

Do we have all the variants, for sure unlikely, do you mean like:

> **[FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision](https://news.ycombinator.com/item?id=40938577)**
>
> 287 points —
> 60 comments —
> jhshah —
> 5:06 PM - 11 Jul 2024

> **Better performance with lower precision** : FlashAttention-3 can work with lower precision numbers (FP8) while maintaining accuracy.

Such lower precision requires recent hardware (GPUs), and I don’t think we can compete if we do not use/target such. We are also behind in even more quantized models. We have SafeTensors.jl but that format and code is limited to bfloat16 and FP8 smallest (or so it seems, maybe not inherently and will support smaller?). It uses DLFP8Types.jl so seems implemented in software (then slowly).

Did you mean we need Grouped Query Attention, I’m not sure we might have it already?

> **[Demystifying GQA — Grouped Query Attention for Efficient LLM Pre-training](https://towardsdatascience.com/demystifying-gqa-grouped-query-attention-3fb97b678e4a)**
>
> The variant of multi-head attention powering LLMs like LLaMA-2, Mistral7B, etc.

> **[FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision](https://news.ycombinator.com/item?id=40938577)**
>
> 287 points —
> 60 comments —
> jhshah —
> 5:06 PM - 11 Jul 2024

> > … Tri’s publication history has been leaning toward SSM and Mamba style architectures recently. Unlike Flash Attention which has quadratic time complexity wrt sequence length, these latest algorithms are subquadratic. Thus they do much less computation, instead of just doing it more efficiently a la Flash Attention.
> 
> > Dao and Gu published a really long paper this year which demonstrated (among other things) how Mamba/SSM can be formulated such that it’s amenable to acceleration using the same hardware primitives that Transformers benefit from. …

> …Until the strong exponential hypothesis is (dis-)proven, the quadratic cost is required or you have to give something up. Just the cost of exhaustive search.
> 
> As (dis-)proving SETH will resolve the P vs NP problem, I wouldn’t hold my breath. …

Maybe Mamba, SSM or Jamba or some linear transformer will take over, but it seems to me just sticking with quadratic transformer is a safe bet (or even if not good enough code can later be changed?).

> [@Is there an implementation of the attention mechanism in Flux.jl?](https://discourse.julialang.org/t/is-there-an-implementation-of-the-attention-mechanism-in-flux-jl/26959/6):
>
> Maybe this links are interesting for you:

We have many (most?) optimizers here, e.g. many Adam variants:

> **[API · Optimisers.jl](https://fluxml.ai/Optimisers.jl/stable/api/#Optimisers.AdamW)**
>
> Documentation for Optimisers.jl.

> With version 0.4 the default update rule for AdamW has changed to match the pytorch implementation.

We even have [API · Optimisers.jl](https://fluxml.ai/Optimisers.jl/stable/api/#Optimisers.Lion) (once the best I though, no longer?), but clicking on some docs such as it shows strange (placeholder?) text:

> In addition to the main course, you may wish to order some of these condiments:

I thought we redundantly have them at SciML, but its docs actually link over to flux.ml:

> **[Optimisers.jl · Optimization.jl](https://docs.sciml.ai/Optimization/stable/optimization_packages/optimisers/)**
>
> Documentation for Optimization.jl.

> - [`Optimisers.RMSProp`](https://fluxml.ai/Optimisers.jl/dev/api/#Optimisers.RMSProp): **RMSProp optimizer**

> - [`Optimisers.Adam`](https://fluxml.ai/Optimisers.jl/dev/api/#Optimisers.Adam): **Adam optimizer**

---

<div class="post-metadata">

**Author:** ![mantzaris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mantzaris/32/3852_2.png) [@mantzaris](https://discourse.julialang.org/u/mantzaris)\
**Post date:** [November 15, 2024, 3:16pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/12 "2024-11-15T15:16:38Z")

</div>

In terms of ‘competing’, I am thinking of getting into the race track from behind and with a good stack design/architecture be able to have iterative contributions to improve upon it. Almost like an end-to-end scaffolding. As long as the modularity and separation of concerns of the stack components is well thought through it should allow others to plug in their improved components.

At the moment what are the components which don’t exist yet? I would be great if there was a ‘roadmap’ / list of ‘milestones’ and then we could see what needs to be done and get updates. Most importantly though I am thinking that the ‘design’ for seeing how the components can link together is vital for the operation of the stack.

What is the list of ‘components’ we need and how will they interface together?

The distributed learning module is exciting.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [November 15, 2024, 5:28pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/13 "2024-11-15T17:28:14Z")

</div>

Just looking at the diversity and breadth of libraries and packages in the equivalent Python ecosystem, I don’t think there will be a single design or architecture to rule them all here. It might be better to start with a concrete example and use case that includes all the pieces you want to see in the stack, then identify and fill in gaps from there. Llama2.jl is a great example of this: it tried to fill the niche of a llama.cpp equivalent in the Julia ecosystem.

---

<div class="post-metadata">

**Author:** ![mantzaris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mantzaris/32/3852_2.png) [@mantzaris](https://discourse.julialang.org/u/mantzaris)\
**Post date:** [November 15, 2024, 5:52pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/14 "2024-11-15T17:52:46Z")

</div>

You are very right that a single approach is not very ‘organic’ and that only ecosystem will prevail in producing a great wealth of libraries. I totally agree with a ‘concrete example and use case that includes all the pieces’.  
From that, the question is, what is a list of the ‘repos’/‘pieces’ that would be good to have to cover a full stack? A list would be great, or a diagram, with a few high level bullet points, knowing that it is subject to change etc. Then the interface would be good so that others that want to substitute a component can easily do so. Eg. if the new ‘distributed’ approach is sought after to do the training, a good abstraction within the stack and clear defined interface should allow a substitution to try that out. A big monolithic approach may hamper the efforts of specialized development but of course it should not be too ‘shallow’ and scattered. Do you have an idea for a basic outline of a stack we could work on?

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [November 15, 2024, 7:22pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/15 "2024-11-15T19:22:48Z")

</div>

This difficulty of analysis and requirements gathering is precisely why I’d recommend focusing on one (maybe two) concrete examples first. For example, maybe you want to pre-train a new open source LLM on a specific set of datasets, evaluate it on a common benchmark suite and quantize it to run inference. All of those steps currently have gaps in the ecosystem, but the only way to figure out what exactly the gaps are is to start writing code. It’s less work to extract a common stack out from those afterwards than it is to try to come up with an ideal one before any code exists to use it.

---

<div class="post-metadata">

**Author:** ![TheCedarPrince](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thecedarprince/32/17323_2.png) [@TheCedarPrince](https://discourse.julialang.org/u/TheCedarPrince)\
**Post date:** [November 21, 2024, 11:10pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/16 "2024-11-21T23:10:13Z")

</div>

Fun news, seems like there is some action on building some LLM technology with Julia:

> **[Alex Tantos (@alextantos.bsky.social)](https://bsky.app/profile/alextantos.bsky.social/post/3lbgv27iafc2y)**
>
> Inspired by 
> @rasbt
> Build a LLM (from Scratch), I’ve written a Medium article on implementing a simple self-attention mechanism using Python and Julia. It includes practical, theory-driven details to deepen understanding. #LLMs #Python #JuliaLang...

Wanted to share!

---

<div class="post-metadata">

**Author:** ![avikpal](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/avikpal/32/6550_2.png) [@avikpal](https://discourse.julialang.org/u/avikpal)\
**Post date:** [November 22, 2024, 6:29pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/17 "2024-11-22T18:29:03Z")

</div>

[Reactant](https://enzymead.github.io/Reactant.jl/dev/) would be the way to go if you want good performance for these workloads. If you want a starter code, we have some WIP versions scattered across PRs atm

1. [feat: nanoGPT implementation using Reactant by avik-pal · Pull Request #1062 · LuxDL/Lux.jl · GitHub](https://github.com/LuxDL/Lux.jl/pull/1062)
2. [feat: add a Llama2 model by avik-pal · Pull Request #88 · EnzymeAD/Reactant.jl · GitHub](https://github.com/EnzymeAD/Reactant.jl/pull/88)

Even the quantized ops needed for inference exist in the StableHLO land but we haven’t hooked them up yet on the Julia side but it is definitely doable.

---

<div class="post-metadata">

**Author:** ![Alex\_Tantos](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alex_tantos/32/10636_2.png) [@Alex\_Tantos](https://discourse.julialang.org/u/Alex_Tantos)\
**Post date:** [November 23, 2024, 2:34pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/18 "2024-11-23T14:34:42Z")

</div>

Thank you so much for referring to my recent [Medium](https://medium.com/@atantos) article here! 🙏 I’m excited to share that this is part of an ongoing series where I translate the Python code from Sebastian Raschka’s [Build a Large Language (from Scratch)](https://www.amazon.com/Build-Large-Language-Model-Scratch/dp/1633437167) into Julia.

Apart from the recent article on the self-attention mechanism (section 3.3.1 of the book), I’ve also written 7 more Medium articles where the Python and Julia code from the second chapter is translated and explained. You can find the full code in the [GitHub repo](https://github.com/atantos/Build_LLMs_from_Scratch_in_Julia), which I’ve just set up to accompany the article series.

I’d love to connect on [Bluesky](https://bsky.app/profile/alextantos.bsky.social) and [X](https://x.com/alextantos1) to exchange ideas and collaborate further. If you find this work useful, please feel free to star/share the [GitHub repo](https://github.com/atantos/Build_LLMs_from_Scratch_in_Julia) to reach more users who might want to join this effort of translating Python to Julia for building an LLM from scratch.  
P.S.: I always share both the member link and the (free) friend link for the [Medium](https://medium.com/@atantos) articles, in case you’re not a Medium member.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [November 26, 2024, 12:02am UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/19 "2024-11-26T00:02:03Z")

</div>

> **[AWS Trainium](https://aws.amazon.com/ai/machine-learning/trainium/)**
>
> Learn about AWS Trainium an ML accelerator presented by AWS.

> To deliver high performance while meeting accuracy goals, Trainium is optimized for FP32, TF32, BF16, FP16, UINT8, and the new configurable FP8 (cFP8) data type.

Despite seeing AWS has their own chip, I think they’re behind the curve of new research. Nvidia has FP8 and INT4 and their Blackwell’s FP4 Tensor Core (I think it also has FP6), but it’s also outdated.

Something to emulate in Julia, or get access to such software:

### 1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit

[https://arxiv.org/pdf/2408.14267](https://arxiv.org/pdf/2408.14267)

> To explore the ultimate limit of FQT (the lowest achievable precision), we make a first attempt to 1-bit FQT. We provide a theoretical analysis of FQT based on Adam and SGD, revealing that the gradient variance influences the convergence of FQT. Building on these theoretical results, we introduce an Activation Gradient Pruning (AGP) strategy. [..] Additionally, we propose Sample Channel joint Quantization (SCQ), which utilizes different quantization strategies in the computation of weight gradients and activation gradients to ensure that the method is friendly to low-bitwidth hardware. Finally, we present a framework to deploy our algorithm. For fine-tuning VGGNet-16 and ResNet-18 on multiple datasets, our algorithm achieves an average accuracy improvement of approximately 6%, compared to per-sample quantization. Moreover, our training speedup can reach a maximum of 5.13× compared to full precision training. Ours code is available at [GitHub - Gaochang-bjtu/1-bit-FQT](https://github.com/Gaochang-bjtu/1-bit-FQT)

> As the training numerical precision continues to decrease, a natural question arises:
> 
> > **What is the ultimate limit of FQT (i.e., the minimum achievable bitwidth)?**

> Answering this question not only advances our understanding of FQT but also provides a crucial direction for future hardware design strategies. Ideally, if we can push the bitwidth down to 1-bit, the training can be implemented with binary operations, such as XNOR and bitcounting operations Courbariaux et al. [2016], and hardware design might be greatly simplified. Binary computation is already shown possible for _inference_ acceleration, such as XNOR-Net Rastegari et al. [2016], but 1-bit _training_ remains unexplored.

> Reducing the bitwidth for FQT is challenging because of (1) the lack of theoretical understanding, especially how gradient quantization affects the convergence; (2) the large quantization error of gradients, which causes a sharp performance drop or even divergence when reducing gradient bitwidth lower than 4-bit (Fig. 1).

> Due to these challenges, current research frontier is still 4-bit FQT. In this work, we make a first attempt towards achieving 1-bit FQT.

> Specifically, our analysis reveals that Adam is more suitable for FQT than SGD in the low-bitwidth regime, due to their different sensitivity to gradient  
> variance.  
> Inspired by the above theory, **we propose a hardware-friendly algorithm for 1-bit FQT.** [..]  
> We examine the **potential of 1-bit FQT on transfer learning tasks in both vision and NLP domain**. [..]  
> On all the datasets, our 1-bit FQT algorithm can successfully _converge_ and demonstrate significantly superior performance compared to directly applying the previous FQT method to the task. The average accuracy drop on visual classification datasets is approximately 5%, compared to training the binary model with full-precision gradients. Notably, the average accuracy loss is  
> negligible (less than 1%) on Flowers Nilsback and Zisserman [2008] dataset and Pets Parkhi et al. [2012] dataset, indicating that 1-bit FQT might indeed be useful in some cases. We implement our algorithm on Hygon and Raspberry Pi devices as a PyTorch-based library binop. Accelerated on-device training can be achieved with simple layer substitution, e.g., replace torch.nn.Conv2d with binop.Conv2d. In practice, our method can achieve up to 5.13× speedup, compared to FP32 PyTorch. These results indicate that in some specific tasks, FQT precision can be pushed to the ultimate 1-bit  
> [..]  
> **3.1 Quantized Training**  
> Here, we describe Quantization-Aware Training (QAT) and Fully Quantized Training (FQT). QAT is employed to accelerate _inference_, while FQT is designed to accelerate both inference and _training_.  
> [..]  
> **8-bit PSQ vs. Ours.** To demonstrate the advantages of our method over other high-bit-width FQT methods, we compare our approach with 8-bit PSQ in terms of both speedup and classification performance ( **there is no 4-bit format among the standard data types** ).

[https://github.com/mobiusml/hqq](https://github.com/mobiusml/hqq)

> This repository contains the official implementation of Half-Quadratic Quantization ( **HQQ** ) presented in our articles:
> 
> - HQQ: [https://mobiusml.github.io/hqq\_blog/](https://mobiusml.github.io/hqq_blog/)
> - HQQ+: [https://mobiusml.github.io/1bit\_blog/](https://mobiusml.github.io/1bit_blog/)
> 
> ### What is HQQ?
> 
> **HQQ** is a fast and accurate model quantizer that skips the need for calibration data. Quantize the largest models, without calibration data, in just a few minutes at most 🚀.
> 
> FAQ  
> **Why should I use HQQ instead of other quantization methods?**
> 
> - HQQ is very fast to quantize models.
> - It supports 8,4, **3,2,1 bits**.
> - You can use it on any model (LLMs, Vision, etc.).  
> [..]
> 
> **What is the quality of the quantized models?**  
> We have detailed benchmarks on both language and vision models. Please refer to our blog posts: [HQQ](https://mobiusml.github.io/hqq_blog/), [HQQ+](https://mobiusml.github.io/1bit_blog/).
> 
> **What is the speed of the quantized models?**  
> 4-bit models with `axis=1` can use optimized inference fused kernels like torchao’s int4\_gemm. This is the same kernel used in [gpt-fast](https://github.com/pytorch-labs/gpt-fast) and based on our benchmarks, it’s the fastest kernel available right now. We also support the [Marlin](https://github.com/IST-DASLab/marlin/tree/master/marlin) kernel. Moreover, we focus on making hqq fully compatible with `torch.compile` which speeds-up both training and inference.

In case 1-bit (up to 3-bit) isn’t optimal just yet, then standard floating point is neither or FP4, but:

### Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs

[https://arxiv.org/pdf/2405.03103](https://arxiv.org/pdf/2405.03103)

> … Yet recently, alternative formats such as Normal Float (NF4) have increased model accuracy at the cost of increased chip area. In this work, we first conduct a large-scale analysis of **strong text** LLM weights and activations across 30 networks and conclude that most distributions follow a Student’s t-distribution. **We then derive a new theoretically optimal format, Student Float (SF4)**, that improves over NF4 across modern LLMs, **for example increasing the average accuracy on LLaMA2-7B by 0.76% across tasks**

See also:

> **[1-bit Adam: Up to 5x less communication volume and up to 3.4x faster training](https://www.deepspeed.ai/tutorials/onebit-adam/)**
>
> Note: On 03/07/2022 we released 0/1 Adam, which is a new communication-efficient Adam optimizer partially following the 1-bit Adam’s design. Compared to the 1-bit Adam described below, 0/1 Adam provides better communication efficiency and the same...

Note also from earlier this year:  
[https://arxiv.org/pdf/2402.17764](https://arxiv.org/pdf/2402.17764)

[https://github.com/microsoft/BitNet](https://github.com/microsoft/BitNet)  
see there for a string of related papers up to this month:

> **[BitNet a4.8: 4-bit Activations for 1-bit LLMs](https://arxiv.org/abs/2411.04965)**
>
> Recent research on the 1-bit Large Language Models (LLMs), such as BitNet b1.58, presents a promising direction for reducing the inference cost of LLMs while maintaining their performance. In this work, we introduce BitNet a4.8, enabling 4-bit...

[https://medium.com/@zaiinn440/llama-bitnet-training-a-1-58-bit-llm-3831e517430a](https://medium.com/@zaiinn440/llama-bitnet-training-a-1-58-bit-llm-3831e517430a)

> **[Why 1-Bit Transformers Will Change the World | HackerNoon](https://hackernoon.com/why-1-bit-transformers-will-change-the-world)**
>
> Why are 1-bit transformers such a big deal? Read this article to find out!

> They can run on desktop GPUs without any expensive hardware requirements.
> 
> > **Where 175 GB was required, now only 1.2 GB is required!**

---

<div class="post-metadata">

**Author:** ![mantzaris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mantzaris/32/3852_2.png) [@mantzaris](https://discourse.julialang.org/u/mantzaris)\
**Post date:** [November 26, 2024, 3:13am UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/21 "2024-11-26T03:13:55Z")

</div>

> [@mantzaris](#):
>
> @Palli

That is incredible research. I was unaware of that progress. Thanks for sharing.  
It should be possible to allow the flexibility to work with the 1bit model and the typical FP32/16? Will that be possible with the Flux setup currently? This approach would definitely make it possible for training to be done without serious funding or corporate support that might be necessary for anything beyond a gpt2 version, produced just for demonstrative purposes.

[Next page](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796.md?page=2)
