# Community Interest Check: LLMs from Scratch in Pure Julia

**URL:** <https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796>\
**Category:** Offtopic\
**Tags:** package\
**Created:** [October 26, 2024, 7:10pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796 "2024-10-26T19:10:38Z")\
**Posts on this page:** 1\
**Showing post:** 11

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [November 15, 2024, 10:07am UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/11 "2024-11-15T10:07:55Z")

</div>

If you’re asking me, there so much I can write on this, but I want to be confident in my answers, so I let others answer.

About training, I thought, that training scratch, would be basically an impossible task, and also missing infrastructure code (in Julia if we can not use already available). This seems like a huge deal from August:

> **[GitHub - NousResearch/DisTrO: Distributed Training Over-The-Internet](https://github.com/NousResearch/DisTrO)**
>
> Distributed Training Over-The-Internet

> This is the repository for DisTrO (Distributed Training Over-The-Internet), a family of low latency distributed optimizers that reduce inter-GPU communication requirements by three to four orders of magnitude.

This means people could help out together, I think with their home GPUs, but it’s unclear it lowers GPU requirements though, so we can likely not get 300,000 GPUs or something (or that many people) to help.

Either you fine-tune (doable) or from scratch, but I’m thinking do only those extreme exist? I suppose you do not start totally from scratch all the time, would be best to start from some early checkpoint.

I’ve not looked into DisTrO closely, is it a replacement for Adam, Lion etc. (likely not) or builds on such? Probably at least replaces Deepspeed.

DoReMi is also a very intriguing development, training a smaller model, then larger based on it:  
[https://neurips.cc/virtual/2023/poster/70588](https://neurips.cc/virtual/2023/poster/70588)

> DoReMi improves perplexity across _all_ domains

Not to be confused with DoReMi there (looking it up again, got me confused…), this paper is on robotics:

> **[DoReMi](https://sites.google.com/view/doremi-paper)**
>
> Abstract: In this paper, we propose DoReMi, a novel language model grounding framework that enables immediate Detection and Recovery from Misalignments between plan and execution. Specifically, we leverage LLMs to play a dual role, aiding not only in...

On this:

> [@mantzaris](#):
>
> 1. Creating efficient Julia-native attention mechanisms and optimizers

What do you mean? I’m not sure I know enough about “attention”, it basically means transformers, right and we have them already?

[https://h2o.ai/wiki/self-attention/](https://h2o.ai/wiki/self-attention/)

[https://h2o.ai/wiki/attention-mechanism/](https://h2o.ai/wiki/attention-mechanism/)

Do we have all the variants, for sure unlikely, do you mean like:

> **[FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision](https://news.ycombinator.com/item?id=40938577)**
>
> 287 points —
> 60 comments —
> jhshah —
> 5:06 PM - 11 Jul 2024

> **Better performance with lower precision** : FlashAttention-3 can work with lower precision numbers (FP8) while maintaining accuracy.

Such lower precision requires recent hardware (GPUs), and I don’t think we can compete if we do not use/target such. We are also behind in even more quantized models. We have SafeTensors.jl but that format and code is limited to bfloat16 and FP8 smallest (or so it seems, maybe not inherently and will support smaller?). It uses DLFP8Types.jl so seems implemented in software (then slowly).

Did you mean we need Grouped Query Attention, I’m not sure we might have it already?

> **[Demystifying GQA — Grouped Query Attention for Efficient LLM Pre-training](https://towardsdatascience.com/demystifying-gqa-grouped-query-attention-3fb97b678e4a)**
>
> The variant of multi-head attention powering LLMs like LLaMA-2, Mistral7B, etc.

> **[FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision](https://news.ycombinator.com/item?id=40938577)**
>
> 287 points —
> 60 comments —
> jhshah —
> 5:06 PM - 11 Jul 2024

> > … Tri’s publication history has been leaning toward SSM and Mamba style architectures recently. Unlike Flash Attention which has quadratic time complexity wrt sequence length, these latest algorithms are subquadratic. Thus they do much less computation, instead of just doing it more efficiently a la Flash Attention.
> 
> > Dao and Gu published a really long paper this year which demonstrated (among other things) how Mamba/SSM can be formulated such that it’s amenable to acceleration using the same hardware primitives that Transformers benefit from. …

> …Until the strong exponential hypothesis is (dis-)proven, the quadratic cost is required or you have to give something up. Just the cost of exhaustive search.
> 
> As (dis-)proving SETH will resolve the P vs NP problem, I wouldn’t hold my breath. …

Maybe Mamba, SSM or Jamba or some linear transformer will take over, but it seems to me just sticking with quadratic transformer is a safe bet (or even if not good enough code can later be changed?).

> [@Is there an implementation of the attention mechanism in Flux.jl?](https://discourse.julialang.org/t/is-there-an-implementation-of-the-attention-mechanism-in-flux-jl/26959/6):
>
> Maybe this links are interesting for you:

We have many (most?) optimizers here, e.g. many Adam variants:

> **[API · Optimisers.jl](https://fluxml.ai/Optimisers.jl/stable/api/#Optimisers.AdamW)**
>
> Documentation for Optimisers.jl.

> With version 0.4 the default update rule for AdamW has changed to match the pytorch implementation.

We even have [API · Optimisers.jl](https://fluxml.ai/Optimisers.jl/stable/api/#Optimisers.Lion) (once the best I though, no longer?), but clicking on some docs such as it shows strange (placeholder?) text:

> In addition to the main course, you may wish to order some of these condiments:

I thought we redundantly have them at SciML, but its docs actually link over to flux.ml:

> **[Optimisers.jl · Optimization.jl](https://docs.sciml.ai/Optimization/stable/optimization_packages/optimisers/)**
>
> Documentation for Optimization.jl.

> - [`Optimisers.RMSProp`](https://fluxml.ai/Optimisers.jl/dev/api/#Optimisers.RMSProp): **RMSProp optimizer**

> - [`Optimisers.Adam`](https://fluxml.ai/Optimisers.jl/dev/api/#Optimisers.Adam): **Adam optimizer**

---

_[View the full topic](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796)._
