# Community Interest Check: LLMs from Scratch in Pure Julia

**URL:** <https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796>\
**Category:** Offtopic\
**Tags:** package\
**Created:** [October 26, 2024, 7:10pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796 "2024-10-26T19:10:38Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [October 26, 2024, 9:00pm UTC](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796/4 "2024-10-26T21:00:10Z")

</div>

It’s probably not better to go _pure_ Julia (rather than use state-of-the-art code and algorithms), unless as a learning exercise. But if you do, consider “1-bit networks” (from 2023 and from this week):

[https://arxiv.org/pdf/2410.16144](https://arxiv.org/pdf/2410.16144)

It’s very likely if you redo some software, you reimplement an outdated way. E.g. transformers are likely going away in current form.

> **[BitNet: Scaling 1-bit Transformers for Large Language Models - Microsoft...](https://www.microsoft.com/en-us/research/publication/bitnet-scaling-1-bit-transformers-for-large-language-models/)**
>
> The increasing size of large language models has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. In this work, we introduce BitNet, a scalable and stable 1-bit Transformer architecture...

We’ve likely gone to the end of the line with quantization with such 1-, 2-bit networks, and it helps keep the size down. To stay competitive with training you need thousands of GPUs, and software that can target so many, so it seems out of the question to use pure Julia. But maybe you can go half way there, leave out some parts like distributing to many GPUs, use DeepSpeed or something for that.

Training from scratch is still very costly, so no need to, since you can finetune a model for Julia use. But then you need to choose the best model to start from and formats/quantization as in llama.cpp or this new bitnet.cpp from Microsoft. See on the former (and relation to Llama2.jl):

> [@Is Rust (and/or Python) the new high-level API (here ANN/GGUF/LLama example)](https://discourse.julialang.org/t/is-rust-and-or-python-the-new-high-level-api-here-ann-gguf-llama-example/121721):
>
> This state-of-the-art library has Rust and Python API, but ironically no (C++) high-level one despite written in C++: I’m guessing C-like C++ is kinda necessary for bindings to other [languages] [TensorFlow, also written in C++, only had officially stable Python API (while Julia’s API better until unmaintained).] If we want to call this library then we can for sure, using PythonCall.jl, also an option to call Rust’s API. So which would you prefer? You might think why not use: GGUF mod…

KAN networks (they can be drop-in replacement into MLP part of transformers, if I recall) are worth-while to reimplement in Julia:

> [@Has Julia met your need for AI and ML?](https://discourse.julialang.org/t/has-julia-met-your-need-for-ai-and-ml/119818/5):
>
> The reason Python is used, is because it’s a good interface to high-speed code written in other fast languages. Julia can also be a good interface, and that faster language. Training massive AI/ML models relies on libraries like Microsoft’s DeepSpeed (“[Python 69.0%](https://github.com/microsoft/DeepSpeed/search?l=python) [C++ 20.3%](https://github.com/microsoft/DeepSpeed/search?l=c%2B%2B)[Cuda 9.8%](https://github.com/microsoft/DeepSpeed/search?l=cuda)”, Julia could theoretically have replaced all the languages, the numbers are for source code, for runtime about 0% is Python), for distributed training, this is only of interest to big companies that can afford to…

KAN networks are likely not compatible with 1-bit networks, I mean their weights larger, but might still be a good thing, if you get away with fewer. Also I think not intirely contradictory, since you can still have a transformer and other parts with 1-bit weights, where KAN is not replaceing the MLP part. But isn’t the MLP part the largest part of the total?

I think also worthwhile to help with this:

> [@Julia implementation of Liquid Neural Networks - anyone wanna help make one?](https://discourse.julialang.org/t/julia-implementation-of-liquid-neural-networks-anyone-wanna-help-make-one/100953):
>
> Following a thread on #machine-learning on Julia slack, i’m interested in putting together a Julia implementation of [“Liquid Neural Networks”](https://cap.csail.mit.edu/sites/default/files/research-pdfs/Robust%20flight%20navigation%20out%20of%20distribution%20with%20liquid%20neural%20networks.pdf). These are time-adaptive neural networks, inspired by neuroscience models, where each neuron’s activity is a differential equation. The synapses are also plastic and continue to adapt after training. I think this is very interesting, and Julia has a great ecosystem for automatic differentiation with differential equations. Anyone interested in teaming toge…

Best models will likely use new ways of multiplying not yet in software (but you could emulate slowly(?) for compatibility until hardware catches up, or maybe just use Float8, of bflot, I don’t recall, might be compatible with it):  
[https://arxiv.org/html/2410.00907v2#S2](https://arxiv.org/html/2410.00907v2#S2)

> [@In Binary Floating Point, you can multiply without multiplying](https://discourse.julialang.org/t/in-binary-floating-point-you-can-multiply-without-multiplying/121218/21):
>
> This is very intriguing (though I think the benefit of the methods doesn’t translate to 4-bit floating point, it may not matter as much, since likely useless for training neural networks; while useful for inference, though even then 2/3-bit integers taking over). The proposed ℒ-Mul method will lead to a significantly reduced energy consumption for both model training and inference. […] multiplying two 32-bit floating point numbers (fp32) costs four times energy as adding two fp32 numbers, and …

---

_[View the full topic](https://discourse.julialang.org/t/community-interest-check-llms-from-scratch-in-pure-julia/121796)._
