# Proof of concept LLM chatbot built with Julia: KeemenaLM.jl

**URL:** <https://discourse.julialang.org/t/proof-of-concept-llm-chatbot-built-with-julia-keemenalm-jl/137961>\
**Category:** Package Announcements\
**Created:** [July 5, 2026, 8:37pm UTC](https://discourse.julialang.org/t/proof-of-concept-llm-chatbot-built-with-julia-keemenalm-jl/137961 "2026-07-05T20:37:34Z")\
**Posts on this page:** 1\
**Showing post:** 32

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [August 2, 2026, 11:58pm UTC](https://discourse.julialang.org/t/proof-of-concept-llm-chatbot-built-with-julia-keemenalm-jl/137961/32 "2026-08-02T23:58:42Z")

</div>

Consider using Muon (PR seems ready, or stuck?):

> <https://github.com/FluxML/Optimisers.jl/pull/203>
>
> This adds Muon (https://kellerjordan.github.io/posts/muon/), which uses an appro…ximate orthogonalization before the update. There isn't a publication, but it gave a key improvement in the nanoGPT training "speedrun" attempt: https://github.com/KellerJordan/modded-nanogpt
> 
> This was started by @mashu 
> 
> \### PR Checklist
> 
> \- \[\] Tests are added
> \- \[\] Documentation, if applicable

> **[To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters](https://arxiv.org/abs/2603.00742v1)**
>
> While Adam has long been the ubiquitous default optimizer for deep neural networks, Muon has recently seen rapid adoption due to its superior training speed. Although much of the literature focuses on validating the benefits of Muon, our work...

or (I didn’t know of this, newer/better variant?):

> [@Has anyone tried implementing the LiMuon Optimizer?](https://discourse.julialang.org/t/has-anyone-tried-implementing-the-limuon-optimizer/132742):
>
> Has anyone tried implementing the [LiMuon optimizer](https://arxiv.org/abs/2509.14562)? I made an issue [here](https://github.com/FluxML/Optimisers.jl/issues/213) in the Optimisers.jl package but thought I’d ask here as well if someone has implemented it elsewhere? Also any hands on experience with the method itself would be interesting to hear about. According to the paper it beats AdamW in a few benchmarks both on training/testing error and convergence speed.

You currently use Adam, why not AdamW (or Muon etc)?

I suppose you’re using SwiGLU (SwiGLU-Clip was new to me and), newer to consider:

> **[PowLU: An Activation Function for Stable Pre-Training of LLMs](https://arxiv.org/abs/2605.25704)**
>
> In contemporary large language models (LLMs), the swish-gated linear unit (SwiGLU) activation function is widely adopted to regulate the information flow and introduce non-linearity. For large positive inputs, SwiGLU approximates the quadratic...

I suppose your problem is the 512 token context window, one reason and/or small/bad training dataset(?), for not being better. I think it might be intentional because of quadratic attention. Too costly for you to increase? I suppose SubQ is too hard to implement. You can have some linear layers, if you don’t already (and/or Mamba layers).

This paper seems awesome, but I guess a premature to add to your design:

[https://papers.ssrn.com/sol3/papers.cfm?abstract\_id=7170898](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7170898)

> We propose Hierarchical Adaptive Memory (HAM), a memory layer that tiers information into working, episodic, and semantic stores by a utility score (frequency, reuse, recency, novelty, predictive utility, stability) and consolidates frequent, redundant content into fewer prototypes while retaining rare content at higher fidelity. This predictive-coding-inspired compression lowers byte footprint at a similar fine-tuning budget. Unlike prior tokenization-level hierarchical memory, HAM operates at the prototype level and so can serve as an external store, a KV-cache compressor, or an internal memory block.

https://embed.reddit.com/r/LocalLLaMA/comments/1v3nw3j/cactus_hybrid_we_taught_gemma_4_to_know_when_its/?embed=true&ref_source=embed&ref=share

AVQ-Attention: Adaptive Vector-Quantized Attention  
[https://arxiv.org/pdf/2607.12789v1](https://arxiv.org/pdf/2607.12789v1)

seems like a very important paper (for you/LLMs too? It’s validated on image classification, AI suggested to me it would work for LLMs, and I challenged that, since not mentioned, only LLMs in sources, it thinks it translates to LLMs). If the I understand correctly and the codebook, M, is small, then much better than the quadratic attention I suppose you use, i.e. it’s O(MN).

CommVQ paper is also intriguing, maybe it can combine with above:

> .. FP16 KV cache size by 87.5% with 2-bit quantization, while outperforming state-of-the-art KV cache quantization methods. Notably, it enables 1-bit KV cache quantization with minimal accuracy loss, allowing a LLaMA-3.1 8B model to run with a 128K context length on a single RTX 4090 GPU. The source code is available at: [GitHub - UMass-Embodied-AGI/CommVQ: [ICML 2025] CommVQ: Commutative Vector Quantization for KV Cache Compression · GitHub](https://github.com/UMass-Embodied-AGI/CommVQ).

In case you or anyone wants to do 2-bit quantized models (a bit premature until your generated model is better):

> **[GitHub - AI9Stars/UniSVQ: \[ICML 2026\] Codes of paper UniSVQ: 2-bit Unified...](https://github.com/AI9Stars/UniSVQ)**
>
> \[ICML 2026\] Codes of paper UniSVQ: 2-bit Unified Scalar-Vector Quantization and LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization.

> Experiments across diverse LLMs demonstrate that LC-QAT consistently outper-  
> forms state-of-the-art QAT methods while using only 0.1%–10% of the training data.

They use AdamW in the paper (I suppose your Adam would also work, and probably Muon or similar too).

https://embed.reddit.com/r/singularity/comments/1odklk9/meta_the_free_transformer_an_improvement_to/?embed=true&ref_source=embed&ref=share

Note, the paper was rejected (it seems very intriguing to me, found it reference in breakthrough paper below, your call if you want to implement it), but read arguments and answers, still starting this way in the metareview:

> **[Verifying your browser | OpenReview](https://openreview.net/challenge?redirect=%2Fforum%3Fid%3DRV43b2B1SC)**

> I agree with the reviewers that the idea of the paper is simple, elegant, and appealing. Adding a latent variable to the decoder with very small overhead is attractive, and the paper shows improvements on several benchmarks. In this sense, the work is a tangible contribution and shows promising results.

> **[paper.pdf](https://explorative-modeling.github.io/static/pdfs/paper.pdf)**
>
> 4.49 MB

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

> Autoregressive LLMs have proven the hardest case, for the reasons discussed in our limitations (Section 7)  
> [..]  
> The Free Transformer [68] conditions a decoder on a latent variable inferred by a VAE, which is exactly the kind of latent exploration searches over, and training it with Explorative Modeling instead would remove the VAE entirely, along with the exposure bias of training on inferred latents (Section 3.1). Beyond language models, few-step models such as MeanFlow [15] are a natural fit for exploration as well, since exploration can supply the generative expressivity their shortened trajectories amortize. We also believe in combining XMs with Energy-Based Transformers (EBTs) [69], where the biggest documented challenge with EBTs has been end-to-end generation and handling highly multimodal distributions, which is exactly what XMs enable. Paired together, XMs and EBTs could enable more dynamic reasoning, search, and generalization over entire sequences.

> **[Alexi Gladstone | Explorative Modeling -- Unlocking a Third Pretraining Axis...](https://alexiglad.github.io/blog/2026/explorative_modeling/)**
>
> What in the world is Explorative Modeling?

[Note e.g. there Figure 9 needs light mode, broke in dark mode]

> **[GitHub - alexiglad/XM: PyTorch Code for Explorative Modeling: Unlocking a...](https://github.com/alexiglad/XM)**
>
> PyTorch Code for Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

> The longer a model generates, the worse things get—you’ve probably seen this firsthand with video generation models that melt into mush after ten seconds, or LLMs that get less coherent over really long generations.

> **[Alexi Gladstone | Energy-Based Transformers are Scalable Learners and Thinkers](https://alexiglad.github.io/blog/2025/ebt/)**
>
> What are Energy-Based Transformers and why should I care about them?

> **TLDR** : We **outscale** (feed-forward) transformers while **generalizing** reasoning/system 2 thinking to any modality/problem **without** requiring verifiable rewards​😮. Energy-Based Transformers are the **first approach** to outscale feed-forward transformers across modalities and with respect to several axes including data, depth, parameters, FLOPs, etc. Energy-Based Transformers can think over every single prediction (i.e. every token in language modeling) and generalize better than existing models.

[https://arxiv.org/pdf/2507.02092](https://arxiv.org/pdf/2507.02092)

> Specifically, we train **Energy-Based Transformers (EBTs)**—a new class of Energy-Based Models (EBMs)—to assign an energy (un-  
> normalized probability) value to every input and candidate-prediction pair, enabling  
> predictions through gradient descent-based energy minimization until convergence.  
> This formulation enables System 2 Thinking to emerge from unsupervised learn-  
> ing, making it modality and problem agnostic. Across both discrete (text) and  
> continuous (visual) modalities, we find EBTs scale faster than the dominant Trans-  
> former++ approach during training, achieving an up to 35% higher scaling rate with  
> respect to data, batch size, parameters, FLOPs, and depth. During inference, EBTs  
> improve performance with System 2 Thinking (i.e., extra computation) by 29%  
> more than the Transformer++ on language tasks, and EBTs outperform Diffusion  
> Transformers on image denoising while using fewer forward passes

> **[Mean Flows for One-step Generative Modeling](https://arxiv.org/abs/2505.13447)**
>
> We propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity...

> **[MODUS: Decoder-only Any-to-Any Modeling of Diverse Modalities](https://modus-multimodal.epfl.ch/)**
>
> MODUS unifies any-to-any multimodal generation in a single decoder-only model.

---

_[View the full topic](https://discourse.julialang.org/t/proof-of-concept-llm-chatbot-built-with-julia-keemenalm-jl/137961)._
