# Which is the best open model for \[Julia\] coding? And practical, and still reasonably good

**URL:** <https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640>\
**Category:** General Usage\
**Created:** [August 6, 2026, 12:35pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640 "2026-08-06T12:35:21Z")\
**Posts on this page:** 1\
**Showing post:** 10

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [August 7, 2026, 2:18am UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/10 "2026-08-07T02:18:20Z")

</div>

As for AirLLM and Mixture-of-Kittens: I’ll be honest with you, I’m not even planning to touch them. A lot of this is new to me. For example, I initially tried to run GLM 5.2, and it took me about five days to realize it’s probably not possible. The most success I had was with K 2.7, where the best I could achieve was about 8 t/s. However, at 70k tokens, the model starts producing garbage, and as for now, I don’t know why. To get it running, I had to build vLLM, Triton, Flash-Attention, and the rccl-cxi-plugin (based on aws-ofi-rccl) and write a run script. It wasn’t a huge deal, but it wasn’t easy for me either. While I know a bit about Conda, I know almost nothing about Python. On top of that, I’m on a distributed filesystem, so loading these models feels like loading a game from a tape. Imagine how that affects every single iteration.

> [@Palli](#):
>
> What do you mean waitlist? It’s good to see “large speedups relative to PyTorch (2x to as high as 10x on the same hardware).”

You know, with a 10x speedup, I’m busy shorting NVIDIA.

> [@Palli](#):
>
> but I see it’s for ML however for “non-LLM” models …

Well, as far as my current understanding goes, it is possible to run LLMs as well:

> [@ReactantServer.jl: serve more models per GPU with Reactant.jl and XLA (gauging interest)](https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358/14):
>
> I also looked into serving LLMs. The main challenges are potentially needing many different compiled program sizes with padding / KV cache using StableHLO semantics / wrapping all of that up together in a way that can be effectively batched. In principle, there shouldn’t be any reason why you couldn’t serve a dense model like [Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) today. I don’t have time to look into this currently, but serving these sorts of models which can reasonably fit on a single workstation class GPU is something I would like to support in the future.

---

_[View the full topic](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640)._
