# Which is the best open model for \[Julia\] coding? And practical, and still reasonably good

**URL:** <https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640>\
**Category:** General Usage\
**Created:** [August 6, 2026, 12:35pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640 "2026-08-06T12:35:21Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [August 6, 2026, 12:35pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/1 "2026-08-06T12:35:21Z")

</div>

Continuing the discussion from [Anthropic gave me 6 months of Claude Max for my contributions to the Julia ecosystem 🙂](https://discourse.julialang.org/t/anthropic-gave-me-6-months-of-claude-max-for-my-contributions-to-the-julia-ecosystem/138591/28):

> [@Anthropic gave me 6 months of Claude Max for my contributions to the Julia ecosystem slight\_smile](https://discourse.julialang.org/t/anthropic-gave-me-6-months-of-claude-max-for-my-contributions-to-the-julia-ecosystem/138591/28):
>
> Proliferation of vibe-coded packages will make the ecosystem unmaintainable without LLM subscription.

I was just trying to point out an “LLM subscription” is not needed. Local models exist. Qwen was quite good at the time and quantized versions (to extreme binary, or ternary) of it like Bonsai 27 B works fast on even an iPhone:

https://embed.reddit.com/r/LocalLLaMA/comments/1uwhukq/bonsai_27b_the_first_27bclass_model_to_run_on_a/?embed=true&ref_source=embed&ref=share

> **[ProgramBench Benchmark Scores & AI Model Leaderboard | BenchmarkList](https://benchmarklist.com/benchmarks/programbench/)**
>
> A benchmark where software engineering agents rebuild complete programs from compiled binaries and documentation, then are scored against hidden behavioral tes…

Kimi K3 is likely best but slowest (not that slow?). Is best on the intriguing ProgramBench benchmark, at 77.8%, just edging out best closed models GPT-5.6 Sol at 77.6% and Claude Fable 5 (max, with fallback) 76.8% and all using KimiCode harness. [GLM 5.2](https://benchmarklist.com/models/z-ai-glm-5.2/ "GLM 5.2") is next best open _on that metric_ (some say not good for Julia, using best harness?), then [Kimi K2.7 Code](https://benchmarklist.com/models/moonshotai-kimi-k2.7-code/ "Kimi K2.7 Code")

GLM 5.2 is likely one of the best, I thought impossible to run locally, or the larger DeepSeek V4 [Flash or Pro] but you CAN run the largest Kimi K3 locally, yes slowly:

> **[Best LLM for Coding 2026 | AI Coding Model Rankings & Benchmarks](https://onyx.app/best-llm-for-coding)**
>
> Which AI model writes the best code? We rank every major LLM — open and closed source — across SWE-bench, HumanEval, LiveCodeBench, and Terminal-Bench coding benchmarks. Compare the best LLMs for coding, software engineering, and programming.

> **[Best LLM for Coding](https://www.vellum.ai/best-llm-for-coding)**
>
> Compare all proprietary and open source models across programming benchmarks, and see which one is the best.

> **[Can Kimi K3 Run in 8GB RAM? C99 Claim Audit](https://kingy.ai/blog/kimi-k3-8gb-ram-c99-1-7tb-ssd/)**
>
> A C99 engine reports Kimi K3 in 8.24GB peak RAM with a 1.7TB storage bill and 32.69 seconds per token. Kingy.ai audits the evidence.

https://embed.reddit.com/r/LocalLLaMA/comments/1vd874t/i_pushed_kimi_k3_onto_one_cpu_with_8_gb_of_ram/?embed=true&ref_source=embed&ref=share

Note the metric there is s/token, not the usual tokens per second; I’ve seem some claim Kimi K3 is fast (on the CPU!), I think they just read the numbers, and I was careful to point out it slow. It just got stuck in my memory it’s possible at all, I didn’t recall if it was quantized, then could be sped up. Or if using BLAS and/or GPU. It’s memory limited, so if not quantized 16x larger than it strictly needs to be. Could be sped up by that amount roughly then.

---

<div class="post-metadata">

**Author:** ![Tortar](https://avatars.discourse-cdn.com/v4/letter/t/6bbea6/32.png) [@Tortar](https://discourse.julialang.org/u/Tortar)\
**Post date:** [August 6, 2026, 12:52pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/2 "2026-08-06T12:52:06Z")

</div>

I think it is really not possible to run Kimi K3 or other really big LLMs locally without proper hardware. Look here: [Release AirLLM v3.1.0 — Kimi K3 (2.8T) on a single card · lyogavin/airllm · GitHub](https://github.com/lyogavin/airllm/releases/tag/v3.1.0), it is incredible, yes, that one can run it with only 4GB of VRAM, though the speed is 292 s/token, which means it isn’t practical for **any** task, you would wait 10-100 days for a single response 😅…maybe there are better implementations but I think it’s kind of impossible that without dedicated hardware one can go anywhere with big local models (for now). The only one which seems maybe worth it is DeepSeek V4 Flash, in the sense, that the necessary hardware is less demanding since it is 281B, but one still needs to have a proper use case to justify not to pay the API which to me it seems at electricity costs (this would an interesting question for me: what is the energy costs and speed in current LLMs with different propor setups).

EDIT: saw that your link says

> Is 32.69 seconds per token a verified decode speed? No. It is one author-reported eight-token average that includes first-step prompt work and cold-start effects. It is not a separately measured steady-state decode rate

32 seconds seems better (still way impractical). And doesn’t seem to be verified anyway. I think the same for the other links.

In any case I would bet on Qwen 3.8 27B for the best locally runnable, still not available but will be soon from what I get. Surely though significantly worse than bigger models, at least this seems the case for now, and I don’t think new releases will change this. Though, who knows? Maybe it works fine for some things. Scale seems one dimension which still counts.

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 6, 2026, 1:50pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/3 "2026-08-06T13:50:18Z")

</div>

IMHO, given the price of Deepseek Flash v4 ($0.084 / $0.168 per 1M tokens), you will spend a fraction of what you need considering the hardware, energy, etc. to run locally an acceptable model.

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [August 6, 2026, 4:13pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/4 "2026-08-06T16:13:12Z")

</div>

I’m afraid @Tortar and @Ronis_BR are right. Unless you’re doing a massive amount of coding and running concurrent requests, it’s probably not economically justifiable to run the latest large models locally on your own hardware. Or at least this is my current understanding.

NVIDIA is offering a decent selection of open-weights models with substantial limits on their free endpoints. Hyper for Charm, among others, offers V4 Flash, and you can get a lot done on their free tier (for workloads suitable for this model).

P.S. BTW, I think we are significantly underestimating Julia in this field. The stack is improving every month. It’s just that the level of entrepreneurship in our community is extremely low, at least in my experience. Moreover, the level of negativity toward this new technology is also very surprising.

---

<div class="post-metadata">

**Author:** ![joa-quim](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joa-quim/32/227_2.png) [@joa-quim](https://discourse.julialang.org/u/joa-quim)\
**Post date:** [August 6, 2026, 5:09pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/5 "2026-08-06T17:09:17Z")

</div>

> [@Palli](#):
>
> GLM 5.2 is likely one of the best,

I tried it inside Claude Code and basically it is a piece of … Not to mentioned that it consumed the credits some 10x faster that even Opus.

Tried also GPT “Sol”. Only slightly better than GLM. And I almost always had to launch Opus to correct the … it did. A huge disappointment, No competition at all _vs_ Claude Code.

My current conclusion is that for more elaborated works there is only Claude Code Opus and above. Sonnet is good sometimes, but the time and limitations makes that Opus was the best choice most of times.

_PS_ I am **very** jealous of Ronan’s gift (muitos parabéns).

---

<div class="post-metadata">

**Author:** ![Ronis\_BR](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ronis_br/32/50999_2.png) [@Ronis\_BR](https://discourse.julialang.org/u/Ronis_BR)\
**Post date:** [August 6, 2026, 6:30pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/6 "2026-08-06T18:30:05Z")

</div>

> [@joa-quim](#):
>
> _PS_ I am **very** jealous of Ronan’s gift (muitos parabéns).

Thanks! (Obrigado!) 🙂

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [August 6, 2026, 9:18pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/7 "2026-08-06T21:18:45Z")

</div>

Kimi k3 seems to do pretty well compared to opus!

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [August 6, 2026, 10:09pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/8 "2026-08-06T22:09:28Z")

</div>

> [@sdanisch](#):
>
> Kimi k3 seems to do pretty well compared to opus!

_Locally_, I guess those kinds of things are _only possible_ in Germany. :- ) But seriously, how are you running it, @sdanisch? The `vLLM` requirement seems to be v0.27.0, but v0.26.0 is currently the latest official release.

@Palli, one thing I’d like to add: I have `ReactantServer.jl` on my waitlist. This looks like a huge deal:

> [@Wired: Python Is So Slow. Can Julia Solve the Two-Language Problem?](https://discourse.julialang.org/t/wired-python-is-so-slow-can-julia-solve-the-two-language-problem/138165/3):
>
> Reactant has gotten to the point where I’m using it for all of my ML training and serving, often seeing large speedups relative to PyTorch (2x to as high as 10x on the same hardware).

EDIT: Just to be clear, to mitigate any problems. The first sentence is harmless. It’s positive. A kind of a regional “smile” / “joke” / “pun”.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [August 7, 2026, 12:10am UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/9 "2026-08-07T00:10:26Z")

</div>

Great, you mean Kimi K3 for coding Julia? Good to know [then]. Not just for general programming, or some other specific language? And what harness, if any (I think they are critical, and not all as good) do you use with it?

It’s a good thing AirLLM added:

> **[GitHub - lyogavin/airllm: AirLLM 70B inference with single 4GB GPU](https://github.com/lyogavin/airllm)**
>
> AirLLM 70B inference with single 4GB GPU

> [2026/07] **Kimi K3 (2.8T)** support: the largest open-source model runs on a single card in **3.72GB** of VRAM, measured end to end on one RTX 6000 Ada. Per-expert streaming loads only the experts a token actually routes to.

> **[Speedups - Mixture-of-Kittens: our open-source MoE megakernel for NVL72s ·...](https://cursor.com/blog/mixture-of-kittens#speedups)**
>
> We're open-sourcing Mixture-of-Kittens, a deterministic MoE training megakernel for NVL72s that fuses communication and computation into a single kernel.

> Today, we’re open-sourcing [Mixture-of-Kittens (MoK)](https://github.com/cursor/mixture-of-kittens), our production MoE training megakernel for NVL72s.
> 
> As we have scaled the training and inference of [Composer](https://cursor.com/blog/composer-2), our agentic coding model, the mixture-of-experts layer has consistently remained the major bottleneck.

It’s twice as fast for e.g. Kimi K2.7 tested there (I suppose would work for K3 too, just wasn’t tested since quite new). “ **forward throughput** ” implies to me inference, but I’ve not read yet through all of this and what refers to it or what to training:

> | | DeepEP-based | MoK |
> | --- | --- | --- |
> | Tokens / second / GPU | 760.9 | 1,070.2 (1.41x) |
> 
> Overall, MoK delivered an approximately 41% tokens-per-second speedup over our previous DeepEP-based production setup, allowing us to train our models more efficiently across our GB300 NVL72 infrastructure.

> [@j\_u](#):
>
> `ReactantServer.jl` on my waitlist.

What do you mean waitlist? It’s good to see “large speedups relative to PyTorch (2x to as high as 10x on the same hardware).” but I see it’s for ML however for “non-LLM” models:

> ### [Who this is not for](https://enzymead.github.io/ReactantServer.jl/dev/design/philosophy/#Who-this-is-not-for)
> 
> …  
> LLM serving at scale. vLLM, TGI, TensorRT-LLM, and similar projects are purpose-built for that domain and do it well. This project does not compete in that market.

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [August 7, 2026, 2:18am UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/10 "2026-08-07T02:18:20Z")

</div>

As for AirLLM and Mixture-of-Kittens: I’ll be honest with you, I’m not even planning to touch them. A lot of this is new to me. For example, I initially tried to run GLM 5.2, and it took me about five days to realize it’s probably not possible. The most success I had was with K 2.7, where the best I could achieve was about 8 t/s. However, at 70k tokens, the model starts producing garbage, and as for now, I don’t know why. To get it running, I had to build vLLM, Triton, Flash-Attention, and the rccl-cxi-plugin (based on aws-ofi-rccl) and write a run script. It wasn’t a huge deal, but it wasn’t easy for me either. While I know a bit about Conda, I know almost nothing about Python. On top of that, I’m on a distributed filesystem, so loading these models feels like loading a game from a tape. Imagine how that affects every single iteration.

> [@Palli](#):
>
> What do you mean waitlist? It’s good to see “large speedups relative to PyTorch (2x to as high as 10x on the same hardware).”

You know, with a 10x speedup, I’m busy shorting NVIDIA.

> [@Palli](#):
>
> but I see it’s for ML however for “non-LLM” models …

Well, as far as my current understanding goes, it is possible to run LLMs as well:

> [@ReactantServer.jl: serve more models per GPU with Reactant.jl and XLA (gauging interest)](https://discourse.julialang.org/t/reactantserver-jl-serve-more-models-per-gpu-with-reactant-jl-and-xla-gauging-interest/137358/14):
>
> I also looked into serving LLMs. The main challenges are potentially needing many different compiled program sizes with padding / KV cache using StableHLO semantics / wrapping all of that up together in a way that can be effectively batched. In principle, there shouldn’t be any reason why you couldn’t serve a dense model like [Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) today. I don’t have time to look into this currently, but serving these sorts of models which can reasonably fit on a single workstation class GPU is something I would like to support in the future.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [August 7, 2026, 3:40pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/11 "2026-08-07T15:40:45Z")

</div>

5 posts were split to a new topic: [What is your favourite font for coding in Julia?](https://discourse.julialang.org/t/what-is-your-favourite-font-for-coding-in-julia/138662)

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [August 7, 2026, 9:26pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/12 "2026-08-07T21:26:35Z")

</div>

> [@sdanisch](#):
>
> Kimi k3 seems to do pretty well compared to opus!

Good to know, I suppose it’s the best open weight model, because of its enormous size, so I would also want to know what is second best or by size class. And how slow Kimi K3 needs to be:

For Kimi K3 from sqlite.ai, a token generated every other second (vs “[AirLLM](https://github.com/lyogavin/airllm) running K3 at ~5 minutes per token”) and 10.65 tokens per second for Kimi-Linear model:

> **[WASTE: Run Kimi K3 by Streaming Experts from NVMe – ToKnow.ai](https://toknow.ai/posts/waste-streaming-kimi-k3-nvme/)**
>
> WASTE is a dependency-free C inference engine that runs the full 2.78-trillion-parameter Kimi K3 from a 982 GB container. It needs 29.06 GB to open the model and about 64 GB for practical use.

> recommends 64 GB plus a fast internal SSD. WASTE’s [efficiency documentation](https://github.com/sqliteai/waste/blob/main/docs/EFFICIENCY.md) reports 0.45 to 0.62 tokens per second on a 64 GB M5 Pro. WASTE uses 3-bit residual vector quantization for experts, keeps more sensitive shared weights at 4 or 8 bits, and uses router lookahead to begin reads before the next layer needs them. At 4K context, K3’s compressed KV cache is about 0.21 GB.
> 
> .. The documentation also measures Kimi-Linear 48B at 10.65 tokens per second in a 19 GB container with a 1.28 GB minimum. K3 remains the main and best-tested target. .. Conversion takes about 4.7 hours with three workers and needs another 1.42 TB of temporary space.

Another option slightly slower (is it meant for lesser hardware?): [GitHub - gavamedia/deltafin: Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents. · GitHub](https://github.com/gavamedia/deltafin)

This one seemed interesting: [Maple-Preview, an AI that runs on iPhones and boasts performance equivalent to Bonsai 27B while being 13 times faster, represents another step forward for local AI. - GIGAZINE](https://gigazine.net/gsc_news/en/20260805-maple-preview/#gsc.tab=0)

> 13 times faster than Bonsai 27B [quantized Qwen], which also runs on an iPhone. .. Maple-Preview was reportedly able to accurately solve problems from the International Mathematical Olympiad. [at 281.5 tokens/s on Macbook Pro (M5 Pro)] Maple-Preview consumes only 7.69GB of memory even when handling 131,000 tokens.

I do have only 2 GB VRAM not sure I’m even using the GPU and currently only 32 GB of my 128 GB regular RAM installed, i.e. only one DIMM, would have more channels bandwidth otherwise; 16 cores (hyperthreaded I guess), I changed default from `--threads 16` shown in the (modified) llama.cpp github to half the threads, and got more speed, I was happy before:

```julia-auto
$ ./build/bin/llama-completion -m models/maple-preview-TQ2_0-head-Q4_K.gguf --threads 8 --temp 1.0 --top-p 0.95 --jinja --conversation
..
0.32.477.396 I common_perf_print: prompt eval time = 268,39 ms / 24 tokens ( 11,18 ms per token, 89,42 tokens per second)
0.32.477.404 I common_perf_print: eval time = 18496,73 ms / 428 runs ( 43,22 ms per token, 23,14 tokens per second)

```

I tested it with my standard question “What is the Julia language, and can you show me example code?” [first] at:

> **[DeepGrove](https://chat.deepgrove.ai/)**
>
> DeepGrove — frontier intelligence, on any device.

Neither that model nor MiniMax claim to support coding, maybe they do and I would like to know if good for:

> **[MiniMaxAI/MiniMax-H3 · Hugging Face](https://huggingface.co/MiniMaxAI/MiniMax-H3)**
>
> We’re on a journey to advance and democratize artificial intelligence through open source and open science.

> **[maternion/strand-rust-coder](https://ollama.com/maternion/strand-rust-coder)**
>
> The model fine-tunes Qwen2.5-Coder-14B for Rust-specific programming tasks using a 191K-example synthetic dataset built via multi-model generation and peer-reviewed validation.

> The model fine-tunes Qwen2.5-Coder-14B for Rust-specific programming tasks using a 191K-example synthetic dataset built via multi-model generation and peer-reviewed validation.

I suppose we could do similar for any of [those] models if needed.

---

<div class="post-metadata">

**Author:** ![csvance](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/csvance/32/218927_2.png) [@csvance](https://discourse.julialang.org/u/csvance)\
**Post date:** [August 8, 2026, 8:16pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/13 "2026-08-08T20:16:04Z")

</div>

When I ran out of credits on my Anthropic plan, I tried out Deepseek Flash v4 with pi.dev as the agent. The model was capable enough to port additional ViT backbones for Luximm.jl and also refactor my ViT code to add support “multiscale” outputs that are needed for problems like semantic segmentation, object/keypoint detection, etc. I spent around 50 cents total and it knocked out many other tasks as well. Apparently you can run Flash v4 on two RTX 6000 Pro Blackwell at impressively high token rates, trying to convince my boss to move a few things around so I can have it working on Julia open source stuff full time 😅

EDIT: If you are interested in seeing my setup which has extensive sandboxing via bwrap + supply chain hardening, seamless kaimon integration, herdr to manage multiple agents easily, you can see it here: [GitHub - csvance/pi-sandbox · GitHub](https://github.com/csvance/pi-sandbox)

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [August 8, 2026, 10:09pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/14 "2026-08-08T22:09:17Z")

</div>

@Palli, to sum up my part of the recent discussions. I think you’ve done a tremendous job with the research and in general popularizing this topic, not only recently but over the course of the last several months. As I indicated in the other threads, I am sustaining my preliminary interest. I believe this is an important topic for the development of Julia (however, I might be very wrong).

At the same time, I think this is potentially a substantial project. For context, I read that training Poolside Laguna S 2.1 (118B parameters) required 4K Nvidia H200 GPUs over a 4-week timeframe. Similarly, Arcee AI Trinity Large (400B parameters) required 2K Nvidia B300 GPUs with a 33-day pretraining window. Furthermore, there is unconfirmed speculation that Moonshot had a computing agreement for around 20K H200 GPUs to train Kimi K3 (2.8T parameters).

While I’m a bit more skeptical than @csvance regarding the capabilities of some of the models mentioned here, I share @Tortar’s concerns about energy and dynamic regulatory environment (I think both are important aspects to monitor), and I certainly share the enthusiasm expressed by @Ronis_BR, @joa-quim, and @sdanisch. (Moça do corpo dourado, do sol de Ipanema, o seu balançado é mais que um poema).

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [August 8, 2026, 11:45pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/15 "2026-08-08T23:45:16Z")

</div>

> [@j\_u](#):
>
> I read that training Poolside Laguna S 2.1 (118B parameters) required 4K Nvidia H200 GPUs over a 4-week timeframe.

I’m not suggesting we train such a model or larger _from scratch_. I mentioned so-called _“fine-tuning” (i.e. postraining) of models, i.e. for Julia_ (if that is actually needed at all…); such could be done in a day way back, on one GPU. I don’t recall how large such a model it was or exactly when, and I suppose current, larger, models could require longer finetuning. Though that probably needs not be the case, since often you freeze some layers, and only finetune others. Also this thread was mostly about what is the best model now, without too many hacks or finetuning.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [August 9, 2026, 12:10am UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/16 "2026-08-09T00:10:41Z")

</div>

Flash is on the left. The price has already been raised once.

| **[FIM Completion（Beta）](https://api-docs.deepseek.com/guides/fim_completion)** | **Non-thinking mode only** | **Non-thinking mode only** | |
| --- | --- | --- | --- |
| **1M INPUT TOKENS (CACHE HIT)** | **$0.0028** | **$0.003625** | |
| **1M INPUT TOKENS (CACHE MISS)** | **$0.14** | **$0.435** | |
| **1M OUTPUT TOKENS** | **$0.28** | **$0.87** | |

> (2) We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected.

So use it while you can… I’m not sure if the announced increase is not yet in effect for those with subscription.

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [August 10, 2026, 2:13am UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/17 "2026-08-10T02:13:23Z")

</div>

> [@Palli](#):
>
> Also this thread was mostly about what is the best model now, without too many hacks or finetuning.

My apologies. I was biased by the combined impression of your previous posts, including the other threads. Additionally, your initial post made several references to some of the largest models currently available.

I have not tested every one of them. I was somewhat disappointed with `GLM 5.2` and consequently didn’t spend much time with it. `Kimi K 2.7 Code` is decent, though its extensive thinking process is not always accurate. I occasionally use `DeepSeek V4 Flash`, which is impressive, particularly for the initial analysis of large logs and documents. `Mimo Pro 2.5` also impressed me, although I suspect the CLI contributed significantly to that experience. My primary model is `Claude Sonnet 4.5`, and I turn to `Claude Sonnet 5` and `GPT 5.6 Terra` from time to time in case of trouble.

Regarding training, as I mentioned in the other thread, I plan to use `MaxText` to train `Qwen3.6-27B` for deployment via `ReactantServer`. I am also researching `Trinity Large TrueBase` and `Base`, as I believe there are few recent public examples of such implementations. This is currently on my waitlist, as I have yet to verify the legal, software, and hardware requirements and as many others I have limited time and a few other things on the list that I planned to do.

In general, looking ahead, in the medium to long term, I’m expecting some form of specialization, however, I believe the prospect of running a truly useful model on a mobile device or with only a few GBs of VRAM seems to be a bit too optimistic to me. [P.S. I hope this reply is relevant to the topic. :- )]

---

<div class="post-metadata">

**Author:** ![yvikhlya](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yvikhlya/32/3753_2.png) [@yvikhlya](https://discourse.julialang.org/u/yvikhlya)\
**Post date:** [August 10, 2026, 10:48am UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/18 "2026-08-10T10:48:21Z")

</div>

Modern LLMs can generate code for pretty much any language, not just those languages, which they were trained on, if you provide enough context and proper tools. You can even make your custom DIY language, then feed sources/specs to LLM and it will generate correct code without re-training (I tried this). Small local models, below 100B params, are good only for simple things like auto-completion. For any serious coding and planning you need real big LLM, but not necessarily the top tire one. Cheaper models do good as well, they just take more attempts and more tokens to solve the task. Some providers even give you an estimate of $$/real-world-task for their models. I like Xiaomi MiMo-2.5, which is probably the cheapest one out of capable models.

---

<div class="post-metadata">

**Author:** ![j\_u](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j_u/32/219081_2.png) [@j\_u](https://discourse.julialang.org/u/j_u)\
**Post date:** [August 10, 2026, 3:13pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/19 "2026-08-10T15:13:02Z")

</div>

@Palli 1. Reflecting, I wondered if you were referring to distributed computation when you mentioned mobile devices. 2. I recall you sharing a research paper a few months ago regarding distributed training or inference. I haven’t been able to locate the link, but I believe it would be valuable to revisit it, especially in this context. 3. As an addendum to my summary post above, I would like to confirm my preliminary interest again. I currently have nominal access to 256 GPUs for approximately six months, with potential for expansion (TBC - I can’t promise anything). I am not a professional coder / software engineer (Julia, C, q, Redpanda, Oracle ADB, Oracle HeatWave, QuestDB, kdb-X, and BeeGFS). I can dedicate some time to this potential project, however, certainly not on a full-time basis for now.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [August 13, 2026, 7:06pm UTC](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640/20 "2026-08-13T19:06:49Z")

</div>

People are abandoning Claude Code for that (open source) agent harness (built on Pi and [Recursive Language Models: the paradigm of 2026](https://www.primeintellect.ai/blog/rlm)), since it’s more token efficient for coding (and more like [MazeBench](https://mazebench.com/blog?post=maze-bench-results) is an open-world 3D spatial reasoning environment .. Frontier models are shown to greatly struggle on this task, expending billions of tokens to solve only a fraction of the overall world.") please try whatever model, e.g. GLM 5.2, again with this harness:  
[GitHub - PrimeIntellect-ai/prime-agent: A self-improving RLM agent for coding workflows and long-running autonomous tasks. · GitHub](https://github.com/PrimeIntellect-ai/prime-agent)

> Prime Agent serves as both a coding agent to be used, and a _harness design to be evaluated_ for research. We make special note that while many modern frontier models are trained around a specific harness, currently no model has been trained around Prime Agent or its core feature set.
> 
> **ARC-AGI 3.** [ARC-AGI 3](https://arcprize.org/arc-agi/3) is a popular intelligence benchmark  
> ..  
> Our best results use Opus 5 in Prime Agent to achieve **95.5% RHAE Best@1** , **which surpasses the ARC reported human expert baseline of 95.4%**. Across three runs, we find that Prime Agent consistently performs well [95.0, 95.2, 95.5] and 99.97% Best@3 with all 183/183 levels complete. Our median score card action replay (95.2%) for ARC-AGI-3 can be found [here](https://arcprize.org/scorecards/2af780b4-f2a1-43e9-a794-b23da3cd3f9f).

GLM 5.2 gets better benchmark scores than Opus 5 on EmulatorBench _long coding_, than Opus 5 (it is boosted still on most of the benchmarks vs with Claude code) using this agent, while GPT-5.6 Sol (high) is boosted more on this benchmark:

> We report preliminary results on this long-context coding benchmark averaged over 16 emulator reconstructions, as well as two emulators, the SEGA Genesis and Nintendo Game Boy Color, that Prime Agent successfully reproduces. For Opus, our runs surprisingly failed to solve the tasks despite successful tool-call responses.

I suppose this will also work great with Kimi K3, but will likely boost lesser models more.

[![](https://global.discourse-cdn.com/julialang/original/3X/7/3/738a3c965cd5420ddb027846b8e1473f01f86d63.jpeg "Prime Agent: The Self-Improving AI Is Here") ](https://www.youtube.com/watch?v=BVUauHOp89Y)

I was skeptical of that guy, but that video “Prime Agent: The Self-Improving AI Is Here” gat dood info, I’ve been intriguing to try out Hermes Agent, haven’t watched yet his “Prime Agent VS Hermes: Which AI Agent Wins?” (I think I know the answer…) nor “Prime Agent AI Just Changed Agents Forever”.

Apparently 100,000 is “godlike” in this game, and people didn’t believe, so they run it live:

[![](https://global.discourse-cdn.com/julialang/original/3X/8/3/83ebc934c7a81cbeba3b49b73bcc276fb7d1f0e2.jpeg "Prime Agent Plays Factorio") ](https://www.youtube.com/watch?v=5oWKw1V9gUg)

Launching today, beating Kimi K3 at least on one benchmark, Cybergym, and all other models, greatly improved from its preview:

> <https://x.com/deepseek_ai/status/2087864585504305397>

[Those stats are likely without any harness, at least without Prime Agent? So could be improved further, most interested how it work in Julia.]

> **[GitHub - deepseek-ai/deepseek-harness: DeepSeek Harness: Everything is a Plugin.](https://github.com/deepseek-ai/deepseek-harness)**
>
> DeepSeek Harness: Everything is a Plugin.

From today referencing many papers, including from this month;

> **[Introducing NAC, an Open-Source Harness for Long-Running Agent Work](https://www.arcee.ai/blog/nac)**
>
> An open-source runtime built for complex engineering work. Nac coordinates parallel agent workers and persistent state across long tasks.

> This blog explains the implementation, the ideas it draws from, and why agent harnesses are becoming a new kind of inference runtime.  
> ..  
> Recursive Language Models make the trend especially clear: a long prompt lives in an external environment that the model can search, slice, and transform programmatically, including by invoking further model calls over selected pieces.[[16]](https://arxiv.org/abs/2512.24601)[[17]](https://www.primeintellect.ai/blog/prime-agent) Other systems take a different path, invoking many fresh sessions over the course of a run. Anthropic’s long-running harnesses[[18]](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents)[[19]](https://www.anthropic.com/engineering/harness-design-long-running-apps), the Ralph loop[[20]](https://ghuntley.com/ralph/), and Engram[[21]](https://arxiv.org/abs/2603.21321) carry progress through files and other external systems rather than one uninterrupted context. Coding agents do something similar whenever they preserve progress by editing the codebase itself.

> **[LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks](https://arxiv.org/abs/2608.01964)**
>
> Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and...

> .. LongHorizon-Harness improves Qwen~3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench~2.1, and from 2.8% to 8.3% on OSWorld~2.0. It also raises Claude Opus~4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.

> **[Qwen3.8-27B Open Model Release Set for Tomorrow](https://x.com/i/trending/2087919337394819573)**
>
> The Qwen3.8-27B, a dense 27-billion-parameter model billed as a workhorse with flagship-level smarts, releases open weights on August 14, 2026, on Hugging Face and ModelScope. It follows the recent launch of the massive 2.4-trillion-parameter...

There’s also (YouTube video) from someone else: “Introducing Granular: The Agentic Terminal for Non-Developers”.

I can’t resist posting this (AGI coming, since ARG-AGI-3 saturated?); not final answer, but better lower bound: [Learning more about Claude's mathematical capabilities \ Anthropic](https://www.anthropic.com/research/riemann-zeta)

> Drawing on extensive prior research by mathematicians over the past decades, it has increased this bound from 41.6% to 67.2%.

Older what I had in draft:  
Claude Code might not be the best harness, and those are also important, there OpenCode is claimed best:

> **[AI Coding Agent Benchmarks & Leaderboard | Artificial Analysis](https://artificialanalysis.ai/agents/coding-agents?coding-agents-harness-comparison-chart=harness-deep-swe)**
>
> We measure real-world performance of coding agents on software engineering tasks, including cost, token usage, and execution time. We compare how performance changes across agents, models, and execution settings.

> This chart holds the underlying model constant at Claude Opus 4.7 and compares how it performs across different coding-agent harnesses, including Cursor, Claude Code, and OpenCode.

It and e.g. Grok Build harness (since middle of last month, and it gets almost highest score) are open source: [GitHub - xai-org/grok-build: SpaceXAI's coding agent harness and TUI. Fullscreen, mouse interactive, extensible. · GitHub](https://github.com/xai-org/grok-build)

ext install moonshot-ai.kimi-code

The fastest hardware by far, from Taalas (the company just bought by AMD), at over 14000 tokens per second, answering my standard question (i.e. a full page appeared instantly), usable here (they put up a website, since they were not believed about the speed):

> **[chat jimmy](https://chatjimmy.ai)**
>
> chat jimmy LLM web interface

It hardcodes the model in ROM (now old Llama 3.1), so companies would have to change silicon e.g. every 5 months, but would still profit that way…

[Next page](https://discourse.julialang.org/t/which-is-the-best-open-model-for-julia-coding-and-practical-and-still-reasonably-good/138640.md?page=2)
