# Julia is one of the most token-efficient programming languages

**URL:** <https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996>\
**Category:** Offtopic\
**Created:** [January 12, 2026, 4:29am UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996 "2026-01-12T04:29:13Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaoxi](https://avatars.discourse-cdn.com/v4/letter/x/a9adbd/32.png) [@xiaoxi](https://discourse.julialang.org/u/xiaoxi)\
**Post date:** [January 12, 2026, 4:29am UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/1 "2026-01-12T04:29:13Z")

</div>

Among the most popular programming languages, Julia is one of the most token-efficient programming languages​.

 ![token-efficiency-chart](https://global.discourse-cdn.com/julialang/original/3X/2/b/2b2e8999de55530ec98bcc9db1e61c2cf21d9ef8.png)

> **[Which programming languages are most token-efficient?](https://martinalderson.com/posts/which-programming-languages-are-most-token-efficient/)**
>
> Comparing token efficiency across 19 popular programming languages using RosettaCode data - from Clojure to C, there's a 2.6x difference.

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [January 12, 2026, 3:13pm UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/2 "2026-01-12T15:13:16Z")

</div>

awesome! I instantly added this to “Why Julia - A Manifesto!”

---

<div class="post-metadata">

**Author:** ![xiaoxi](https://avatars.discourse-cdn.com/v4/letter/x/a9adbd/32.png) [@xiaoxi](https://discourse.julialang.org/u/xiaoxi)\
**Post date:** [January 13, 2026, 7:36pm UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/3 "2026-01-13T19:36:49Z")

</div>

> [@Datseris](#):
>
> “Why Julia - A Manifesto!”

Is there any possibility that it will be published on Julia-lang website?

---

<div class="post-metadata">

**Author:** ![greatpet](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Post date:** [January 13, 2026, 9:31pm UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/4 "2026-01-13T21:31:04Z")

</div>

> [@xiaoxi](#):
>
> Is there any possibility that it will be published on Julia-lang website?

The “official” blog collection is here:

> **[The Julia Language Blog](https://julialang.org/blog/)**
>
> The official website for the Julia Language. Julia is a language that is fast, dynamic, easy to use, and open source. Click here to learn more.

There are just a dozen blog entries per year. Is there a way to nominate a blog article to appear there?

---

<div class="post-metadata">

**Author:** ![sairus7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sairus7/32/10816_2.png) [@sairus7](https://discourse.julialang.org/u/sairus7)\
**Post date:** [January 14, 2026, 1:32pm UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/5 "2026-01-14T13:32:37Z")

</div>

Isn’t that result the same as code length? Just counted in tokens instead of letters?

---

<div class="post-metadata">

**Author:** ![Tortar](https://avatars.discourse-cdn.com/v4/letter/t/6bbea6/32.png) [@Tortar](https://discourse.julialang.org/u/Tortar)\
**Post date:** [January 14, 2026, 2:43pm UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/6 "2026-01-14T14:43:08Z")

</div>

Probably yes, the two measure should be highly correlated I suppose, maybe `end` in Julia is a single token often, which could explain why Julia does so well

---

<div class="post-metadata">

**Author:** ![ffevotte](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ffevotte/32/6587_2.png) [@ffevotte](https://discourse.julialang.org/u/ffevotte)\
**Post date:** [January 14, 2026, 3:07pm UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/7 "2026-01-14T15:07:52Z")

</div>

Yes, but it looks like _token_ should be taken to mean a very specific thing in this LLM-related context. (Definitely not what I expected from my mental model related to lexing & parsing)

From [https://christophergs.com/blog/understanding-llm-tokenization:](https://christophergs.com/blog/understanding-llm-tokenization:)

> In the machine learning context, a token is typically not a word. It could be a smaller unit, like a character or a part of a word, or a larger one like a whole phrase. The size of the tokens vary from one tokenization approach to another, as we will see in this post.

---

<div class="post-metadata">

**Author:** ![Diego\_Javier\_Zea](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/diego_javier_zea/32/1858_2.png) [@Diego\_Javier\_Zea](https://discourse.julialang.org/u/Diego_Javier_Zea)\
**Post date:** [January 16, 2026, 6:51am UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/8 "2026-01-16T06:51:40Z")

</div>

You can enter some code here to see how it’s tokenized (although this is probably a different tokenizer): [https://platform.openai.com/tokenizer](https://platform.openai.com/tokenizer)

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [January 22, 2026, 12:34pm UTC](https://discourse.julialang.org/t/julia-is-one-of-the-most-token-efficient-programming-languages/134996/9 "2026-01-22T12:34:09Z")

</div>

Vectorized language(s) J (and APL) are even better. Leads me to believe Julia could also be better. You can code vectorize or not (loopy), even mix. And maybe generating Julia defaults to non-vectorized. Possibly we should ask for vectorized by default, or at least let the AI consider it.

> **Update:** A lot of people asked about APL. I reran on a smaller set of like-for-like coding tasks - it came 4th at 110 tokens. Turns out APL’s famous terseness isn’t a plus for LLMs: the tokenizer is badly optimised for its symbol set, so all those unique glyphs (⍳, ⍴, ⌽, etc.) end up as multiple tokens each.

> **Update 2:** A reader reached out about J - a language I’d never heard of. It’s an array language like APL but uses ASCII instead of special symbols. It dominates at just 70 tokens average, nearly half of Clojure (109 tokens). Array languages can be extremely token-efficient when they avoid exotic symbol sets. If token efficiency turns out to be a key driver, this is perhaps a very interesting way for languages to evolve.

[For those that don’t know, J is an APL-variant; I’m not sure what the “unique glyphs (⍳, ⍴, ⌽, etc.)” map the there exactly, I assume more than a single ASCII letter. In APL code-page those are I believe only one byte, in Unicode more; still just one token? It seems J might might actually need more tokens for final code? Maybe the reatively less efficient APL is just about not finetuning on it more? I still guess J is more or much finetined on, or a higher-resource language?]

> Unsurprisingly, dynamic languages were much more token efficient (not having to declare _any_ types saves a lot of tokens) - though JavaScript was the most verbose of the dynamic languages analysed.

I tended to believe JavaScript was one of the best languages for genAI, likely because a high-resource language, like Python; and also just much used (many do not even consider there being any alternatives, for web use). This test of token efficient, is though maybe too limited ignores library/framework use that I guess the AI knows about and exploits e.g. for web use. Also this is only about completing to a correct solution, not taking runtime into account, so there are also other metrics.

I would also like to know how well another APL variant, BQN, works. And Forth and other concatenative languages, like Factor and Kitten. Anyone have a good feeling or bad on using those? AI thinks such Concatenative/tacit are genAI hostile, but not APL tacit languages…
