Possible applications of Jev (ai model)

Hey everyone,

so a few days ago a new AI model was released under the name of Jev. I signed up to the beta, because I thought it was an interesting new take on the entire market and the reason why I am posting this at all (this is not propaganda, I am merely a user).

What they offer is not a frontier model, that’s not the purpose. Instead Jev is a low-cost & highly specialized model that only does one thing: decisions. You feed it text context and it will either a) answer your yes/no question b) select from a list of options (categorize) c) score the result.

It is incredibly fast for an llm (nowhere near functional designs, but much faster than your usual model) and only costs USD 0.04 for 1 Mio input tokens - the output tokens are free.

To me this opens up many many options that were difficult to solve with LLMs in the past and something I believe also fits into the Julia world. Since I got access to it, I also wrote up a quick article with my first impressions that you can find on substack: My Thoughts on Jev, after the Private Beta - Paul Piper

Anyway, would love to hear your thoughts. But I am already going over potential applications in my head. It is defintely the right step in the right direction, as I hated the idea of being dependend on overly complex and heavy frontier models…

It’s fast, but not a large language model, LLM (i.e. autoregressive, based on RLHF), at least as usually understood, GenAI. I understand it to takes text in yes, and give only good decisions out, i.e. it’s for reliable automation. From their FAQ:

Is Jev just a smaller LLM?

Jev is neither small nor an LLM, hence being off the intelligence Pareto curve. Jev’s efficiency comes from optimizing for a different task. It’s built for structured decisions inside software, with an interface and training approach designed for that purpose. It understands language, but doesn’t generate free-form text or function as a chatbot.

I liked the talk from the founder, co-created or ChatGPT and one of the main researchers behind reinforcement learning with human feedback (RLHF), ChatGPT and other LLMs are about. He goes into it and the flaws and what he’s now did better.

I only learned of Jev yesterday from this video:

I recommend watching all of it (see does good videos), and the claim at the timestamp I point to. The founder is clearly not Sam Altman, but a researcher, and I like him and his talk, but it was clear he didn’t want to go into some details then pre-launch, so her talk may be more practically informative.

Important distinction and yes, that is accurate- It does not chat. It merely judges. Thanks for the clips too - interesting!

btw, can we appreciate that for once a woman is one of the core drivers? I think that’s awesome and doesnt happen often enough

That would be funny if we circled back to DAG style agents with conditional edges and flow control, basically what we had to do before tool calling got good. But of course, open ended agnetic loops have far fewer structural guarantees about “staying on the rails” so to speak. I’d be interested in seeing how the DAG approach works with a much more capable model like I suppose Jev is supposed to be, and how far this approach can scale up, how crazy we can get with the number of states, etc.

Seems like a positive direction to me.

EDIT: models like Jev should also be compute bound instead of memory bandwidth bound. It’s much easier / cheaper to scale up compute than it is memory bandwidth for modern accelerators IIRC. So from a cost standpoint, it makes sense it would be cheaper since prefill is compute bound instead of memory bandwidth bound like decoding.

Some months ago I did some RE (Reverse Engineering) tasks to build an USB Linux driver for fingerprints driver with the help of 2 agents in a collaboration (feeding each other in turn but under two differents OS)

When working on this task I noticed a very repetitive behavior of AI which is in fact inpired by a very old scientific paper Strong Inference method (Platt, 1964),

Maybe such tasks could be simplified and distributed to agent this way:

Task Tool Why?
Generate hypotheses LLM Good at brainstorming (e.g., “List possible roles for opcode 0xA3”).
Test hypotheses Jev Specialized in structured decisions (Yes/No, classification, scoring).
Rank hypotheses Jev Can score hypotheses (1-10) based on context (e.g., payload structure, device responses).
Refine explanations LLM Can synthesize results into human-readable insights.

Jev seems to be cheap and fast for validation, while LLMs are better for creative tasks.

My 2cts

I left off reading about Jev at this article, the table of use cases sounded reasonable:

I don’t have any personal ideas because I don’t have a use for automated fuzzy decisions right now. The marketing sounds too good to be true, and there’s not much said about the training data and the sort of decisions therein, so I’d wait and see how other people do with it if I did need it.

there is not anything said about the architecture, and being so close is not a good direction of progress in my opinion. Many think that it is a repackage of some known architecture which has been there for years, and most claims are exaggerated, the best take to me seems that it is maybe a version of fastino/gliner2.5-multi-v1 · Hugging Face which is indeed very nice and open, and in the BERT family. So Jev “success” could be really just marketing, and not giving any detail seems to indicate that this could be even more probable.

There are at least more projects similar (or claimed to be), though hard to know without a paper on Jev:

I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset.

The RLCD(unofficial, btw) I did is a policy-gradient reinforcement learning approach that kinda optimizes decision models against strictly proper scoring rules, ensuring maximum reward is achieved only when outputting true, mathematically calibrated probabilities.

NB: It can be run on low end PC as its a small 421M model, cheers

HF space to try: Laya Demo - a Hugging Face Space by convaiinnovations

I don’t know if this is comparable, at least it can also Doom too (like Jev and Laya can, but unclear if they too can play from pixels only):

openjev is Qwen3.5 turned into a jev model: a single cross-encoder that reads a premise and a hypothesis and answers with entailment, contradiction or neutral. That one primitive is enough to rerank answers, grade them against a reference, guard content, and play games in real time: hand it the game state and a few statements about it, and the argmax entailment is the move. Nothing is trained per task. ..

  • Doom straight from the pixels (first video): 10.4 kills per episode, twice the v1 model (5.2); random play gets 1.
  • Crafts an iron pickaxe from nothing in real Minecraft (second video)

Traditional LLMs have prefill compute bound (always (for sure true)? according to e.g. Prefill vs. Decode Bottlenecks: SRAM–Frequency Tradeoffs and the Memory-Bandwidth Ceiling), and decode phase memory-bound. And that latter decode phase, producing more costly tokens. While for Jev “Output tokens: FREE (too cheap to meter)” and I understand the decode stage absent, not just short output. Input is also very cheap and fast.

I note input for Jev is “Unstructured data (e.g. text) with an emphasis on structured program state.” (vs for LLMs, claimed “Unstructured data (e.g. text) with an emphasis on sequential messages.”):

[because of] training method we call Reinforcement Learning for Calibrated Decisions (RLCD).
..
While Jev gives up string generation, it’s optimized for structured outputs and can’t hallucinate.
..
Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.
Extraordinary claims require extraordinary evidence so see below for the receipts. :nail_polish:
..

Fun Demos

Perhaps the most exciting part of our work is enabling new use cases. We have a lot more to show you, but here are a couple of the team’s favorites:

Doom

We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI. The engineer behind it was worried about making 10 queries a second (which ends up costing ~$7/hour), but the rest of us agreed that was lower than expected! This is so fun we intend to not only release an in-depth walkthrough, but also host some events to hack on this.

[playing doom video demo]

We deliberately chose not to publish performance against public benchmarks. In fact, we plan to only have one-off evals when we make product updates.

Given that we are opening up a new frontier for models, we’re pushing for more useful best practices:

  • Put no weight on public benchmarks.

  • Encourage users to create their own evals for their use cases (System One tasks are much easier to evaluate).

  • Disclose the nuance in your evals.

  • De-emphasizing benchmarks even when you’re ahead.

See this blog post about our philosophy around optimizing for benchmarks.

We wouldn’t train on your data even if you asked us to (no offense). We do some pretty sophisticated stuff, but if you want to find out more, we’d have to hire you.

dev.to/gabrielanhaia/he-says-he-co-invented-chatgpt-his-new-ai-jev-wont-write-a-word-e3c

Noul is a yes/no question answered with the probability of true. The AI SDK calls it boolean.

[AI tells me, not a real English word and Phonetic Mashup: It sounds like a blend of “Null” (zero/void) and “Boolean” (true/false). It thinks null result is not one of the output options.]

Anyone actually use the Jev model yet, its only over the cloud and no way to run locally, I had to sign up for a wishlist as well and then build it all with Go, then spent way longer than I wanted to structure its outputs and so on because the model doesnt talk at all. Trying to use it as a pre-parser model for type saftey since this model is ridiculously fast it can possibly be part of a parsing pipeline or lsp server agent is my ideas.

Also they ask alot of questions in the sign-up and want you to describe your work and link your repo, was kinda odd but seems harmless.

Edit: Just and idea but a language parser is good at parsing its language but what if I want a LSP that load another languages info and helps convert the types to Julia as I write the ccall that would be really nice.

They are currently trying to replicate the jev model locally, I found a few repos Ill drop the links.

Official dev jev docs…

After alot more testing these system 1 models are like JIT agents, they cant hallucinate because you make its choices and why. It just hurrys up and chooses it when you want it like a JIT. meaning anything deterministic in its output can be used im sure.

I still have a lot of information about Jev that I haven’t caught up with, but I found a tool that looks useful, so I’m sharing it here.

No idea how rigorous this is, but here’s a collection of similar tools with estimated costs and quality: Jev does come in 1st overall, but just barely and not in every aspect.

JevBench by Benchmark Heaven — Jev-class model benchmark | Benchmark Heaven

So I’ve not spent very much time at all reading about this and it appears that they’ve had to pause signups as a result of the hype/demand, but is this just intended to be used as a general-purpose classification tool? And is the primary advantage that you don’t have to train it on your data for it to work well? I can see this being very useful where you don’t have any/enough training data to train your own classifier, but what are the other motivations for using something like this?

I may have enough training data to train my own LLM, too, but I won’t do that. Jev seems to be a classifier that’s trained on the entire Internet’s knowledge, and it’s not necessarily easy to replace it with a specialized model (assuming Jev lives up to its hype).

TypeSafe AI claims Jev is trained on only synthetic data. I assume that involved models trained on the internet in some chain, but they didn’t provide any details.

I believe the primary advantage is that people typically use mainstream generative transformers for processing natural language prompts, and that is fairly costly for the answer being multiple choice with measures of confidence. In the JevBench I linked earlier, GPT-6 Luna and DeepSeek 4.1 Flash blow Jev out of the water in terms of Intelligence (chance-corrected accuracy of answers) and Calibration (accuracy of confidences) in the 90s out of 100 versus Jev’s 53 & 76, but the respective cost per 1000 decisions at $0.127 and $0.59 compared to Jev’s $0.04 offset the performance so much in the benchmark that they’re way below in the rankings. Obviously different people would prefer different benchmarks and weights for their more specific expenses; if each correct decision made me $0.00021, I’d make more money with GPT-6 Luna.

This compiles with ggmlc and can be quried from localhost from julia through JSON and HTTP.

Call laya and point at the input and output files, also laya has a server with laya serve.

Also laya.cpp is a thing as well.

This model smokes jev for almost everything except larger context decisions. The cloud jev is faster, also im unsure of jevs actual size, but laya is really small. 1 decision head and a encoder.

This mostly depends on how well you setup the input and output files. There is also something with these models nobody is talking about and thats the nested harness stuff, how a system 1 can iterate through a loop and feedback and tool call with custom guardrails, Your writing the if statement style control flow guards so the agent can use that to call things dynamically as they happen. A agent that can chain and decide specific events within 100ms that you can determine without any retraining. Also its impact on the gpu(local Laya model) is almost non existant. Its not supposed to compete with gpt-astra thats a different type of model that can do jevs job but why when we engineer purpose built agents that can be better for what they do. Build the app with gpt-astra or whatever then embed the apps decisions with jev is the right mental model for this stuff.

A.

Julia 1 is the first model in the Julia family and the first released test of our training system. At 144.3M parameters, it turns a state, a question, and possible answers into one clear decision.

I posted this since a relevant similar model, beating Jev on their benchmarks (except one), but also because of the name.. just coincidence.

B.
The above is on-topic I guess, but this here not, brand new Context Language Models (maybe I should have forked the discussion, feel free to do that if answering), or only in general AI related. I think this might though be the next big thing. Apparently harnesses (like Claude Code) will get less important, though the current way to run is with Pi harness supporting:

$ pi install npm:@lolipopshock/pi-clm

We introduce Context Language Models (CLMs), language models that natively manage their own context. .. Building CLMs zero-shot with existing models outperforms SOTA context-management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies

Context-as-a-file implementation for CLMs. To implement CLMs, we mirror the LM’s live context as a directly editable file and provide its path in the system prompt. The LM can edit this file using general Bash commands, just as it would edit other files in storage. Unlike ordinary files, edits to the context file are automatically synchronized with the LM’s context and sent to the LLM server for continued generation. When the LM does not edit the context file, the generated tokens are appended to the existing context by default. This implementation balances context reuse with the flexibility to edit the context.

5.1.1 Coding and Deep Research Tasks
We first evaluate CLMs on two terminal-coding benchmarks, TerminalBench 2.1 (TB2.1) (Merrill et al., 2026) and TBLite (OpenThoughts-Agent team, 2026), and on the deep-research benchmark BrowseComp-Plus (BCP) (Chen et al., 2025). ..

CLMs outperform harness-defined and action-based baselines. On BCP, CLMs outperform all baselines

Different retrievers matter a lot: