2026-10-11 18:31 UTC
DANGMUAAI & Developer Tools, Decoded
BackAI Models

Jev Bets $870M That Agents Shouldn't Ask an LLM to Decide

TypeSafe AI raised $870M at a $7.5B valuation for Jev, a model that returns typed decisions and probabilities instead of text. Where it fits in an agent graph.

DangMua EditorialOct 11, 20263 min read
Jev Bets $870M That Agents Shouldn't Ask an LLM to Decide

TypeSafe AI raised $870 million at a $7.5 billion valuation for a model that never writes a sentence. Its model, Jev, returns only classifications and probabilities.

The round was led by a16z, according to the report on the raise, and the company was founded by ex-OpenAI researcher Diogo Almeida. The pitch is narrow on purpose: the decisions inside an agent loop do not need a chat model.

What Jev returns instead of text

Jev takes a state plus a set of typed questions and answers them with calibrated probabilities, generating no text at all. It exposes three question types: Choice picks from predefined options, Score rates an ordered scale, and Noul returns a probability for a yes/no statement.

That last one is the useful primitive. A Noul answer is a value between 0 and 1 representing the probability that the statement is true, which application code can threshold however it wants rather than parsing out of a paragraph.

Where it slots into an agent graph

A public LangGraph.js example, JevSample, shows the shape. Its graph wires up nodes named router, generator, prefilter, safety, dispatch, human_review, executor and verifier — a state machine, not a linear pipeline.

Two of those nodes are where the typed model earns its place. The router asks a Choice question whose options are already the graph's transitions, so the dispatcher switches on an explicit value instead of interpreting model prose. The safety node runs after a deterministic prefilter and turns a Noul probability into policy:

if (safety.noul >= 0.9) { return "execute"; }
if (safety.noul >= 0.6) { return "review"; }
return "blocked";

The sample's author draws the line clearly: Jev provides the judgment and does not execute anything, and the application stays responsible for what the judgment means operationally. Null checks and length limits stay in ordinary code — there is no reason to spend a model call on if (!input).

The price argument

Jev is priced at $0.042 per million input tokens with free output, which is the whole economic case: micro-decisions you would never route through a chat model become cheap enough to automate at high frequency. The report on the raise says Jev launched September 15 and credits it with pushing OpenAI to ship a competing Decisions API within two weeks, citing an acknowledgment from OpenAI's API lead — that account comes from a single write-up, not a confirmed statement.

An open-weights challenger, by its own claim

A separate team, VIDRAFT, is pitching the open version of the same idea. In its own post it claims its model Darwin-27B-ZTC-v2 ranks number one of 102 models on the System One Mosaic Benchmark decision-engine leaderboard, with a Borda score of 89.58 and a task average of 66.46, released under Apache-2.0.

The method it describes, Zero-Token Confidence, reads the problem in one forward pass, takes the final-layer hidden state and applies a calibrated probe to produce the decision — nothing is sampled, so the same input always yields the same decision. Treat the ranking as a vendor claim on a vendor-adjacent benchmark until someone outside the team reproduces it.

Worth trying or not

If your agent calls an LLM to answer "which branch" or "is this safe," you are paying generation prices for a classification, and you are parsing text that can drift between runs. A typed decision call removes both problems at the routing layer while leaving generation where it belongs.

What to watch next: whether OpenAI's Decisions API prices anywhere near $0.042 per million input tokens, and whether an independent run of the S1MB leaderboard backs the open challenger's numbers.

More from DangMua