A closed model from an ex-OpenAI founder, and open rebuilds within a week
On 15 September 2026, TypeSafe AI launched Jev, a model that does not chat. It takes the state a program already holds and a set of typed questions, and returns a probability for every allowed answer in a single call. Founder Diogo Almeida introduced it in a launch post as the first “System One” model, and wrote that at OpenAI he “helped build the methods” behind ChatGPT. Jev shipped closed, as an early-access API, and within a week it was also reachable through third-party gateways including OpenRouter and Vercel’s, according to a roundup by the gateway company Requesty.
The next day, an open imitation was already on Hugging Face. Nine days after launch, an independent benchmark ranked an open 4B rebuild above Jev on its overall score. It did not out-reason Jev. It was faster and cheaper, and on the same benchmark it came close enough for that to decide the ranking.
Almost all of them run on a card you own, which raises the question this page ends on: what a decision model costs when it shares the card with the model doing the thinking.
What Jev is, in TypeSafe’s words
TypeSafe describes two kinds of model. Chat models generate text one token at a time. A System One model reads a state and answers questions whose possible answers are fixed in advance: yes or no, a choice between options, or a score. TypeSafe says Jev answers all of them in a single query and cannot make a type error, because it only ever returns one of the answers you defined, each with a calibrated probability.
The claims on the launch post, all TypeSafe’s own:
| Claim | TypeSafe’s figure |
|---|---|
| End-to-end response time | 70 to 500 ms |
| Speed against frontier models on System One tasks | 40x to 200x faster |
| Input price | $0.042 per million tokens |
| Output price | free |
| Training method | Reinforcement Learning for Calibrated Decisions (RLCD) |
TypeSafe is candid about its evidence: the launch post calls the 193.6x and 444.6x figures on its home page “on the higher end of real world gains”, and notes that its own team wrote the evaluation workflows. The uses it lists are routing, classification, scoring and guardrails: “smart if-statements” inside ordinary software.
The week that followed
At 05:36 UTC on 16 September, the day after launch, Almeida replied on X that open-source models “would have a hard time being competitive because the bottleneck for a new task/north star is data and not an architecture”. In the same post he added that competitive open models would most likely “come from us or people like us”, as TypeSafe makes “even smaller models that could run locally”. At 07:59 UTC, less than two and a half hours later, the first rebuild in the table below was on Hugging Face.
Each date below is the day the repository appeared on Hugging Face, and each claim is the author’s own.
| Date | Model | Built on | What its author reports |
|---|---|---|---|
| 16 Sep | decider-2b, Mapika | Qwen3.5-2B-Base | “An open reproduction” of the System One model class, up by 08:00 UTC; the first of the family that produced decider-4b |
| 16 Sep | System One scorer, pngwn | Qwen3.5-4B-Base | One forward pass, 112.3 ms per question at 4 options on an A100. Non-commercial licence, inherited from its data |
| 17 Sep | kev, Jared Palmer | Qwen3.5 and Qwen3.8 | “A Jev-like family of decision models” you can train and run yourself; about 7,200 GitHub stars by 26 Sep |
| 18 Sep | open-jev, Kotoba Labs | DeBERTa-v3-large, 434M | Public gold labels only, no teacher model; 28 ms for 10 questions on an H100 |
| 18 Sep | Laya, ConvAI Innovations | ModernBERT-large, 421M | 32.8 ms median per question on a Tesla T4, against published figures of 236 to 276 ms for Jev that it did not measure itself; higher accuracy on its own decision set, Jev leads on Banking77; about 25,500 GitHub stars by 26 Sep |
| 20 Sep | Winnow-12B | Gemma 4 12B IT | Ties Jev at 85.71% on the 231 public JevBench items its card tests, at Q8; tested with 64K context and vision on a 16 GB RTX 5070 Ti |
| 20 Sep | OpenThai-SystemOne, iApp | Qwen3.5-0.8B-Base | Thai and English decisions, about 40 ms on an H100, 154 ms on an M3 Max |
| 21 Sep | Jev-Style 2B | Qwen3.5-2B-Base | 77 ms per decision on an M1 Max; a 0.8B v3 followed on 24 Sep |
| 21 Sep | CLM-v0.1-8B, Contrastive-LM | frozen Qwen3-8B plus a 20M head | “On par with Jev” on computer-use, gaming and tool-calling tasks, “up to 9x lower latency”; code published 23 Sep |
| 22 Sep | decider-4b, Mapika | Qwen3.5-4B-Base | “Nothing was distilled from Jev”; trained on public data and labels from local Qwen 27B models |
| 22 Sep | JevK5 | Qwen3.5-4B | About 13 ms per decision on an H100 with its own runtime; 2B and 9B versions followed |
Most are small, from 421 million to 12 billion parameters, and most start from Qwen3.5. None names Jev as a teacher: decider-4b says nothing was distilled from Jev, open-jev uses no teacher, and JevK5 names Qwen3.6-27B, run locally, and OpenAI’s GPT-6 Luna through its API. Winnow-12B keeps its training data private.
Did the open ones beat it?
On one benchmark’s overall score, yes. On reasoning, not yet.
JevBench is run by Benchmark Heaven, which describes itself as an open-source hobby project under the MIT licence. Its authors also run jev-router.com, a service for self-hosted open decision models; the page discloses this and says the service gets no scoring advantage and is not a ranked entrant. Version 1.4.2, scored on 24 September, lists 93 systems and ranks 89 of them on four axes, Intelligence, Calibration, Speed and Cost, combined with equal weight in a harmonic mean, so one weak axis drags the whole score down.
| JevBench v1.4.2 | Official score | Intelligence | Estimated cost per 1,000 decisions |
|---|---|---|---|
| decider-4b v2 (open) | 64.1 | 49.4 | $0.020 |
| Jev 1.13.0 (closed) | 63.3 | 53.1 | $0.040 |
| JevK5 v0.2.0 (open) | 62.0 | 48.9 | $0.022 |
| Cygnet (open) | 61.8 | 49.5 | $0.037 |
Jev’s cost is its list price; the others are the benchmark’s estimates.
The page says it plainly: Jev out-reasons decider-4b v2 and is better calibrated, while decider-4b v2 leads on speed and cost. On the page’s separate Capability ranking, which leaves speed and cost out, Jev is first.
JevBench also attaches a note to its winner. decider-4b’s author disclosed that 8,000 of the v2 training rows came from generators written from the published names of the benchmark’s ten sealed question families, without reading any item. JevBench checked the result and marked it “LEGIT” on 24 September, and says the author’s private training rows could not be audited for overlap.
Sealed questions show the gap most honestly. On the 308 that no entrant has seen, Jev scores 36.7% and decider-4b v2 34.7%, against 29.3% for guessing. On the public questions the two score 86.6% and 83.5%.
TypeSafe does not accept the scoreboard. The show notes of the Latent Space podcast episode with Almeida, published on 21 September, list a generic JevBench benchmark among things he “has rejected publicly”, and one chapter is titled “Why TypeSafe rejects public benchmarks”.
A second, larger scoreboard reads the same way. On 22 September apolinario (@multimodalart) published the Decision Index 0.1, which Hugging Face reposted: 30 or more open decision models against Jev, 35 benchmarks, 130,000 questions each. His summary in the launch thread: “jev still leads open models, with some margin in case of overall knowledge (probably a bigger model), however, on tool use/automation/retrieval and classification, it’s close!” And JevBench’s author, Florian S, wrote on 24 September that he is considering giving intelligence more weight in the next version, which could reorder the top.
So a closed model was matched on JevBench’s overall score within nine days, by independent developers with 4B open base models, on a benchmark its maker rejects. The reasoning gap is still there, and it is small.
How r/LocalLLaMA took it
Not quietly. The thread that launched Laya on 17 September was titled “I literally built the Jev architecture one year back and completely open-sourced it“. The public record is shorter: Laya’s GitHub repository was created on 18 September, and its first commit, “Initial commit for Laya package”, is from that day. By 26 September it had about 25,500 stars, and ports to Rust, ONNX and a vision model had appeared around it.
Within a week the forum was pushing back on the noise. On 23 September one thread asked the moderators to act on “half the forum getting filled with these advertising posts for Jev” On 24 September the next contender arrived under the title “JEV almost dead: CLM vs JEV“. In the comments, verdicts on which model is actually better split in every direction, which matches what the two scoreboards show: close, and not settled.
System 1 beside System 2, on one card
The practical idea behind all of this is old: a fast model decides, a slow model thinks. Almeida uses the same frame for Jev: on launch day he wrote that it could take on open-ended tasks “in the right harness (i.e. code/workflow is the system 2)”. It is already a product: on 25 September OpenRouter introduced typesafe/jev-router, “a cache-aware model router powered by Jev” that picks the model and the reasoning effort for each request. Locally, both have to fit on the same card, and the decider’s memory comes straight out of the reasoner’s context window.
We sized it with our engine. Each decider is its real GGUF file from its repository, running with an 8K context and an F16 cache. The System 2 model keeps a Q8_0 cache, and each model runs as its own llama.cpp server, so each carries its own overhead. We test fixed steps from 4K to 256K, and each cell is the largest step that fits. Context is in tokens, with K = 1,024.
| Card | System 2 model | Alone | + 0.8B decider, Q8_0 | + 4B decider, Q4_K_M | + 4B decider, Q8_0 |
|---|---|---|---|---|---|
| RTX 3060 12 GB | Qwen3.5 9B, Q4_K_M | 256K | 128K | 64K | under 4K |
| RTX 4060 Ti 16 GB | Qwen3.5 9B, Q6_K | 256K | 256K | 128K | 64K |
| RTX 4060 Ti 16 GB | Qwen3.8 27B, Q3_K_M | 32K | 4K | under 4K | under 4K |
| RTX 4090 24 GB | Qwen3.8 27B, Q4_K_M | 128K | 128K | 64K | 16K |
| RTX 5090 32 GB | Qwen3.8 27B, Q6_K | 192K | 192K | 128K | 96K |
CLM sits outside the table by design: its decider is a full Qwen3-8B encoder, so it costs a card what an 8B model costs, far more than any decider above.
The 4B decider files are 2.71 GB at Q4_K_M and 4.48 GB at Q8_0. The 9B and 12B deciders leave no usable room for a 27B on any of these cards.
Three readings of the table:
- A 0.8B decider is almost free. On a 24 GB or 32 GB card the reasoner keeps all the context it had alone.
- The decider’s quant matters more than its size. On a 4090, the same 4B decider takes the 27B from 128K to 64K at Q4_K_M, and to 16K at Q8_0.
- On 16 GB, pair with a 9B. Qwen3.8 27B needs Q3_K_M or lower to fit a 16 GB card at all, and a decider beside it leaves almost nothing. A Qwen3.5 9B keeps all 256K next to a 0.8B decider.
Winnow-12B points the other way: its card sells decisions, chat and vision from one set of weights, one model doing both jobs on a 16 GB card.
The Decision Index points the same way from the top end. The open entries right behind Jev there, per apolinario, are “inference techniques that allow for one-pass inference of Qwen3.8-27B and diffusion gemma, no fine tuning”. If a card already runs Qwen3.8 27B as its System 2, one-pass scoring on the same weights is a decider with no second model to fit.
Speed: a decision is mostly reading
A decision model generates almost nothing. It reads the state and the questions and scores the options in one pass, so its speed is prompt processing, not the token-by-token writing speed most benchmarks quote. The authors’ own figures, on their own hardware:
| Model | Reported time per decision | Hardware |
|---|---|---|
| JevK5 | about 13 ms | H100, its own runtime |
| Laya | 32.8 ms median | Tesla T4 |
| OpenThai-SystemOne | about 40 ms, 154 ms | H100, M3 Max |
| Jev-Style 2B | 77 ms | M1 Max, MLX |
| Jev, hosted | 70 to 500 ms end to end | TypeSafe’s API, network included |
The hosted figure includes a round trip over the internet and the local ones do not, which is the real choice between an API and a card. Our tokens per second checker estimates generation only, and its table of measured runs lists prefill separately, because the two behave nothing alike.
What to run
- 12 GB: a 0.8B decider takes a Qwen3.5 9B at Q4_K_M from 256K to 128K, and a 4B at Q4_K_M leaves it 64K.
- 16 GB: a 4B decider at Q8_0 plus Qwen3.5 9B at Q6_K keeps 64K of context. See what else fits with the calculator.
- 24 GB and up: with a Q8_0 cache, a 4B decider at Q4_K_M beside Qwen3.8 27B keeps 64K on a 4090 and 128K on a 5090.
Check the runtime before the file. decider-4b’s reference runtime is transformers, JevK5 ships its own, and Winnow-12B ships a llama.cpp-based server. The GGUF builds run in llama.cpp, but the typed-question interface each author built may not come with them.
FAQ
What is Jev?
Jev is TypeSafe AI’s first “System One” model, launched on 15 September 2026. Instead of generating text, it reads a state and returns a calibrated probability for each allowed answer to typed questions: yes or no, a choice, or a score. It is available as a closed API, from TypeSafe and through third-party gateways.
Is there an open-source Jev?
Several. decider-4b, JevK5, Winnow-12B, Jev-Style, Laya, open-jev and OpenThai-SystemOne all appeared on Hugging Face within ten days of Jev’s launch, and all seven are released under Apache 2.0. Jared Palmer’s kev and Contrastive-LM’s CLM followed the same pattern. None of them is Jev’s weights: they are independent models built for the same kind of task.
Did an open model beat Jev?
On JevBench v1.4.2’s official score, decider-4b v2 ranks first at 64.1 against Jev’s 63.3, because it is faster and cheaper. Jev still leads on the benchmark’s Intelligence axis, 53.1 against 49.4.
Can I run a Jev-style model next to a big model on one GPU?
Yes, and the decider’s quant decides how much context the big model keeps. On a 24 GB card with Qwen3.8 27B, a 4B decider at Q4_K_M leaves 64K of context, the same decider at Q8_0 leaves 16K, and a 0.8B decider leaves all 128K it had alone.
How much VRAM does an open Jev alternative need?
The weights run from 0.81 GB for a 0.8B decider at Q8_0 to 12.67 GB for Winnow-12B at Q8_0. The popular 4B ones, decider-4b and JevK5, are 2.71 GB at Q4_K_M and 4.48 GB at Q8_0.