Over the last couple of weeks I’ve had more questions about decision models than about almost anything else in AI. What are they? Are they just small LLMs? Which one should I use? So I sat down with every launch post, model card and leaderboard I could find, and this is the guide I wish I’d had on day one. We’ll cover how they work, how to read their confidence numbers, a close look at each option, and some practical advice on picking one for your own project.

What a decision model does
A decision model takes two things: some state, such as a support ticket, a block of JSON, or a proposed tool call, and one or more questions you define. For each question you also supply the possible answers. The model scores every answer you supplied and returns typed results with probabilities attached. Because it can only pick from your list, the output always fits your schema.
TypeSafe’s API, which most of the field now follows, defines three question types:
- Choice: pick one option from your list, with a probability for each option. Jev supports up to 255 options.
- Score: place something on an ordered scale, with a distribution across the levels.
- Noul: the probability that a stated condition is true.
That shape makes these models a natural fit for routing a request to the right model or queue, checking a message against a policy, deciding whether an agent should call a tool, triaging tickets, and cheaply judging another model’s output.
Two ways to build one
Under the hood, I see two main designs.
The first keeps a pretrained language model’s body and swaps its text output layer for a small head that points at one of the supplied options. AWS’s Strands Decider uses a pointer head of “just over a million total parameters” in place of the language model head. Cloudflare’s Clef models keep a frozen Qwen backbone and add a joint schema head that scores the valid choices in parallel. TypeSafe describes Jev as a new architecture with a parallel sampler, and it hasn’t published the details.
The second design keeps the standard next-token head and fine-tunes the model to answer with a single letter or code. The probability the model puts on each answer token becomes the score for that option. Together AI’s Tev1 and Bespoke Labs’ Nimble work this way. It’s a wonderfully accessible approach. Together’s walkthrough is titled “How to train your own Jev for $17”, and Bespoke’s repository says, “We built Nimble in one day, so expect some rough edges.” The quality of the probabilities depends heavily on the fine-tune.
Calibration in plain language
A decision model earns your trust when its “90 percent sure” really means it’s right about nine times out of ten. That property is called calibration, and two numbers show up everywhere.
The Brier score is the average squared gap between the probabilities the model gave and what actually happened. Lower is better, and zero is perfect. Confident wrong answers get penalized heavily. The exact scale depends on how a benchmark computes it, so compare Brier scores only within the same leaderboard.
Expected Calibration Error (ECE) sorts predictions into confidence buckets, such as 70 to 80 percent, and compares the average confidence in each bucket with how often the model was actually right. ECE is the weighted average of those gaps. The community Decision Index gives a lovely real example. Across 72,594 sampled decisions, Jev’s average confidence was 0.81 and its accuracy was 0.74, giving an ECE of 0.074. In other words, Jev runs about seven points overconfident on that board, which is useful to know when you set thresholds.
How the wave started
TypeSafe AI introduced System One Models and Jev on September 15. Founder Diogo Almeida called Jev “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” TypeSafe trains it with a method it calls Reinforcement Learning for Calibrated Decisions, prices input at $0.042 per million tokens with free output, and quotes end-to-end response times of 70 to 500 ms.
What happened next was remarkable. By October 2 the community Decision Index listed 70 open entrants, and Together AI, Bespoke Labs, AWS, Cloudflare and Perplexity had all released decision models. Ollama added a Jev-compatible /v1/systemone endpoint in v0.35.0, which I wrote about in my September 30 post. That shared API shape is a gift for builders, because it lets you swap models with very little code.
A close look at the options
Here’s the landscape, with the source of each number labeled. “Board” means the independent Decision Index, which runs every entrant on the same hardware. “Vendor” means the maker ran it.
| Model | Maker | Size | Head type | License | Runs | Notable number |
|---|---|---|---|---|---|---|
| Jev 1.13 | TypeSafe AI | Undisclosed | Own architecture | Closed | Hosted API | 57.91 board, ECE 0.074 |
| Tev1 4B / 0.8B | Together AI | 4B / 0.8B | Letter scoring | Weights license not yet stated | Together, Ollama, local | 29.24 / 12.85 board |
| Nimble v1, v2 | Bespoke Labs | 9B | Letter scoring | Apache-2.0 | Ollama, local | v2: 39.57 board, ECE 0.024 |
| Nimble v3 | Bespoke Labs | 9B | Letter scoring | CC BY-NC 4.0 | Local | 56.88 self-run |
| Strands Decider 2B | AWS | 2B | Pointer head | Apache-2.0 | Local CPU, GPU, Mac | 167/231 JevBench public (vendor) |
| Clef-Flash | Cloudflare | 9B | Schema head | Apache-2.0 | Workers AI, local | 38.8 ms median (vendor, hosted) |
| Clef | Cloudflare | 27B | Schema head | Apache-2.0 | Workers AI, local | 94.20 BANKING77 (vendor) |
| pplx-decider-v1-27b | Perplexity | 27B | Fine-tune | Apache-2.0 | Local | 56.40 board, ECE 0.018 |
Jev 1.13 is the reference point. It leads the Decision Index at 57.91 on a chance-corrected scale where zero is random and 100 is perfect. It’s hosted only, in early access, and the board measured a 524 ms median including the network round trip.
Tev1 from Together AI comes as 4B and 0.8B fine-tunes of Qwen3.5. The 0.8B version is an 812 MB download in Ollama. Together labels both as experimental, and its model card lists calibration among the areas that “have not been comprehensively evaluated.” The Hugging Face cards also carry no weights license tag yet.
Nimble from Bespoke Labs is a LoRA on Qwen3.5-9B. Versions 1 and 2 are Apache-2.0, and the v2 board score of 39.57 comes with excellent calibration. The new Nimble-9B-v3 moved to CC BY-NC 4.0, which rules out commercial use, so check your version carefully. Bespoke reports 56.88 for v3 from its own run of the Decision Index kit.
Strands Decider 2B from AWS is the one I’m most excited to run on a laptop. AWS reports a median of around 115 ms per decision on an RTX 3090 and around 153 ms for small tasks on an M3 MacBook. AWS calls it “a small, open source, decision model,” the card lists Apache-2.0, and a community ONNX export already exists.
Clef-Flash and Clef from Cloudflare are the multimodal pair. The Clef-Flash card says it reads state “as text, JSON, images, or video,” and Cloudflare quotes a 64K context. Cloudflare’s own table is honestly mixed: Clef-Flash scores 98.76 on BFCL and 93.11 on API-Bank, and 65.58 on When2Call against Jev’s 80.97.
pplx-decider-v1-27b from Perplexity is a Qwen3.8-27B fine-tune. On the board it’s listed with the engine name autojev-27b, and it scores 56.40 there with an ECE of 0.018, the best calibration among the leaders. The card asks for a CUDA GPU with room for about 49 GiB of weights.
A few more board entrants are worth knowing. Surogate Rune 26B-A4B v3 scores 57.44, the closest anyone has come to Jev on that board. Jebadiah 27B scores 54.67 with an ECE of 0.014. InternLM’s Intern-Decision family covers 4B, 2B and 0.8B, with the 0.8B model showing strong calibration at an ECE of 0.026. I haven’t checked the licenses for these, so please read their cards before you build on them.
Which one fits which job
- On-device or laptop triage: start with Strands Decider 2B. If you’d like Ollama’s ready-made endpoint today, try Tev1 4B or Nimble v2 there. Expect a real accuracy step down from Jev; on the board, the best 4B entries score in the low 40s.
- CPU-only routing: Strands Decider 2B is built to run on CPU as well as GPU, and the ONNX export opens up more runtimes.
- Multimodal guardrails: Clef-Flash for speed, or Clef for more accuracy, when your state includes screenshots, images or video.
- Highest accuracy, hosted: Jev 1.13 leads the independent board. Clef on Workers AI may belong here too once it has an independent score.
- Highest accuracy, open weights: Surogate Rune v3 or pplx-decider-v1-27b, with the Perplexity model offering the best-calibrated probabilities.
- Commercial-safe licensing: Strands Decider 2B, Clef, Clef-Flash, pplx-decider-v1-27b and Nimble v1 or v2 are Apache-2.0. Wait on Tev1 until Together states a weights license, and leave Nimble v3 for non-commercial work.
Honest caveats
Most numbers in this space come from the makers. The Decision Index is the only independent cross-model board I know of, and it hasn’t scored Strands Decider, Clef, Clef-Flash or Nimble v3 yet. Its latency figures also come from a single NVIDIA RTX PRO 6000 card, and Jev’s include a network round trip, so the board itself notes those aren’t comparable.
Watch out for the name “JevBench,” because three unrelated things use it. jevbench.dev ranks models by verified wins in games like StarCraft II. Benchmark Heaven runs a separate composite “JevBench” on Hugging Face. And the “JevBench public” set that AWS and Perplexity cite is a third thing, whose maintainer I couldn’t confirm.
There are no independent phone or NPU latency numbers for any of these models yet. Every figure I found comes from a desktop GPU, a Mac or a server-class card.
Finally, know the weak spots. InfoQ, summarizing TypeSafe’s own documentation, says Jev is unreliable at counting, arithmetic and date comparison, and recommends keeping that math in your code. That advice is sensible for every decision model on this list.
Try it yourself
- Pick one real decision you make often. Ticket routing, a policy check, or “should the agent call this tool?” are perfect starters.
- Collect 100 to 200 labeled examples from your own logs, with the right answers attached.
- Run two models side by side. Something like
pip install strands-deciderfor AWS’s model, andollama pull tev1:0.8bfor the smallest rung, gets you going in minutes. - Check calibration on your data. Group answers by confidence and see whether 80 percent really means 80 percent.
- Keep a bigger LLM for the hard cases. Send low-confidence decisions up the chain to a larger model or a person.
These models are small, fast and refreshingly honest about their uncertainty, and the field is moving at a lovely pace. Pick one, point it at a real decision, and have fun seeing what it can do for you.