Small Decision Models Go Local

A lot of what we ask AI agents to do is make small calls: which team gets this ticket, which model should answer this prompt, is this tool call safe to run. Those decisions work best with a quick, typed answer and a probability attached, and a small model can often handle them well. On September 28, Ollama made that much easier to do on your own machine, and I’m genuinely excited about it.

A laptop running several tiny glowing decision nodes that sort incoming tickets into labeled lanes, feeding a larger AI agent in the background

What Ollama added

Ollama v0.35.0 introduces decision models through a new /v1/systemone endpoint, based on TypeSafe’s Jev API. In the release notes’ words, decision models “return choices, probabilities, and scores instead of text.” Ollama suggests using them for ticket triage, model routing and content classification.

There are three question types. A choice question picks an option and returns a probability for each one. A noul question returns the probability that a condition is true. A score question places something on an ordered rubric. You put the text you want judged in state, name your questions, and get structured answers back.

I wrote about this style of typed decision model when I looked at OpenThai-SystemOne on September 23. What’s new here is that it now lives inside a runtime a lot of you already have installed.

The launch models

Two families launched alongside the endpoint. Nimble from Bespoke Labs is a 9B decision model fine-tuned from Qwen3.5-9B. Tev1 from Together AI comes in 4B and 0.8B sizes, both fine-tuned from Qwen3.5, and Together describes the family as experimental. The 0.8B version is an 812 MB download, small enough for almost any laptop.

Ollama’s pages share one comparison: mean accuracy across Bespoke Labs’ 13 public datasets with human labels, 3,880 decisions in all. Nimble 9B scores 75.7%, Tev1 4B 73.3%, and Tev1 0.8B 63.5%. For reference, TypeSafe’s hosted Jev 1.13 scores 76.0% on Bespoke’s own run of the same decisions. Those are vendor and partner numbers, so please treat them as a starting point.

The edge pattern I like

Here’s why I’m excited. The shape that keeps showing up is lots of cheap local judgments feeding a bigger agent. A tiny model decides whether a message is billing or a bug, whether a request needs the large model at all, or whether a proposed shell command looks harmful. Only the interesting cases travel further up the chain. That saves money and keeps more of your data on your own hardware.

Samsung described a similar design at its AI Forum. According to Edaily’s coverage, Samsung’s Personal Data Engine uses small on-device models to turn phone, sensor and wearable data into structured knowledge, and pipelines a signal-filtering small model into an on-device LLM for deeper analysis. Samsung’s Lee Yoon-soo said the company adheres “strictly to 100% on-device processing” for that personalization data. Samsung didn’t share model sizes, but the shape is familiar.

Quantization is the other half of this story. Fermion Research’s Phonon-2 is a 164 MB English speech recognition model built from NVIDIA’s 2,508 MB Parakeet TDT 0.6B v3, with its encoder at about 2.1 bits per weight. Fermion reports a 5.21% average word error rate against 4.96% for the original, and about 174x realtime on an M5 MacBook Air. Those are Fermion’s own numbers, and they show how small a genuinely useful model can get.

Reasons to stay measured

These are narrow models, and nobody has benchmarked them independently yet. The base models also predate the Ollama integration. Nimble’s original release came out earlier in September, and Bespoke’s own repository says “We built Nimble in one day, so expect some rough edges.” Together’s model card notes that calibration hasn’t been comprehensively evaluated, and Ollama’s pages remind you that a confidence score “isn’t the chance that the answer is right.”

Licenses are shifting too. Ollama lists its Nimble package as Apache 2.0, but the new Bespoke-Nimble-9B-v3, posted September 29, is CC BY-NC 4.0 and can’t be used commercially. Together’s Hugging Face card says the license for Tev1’s fine-tuned weights “is being finalized.” Check the terms for the exact version you pull before you build a product on it.

Finally, decision models are available through the API only for now. They aren’t in the Ollama CLI or the Ollama Python and JavaScript libraries yet.

Try it today

  1. Update Ollama to 0.35 or later.
  2. Pull the small one with ollama pull tev1:0.8b.
  3. Send it a ticket. This follows the shape in Ollama’s release notes:
curl http://localhost:11434/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "tev1:0.8b",
    "state": "Our checkout has returned 500 errors since 9am.",
    "questions": {
      "label": {
        "type": "choice",
        "instructions": "Which label fits this ticket?",
        "criteria": {
          "billing": "Payments and refunds",
          "bug": "Software errors",
          "account": "Login and account access"
        }
      }
    }
  }'
  1. Look at the probabilities, then try a few of your own real messages and add a none option in case nothing fits.
  2. Run the same questions against tev1:4b or nimble and compare. Pick thresholds from your own data.

Start with one small decision you make a hundred times a day, and let a local model take a first pass at it. I think you’ll be surprised how much it helps.