Ollama adds local decision models for routing and moderation
The September 29th release brings Bespoke Labs' Nimble and two Together AI models to Ollama's new /v1/systemone API, which returns typed answers instead of generated prose.
By Ryan Merket · Published
Primary source: Ollama on X
Why it matters
Ollama is bringing typed classification and scoring into its local model workflow. That gives developers a way to test lower-latency, on-device decisions, while the published moderation results show why teams still need their own evaluations and human review thresholds.

Ollama's September 29th announcement adds locally run decision models to its model runner, giving developers an API for classification and routing tasks that returns constrained answers instead of chat text. The new /v1/systemone endpoint is available in Ollama 0.35 and accepts text or structured data as a state, plus named questions that request a choice, a yes-or-no result, or a score.
https://x.com/ollama/status/2105152056382345544
The release extends the kind of developer tool co-founder Jeffrey Morgan has built before. Morgan and co-founder Michael Chiang met at the University of Waterloo and started Kitematic, a graphical tool for running Docker containers. Docker acquired Kitematic in 2015; Morgan and Chiang say their work there became part of Docker Desktop. They later founded Ollama, which launched publicly in 2023 to make running open models on personal computers simpler.
That history helps explain the product decision: Ollama is making a model capability available through the local workflow developers already use, rather than asking them to send every classification request to a separate hosted service. A developer can send a support ticket and ask which team should handle it, whether it requests a refund, and how urgent it is. The endpoint returns those answers in defined fields, with probabilities for the allowed choices. Software can branch on the returned values without parsing a paragraph of generated text.

Ollama lists three models for the endpoint: Nimble, a 9-billion-parameter model from Bespoke Labs, and two experimental models from Together AI, Tev1 at 4 billion parameters and a 0.8-billion-parameter version. Nimble is fine-tuned from Qwen3.5-9B. Its model page says it scores answer tokens directly rather than generating a reasoning trace, and supports up to 64 questions about the same input in one call. The listed model is Apache 2.0 licensed.
Ollama reports that Nimble averaged 91 milliseconds per decision in a Pac-Man demonstration on a MacBook Pro with an M5 Max. That is a result from one model on one named machine, not a general speed guarantee. The local route avoids network transit and per-request charges for hosted inference, but running it still requires a computer with enough memory and processing capacity. Ollama's listing shows a 9.5 GB default download for Nimble, with a smaller 5.6 GB quantized option.
The benchmark evidence also puts limits on the use cases Ollama highlights. A comparison published by Bespoke Labs and reproduced on Ollama's Nimble model page covers 3,880 examples across 13 public datasets. Nimble's macro-average accuracy was 74.8%, compared with 76.0% for Jev, the hosted decision model from TypeSafe whose API shape the Ollama endpoint follows. On the Civil Comments moderation dataset, Nimble's reported agreement with human labels was 70.3%, while Jev scored 81.0%. Those figures come from the model maker's evaluation, and they are agreement with labels in those datasets, not proof of performance on a company's own moderation queue.
That qualification matters for systems that automatically block content, route high-priority customer cases, or make safety decisions. Nimble returns probabilities and a confidence measure, but Ollama's documentation cautions that confidence reflects how concentrated the model's choice is; it is not a promise that the answer is correct. Developers still need to test their own inputs and decide when low-confidence cases should go to a person. The decision capability documentation describes the supported answer types and API behavior.
Ollama is adding the feature after raising $65 million in a Series B led by Theory Ventures in July. The founders' announcement said the financing brought total funding to $88 million and named Benchmark, 8VC, Y Combinator and other investors. The decision API does not, by itself, establish a new revenue line: the launch post frames local use as having no additional inference cost, and gives no separate price for the endpoint. Its immediate bet is distribution. If developers can use Ollama for narrow decisions as well as text generation, the software can sit in more application workflows, where low latency and keeping inputs on-device are practical advantages.
For now, Ollama describes the launch as an initial step and says it plans to add more decision models, including cloud-served options. That would broaden the same API from a local machine to hosted inference; the present release centers on local execution and three models. The technical premise is straightforward: many application steps need a reliable category, score or gate, not another block of prose. Whether that premise holds in production depends on task-specific accuracy, which the published benchmark leaves developers to establish for themselves.