Decision Models: Ollama and OpenAI Ship Jev APIs in 24 Hours

Ollama 0.35 adds Jev-style decision models locally the same day OpenAI previews its Decisions API for Luna. Same primitive, two stacks.

If you have ever watched an AI agent run, you have seen the same pattern. It spends most of its time answering questions that have a yes-or-no, a category, or a number: should I retry, is this a billing email or a tech-support email, how urgent is this ticket on a 0-to-2 scale. None of those need a paragraph of prose. They need a typed, calibrated answer, fast. Two AI labs - one open-weight, one frontier - shipped the same primitive this week to fix exactly that mismatch.

What a decision model actually is

TypeSafe AI calls the primitive a System One decision: “decisions, not strings.” You hand the model a state object and a set of named questions; it answers with typed fields and a confidence number, in a single round trip. The decisions come back in three shapes. choice returns a label plus a vector of probabilities, with a confidence. noul is a yes-or-no numeric probability. score returns a number mapped to a legend of strings. There is no streaming prose, no retries, no chat UI to render.

TypeSafe’s first public model is Jev, built on a new training method the company calls RLCD - “Reinforcement Learning for Calibrated Decisions.” The claim TypeSafe makes is that every Jev decision includes a calibrated confidence, so a downstream agent can route, escalate, or abstain. On TypeSafe’s published numbers, Jev runs at 0.114 seconds per workflow at $0.000081 per workflow, against 8.566 seconds and $0.013880 for an LLM baseline. The page frames that as 193.6x faster and 444.6x cheaper on its own System One workloads, with a published input price of $42 per billion tokens - 238x lower than Claude Fable 5.1’s input price. The TypeSafe docs (docs.typesafe.ai) describe Jev as TypeSafe’s first System One model and confirm the three primitives: choice, noul, and score.

The pragmatic framing is that this is what classifiers were always supposed to be: typed, calibrated, machine-consumable, and fast enough that an agent can call them inside an inner loop.

Two ships in twenty-four hours

On September 29, 2026, Ollama 0.35 shipped with “decision models, based on TypeSafe’s Jev API for fast, typed decisions.” Three open-weight models ship at launch: nimble (a 9B model from Bespoke Labs), tev1 (a 4B experimental model from Together AI), and tev1:0.8b (a 0.8B Together AI experimental). The Ollama blog reports that “Nimble 9B averaged 91ms per decision in the Pac-Man example below” when running locally on a MacBook Pro M5 Max. The endpoint is /v1/systemone; the SDK is the open-source typesafe-sdk on PyPI.

The same day, at OpenAI’s DevDay in Fort Mason, Simon Willison’s live blog recorded OpenAI previewing a Decisions API that gives the Luna model “a predefined set of options to choose from” and answers “in a fraction of a second.” Willison called the Decisions API “OpenAI’s response to ‘Jev,’ which emerged from stealth ‘less than two weeks ago.’” The Decisions API sat in the same keynote as Dots, GPT-6.1 Sol, Codex going cloud-only, and the OpenAI Marketplace launch.

The next morning, TechCrunch reported the broader read: OpenAI built the Decisions API “after a series of incidents where its agents misbehaved on the open internet,” and TypeSafe’s CEO Diogo Almeida (a former OpenAI engineer who co-invented reinforcement learning) was framed as the founder whose work OpenAI was effectively copying. Almeida told TechCrunch: “Fast and cheap is very easy, you know. If you want it really fast and cheap, use dice, right? Intelligence is the hard part, and my North Star is always pushing the intelligence-per-dollar Pareto curve.” He also joked that OpenAI’s clone could mean “the beginning of the clone wars.”

What the API actually looks like

The Ollama example is concrete enough to copy. The request is a JSON POST to /v1/systemone with a model field, a state object, and a questions array. Each question has a name, an instructions string, and a criteria payload. For a support-ticket triage:

  • a choice question named team with criteria {billing: "...", technical: "...", other: "..."}
  • a noul question named refund with criteria asking whether the customer is asking for one
  • a score question named urgency with a legend {0: "Routine", 1: "Soon", 2: "Urgent"}

A single response returns all three, each with a probability or value and a confidence. The Ollama example reports {"input_tokens": 841, "output_tokens": 4} for one of these calls - four output tokens because the answer is a structured payload, not a paragraph of reasoning.

The pricing difference is the part that should land. The TechCrunch piece noted that Shapor Naghibzadeh of QueryStory built a Jev-based monitoring demo, and framed the cost math as a comparison: monitoring agentic actions with Jev cost $2.94; the same monitoring with a frontier LLM cost $372. Sam Altman, at the same event, framed the cloud version of the same trade-off: “By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections.”

Why This Matters

The local-AI angle is the part that matters for readers who run their own models. You can now pull a 9B classifier, point a Python SDK at localhost:11434, and get a 91-millisecond typed decision in a hot path. For local AI, three things change. First, latency: a decision in the inner loop of an agent no longer needs to go to the cloud, so a local agent can stay fully local. Second, cost: four output tokens per call, on commodity hardware, is dramatically cheaper than even the cheapest API. Third, the open-weight SDK is the same one TypeSafe ships for the cloud, so the same code can swap between a local Ollama server and TypeSafe’s hosted Jev without rewriting the agent.

The frontier-lab angle is also worth naming. OpenAI’s Decisions API exists because the agent-control problem is now concrete enough that a frontier lab is willing to ship a specialized endpoint rather than ask customers to keep using a general chat model for classification. The Decisions API sits in the same DevDay keynote as Dots, GPT-6.1 Sol, Codex going cloud-only, and the OpenAI Marketplace - a set of launches that read as a coordinated move toward specialized primitives underneath general agents.

The catch is that decision models are not a substitute for a reasoning model. They are a complement. A reasoning model is the right tool for “find the right answer”; a decision model is the right tool for “given a small set of options, pick one with calibrated probability, fast.” Mixing them - reasoning for planning, decisions for gating - is where the actual workload savings show up.

The Bottom Line

Two labs - one open-weight, one frontier - shipped the same primitive within twenty-four hours. Ollama 0.35 added the /v1/systemone endpoint with three open-weight decision models from Bespoke Labs and Together AI; OpenAI previewed the Decisions API for Luna the same day. Both return typed, calibrated decisions, fast. For local-AI readers, the practical line is: pull nimble from the Ollama library, point the typesafe-sdk at your local server, and replace most of the classification calls in your agent’s inner loop with a typed, sub-100ms, locally hosted decision.