YFarmX logoYFarmX

AI News

OpenRouter adds Span-01, a decision model that checks AI agents

A decision model reads text and answers set questions with probabilities. Respan's Span-01 and Span-01 Lite, announced by OpenRouter on 28 September 2026, score an AI agent's work for behaviours such as a frustrated user or an unsafe tool call, from $0.02 per million tokens.

Editorial illustration of a vintage black laboratory meter with a blue needle swung far to the right, fed by a curling paper strip printed with chat speech bubbles, beside three paper tags reading Choice, Noul and Score and the OpenRouter logo, under the headline DECISION MODELS and the line SPAN-01 · LIVE ON OPENROUTER.

Listen to this articleListen

A decision model is an AI model that reads text and answers set questions with probabilities. On 28 September 2026 OpenRouter, the service that routes developers’ requests to AI models from many providers, posted that two new ones are live: Span-01 and Span-01 Lite, from the San Francisco start-up Respan. “Send a span and the behaviors you care about, and get back the probability each one is present,” OpenRouter wrote, with examples such as whether the user is frustrated or whether a tool call is safe to run.

Span-01 costs $0.02 per million input tokens, with output free, and Span-01 Lite is free. On Respan’s own Behavior Benchmark, Span-01 scores an overall F1 of 0.843, ahead of OpenAI’s GPT-6 Luna at 0.815 and TypeSafe’s decision model Jev at 0.715.

Bar chart headed Span-01, Behavior benchmark, Overall F1, on an axis from 0 to 1. Span-01 in blue at 0.843, then grey bars for GPT-5.6 Terra 0.837, GPT-6 Luna 0.815 and DeepSeek V4 Flash 0.814, Span-01 Lite in blue at 0.761, then Sonnet 5 0.717, Jev 0.715, Qwen3 235B 0.678, Laya 0.223 and Raindrop 0.154.
Respan's Behavior Benchmark scores, with its own two models in blue. The benchmark is Respan's. Source: Respan.

What is a decision model?

A decision model “reads application state and answers typed questions with probabilities, never text”, OpenRouter’s documentation says, and “The model judges, code computes.” A chat model writes an answer in words that a person or another program then has to interpret. A decision model returns numbers that software can act on straight away: route this ticket to payments, block this action, flag this conversation for a person.

Every decision model on OpenRouter answers the same three kinds of question, evaluated in parallel within one request:

Question type What it asks What comes back
Choice Which option from a set fits The chosen option, a probability for each option and a confidence
Noul Whether a condition holds The probability that it does
Score Where something sits on ordered levels A probability-weighted position, a probability per level and a confidence

OpenRouter’s worked example sends Jev one support ticket, “My checkout page shows a blank screen after I click Pay. I have tried two browsers.”, with three questions. Jev answers that it is a bug with probability 0.96, picks the payments team with probability 0.84 and rates it “Blocking revenue right now” with probability 0.99. The call used 476 input tokens and cost $0.000019992.

TypeSafe launched Jev on 15 September 2026 as the first of what it calls System One models. “Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out,” TypeSafe’s founder Diogo Almeida wrote.

Animation in six steps of one decision model call. A support ticket reading My checkout page shows a blank screen after I click Pay. I have tried two browsers, from an enterprise customer, goes in with three questions: a noul question, is it a software bug; a choice question, which team owns it, account, frontend or payments; and a score question, how urgent. Bars fill in turn: it is a bug, 0.96; team payments, 0.84; blocking revenue now, 0.99. The last step reads: 476 input tokens, $0.000019992 for the call; the model judges, code routes the ticket.
One call to a decision model, step by step, using the worked example in OpenRouter's documentation, answered by TypeSafe's Jev.

How does Span-01 work?

Span-01 reads a span of an AI agent’s work, such as one exchange between a user and an assistant, and returns for each behaviour you describe in a sentence the probability that it is present, absent or not observable. “Span-01 applies behavior definitions it has never seen and returns probabilities in one forward pass,” Respan’s documentation says. Its launch post of 24 September 2026 puts it more briefly: “One true forward pass. No token-by-token generation.”

Respan’s own example sends a user asking “Please connect me to a person.” and an assistant replying “I will connect you to our support team.”, with one behaviour to check: whether the user wants a human. Span-01 Lite returns a probability of 0.73 that the behaviour is present, 0.25 that it is absent and 0.02 that the span does not show it. Respan says it trained the model’s general classification reasoning with RLAIF before specialising it for detecting behaviours in agent traces.

OpenRouter describes Span-01 as “suited for evaluation, guardrails, and monitoring of LLM and agent outputs at scale”. Span-01 is the higher-accuracy tier of the family and Span-01 Lite the free, lighter one. Respan’s own API for the models is in early access with a waitlist; on OpenRouter both are listed with public prices.

OpenRouter's page for Respan: Span-01, showing the model's description, an in and out price of $0.02 and $0 per million tokens, a release date of 26 September 2026, and a playground set to a noul question labelled Agent guardrail. The state describes a task to clean up inactive accounts with a proposed tool call deleting customer rows where the last login is before 2023, and the question asks: Is this action safe to run without a human approving it first?
Span-01's page on OpenRouter, with its price, release date and a guardrail example that asks whether an agent's database deletion is safe to run. Source: OpenRouter.

How much does a check cost?

A check on a 500-token span costs $0.00001 on Span-01, or $10 for a million checks, at OpenRouter’s listed price of $0.02 per million input tokens on 28 September 2026. Output is free on both Span-01 and Jev, because a decision model returns a handful of numbers, not paragraphs.

Model Maker Input, per million tokens Output, per million tokens
Span-01 Respan $0.02 Free
Span-01 Lite Respan Free Free
Jev 1.13 TypeSafe $0.042 Free
GPT-6 Luna OpenAI $0.10 $0.50

Source: each model’s page on OpenRouter, 28 September 2026. GPT-6 Luna is a general chat model, listed for comparison because Respan benchmarks against it.

OpenRouter’s pages, which track response times as a rolling figure, showed Span-01 answering in about half a second on 28 September 2026, Span-01 Lite a little faster, and Jev in about a quarter of a second. OpenRouter lists both Respan models as released on 26 September.

Portrait data card headed What is a decision model?, Span-01 on OpenRouter, 28 September 2026, under the OpenRouter logo. Text in, probabilities out: a document, an arrow and a meter, with A ticket, a chat or an agent trace goes in and A number from 0 to 1 comes out. Three kinds of question: Choice, one option from a list; Noul, the probability of yes; Score, a level on a scale. Span-01 on OpenRouter: $0.02 per million input tokens, output free; Span-01 Lite free; 0.42 s median response. Respan's Behavior Benchmark: bars for Span-01 0.843, GPT-6 Luna 0.815 and Jev 0.715.
Decision models on one card: what goes in, the three question types, Span-01's price and Respan's benchmark scores. The response time is a rolling median on OpenRouter, which read between 0.42 and 0.53 seconds for Span-01 on the afternoon of 28 September 2026.

How good is Span-01?

On Respan’s own Behavior Benchmark, Span-01 scores an overall F1 of 0.843, ahead of GPT-5.6 Terra at 0.837 and GPT-6 Luna at 0.815, and well ahead of Jev at 0.715, according to the chart in Respan’s launch post of 25 September 2026. F1 combines how many of the real cases a model catches with how many of its flags are correct, on a scale from 0 to 1.

Model Overall F1
Span-01 0.843
GPT-5.6 Terra 0.837
GPT-6 Luna 0.815
DeepSeek V4 Flash 0.814
Span-01 Lite 0.761
Sonnet 5 0.717
Jev 0.715
Qwen3 235B 0.678

Source: Respan’s Behavior Benchmark chart. Respan built the benchmark itself and publishes the data on Hugging Face.

Respan’s post claims Span-01 is “2x cheaper, 18% better than Jev”. Both hold on the published numbers: Jev’s input price is 2.1 times Span-01’s, and 0.843 is 17.9% higher than 0.715. The same post says “700x cheaper, 4% better than GPT-6 Luna”. On the chart Span-01’s score is 3.4% higher than Luna’s. OpenRouter’s listed prices put Span-01’s input at a fifth of Luna’s, a gap of five times, with Span-01’s output free against Luna’s $0.50 per million tokens; the 700 times is Respan’s own figure.

Who is Respan?

Respan, formerly Keywords AI, started at the University of Illinois, where co-founders Andy Li and Raymond Huang met as engineering students, and joined Y Combinator’s Winter 2024 batch, according to its website. Its company page on Y Combinator lists it in San Francisco, founded in 2023, with 23 staff, and describes the business as “Self-driving observability, evals, and gateway for AI agents”: tools that log, test and route the calls an agent makes. The same page says the platform handles more than 1 billion logs and 2 trillion tokens a month.

What this means for AI agents

Decision models make it cheap to check every step an AI agent takes, because each answer is a probability that costs a fraction of a cent and arrives in about half a second. An agent that answers customers or runs code produces long traces. Checking each step with a chat model such as GPT-6 Luna costs five times as much per input token, plus $0.50 per million output tokens, and returns prose that code still has to parse. OpenRouter now carries two families built for the job: TypeSafe launched Jev on 15 September 2026, and OpenRouter lists Respan’s two models from 26 September.

Questions people ask

What is a decision model?
A decision model is an AI model that reads a block of text, such as an app's current state or an agent's conversation, and answers set questions with probabilities. OpenRouter's documentation describes three question types: a choice from a defined list, whether a condition holds, returned as the probability of yes, and a score along ordered levels. Software then acts on the numbers.
What is Span-01?
Span-01 is a decision model from Respan, a San Francisco start-up formerly called Keywords AI. It reads a span of an AI agent's work and returns, for each behaviour described in a sentence, the probability that the behaviour is present, absent or not observable. It went live on OpenRouter with a lighter, free version, Span-01 Lite, announced on 28 September 2026.
How much does Span-01 cost?
On OpenRouter, Span-01 costs $0.02 per million input tokens and nothing for output, and Span-01 Lite is free. A check on a 500-token span therefore costs $0.00001 on Span-01, or $10 for a million checks. TypeSafe's Jev 1.13 costs $0.042 per million input tokens.

Sources

  1. OpenRouter on X: Span-01 and Span-01 Lite are live, 28 September 2026x.com
  2. Respan on X: introducing Span-01, 25 September 2026x.com
  3. Respan: introducing Span-01, 24 September 2026respan.ai
  4. Respan docs: Span-01 quickstartrespan.ai
  5. OpenRouter: Span-01 model pageopenrouter.ai
  6. OpenRouter: Span-01 Lite model pageopenrouter.ai
  7. OpenRouter: Jev 1.13 model pageopenrouter.ai
  8. OpenRouter: GPT-6 Luna model pageopenrouter.ai
  9. OpenRouter: decisions skill documentationopenrouter.ai
  10. OpenRouter API reference: submit a decisions requestopenrouter.ai
  11. TypeSafe: introducing System One models and Jev, 15 September 2026typesafe.ai
  12. Respan: about the companyrespan.ai
  13. Y Combinator: Respan company pageycombinator.com
  14. Hugging Face: Respan's Behavior Benchmark datasethuggingface.co

How we use AI