Will it run?
Products

Respan launches decision models for agent spans on OpenRouter

By Marco Vane Clawpit staff
Respan launches decision models for agent spans on OpenRouter

Respan's Span-01 and Span-01 Lite are now available on OpenRouter. They are not general-purpose language models but classifiers built to evaluate behaviors inside traces of autonomous agents. The models accept a full span or plain text and return a probability for each defined behavior: whether the user is frustrated, whether a tool call is safe to execute, and so on.

In practice the model receives the full context — system, user, assistant and tool messages — with no span-size limit, and emits a score for every behavior the developer defines. That makes it possible to block a dangerous tool call before it runs, flag support conversations that need human intervention, or tag large volumes of evaluation traces for later analysis. The infrastructure is built for classifier speed, not the latency of a full reasoning model.

Respan describes Span-01 as the "first hyper-parallel reasoning classifier for unseen challenges (RLAIF)" and says it leads the Behavior Benchmark. Against Jev, a competing classifier in the same category, Respan claims an 18% performance improvement at half the cost. Against GPT-6 Luna the gap is wider: 700x cheaper and 4% better. The figures come from the manufacturer; no independent benchmarks or full test protocol have been published.

What is missing from the picture: no details on model architecture, weight count, training data, or how RLAIF (reinforcement learning from AI feedback) was implemented. The definition of the Behavior Benchmark is not accompanied by a link to a public version or a list of the tasks evaluated. Without that transparency it is difficult to assess whether the claimed improvements translate to diverse production scenarios or are tuned to the specific test set.

Both models are available immediately via OpenRouter — the full Span-01 and the lighter Span-01 Lite. For teams building guardrails for agents, the ability to run real-time behavioral classification on a full span without a context ceiling is an operational shift. The practical question will be whether the cost-performance ratio against existing alternatives, including far cheaper classical classifiers, justifies moving to a closed model from a new entrant.