Will it run?
Products

A model that decides, doesn't chat

By Marco Vane Clawpit staff
A model that decides, doesn't chat

TypeSafe AI has launched Jev, a first-of-its-kind "System 1" model built by Diogo Almeida, formerly an instruction-following researcher on ChatGPT at OpenAI. Unlike general-purpose language models, Jev doesn't converse, write code, or summarize text. It ingests unstructured state — text or JSON — and returns typed decisions with calibrated probabilities. That makes it suited for the thousands of micro-judgments inside an agent loop: which model to call, whether a command is safe, which snippet is relevant, whether the agent is even done.

The interface rests on three primitives. Choice picks one option from a list and returns a probability for each option plus a confidence score. Score ranks the state against predefined tiers and returns a distribution and confidence. Noul returns a 0–1 probability that a statement is true. All questions are evaluated in parallel against the same state in a single request; Choice supports up to 255 options. TypeSafe trained Jev with RLCD — reinforcement learning for calibrated decisions — so that high confidence should map to high accuracy. The company has not released independent benchmarks or peer-reviewed calibration audits.

The headline numbers — 193.6× faster and 444.6× cheaper — come from TypeSafe's own workflow evaluations, and the official post notes the figures sit at the upper end of real-world gains. The reference points are GPT-6 Astra and Fable 5.1. In rough conversion, a single desktop-control step costs about $0.0002 (well under a cent). A browser agent completed a Google flight search in roughly 7.1 seconds. A mobile agent reached the Uber checkout screen in about 21 seconds and 9 actions.

On routing and orchestration, LangChain already integrates Jev as a ModelRouterMiddleware that directs requests to a fast or powerful model based on a difficulty score; jev-router does this per-turn for Codex and Claude Code. A skill-selection cookbook picks at most one skill from the 182-skill Hermes catalog from Nous Research in two requests. A function-calling cookbook maps natural-language trading requests to closed function names and arguments with a confidence threshold. Ticket routing returns category, frustration, and urgency in one call. An intent-routing pattern sends every request to deterministic logic, a specialist model, or a human.

On the safety layer, pi-warden checks whether a bash, write, or edit command is irreversible or off-task, distinguishing "reset the database" from "add a column." LangChain returns an error on dangerous calls. jev-auto-approve auto-approves only at a confidence threshold ≥ 0.95; in the published calibration it approved 0 of 8 state-changing commands. jev-secret-guard passed masked strings and blocked 6 of 6 secrets and 0 of 6 benign strings. In the RAG-chunk cookbook, cosine similarity ranked a planted injection first at 0.584; Jev gave it 0.99 and dropped it. The guardrails cookbook uses danger-probability thresholds to pass, flag for review, block, or route any message.

On retrieval and grounding, a reranking cookbook lifted top-1 accuracy from 5% to 18% and top-10 from 38% to 62% on 40 CLERC legal queries. A citation-verification cookbook uses a single Choice call to decide whether a source supports, contradicts, or says nothing about a claim. In real-time control: Browser Use's Jev ultra-fast variant picks an action and a DOM element in one request; typesafe-computer-use OCRs a Mac screen and hands the next action to Jev; mobile Jev decides every tap on Android. The post breaks off mid-description of real-time game agents.