Will it run?
Models

OpenAI posts incomplete GPT-Live-1 voice-agent benchmarks

By Rae Whitlock Clawpit staff
OpenAI posts incomplete GPT-Live-1 voice-agent benchmarks

OpenAI's developer account on X published a thread that repeats the same prompt three times: how GPT-Live-1 performs on the benchmarks that matter for production voice agents. The post says testing covered four areas — task completion, two-way conversation and turn-taking, response speed, and tool use — but supplies only a single figure: a score on the Tau3 benchmark. The other metrics the company itself flagged as relevant are absent, and the text cuts off mid-comparison sentence.

Tau3 measures whether a voice agent can finish real customer-support tasks across three scenarios: aviation, retail, and telecommunications. It is not a general language test but a full-workflow simulation — intent recognition, information retrieval, tool invocation, and case closure. According to the post, GPT-Live-1 was paired with the GPT-6 Astra reasoning model at "medium reasoning effort", and the pair completed 83.6% of tasks on the first attempt.

An 83.6% first-attempt success rate suggests meaningful ability to handle inquiries without escalating to a human or forcing the customer to repeat themselves. The post provides no basis for comparison, however — not against earlier model versions, not against competitors, and not against any internal threshold the company defines as production-ready. Without that context the number remains a single data point that is hard to interpret.

The three other benchmark areas the post highlights — natural conversation management, latency, and function-calling success — carry no numbers at all. No latency figures, no turn-taking success rates, no function-calling accuracy. The fact that the company listed them as the focus of testing and then published no results leaves a gap that every developer building on the platform will have to fill with independent evaluation.

The thread ends on the word "compared" without completing the comparison — compared to what? GPT-4o? An open model? A human baseline? The cutoff, combined with the tripled introductory prompt, reads like a draft published by mistake or a marketing rollout that was not finished. Until OpenAI releases the remaining data, developers should treat the 83.6% figure as an early indication only, not a binding specification.