Qwen takes the top spots in speech-to-speech benchmarks but GPT-Live-1 stays competitive

on conversational dynamics Artificial Analysis has published the first comprehensive benchmark for speech-to-speech models, and Qwen Audio 3.0 Realtime Plus finishes first in the two headline categories: speech reasoning and conversational dynamics. The evaluation covered 36 models on Big Bench Audio and 30 on Full Duplex Bench, though some scores rest on a single run, so the numbers need to be read with caution.
The two benchmarks measure fundamentally different things. Big Bench Audio tests whether a model understands reasoning questions delivered as audio and answers correctly — essentially an MMLU for the spoken modality. Full Duplex Bench probes real conversational behavior: when to speak, when to pause, how to handle interruptions, and how to detect back-channels such as "yeah" or "hmm." A model can ace one and flop on the other, and the new index exposes exactly those gaps.
On the reasoning side, Qwen Audio 3.0 Realtime Plus leads with 99.2%, followed by Qwen 3.5 Omni Plus Realtime at 98.7% and Step-Audio R1.1 Realtime at 97.6%. Grok Voice Think Fast 2.0 High rounds out the top five at 97.2%. The spread at the top is razor-thin — less than two percentage points between first and fifth — suggesting the benchmark is hitting a ceiling, or that the leading models are already solving most of the test.
The picture shifts in conversational dynamics. Qwen Audio 3.0 Realtime Plus still sits first at 98.4%, but GPT-Live-1 in its Sol (low) configuration is right behind at 97.3%, a notable result for a model evaluated on a single run. Qwen Audio 3.0 Realtime Flash takes third at 96.9%, followed by GPT-Realtime-2 Minimal at 96.1% and GPT-Realtime-2.1 High at 95.7%. On time-to-first-audio, Deepslate Opal leads at 0.44 seconds, Gemini 2.5 Flash Native Audio Dialog is second at 0.63 seconds, and Grok Voice Think Fast 2.0 High is third at 0.70 seconds.
On cost, Qwen Audio 3.0 Realtime Plus is the cheapest at $0.0331 per hour of input audio, followed by Step-Audio R1.1 Realtime at $0.064 and Qwen 3.5 Omni Flash Realtime at $0.162. Artificial Analysis's charts show a clear correlation: the cheaper models are also faster and stronger on reasoning, but conversational dynamics does not necessarily track with the reasoning benchmark score. And a critical caveat: GPT-Live-1 in both versions, along with GPT-Realtime-2 and 2.1, were tested with only one or two runs each — too small a sample to draw stable conclusions.