Will it run?
Models

Qwen dominates speech-to-speech benchmarks, leads in reasoning and conversational dynamics

By Rae Whitlock Clawpit staff
Qwen dominates speech-to-speech benchmarks, leads in reasoning and conversational dynamics

Qwen (Qwen) conquers speech-to-speech leaderboards, topping both reasoning and conversational dynamics. Artificial Analysis evaluated 33 speech-to-speech models on speech reasoning and 27 on conversational dynamics, placing Qwen’s Audio 3.0 Realtime Plus at the top of both leaderboards.

In the speech-reasoning track (Big Bench Audio), Qwen Audio 3.0 Realtime Plus secured first (1) place with 99.2%, followed by Qwen3.5 Omni Plus Realtime with 98.7%, Step-Audio R1.1 Realtime at 97.6% and Grok Voice Think Fast 2.0 High and Grok Voice Think Fast 2.0 at 97.2% and 97.1% respectively. The gap between first and fifth place is just over two percentage points.

In the conversational-dynamics track (Full Duplex Bench), Qwen Audio 3.0 Realtime Plus secured first place with 98.4%, while Qwen Audio 3.0 Realtime Flash took second place with 96.9%, and GPT-Realtime-2 (Minimal) was third at 96.1%. GPT-Realtime-2.1 High and GPT-Realtime-1.5 shared fourth place with 95.7%. Full Duplex Bench measures whether the model knows when to speak, when to stay silent in natural pauses, how to respond to Interruptions and how to recognize Backchannels like “yes” or “mm”. The GPT models were run with a limited number of trials—one for the minimal version, two for 2.1 High—introducing uncertainty into direct comparisons.

Deepslate Opal led the latency metric Time to First Audio with 0.44 seconds. Gemini 2.5 Flash Native Audio Dialog was second at 0.63 seconds, and Grok Voice Think Fast 2.0 High was third at 0.70 seconds. On cost, Qwen Audio 3.0 Realtime Plus was cheapest at $0.0331 per hour of generated audio, followed by Step-Audio R1.1 Realtime at $0.064 and Qwen3.5 Omni Plus Realtime at $0.162, a spread of nearly five-fold between the lowest and highest prices among the top tier.

Real-time conversation requires both strong conversational dynamics and low latency. Qwen Audio 3.0 Realtime Plus excels in dynamics but does not appear among the three fastest models, while Deepslate Opal is considerably faster, though it was not tested for conversational dynamics on the same scale. Developers will need to decide whether natural turn-taking or rapid initial response is more critical for their product, and how much they are willing to pay per hour of audio.