Will it run?
Models

Artificial Analysis publishes methodology for streaming speech-to-text leaderboard

By Rae Whitlock Clawpit staff
Artificial Analysis publishes methodology for streaming speech-to-text leaderboard

Artificial Analysis has released the methodology behind its new leaderboard for streaming speech-to-text models. The headline metric, AA-WER Streaming, blends three datasets: AA-AgentTalk weighted at 50%, VoxPopuli at 25% and Earnings22 at 25%. The score represents the percentage of words transcribed incorrectly; lower is better.

The benchmark measures two distinct transcription points: the first partial transcript that appears immediately after speech-end detection, and the final transcript delivered after additional processing. Each stage receives its own word-error rate and its own latency figure in seconds. Charts plot the Pareto frontier between accuracy and latency, showing which models sit in the most attractive quadrant — high accuracy with low latency.

A third dimension is price: transcription cost in dollars per 1,000 minutes of audio. The metric allows direct comparison of models offering different trade-offs among speed, accuracy and cost — parameters that are critical for real-time applications such as voice bots, live call transcription and automatic captioning.

The methodology page details the calculation and chart structure but does not include the benchmark results themselves: no model names, no numerical scores, no positions on the Pareto curve. Artificial Analysis has not yet published the full table with raw data.