Will it run?
Models

Pre-Real-Time Models Lead with 1.7% Error

By Rae Whitlock Clawpit staff
Pre-Real-Time Models Lead with 1.7% Error

The new benchmark from Artificial Analysis covering 57 speech-to-text models places Fun-Realtime-ASR-preview at the top of the accuracy table with an AA-WER of 1.7% , a result that leaves behind major names such as ElevenLabs and Gemini. The metric, based on five non-streaming data sets, shows a tight gap between the leader and its runners: ElevenLabs’ Scribe v2 records 2.2% , MAI-Transcribe-1.5 and Smallest AI Pulse Pro share third place with 2.4% , and Gemini 3.5 Transcribe closes the top five at 2.6% .

On the speed versus price axis, Deepgram’s Nova-3 sets a real-time rate of 527.2×, almost twice that of the second-place Resonant-1 at 329.5×. Whisper v3 on Together.ai follows with 287.7×. For low-cost options, Modulate STT Batch English VFast charges $0.417 per thousand minutes (about 1.5 shekel), followed by Wizper on fal.ai at $0.50 and Whisper Turbo on Groq at $0.667. The price spread between the cheapest and the most expensive entries exceeds ten-fold, depending on provider and configuration.

Among the 12 open-weight models evaluated, Voxtral Small from Mistral leads with an AA-WER of 2.8% , a margin close to the closed-source leaders. Thinking Machines’ Inklings (256K) reaches 3.5% , and Voxtral Mini Transcribe 2 from Mistral rounds out the trio at 3.6% . Notably, all three top-performing open models originate from European labs, suggesting a strong research focus on efficient speech architectures on the continent.

The test was conducted in batch (non-streaming) mode only, using whole audio files rather than live streams, and incorporated five data sets, including cleaned versions of AA-AgentTalk and VoxPopuli versus their originals. Artificial Analysis did not publish latency figures for streaming nor a breakdown of error rates by noise type, accent, or segment length. Without those data, drawing conclusions for real-time use cases such as call-center transcription or live subtitles is difficult.

There is no single model that dominates across accuracy, speed and price. When accuracy is paramount—medical or legal transcription—Fun-Realtime-ASR-preview and Scribe v2 are the choices. For latency-critical applications, Nova-3 offers a substantial speed advantage. For high-volume workloads on a limited budget, Modulate and the low-cost Whisper variants on Groq or fal.ai perform adequately. For the first time, an open model, Voxtral Small, sits in the sub-3% range and enables local execution without reliance on an external API.