Will it run?
Models

Alibaba's Qwen team releases Qwen-Audio-3.1-Realtime, a bidirectional voice model that decides when to speak

By Ilse Brandt Clawpit staff
Alibaba's Qwen team releases Qwen-Audio-3.1-Realtime, a bidirectional voice model that decides when to speak

Alibaba's Qwen team has released a suite of five audio models covering speech recognition, synthesis and real-time interaction. The flagship, Qwen-Audio-3.1-Realtime, is built as a voice agent that supports function calling and operates in full duplex — it can listen and speak simultaneously and decides for itself when to stop listening and start answering. The service is available as a managed API on QwenCloud over WebSocket; no open weights are being released at this stage. The context window stands at 262 thousand tokens (approximately 245 thousand maximum input, 16 thousand output), with default limits of 60 requests and 100 thousand tokens per minute.

Two models share the same audio encoder and LLM backbone. A bidirectional decision model predicts whether to continue listening, speak, stop or return to listening; a speech-to-text model writes the response content as text; and a context-aware voice renderer converts that text into streaming speech while accounting for conversation history, vocal cues and acoustic context. Alongside the main model, Qwen is also launching Qwen-Audio-3.1-ASR-Flash-Filetrans, a companion for long-form offline audio transcription with support for hotwords, speaker diarisation, punctuation and Chinese dialect detection alongside multilingual capability.

Development is organised in three layers: Think, Act, and Speak and Coordinate. In the Think layer, the team applies M²-OPD, an on-policy distillation method in which a text teacher and a frozen audio model score every student token rather than imitating pre-written responses. Domain experts for empathy, pragmatic intent and acoustic scenes are trained with GRPO and merged into a single per-layer model. In the Act layer, each training domain comes with a tool registry, a JSON state database and a natural-language business policy; tasks are defined as write, reasoned refusal or unsupported request, and scoring checks the final state before fluency. Search training reduced the average queries per call from 4.37 to 1.05, at the cost of a slight drop in activation F1 from 60.87% to 58.61%.

The Speak and Coordinate layer shows measurable gains: on Full-Duplex-Bench v1.5 the rate of responding while the human is speaking to someone else fell from 0.13 to 0.03, and on v3.0 the filler-word rate dropped from 0.759 to 0.296. There are trade-offs: the unwanted resume-after-interruption rate rose from 0.035 to 0.130, and interruption stop latency sits at 1.116 seconds versus 0.383 seconds for GPT-Realtime-2. On general benchmarks, Audio MultiChallenge improved from 47.12 to 52.21, the average BBA score across 14 languages climbed from 81.7% to 88.1%, and FLEURS WER fell from 9.01 to 3.98. The τ-Voice figures are based on semi-duplex alignment and are not comparable to official full-duplex results; in a 50-session human red-team evaluation GPT-Realtime-2 still leads 96% to 92%.

Pricing is set at $6.40 per million audio input tokens, $0.80 per million text input tokens and $24 per million output tokens; text output is not charged. The transcription model costs $0.15 per million input tokens and $0.47 per million output tokens (approximately 0.55 and 1.7 shekels respectively). Prices were verified on 28 September 2026; differing tokenisation rates across providers make direct comparison difficult.