Alibaba launches Qwen3.8-Omni-Flash, first omni-modal model with built-in agentic capabilities

The new model is built on Qwen3.8-Flash-Next, whose open weights were released in August 2026. The context window stands at one million tokens, though QwenCloud caps practical input at 991 thousand and output at 131 thousand; maximum reasoning length reaches 262 thousand tokens. Output is text only — users who need generated speech are directed to Qwen3.5-Omni. Reasoning is enabled by default with reasoning_effort set to xhigh and can be disabled by setting it to none. The API supports both the DashScope and OpenAI protocols, including Chat Completions, Responses API, function calling, web search, structured outputs, context caching and batch requests.
Instead of reading a video file from start to finish, the agent begins with the question, decides which segments to watch and listen to, and gathers evidence in coarse-to-fine rounds. According to the Qwen research team, the approach lifted accuracy on OmniVideoBench from 63.4 to 67.8 while cutting token usage by roughly 45.7%, from 145,736 to 79,117. All figures come from the company; independent results were not available at publication.
Across 29 evaluations, the average score improves by more than 25% over Qwen3.5-Omni-Plus. WildClawBench-MM rises 36.5 points, AgenticVBench 22.3 points, and UniClawBench reaches 69.6. LongAudioSpan adds 8.3 points, OmniVideoBench another 9.6, and OmniCap-IF improves CSR by 8.5 and ISR by 14.1. The official tweet cites an average gain of 19.5 points on the two agentic benchmarks. The company claims audio-video performance is close to Gemini 3.8 Flash and that general audio performance exceeds it, without publishing the comparison protocol.
On QwenCloud the price is $0.15 per million input tokens, $0.47 per million output tokens and $0.016 per million implied cache hits. Compared with the previous generation, the company reports a reduction of more than 98% per hour of audio input, more than 93% for audio-video and roughly 89% for video. Technical limits: video files up to two hours and 2 GB via URL, audio up to three hours, 113 languages and dialects, stable sampling up to 15 frames per second, stereo two-channel and FOA four-channel via use_multichannel. The service is available in six regions: Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia. An API call takes only a few lines with the OpenAI SDK.