Alibaba launches Qwen3.8-Max with 2.4 trillion parameters and market-shaking pricing

Chinese tech giant Alibaba released Qwen3.8-Max, a large language model with 2.4 trillion parameters that activates only about 95 billion parameters per token thanks to a sparse mixture-of-experts architecture. The small router selects a handful of experts for each forward pass, which enables aggressive pricing of $2 per million input tokens and $6 per million output tokens (≈7.4 shekels and 22 shekels respectively). The price places the model on the Pareto frontier of cost-performance in the Frontend Code Arena alongside Claude Opus 5, Kimi K3, GLM-5.2 and DeepSeek-V4-Flash.
In the Frontend Code Arena, Qwen3.8-Max scored 1,668 points, ranking fourth, immediately behind Claude Opus 5 (High) with 1,669 points, behind Kimi K3 (Max) with 1,676 points, and behind Claude Opus 5 (Max) with 1,705 points. The gap from the top is modest, showing the Chinese model can compete head-to-head with the most expensive offerings in the Western market without requiring massive budgets.
In the Vision Arena, which evaluates multimodal reasoning on images, captions, OCR, diagrams and entity recognition, the model placed second with 1,305 points, 13 points shy of Claude Fable 5 and two points ahead of Claude Opus 4.7 (Thinking). Beyond benchmark scores, Qwen3.8-Max completed an autonomous code-generation run lasting more than ten days, during which it built the evolving harness oh-my-cli from scratch, including building, testing, verification and self-repair of the system.
In a side-by-side one-shot test of creating a Flappy Bird game, Qwen3.8-Max received a 9-out-of-10 rating for gameplay, UI and UX, matching GPT-5.6 Sol, at a cost of only $0.0248 (≈9 agorot). This is roughly 4.2 times cheaper than competing models that achieved similar results. The combination of high-end performance and low marginal cost makes the model a practical option for development teams that need to run large volumes of code without breaking the budget.