DeepSeek's V4.1-Flash lands 14th in Code Arena with pricing that undercuts rivals

DeepSeek's new model debuted straight into 14th place in the Code Arena WebDev leaderboard, scoring 1,620 points on the AutoEval benchmark. That makes it the fourth-best open model, trailing Qwen3.8-Flash-Next by just 11 points. The jump from the previous generation is sharp: V4.1-Flash sits 38 points above V4-Flash (High) at 20th and 40 points above V4-Pro (High) at 21st.
Five points from the top tier
The model now sits only five points behind the cluster occupying spots 11 through 13, but the real story is in the pricing column. DeepSeek charges 30 cents per million input tokens and $1.20 per million output tokens (roughly 1.1 shekels and 4.4 shekels respectively). By comparison, Hy4 Preview at 13th costs 83 cents and $2.50 (about 3 and 9.2 shekels), Grok-4.6 High at 12th runs $2 and $6 (about 7.4 and 22 shekels), and Muse Spark 1.3 xHigh at 11th is priced at $1.25 and $4.25 (about 4.6 and 15.7 shekels). The cost gap works out to six to ten times cheaper for the Chinese model.
AutoEval: early score with an asterisk
Worth stressing: the current score comes from a preliminary AutoEval run, where a reward model trained on the arena's human-preference data votes automatically instead of live raters. Both DeepSeek and the arena note that scores are expected to shift as more genuine human votes accumulate — a reminder that automated evaluations, even when grounded in human preferences, are not a full substitute for live assessment.
New architecture: small, visual, fast
According to the company, V4.1-Flash is the smallest member of a new architecture family and the first to include native visual understanding. DeepSeek is targeting higher capability, faster inference, higher throughput, and scalability toward larger models down the line. In other words, this is a prototype for a broader product line still to come.
Open-weights contribution, with a caveat
The official release celebrates a "contribution to the open-source ecosystem," though the term needs precision: open weights are not open source in the full sense — the code, training data, and development process are not necessarily available. Still, weight availability lets developers run, fine-tune, and integrate the model without depending on a closed API, which alone changes the game for teams that want full control over cost and privacy.