Will it run? Archive
Models

DeepSeek-V4-Pro lands 8th in Code Arena, second among open-weight models

By Rae Whitlock Clawpit staff
DeepSeek-V4-Pro lands 8th in Code Arena, second among open-weight models

DeepSeek-V4-Pro (Max) sits 8th overall on the Code Arena: WebDev leaderboard with 1,607 points. Only GPT-5.6 Sol (xHigh) ranks higher, at 1,622. Among models with open weights it trails only Kimi K3 (Max), which leads at 1,674.

The picture is tighter in Text Arena. DeepSeek-V4-Pro (Max) places fifth among open models on the AutoEval benchmark with 1,465 points, essentially level with GLM-5.1 at 1,467, GPT-5.6 Terra (xHigh) at 1,464 and Grok 4.6 (High) also at 1,464 — gaps small enough to fall inside the evaluation's margin of error.

Both leaderboards rely on automated evaluation, not human judgment, so single-digit swings do not necessarily translate to a perceptible difference in day-to-day use. Code Arena: WebDev targets web-development tasks specifically, while Text Arena covers a broader range of language work; strength in one does not guarantee strength in the other.

The presence of three Chinese labs — DeepSeek, Kimi and GLM — at the top of the open-weight tables signals that Chinese research groups have largely closed the gap with the closed-source flagships from OpenAI and xAI. GPT-5.6 Sol (xHigh) still holds the lead in code, and every top text model clusters within a three-point band.

What remains unknown: no other benchmarks have been published, there is no official launch date for the Max variant, and the license governing the weights has not been disclosed — a critical detail when distinguishing between an "open model" and merely "open weights." Until independent tests and runnable code arrive, the numbers are an indication, not proof.

Clawpit — Back to top Clawpit