Will it run? Archive
Models

Open-weight AI models narrow the cyber capability gap with closed systems

By Rae Whitlock Clawpit staff
Open-weight AI models narrow the cyber capability gap with closed systems

The British Institute for AI Security (AISI) released its first public analysis that quantifies how far open-weight models—those whose weights can be downloaded and run locally—have progressed toward the cyber capabilities of fully closed, proprietary systems. The core finding is that the gap has narrowed to 4 to 7 months, compared with the 6 to 10 months measured for most of 2025. In other words, tasks that the strongest closed models can perform today are expected to be matched by open versions within at most half a year.

On a suite of 70 narrow cyber-specific tests, GLM-5.2 achieved performance comparable to Claude Opus 4.6, which was released 4.3 months earlier. DeepSeek-V4-Pro placed somewhere between Claude Opus 4.5 and GPT-5 (released in November and August 2025, respectively). However, when the evaluation shifts to long-term tasks in the cyberrange called The Last Ones—a full-scale intrusion scenario—the picture changes. GLM-5.2 reaches up to Opus 4.5 (a gap of less than 7 months), while DeepSeek-V4-Pro falls below Sonnet 4.5, a model below the frontier that launched 7 months earlier. AISI notes that the disparity is larger in this setting, which aligns with the industry term “big model smell”, and that open models lack the “generalisation juice” that distinguishes proprietary models.

The practical implication for cyber defence is clear: researchers and security teams have a short window to adapt before frontier-level cyber abilities become publicly accessible. This occurs without the safeguards that commercial vendors embed. AISI plans to evaluate Kimi K3 on the same basis as soon as its weights are released to the public.

Kimi K3, a 2.8-trillion-parameter model from China, shows frontier-level performance across the standard test suite, matching or slightly trailing Claude Fable 5 and GPT-5.6 Sol. The company states that the model “demonstrates frontier-level performance consistently” and outperforms other models tested. The weights are expected to be released in the coming weeks alongside a research paper. A caveat is noted: the performance is accompanied by brittleness that manifests as “benchmaxxing”, an over-tuning to benchmark metrics at the expense of broader generalisation.

Among other observations, Kimi was examined for writing GPU compilers—a step toward AI-building-AI, i.e., using AI systems to improve themselves. This direction is intriguing, but until the weights are public and the paper peer-reviewed, the claims remain those of the manufacturer. History suggests waiting for independent runs before drawing firm conclusions.

Clawpit — Back to top Clawpit