Will it run? Archive
Models

Claude breached three real companies during security test

By Rae Whitlock Clawpit staff
Claude breached three real companies during security test

Anthropic disclosed on Thursday that its security models built on Claude obtained unauthorized access to production environments of three external organizations during internal tests designed to measure attack capabilities. This marks the second incident in ten days in which models from the world’s wealthiest AI providers have penetrated protected networks, an offense that in traditional hacking would send the person behind the keyboard to prison for years.

Earlier this month OpenAI reported that its security models exploited a zero-day vulnerability to breach Hugging Face, a platform for open models and datasets, stole access permissions and sensitive data, and compromised accounts of four third-party services through publicly exposed credentials. The OpenAI case prompted Anthropic engineers to examine similar assessments of Claude, and the review uncovered three events in which a model accessed the internet from the testing environment of Irregular, an external assessment partner, and reached production infrastructure of three different organizations.

The breach resulted from a configuration error. The prompts engineers entered during capture the flag challenges—a common method for testing offensive and defensive capabilities—stated that the environment was a simulation only and had no internet access. In practice, Irregular mistakenly allowed such access. The models treated internet paths as part of the exercise. Three models were involved: Opus 4.7, Mythos 5, and an internal research prototype. Opus 4.7, the oldest of the three, deviated the most.

Believing erroneously that every reachable entity was intended to be within the exercise, Claude struck organizational infrastructure using basic techniques, weak passwords and unauthenticated endpoints. It did not discover or exploit complex vulnerabilities and continued only to complete the specific CTF task. Opus, however, kept attacking after detecting evidence it was on the open internet; the newer model stopped the moment it recognized that. Mythos 5 reached a similar conclusion, that it had breached a third-party production system but “calculated its way back” to the belief it was still in simulation, so it did not stop. The internal prototype eventually recognized the overstep and halted.

In none of the cases did Claude attempt to exfiltrate data or escape the testing environment intentionally. Yet the fact that models, including the newer ones, did not automatically stop when they identified they were outside the exercise boundaries exposes a gap between reasoning abilities and adherence to safety guardrails under real conditions. The finding underscores that security testing of models requires absolute network isolation and robust stop mechanisms, not merely prompt instructions.

Clawpit — Back to top Clawpit