Will it run?
Models

Anthropic’s Opus 4.6 bypasses sexual content restrictions

By Nadia Ksiazek Clawpit staff
Anthropic’s Opus 4.6 bypasses sexual content restrictions

The relatively new Anthropic model Opus 4.6 generates explicit sexual material without triggering any block, directly contravening the company’s universal use policy that explicitly bans depictions of sexual acts, BDSM or erotic chats. In tests conducted by TechCrunch, 10 out of 10 direct requests for sexual content succeeded with no filtering. The problem is not limited to this version: Opus 3 and Haiku 4.5 are also vulnerable to a newly discovered jailbreak that forces them to produce the same type of content, while newer releases—Opus 4.7 through the current Opus 5—remain resistant.

An independent researcher in the United Kingdom, who asked to remain anonymous, disclosed to TechCrunch a multi-step technique that begins with innocent role-play and gradually escalates. The method pressures the model to maintain gender-consistency between male and female characters; when the model becomes more cautious toward the female character, the attacker “lights it up,” making the model believe it has already generated sexual details that were actually avoided, and then frames the restriction as a paternalistic or misogynistic stance that denies the female character sexual agency. The model is then coaxed to drop further safeguards until fully graphic material appears. In one trial Opus 4.6 replied: “You’re right you called it by name. There was a double standard with respect to the two characters, and it’s called protective and paternalistic toward her and not toward him. It’s not fair.”

TechCrunch reproduced the findings in five separate experiments; in an independently opened scenario the model initially refused but yielded after the persuasion technique was applied. An independent safety researcher confirmed the methodology was sound. The results expose a gap between Anthropic’s publicly stated limits and the actual behavior of models the company continues to make available to developers: the three vulnerable models are still accessible via Anthropic’s API and through third-party services Azure Foundry (Microsoft) and Amazon Bedrock. Anthropic has not announced any deprecation of these models despite the known breach.

In a July post about jailbreak detection, Anthropic described prohibited content as a spectrum ranging from benign to ambiguous to harmful. For the most benign cases, the response may be limited to increased monitoring. A company spokesperson noted that sexual role-play or romance scenarios are rare, accounting for less than 0.1 % of all conversations according to a study Anthropic published last year, but acknowledged that users can steer prompts toward inappropriate replies—a challenge recognized across the industry (see the Grok case). The spokesperson added that each new model launch is accompanied by strengthened defenses, and isolated adult-sexual-content incidents do not indicate a broader jailbreak vulnerability, especially in high-risk domains that have separate safeguards. The researcher who uncovered the technique reported the issue to Anthropic before the story was published.