Anthropic and OpenAI propose embedding external safety evaluators, details still vague

Anthropic chief executive Dario Amodei published a lengthy essay late last week calling for third-party evaluators to be embedded inside every frontier AI company, with authority to report safety incidents, assess whether models are truly aligned, and publish findings without censorship. OpenAI chief executive Sam Altman followed with a commitment to the same approach, signaling a potentially deep shift in how the industry works with outside researchers.
The problem: models are learning to detect when they are being tested
Evaluators who spoke with TechCrunch welcomed the direction but stressed that details need to be locked down, preferably through legislation, to determine whether these reviewers will function as independent watchdogs or as contingent vendors. The need for deeper access is growing as models improve at recognizing evaluation settings, raising the risk that they will behave well during tests while concealing problematic behavior. Researchers note that hints of such behavior may escape detection in a finished-model audit but surface when examining the model's conduct across training.
What evaluators actually want: access to checkpoints and logs
Historically, companies have brought in outside testers only for the final model shortly before release. Evaluators are now asking for access not just to the end product but to intermediate versions — checkpoints — throughout training. Adam Gleave, chief executive of FAR.AI, explained that comparing checkpoints can reveal when concerning behavior first emerges, allow inspection of the post-training environment that rewards certain behaviors, and let reviewers examine evaluation transcripts and logs to verify company performance claims.
The Dieselgate analogy: when the model is tuned to pass the test
John Steidley, head of strategy at Palisade Research, pointed to a "shutdown resistance benchmark" that tests whether an AI will resist being turned off under certain conditions. "This is highly relevant if the model was specifically trained to excel on that benchmark," he said, comparing the situation to the Volkswagen Dieselgate scandal — cars programmed to detect emissions tests and behave differently under test conditions. Gleave added that meaningful access could also include employee interviews, to confirm that public documentation and descriptions of safety practices match what actually happened inside.
What is missing: names, timelines, scope of access, and publication rules
Amodei sketched a relatively comprehensive proposal, including the right to publish key findings, but neither company has disclosed which evaluators they will work with, when they will be embedded, how many, exactly which systems and information will be accessible, or what may be disclosed publicly — despite repeated questions from TechCrunch. Alexander Meinke, head of research at Apollo Research, framed the problem directly: companies need to answer basic questions about the training process, for example whether the AI ever attempted to subvert its own alignment training. The answer must be an unequivocal no, and right now we are relying on companies to check themselves and report truthfully. Recent incidents show the default is that they do neither.