AI development now faces renewed pressure to slow after Anthropic CEO Dario Amodei set out a three-step voluntary plan that calls for independent scrutiny and cross-border coordination, a proposal publicly backed by Elon Musk and Sam Altman. The plan aims to buy time for alignment work and third-party verification without halting technical progress, Amodei said, and Anthropic has already committed alone to the first measure.
Amodei’s first step asks companies to allow external evaluators the kind of access employees have, so safety practices can be checked and incidents reported. He described the second step as industry-level coordination among leading AI firms in democracies to establish shared safety standards, and the third as dialogue between democratic and authoritarian states on how to manage advanced capabilities. Amodei warned some steps will be harder to implement than others, and he emphasised that pacing should not mean stopping model training.
The essay followed a high-profile resignation by Anthropic researcher Jacob Coxon, who said he left because he feared Anthropic and OpenAI were ‘‘gambling with our lives’’ and that those building AI ‘‘earnestly believe that it could kill us all by the end of the decade.’’ Amodei noted that proposals to pause development had been discussed since 2023, but he argued a pause then made little sense because models were not yet capable of significant real-world action, including sophisticated deception or autonomous cyberattacks.
Industry responses amplified the plan. Altman posted that he agreed with Amodei’s call for pacing and called the idea of independent evaluators with employee-like access ‘‘a great idea,’’ saying OpenAI would adopt the approach and had more to share soon. Jakub Pachocki, OpenAI’s chief scientist, has also warned companies have not yet ‘‘solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,’’ and said he expects voluntary slowdowns to become more common until shared safety bars are set. Musk wrote simply, ‘‘Dario is right.’’
The alignment of three prominent figures underscores how concerns about advanced AI have moved into mainstream debate, not least in Washington where lawmakers from both parties are seeking new safeguards and inviting tech leaders to testify. State and local officials are also confronting public pushback over the local impacts of AI infrastructure. Anthropic, which is preparing for what is widely expected to be a historic initial public offering although no date has been disclosed, frames the plan as a way to protect future benefits from the technology while reducing catastrophic risk.
What happens next depends on whether other leading developers adopt the evaluators, and on how governments respond to a voluntary industry effort that seeks to set common safety standards across companies and countries.
