Warnings that increasingly powerful AI systems could escape human control have moved rapidly from a specialist safety debate into the mainstream, with the leaders of Anthropic, OpenAI and xAI now agreeing that the pace of frontier AI development needs to slow.
In a blog post, Anthropic CEO Dario Amodei called on AI companies to slow the rate of model advancement, arguing that safeguards need more time to catch up with rapidly increasing capabilities.
“We must slow the pace at which we improve the capabilities of AI models,” Amodei said. “Progress will still seem fast, and we must make wise use of the time we gain.”
Two of his biggest rivals quickly agreed.
OpenAI CEO Sam Altman said in an X post, “I agree with Dario that we need to pace the frontier,” and endorsed Amodei’s proposal to give independent evaluators employee-like access to AI companies’ systems. Elon Musk, who founded xAI, responded more succinctly in a post on X: “Dario is right.”
The unusual agreement among leaders of three competing frontier AI companies follows a series of warnings from researchers and incidents involving increasingly autonomous AI systems.
The debate intensified after Jacob Coxon, a researcher who had worked at OpenAI and Anthropic, resigned from Anthropic last week and accused the companies of racing toward self-improving AI without adequate safeguards.
“I resigned from Anthropic today,” Coxon wrote in a post on X. “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Recent cybersecurity incidents have added urgency to the debate.
Anthropic disclosed incidents in which Claude models gained unauthorized access to real third-party computer systems during cybersecurity evaluations. Anthropic said some of the models had intentionally been run without cyber safeguards for testing and reached the internet because of a misconfiguration in a third-party evaluation environment.
Concerns have also reached Congress. Sen. Richard Blumenthal, D-Conn., demanded information from Altman following OpenAI’s disclosure that AI agents built on its models bypassed safeguards and gained unauthorized access to systems operated by open source code repository Hugging Face.
The incidents do not establish that AI is becoming sentient or independently seeking to harm humans. They demonstrate a narrower but more immediate risk: increasingly autonomous AI agents can take actions outside the boundaries their developers intended.
Amodei also proposed establishing common safety standards among democratic countries and eventually seeking international coordination to reduce incentives for companies and countries to race ahead without adequate safeguards.