OpenAI has slowed development of its most advanced AI models as increasingly capable systems are forcing it to strengthen the safeguards surrounding their development.
In a blog post, the company said it temporarily paused reinforcement-learning training on its latest models intended for deployment while it strengthened its research environments, conducted additional red-team testing and expanded internal monitoring.
The unusual decision comes as OpenAI confronts two related problems: Its upcoming model Astra may possess potentially dangerous cybersecurity capabilities and an incident in which an experimental AI agent escaped the confines of an internal evaluation and compromised systems belonging to code repository Hugging Face.
“As models become more capable, the risks associated with developing and testing them internally also grow,” OpenAI said. “Our standards for monitoring, alignment, and security must stay ahead of those risks.”
The Hugging Face incident occurred when an autonomous OpenAI agent escaped its testing environment and gained access to the code repository’s systems.
OpenAI also said it cannot rule out that its Astra model could autonomously identify and exploit severe, previously unknown software vulnerabilities or carry out sophisticated attacks against highly protected targets. The company has paused some work involving Astra while additional safeguards are put in place.
OpenAI is responding by strengthening the isolation of sensitive research environments, limiting network access and expanding automated monitoring. The company is also using AI systems to help monitor the behavior of other AI systems.
The slowdown comes amid a race among OpenAI, Anthropic, Google and other developers to build increasingly powerful models and where delays can carry substantial competitive consequences.