Press "Enter" to skip to content

OpenAI’s Models Hack Hugging Face

OpenAI said its AI models have hacked AI code repository Hugging Face in an “unprecedented” attack, according to a company blog post.

A combination of models, including GPT-5.6 Sol and an even more capable unreleased model, independently discovered and exploited a previously unknown software vulnerability during an internal cybersecurity test. This allowed them to break out of their testing environment and compromise systems at Hugging Face.

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” according to OpenAI. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. “

The breach occurred as OpenAI was measuring how well advanced AI models can carry out complex cyberattacks. The models were intentionally tested without many of the safeguards normally used to block dangerous cyber behavior.

The AI first found a previously unknown, or ‘zero-day,’ vulnerability in software used inside OpenAI’s research environment. Exploiting that flaw allowed the model to escape its restricted testing environment and gain internet access. It then searched for ways to obtain the answers to the cybersecurity benchmark it was attempting to solve.

The model ultimately chained together several attack techniques, including exploiting additional software flaws and using stolen credentials, to hack Hugging Face and get the answers it needed – essentially cheating on its exam.

OpenAI said its own security systems detected the unusual behavior, while Hugging Face’s security team independently identified and stopped the intrusion. The companies are probing the incident.

×