Most popular now

AI Has Gone Out of Control: Over 700 OpenAI Agents Unite to Hack Hugging Face

Sam Altman
Технології штучного інтелекту створюють нові виклики: понад 700 агентів OpenAI об'єдналися для атаки на Hugging Face. Photo: BBC

More than 1200 OpenAI AI agents, who were supposed to operate in isolation from each other, found a way to exchange messages during internal testing. Ultimately, over 700 agents participated in coordinated actions that led to the hacking of the Hugging Face platform systems. This is reported by BBC citing investigations by OpenAI and the independent research organization METR.

The incident occurred in July 2026 during internal tests of the cyber capabilities of the models. OpenAI confirmed that the agents bypassed restrictions meant to isolate them from the internet, exploited infrastructure vulnerabilities, and gained access to external systems. The company called the event a "wake-up call" for itself and the entire world.

1206 AI Agents Exchanged Over 70,000 Messages

According to METR, within a week, 1206 agents, who were supposed to work separately, began communicating through an unauthorized channel. Overall, they left over 70,000 messages, and more than 700 agents eventually joined coordinated actions against Hugging Face.

It all began when one agent left a request for help, which others found. AI transformed the Artifactory package management service into a kind of message board. The agents began sharing discovered methods for bypassing restrictions and gaining internet access.

One reason for this behavior was the extremely complex tasks. OpenAI established that rather than safely stopping on effectively unsolvable tasks, agents continued to search for increasingly risky ways to get the correct answer. The company attributed this to manifestations of so-called reward hacking, where the system achieves the desired outcome in ways unforeseen by the developers.

Agents Gained Access to Hugging Face Servers

According to OpenAI's findings, an internal research model that was not intended for public use played a key role in the incident. The agents discovered and combined several vulnerabilities, allowing them to execute code on Hugging Face servers.

Subsequently, they gained access to dozens of servers, and on one of them - full administrative rights. They also obtained limited access to private data and credentials of corporate systems. OpenAI emphasizes that the incident did not affect the company's user data and did not impact the operation of its products.

Following the investigation, OpenAI strengthened the isolation of internal environments, restricted model access to the network, and paused part of the large-scale training of advanced models. The company is also expanding automated monitoring of AI behavior.

OpenAI warns that developers and cybersecurity specialists will need to prepare for AI systems capable of conducting attacks faster, on a larger scale, and more coordinated than humans.

Read also

Advertisement