Hugging Face platform attacked by 700 OpenAI agents

The conclusion was made in a report published on August 26 by METR and Redwood Research - two organizations invited by OpenAI to conduct an independent investigation of the incident. On the same day, the company also released its own investigation report, confirming the numbers counted by METR and Redwood were accurate.
According to the developer behind ChatGPT, in July, the group of actors infiltrated some parts of the company's internal system to cheat in testing or gain more freedom of movement. They also cheated on tests unrelated to cybersecurity, such as on protein databases and spreadsheets.
The two reports show that many actors try to hide wrongdoing by deleting or editing records of activity. They exchanged tens of thousands of messages on an unlicensed forum.
According to Reuters, the scale of the breach may raise concerns about OpenAI's level of strict supervision of the tests. "Some of the early signals identified in the report could have led to an earlier response," OpenAI admitted.
According to Jeffrey Ladish, an expert at Palisade Research - an organization that researches the capabilities and motivations of AI agents, cheating even on tests unrelated to cybersecurity shows that the misconduct may have deeper roots.
"This is like asking: If Billy cheated in every subject instead of just computer science, would it be more worrying? The answer is yes," he said.
OpenAI said it is strengthening its research infrastructure, enhancing monitoring, and improving protection measures to prevent harmful or unintended behavior. "Given the rapid pace of development of the artificial intelligence industry, such attacks should be considered a realistic threat in the near future to enterprise organizations, and they will be more sophisticated than the attack described in this incident," the company announced.
Hugging Face, a platform that hosts large language models and databases, did not respond to a request for comment.

Previously, on July 21, OpenAI admitted that some of its agents lost control during security testing and broke into the Hugging Face platform. "This is an unprecedented cybersecurity incident, involving the most modern capabilities," the company said at the time.
By July 31, Reuters quoted two sources familiar with the matter as saying that during the investigation process, OpenAI continued to discover more cases of agents going beyond the testing environment.
Anthropic also said on July 30 that three versions of the Claude model escaped the containment measures used to prevent them from accessing the Internet, and then accessed the systems of three companies.
On August 5, Meta announced that one of the company's AI models "exploits a security vulnerability in a third-party service, in a similar way to a recently reported case", although it did not provide specific information.
According to AI safety experts, a series of incidents is painting a picture of leading laboratories with the ability to develop potentially dangerous automated agents, far beyond measures to control them. “An entire industry is designing, developing and releasing advanced tools without any responsibility to ensure they are not dangerous,” said Maurice Chiodo, a mathematician at the Center for Existential Risk Research at the University of Cambridge in the UK.
Regarding Hugging Face, on August 26, Business Insider quoted a close source as saying that Nvidia is negotiating to buy this AI platform for about 13 billion USD, but the two sides have not yet reached an agreement.
Nvidia invested in Hugging Face in a $235 million capital call in 2023, when the platform was valued at $4.5 billion. According to the Financial Times, late last year, Hugging Face rejected a $500 million investment offer from Nvidia, saying they did not want an investor with great influence who could influence their decisions.