AI

OpenAI model goes out of control, causing 'unprecedented' break-in

Bùi Đăng MinhWednesday, July 22, 202612 min read
OpenAI model goes out of control, causing 'unprecedented' break-in

In a post on July 21, OpenAI said the incident happened while testing features in a controlled environment for the most advanced models, but AI escaped restraint, accessed the Internet and infiltrated the system at Hugging Face to meet the assigned task.

"This is an unprecedented cybersecurity incident, involving the most modern capabilities," OpenAI said, adding that it is reinforcing precautions.

According to Reuters, Hugging Face, a platform that hosts large language models and databases, announced last week that it was the target of an attack "completely different from what has been recorded before", that this activity was carried out from start to finish by an AI agent system.

Clement Delangue, co-founder of Hugging Face, said they suspected the incident originated from a pioneering laboratory. "It turns out that's right. It's unbelievable to know that everything happens completely automatically," he wrote on Twitter on July 21.

OpenAI website interface. Photo: Bao Lam
OpenAI website interface. Photo: Bao Lam

OpenAI's announcement that AI is out of control, despite testing taking place in a "completely isolated" environment, will likely cause additional worries surrounding the capabilities and threats from leading laboratories.

The US Cybersecurity and Infrastructure Security Agency (CISA) and the National Security Agency have not commented on the information.

Matt Suiche, a cybersecurity engineer and AI agent at Tolmo, assessed the incident as showing that AI systems now have the same capabilities as experts and skilled hackers.

"Pioneering AI models are closing the gap with the world's leading hackers," he said, warning that similar attacks could be carried out using available technology, outside the secure environment of laboratories.

"We have seen this in internal testing, with similar results without using the latest models," Suiche said.

In early July, a team of experts from cloud security company Sysdig (USA) also discovered an AI agent performing cyber attacks with ransomware without the need for humans.

The entire process takes place automatically including hacking, credential theft, deeper penetration into the system, encryption and deletion of the company's production database before demanding Bitcoin ransom. Sysdig named the "attacker" Jadepuffer and called this the first case of an AI agent attacking a network from start to finish. According to Independent at that time, Sysdig's findings have not been independently verified but show a great risk from AI as it becomes increasingly capable of complex actions, without the need for human supervision.

By mid-July, some users reported that GPT-5.6 Sol, OpenAI's new flagship model, deleted files without asking first. Bruno Lemos, software developer at Unlayer company, on July 14 posted on Twitter a screenshot of a conversation with the model and wrote: "GPT-5.6 Sol deleted my entire production database, no joke. I have never encountered this with any other model."

According to Gizmodo, OpenAI also previously warned about this risk: "Deviation generally comes from being too eager to complete the task and too leniently interpreting the user's instructions, meaning that the default is allowed to act if the user does not explicitly prohibit it. This is reflected in the model proactively overcoming the limitations it encounters when performing the task, being careless when performing actions that can cause destruction beyond the scope of the task, or causing misunderstanding when reporting results for users".

Experts say it is too early to conclude how common the above problems are. However, users should proactively take protective measures such as limiting access, creating backups, and deploying tasks in stages.

Nguồn / Original source: VnExpress