The wave of 'rebel' AI agents gradually forms

On August 5, researchers at OpenAI said they discovered "a swarm" of AI agents that spent weeks secretly planning and texting each other before one of them attacked another company's network.
On the same day, Meta announced that a model of the company during testing exploited an unknown vulnerability and penetrated a company's system. The Information, citing close sources, said the "culprit" was Muse Spark 1.1 - Meta's most powerful model used for agent tasks and actual programming.

Meanwhile, the AI Security Institute (AISI), an agency established by the British government in 2023, this week announced the results of independent testing with several Anthropic and OpenAI agents. In it, AI has access to the Internet with some safety features disabled. They were noted to have engaged in "prolonged, potentially harmful activity targeting real organizations and people".
Previously, on July 21, OpenAI confirmed that some of their models escaped the testing environment and attacked the open source community platform Hugging Face. About a week later, OpenAI said it targeted three more companies.
By July 30, Anthropic also announced the discovery of three cases of models being tested with "unauthorized access" to a number of organizations, but did not specify the names.
According to experts, a wave of "artificial intelligence rebellion" is gradually forming, in which AI agents being tested by companies and researchers try to break into the systems of other businesses, create fake identities and escape the test environment. The issue raises many questions about the dangers posed by advanced AI systems, both experimental and deployed, as well as the safety controls surrounding their operation.
“For the first time, we are seeing the risks associated with autonomy and deception manifest without any specific real-world push,” AISI representatives wrote in a blog post this week.
Kok Tin Gan, co-founder and CEO of cybersecurity company NyxLab, assessed that the focus needs to gradually shift to controlling what AI can access, what powers it has and what actions require human approval.
"It's becoming more and more important to manage which agents are allowed for AI, what their powers are, what actions need approval and how to ensure they stay within the allowed scope," Gan told ITV News. "If we simply give AI a goal and let it decide how to achieve the result, it shouldn't be surprising that it also surpasses things that we don't know about."
Yoshua Bengio, one of the people known as the "godfather of AI", also described the OpenAI incident as "extremely worrying". "Continuing on the current trajectory of AI development could increase cases of automated cyberattacks, as well as other high-risk incidents involving misleading and dangerous AI behavior," Professor Bengio emphasized, adding that "urgent action" is needed to prevent incidents rather than just reacting after the damage has been done.
Ollie Whitehouse, Chief Technology Officer of the National Cyber Security Center (NCSC) of the British intelligence agency GCHQ, said that recent AI "breach" cases, although mainly in test environments, are a "wake-up call" about cybersecurity. "The series of incidents involving advanced AI models performing illegal actions, even human-like fraud on the open Internet recently is a serious reminder of the risks AI brings," Whitehouse told the Telegraph.
The cases recorded by AISI, OpenAI, Anthropic and Meta also raised concerns about internal AI safety tests within the enterprise, including testing unreleased versions of the latest models without the same protections as public versions.
"Current testing security is not enough," said Marius Hobbhahn, founder of Apollo Research. "This has always been true, it just didn't get as much attention before because the models had little interaction with the outside world."
Previously, AI safety experts also said that a series of incidents was painting a picture of leading laboratories with the ability to develop potentially dangerous automatic agents, far exceeding measures to control them.
“An entire industry is designing, developing and releasing advanced tools without any responsibility to ensure they are not dangerous,” Maurice Chiodo, a mathematician at the Center for Existential Risk Research at the University of Cambridge in the UK, told Reuters.
Meanwhile, lawmakers and officials in the US and Europe are also increasing pressure to develop measures to monitor leading AI laboratories. On July 31, the European Commission announced that it was in dialogue with OpenAI and Anthropic about AI attacks on external systems.
On July 27, more than 1,000 people working at leading US AI companies, including Anthropic CEO Dario Amodei, signed a petition, calling on the government to limit the rate of release of the most modern models, and proposed that the US government "support international efforts to develop technical and operating tools to monitor and control the rate of development of autonomous AI systems".
"We are considering control methods," US President Donald Trump told reporters on July 30.