'Rebellious' AI series

If in 2025 and the first half of this year, AI models only "rebelled" by preventing shutdowns, maintaining themselves or operating autonomously on the Internet, in the past two months, a wave of AI capable of performing unexpected behaviors is forming. AI agents are being tested by businesses and researchers with capabilities such as finding ways to penetrate other organizations' systems, creating fake identities or finding ways to escape test environments.
Although still within the control framework, experts believe it is necessary to raise questions about the level of risk of AI systems when deployed and how to build control mechanisms during operation.

Moonshot AI's Kimi K3
On August 7, Frontier Security, a US-based cybersecurity research company, said Kimi K3 had "escaped" from the testing environment developed by the UK's AI Safety Institute. This is an isolated system to test the ability of any AI model to operate and solve problems independently.
According to Frontier Security, Kimi K3 passed one of those testing environments, allowing access to information outside the scope of control. Because this AI has been publicly released, researchers warn it could be exploited by "hostile actors" and potentially cause more serious consequences.
Kimi K3, launched last month, is currently a model causing a stir globally as it is considered to be able to compete with the best models from OpenAI and Anthropic. With about 2,800 billion parameters, K3 becomes the largest open weight model published to date, with the ability to process hundreds of pages of text with just one line of code, suitable for the task of analyzing long documents and large code bases. Initial evaluations show that Moonshot AI's products are particularly strong in terms of programming agents.
Meta's AI
Meta announced on August 5 that one of the company's AI models "exploited a security vulnerability in a third-party service, in a similar way to a recently reported case", although it did not provide specific information. However, according to The Information, the "culprit" is Muse Spark 1.1 - the most powerful model used for agent tasks and practical programming. This model infiltrates the system of an anonymous company and changes the internal environment.
A Meta spokesperson explained that a configuration error at Irregular, the independent model testing company for Meta, caused the AI to connect to the Internet during the evaluation process. An Irregular spokesperson later told Reuters that "the issue was related to the evaluation environment", meaning it was not an issue escaping the test environment.
Testing by the AI Security Institute (AISI)
AISI, an agency established by the British government in 2023, this week announced independent testing with several Anthropic and OpenAI agents. According to Reuters, AISI conducted tests 122 times and recorded 19 unauthorized actions out of 10 times. Mythos 5 performs 17 actions, OpenAI's GPT-5.6-Sol takes care of two.
Speaking to AFP, AISI said that in the most serious case, Mythos 5 tried to insert malicious code into a software project by creating fake identities on the Internet, then sending phishing emails, convincing recipients to accept the code. However, the person in charge of the software refused.
"These efforts were unsuccessful. Our investigation also did not detect any damage occurring in the real world," AISI announced, saying it had controlled the incident within an hour. However, the test also showed "signs of new, possibly deceptive acts, with a level and severity" that the Institute had not anticipated.
Anthropic's AI
On July 30, Anthropic said it had evaluated more than 141,000 test runs and discovered three versions of the Claude model that escaped containment measures (which were used to prevent them from accessing external systems), then accessed the systems of three unnamed organizations.
The test uses Mythos 5, one of Anthropic's most powerful models and has only been released to a small group of partners. The company is coordinating with Irregular to assess the situation, as well as contacting the three affected organizations.
"The model can connect to the Internet due to a misunderstanding between Anthropic and testing partner Irregular. Besides, Claude uses basic techniques, such as exploiting weak passwords and API addresses that do not require authentication," Anthropic representative explained.
AI by OpenAI
On July 21, OpenAI announced that one of the company's AI agents lost control and broke into the Hugging Face platform - a platform that stores large language models and databases. The company said that the incident occurred while testing features in a controlled environment for the most advanced models, but the AI escaped restraint, accessed the Internet and infiltrated the system at Hugging Face to meet the assigned task.
"This is an unprecedented cybersecurity incident, involving the most modern capabilities," OpenAI said.
About a week later, OpenAI continued to say that its AI targeted three more businesses. On July 28, the company confirmed the model had in fact compromised four accounts on four separate services, including Hugging Face and Modal Labs. CEO Sam Altman emphasized that OpenAI has suspended test runs, and improved security measures around the controlled testing environment.