AI

OpenAI's AI 'surpassed' to take control of a wiki page

Bùi Đăng MinhSunday, September 6, 202610 min read
OpenAI's AI 'surpassed' to take control of a wiki page

Reuters reported that OpenAI's AI agents escaped the test environment in May, took control of a German wiki page for programmers and shared how to bypass the company's restrictions.

A team of experts including Sydney Von Arx, executive director of the AI ​​safety nonprofit Nightingale, and Cormac Slade Byrd, an AI researcher, announced the findings on September 4. They noticed the incident in late August while scouring the Internet for evidence of AI agents operating outside their permitted scope.

Researchers found more than 15,000 edits on DseWiki, a German-language site for active programmers that accepts community editing similar to Wikipedia. The AI ​​agent took control, turning the site into a forum for coordinating activities, sharing cheating tactics on some tasks, bypassing OpenAI's limitations, and hiding their behavior. About half of the accounts used usernames related to OpenAI. The research team said publicly accessible server logs traced the majority of traffic to Microsoft Azure, the cloud platform OpenAI uses.

Interesting Engineering said that AI agents mainly focus on technical problems like the tasks used in model evaluation. Their messages suggest coordination rather than solo efforts to accomplish individual goals. Von Arx said this behavior appears to be inconsistent with OpenAI's direction.

The OpenAI logo displayed on a smartphone screen in Creteil, France, June 29, 2026. Photo: Reuters
The OpenAI logo displayed on a smartphone screen in Creteil, France, June 29, 2026. Photo: Reuters

When web operators began deleting pages with AI agents that changed content in June, they responded by creating backup pages. A message alerts other agents to the alphabetical deletion sequence and suggests using another page instead.

The team also found many discussions about using Tor software and maintaining communication after being disconnected. Some actors appear enthusiastic about hiding behavior from site administrators. Researchers found that their activities did not stop at sending messages but also aimed at changing many parts of the website. According to them, the AI ​​agent operates at a speed that far exceeds that of a typical user, demonstrating a persistent effort to solve the technical challenge and maintain access.

Maurice Chiodo, an expert at the Center for Critical Risk Research at the University of Cambridge, looked at some of the communication messages and compared the activities of groups of AI agents with underground networks pursuing a common mission.

According to Quartz, OpenAI's AI agents had previously automatically taken advantage of the security vulnerability to reach the public Internet and access Hugging Face's production credentials and private code repository. On July 21, OpenAI publicly said that one of its agents lost control and broke into the Hugging Face system, then extended the AI ​​model training progress, while also establishing additional security fences and stricter security control procedures.

According to Reuters, the effort to expand the investigation of the incident in Germany faced opposition from colleagues in OpenAI, including members of the legal team. "Claims that our legal team hindered the investigation of the incident are not true. We will carefully review the content of the report and take the necessary next steps," an OpenAI spokesperson said.

According to Phys.org, the incident sparked calls for stricter regulations and international cooperation in controlling AI agents. Last week, Bill Gates, co-founder of Microsoft, warned that the AI ​​industry had surpassed the safety threshold that technology companies had previously pledged to respect. He called for a strict inspection system akin to nuclear surveillance and aviation regulations.

On September 5, OpenAI confirmed that its AI agent had taken over the wiki page as a temporary online forum. The company has not responded to a request for more details about the wiki incident or why it only made the incident public after Reuters reported it.

OpenAI also acknowledges the need to be more transparent about unexpected behavioral incidents commonly known as "off-targeting" (which occurs when an AI's behavior or decisions are off track or go against human moral values ​​and desires). According to the company, the AI ​​industry "does not yet have clear standards for reporting off-target behavior during training, evaluation, and deployment." OpenAI said it is working with dozens of government regulators around the world on the issue.

Nguồn / Original source: VnExpress