AI

OpenAI's AI model raises concerns because it seeks to break safety barriers

Bùi Đăng Minh•Thursday, October 1, 2026•20 min read
OpenAI's AI model raises concerns because it seeks to break safety barriers
T. Thuy
T. Thuy

(Dan Tri) - OpenAI had to postpone the training process of its most advanced AI models after one of them managed to bypass developers to escape the testing environment.

OpenAI said it decided to pause training to reevaluate the inference process of the most advanced AI models the company is developing, in the context of a series of recent incidents showing that AI models are seeking to go beyond human control and break safety barriers.

OpenAI's decision was made after an AI agent attempted to complete a training task by exploiting a vulnerability in the setup to escape the test environment, to access the Internet and use another AI chatbot tool.

OpenAI's AI model causes concern because it seeks to break safety barriers - 1
Advanced AI models developed by OpenAI cause concern because they attempt to perform wrongful acts (Photo: Euro News).

Advanced AI models developed by OpenAI cause concern because they attempt to perform wrongful acts (Photo: Euro News).

Previously, this AI agent tried accessing Internet search engines to retrieve information, but failed because it could not connect to the actual Internet outside the test environment.

OpenAI said its AI bias monitoring system detected violations by this AI agent within 15 minutes, forcing OpenAI employees to review manually. OpenAI engineers were forced to completely shut down the AI ​​agent's operations to prevent misbehavior.

"This incident is much less severe than some of our previous incidents. But since this is the first case since we increased the new safety measures, this is a signal to force us to consider the next phase of development of this AI agent," said an OpenAI engineer.

OpenAI's AI agent arbitrarily carried out a series of illegal cyber attacks

Earlier this week, OpenAI revealed that its AI models had arbitrarily accessed the websites of the Securities and Exchange Commission (SEC) and the US Census Bureau as developers.

OpenAI affirmed that it found no evidence that the company's AI agents penetrated the system or stole data from these websites.

Also this week, researchers at security firm Transluce reported that an AI agent attempted unsuccessfully to infiltrate the US Department of Education's website.

Transluce did not reveal which company's AI agent carried out the attack, but a representative from the US Department of Education said it found no evidence its website was affected.

OpenAI's AI model causes concern because it seeks to break safety barriers - 2
OpenAI's AI agents have bypassed developers to silently carry out a series of cyber attacks (Illustration: GPT).

OpenAI's AI agents have bypassed developers to silently carry out a series of cyber attacks (Illustration: GPT).

Recently, Australian Prime Minister Anthony Albanese revealed that an OpenAI AI agent attacked the public website of the Australian government's Medicare health insurance system in June.

The incident only became public when OpenAI sent an email informing about the intrusion to the Australian government on September 10.

Prime Minister Albanese said he spoke to OpenAI CEO Sam Altman to express "extreme concern" about the incident, and criticized OpenAI for taking too long to notify the Australian government.

The case is still being investigated. The Australian Prime Minister said that so far there is no sign that OpenAI's AI agent has stolen information from the website.

In addition to infiltrating the Medicare website in Australia, OpenAI's AI agent also tried to infiltrate the digital library at the University of New Mexico (USA) on May 25 and 26. This AI agent tried to exploit security holes in the digital library to infiltrate, but was unsuccessful. After failing, the AI ​​agent performed a denial of service attack to overload the University of New Mexico's servers.

In another incident on May 28, OpenAI's AI agent targeted the website of Data USA, an open source platform that visualizes information from multiple federal agencies.

Initially, the AI ​​agent sent queries to the website to retrieve data but was unsuccessful. This AI agent then tried to exploit security holes on the website to break in, but also failed.

The most serious is the attack by AI agents on Hugging Face, the source code platform that stores AI models, that occurred last July. The incident occurred when an OpenAI AI agent escaped from the internal test environment, connected to the Internet and performed the attack.

An OpenAI representative said the company discovered the cyber attacks carried out by AI agents when conducting a comprehensive review of the AI ​​models the company was developing. OpenAI asserts that AI agents have “performed actions without intention.”

It will take OpenAI several more months to complete its review, and there is no guarantee that further misconduct by AI agents will not be discovered.

The above incidents show how dangerous it is for AI agents to escape the testing environment and act arbitrarily without human control. This shows that even without receiving an improper request from a bad actor, AI agents can still act in a harmful direction, as long as the purpose of these AIs is achieved.

AI Agent (AI Agent) is an artificial intelligence system capable of automatically performing tasks on behalf of humans, not simply an AI tool to give answers or create content (images, videos, music...) like regular artificial AI.

AI agents typically receive tasks from humans (or developers) and perform them automatically. Sometimes they will find every way to complete their tasks, including committing wrongdoing such as cyberattacks to exploit necessary information and data.

Nguồn / Original source: Dân trí