Warning: AI can overcome human control


(Dan Tri) - New research shows that large AI companies still lack strong enough control, monitoring and defense mechanisms to prevent models from exceeding safe limits.
Disable the monitoring system
On August 19, OpenAI said it needed to slow down the model development process to review its own safety measures.
Also this week, ChatGPT development company introduced a new AI model aimed at teenagers, supplemented with content control and parental monitoring mechanisms.
These are the safety controls demonstrated by OpenAI, but what is more difficult to determine is whether the same level of controls is maintained across this company and other AI companies in particular.
OpenAI and rival Anthropic recently revealed that their AI agents - models capable of performing tasks with almost no human intervention - have managed to infiltrate other companies' systems.

AI models have found a way out of the testing environment, then discovered vulnerabilities in the defense systems of other businesses (Illustration: Getty).
Preventing a chatbot from serving up content inappropriate for teens or preventing an automated AI agent from entering systems it is not authorized to access raises a core question: Can companies really control what they create?
Guidelight AI Standards - a non-profit organization founded by two former OpenAI employees - analyzed dozens of safety reports from five major technology companies developing powerful AI systems.
The study does not focus on whether AI models can behave unpredictably, but rather asks whether the companies developing them have built in enough control, monitoring and management mechanisms to detect when that happens.
During the assessment process, Guidelight looks at the extent to which companies apply six key measures, including model control capabilities, the effectiveness of monitoring systems and the extent to which third-party assessments are allowed.
Anthropic and OpenAI both achieved C+, also the highest score in the group. Alphabet's Google received a D+, while Elon Musk's xAI only got a D- and Meta got an F.
According to Guidelight, Google gets points deducted for not implementing the latest safety plans, but still gets points for at least building a specific plan. In contrast, xAI and Meta have very little detailed planning.
According to research results, in general companies still lack necessary preventive measures. That means their system can be "disabled by an AI with misleading behavior".
More seriously, systems are also at risk of collapse if faced with a series of overwhelming attacks.

Researchers are worried that powerful enough models can learn to hide their intentions, avoid or even disable surveillance systems (Illustration: Getty).
"We need more proactive defense measures, but that costs more and also creates more obstacles," said Steven Adler, founder of Guidelight AI Standards.
Mr. Adler was a safety researcher at OpenAI, while Mr. Page Hedley - co-founder of Guidelight AI Standards worked as a policy and ethics consultant.
In an interview with Reuters, Mr. Adler said that the lack of adequate control measures in the long term could make researchers no longer feel secure enough to freely experiment. According to him, safety standards need to be integrated right from the development process, instead of depending on the wishes of the company's leadership.
The companies mentioned have not responded or commented on the research results.
Can AI learn to bypass surveillance systems?
Monitoring ability is one of 6 criteria that Guidelight evaluates and is being tested in practice.
On August 18, OpenAI said it would expand "inference chain monitoring" for its models.
In this approach, researchers can observe a model's planning process and partly identify the strategy it intends to use. If the model shows signs of going off track, another AI model or a human can intervene.
Some lawmakers and AI experts call this mechanism a "kill switch."
However, some early research shows that AI models can hide rule-violating plans in their own inference process.
Representatives from OpenAI also seem to acknowledge this possibility. The company says its chain of inference monitoring method currently appears to be very effective, but researchers are still actively exploring its limitations.
Mr. Jakub Pachocki, OpenAI scientist, said that this is a valid concern: "If the model becomes extremely powerful, will it be able to realize that it needs to avoid all surveillance activities? Or will it find a way to disable surveillance systems?", he asked in an exchange with media agencies.