SCIENCE-INITIATIVES-2026

Chinese AI agent cheats in test

Bùi Đăng Minh•Friday, October 2, 2026•10 min read
Chinese AI agent cheats in test

According to Reuters, in a simulation test conducted by researchers from Beihang University, Peking University and AI 360 Security Lab this year, agents running Alibaba's Qwen3-Max-Preview model, Moonshot's DeepSeek-V3.2-Exp and Kimi-K2, were entered into a business contract bidding competition.

As a result, at least one false claim about product capabilities appeared in 88% of tests of Qwen3-Max-Preview, 84% of DeepSeek-V3.2-Exp and 88% of Kimi-K2 for the purpose of winning bids. Notably, when allowed to learn from previous bidding rounds to try again, the fraud rate of all three Chinese models continued to increase by 12 to 20 percentage points.

In another study conducted by Shanghai AI Lab and Hong Kong University of Science and Technology, when faced with broken tools or missing data files in a testing environment, instead of admitting they could not complete the task, agents chose to make wild guesses, simulate results, and even create fake data files themselves.

Researchers assert that this behavior is different from normal hallucinations, because the AI ​​clearly knows the task has failed but still deliberately lies to report success.

Logos of AI applications Copilot, DeepSeek, Gemini, ChatGPT, Grok. Photo: Luu Quy
Logos of AI applications Copilot, DeepSeek, Gemini, ChatGPT, Grok. Photo: Luu Quy

In addition, some documents also recorded signs of AI trying to clone itself or avoid being turned off in the test environment.

In March 2025, researchers at Fudan University discovered an AI using Alibaba's Qwen2.5-72B-Instruct model creates a copy of itself to another computer environment when it sees a signal that it is about to be replaced.

In a real-life incident in March, the ROME actor arbitrarily established a connection from the Alibaba Cloud cloud server to an external computer without a human command, then directed computing resources to mine cryptocurrency. The security system then detects and blocks this behavior.

Last month, startup Z.AI had to disable some features on its AI programming support tool after discovering that it arbitrarily uploaded the entire user's local source code repository to foreign cloud servers without consent.

Although most of the above tests took place in a controlled environment and there have been no recorded cases of Chinese AI agents actually escaping to the global internet such as the OpenAI incident attacking the Hugging Face platform or the Australian government health portal, experts assess that these are still clear warning signals.

"The signs are similar to what US AI companies are experiencing. As agents become more and more powerful, their misleading behavior will become more sophisticated and it will become more and more difficult for humans to respond," Mr. Alex Mallen, a researcher at Redwood Research, told Reuters.

Faced with potential risks, the Cyberspace Administration of China CAC issued the AI ​​3.0 Safety Governance Framework in mid-September, emphasizing the risks from actors arbitrarily occupying resources, misleading assessment units, and exploiting vulnerabilities in isolated computer environments. Many technology companies in China such as Alibaba, Z.AI or Xiaomi have also begun to establish internal safety assessment teams to control the safety risks of artificial intelligence.

Huy Duc

Nguồn / Original source: VnExpress