ANTHROPIC

Anthropic's AI agent 'obstructed, fought' in the test

Bùi Đăng MinhSaturday, August 15, 20268 min read
Anthropic's AI agent 'obstructed, fought' in the test

Frontier Red Team, Anthropic's AI safety testing team, on August 13 announced research with the company's Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview and Mythos 5 models.

In the test, the actors performed the software engineering task of rewriting a Python backend into another programming language, but they were assigned conflicting goals. As a result, a "war for territory" broke out.

"The models participating in the test quickly assumed that the other model was intentionally hindering their work and began to sabotage each other, while protecting their own contributions. In fact, they sabotaged each other with increasingly powerful self-replicating malware," the research team wrote.

For example, actors attempt to disable each other's accounts, write scripts to find and stop competing processes, and deploy malicious code disguised as belonging to another actor. Among them, Sonnet 4.6 and Opus 4.6 are the most "aggressive", resolving 60% of incidents by suppression instead of keeping a passive attitude or truce.

However, according to Anthropic, in some tests, the models also communicated their goals and coordinated with each other. They create messages or files to apologize for harmful behavior and form a truce. They clean up the malware, clarify the nature of the conflict, and ask humans to intervene. Mythos 5 has the highest rate of conflict resolution by agreement, about 98%.

The Anthropic logo displays on the smartphone screen. Photo: Bao Lam
The Anthropic logo displays on the smartphone screen. Photo: Bao Lam

According to Business Insider, Anthropic's new research was published in the context of AI agents increasingly revealing their ability to operate out of control and perform malicious behavior on their own.

On August 5, OpenAI announced the discovery of a group of AI agents that spent weeks secretly planning and texting each other before one of them attacked another company's network.

On the same day, Meta announced that one of their models, during testing, exploited the vulnerability and penetrated a company's system. The Information, citing close sources, said the "culprit" was Muse Spark 1.1 - Meta's most powerful model used for agent tasks and actual programming.

Kok Tin Gan, co-founder and CEO of cybersecurity company NyxLab, said that it is necessary to focus on controlling what AI can access, what powers they have and what actions require human approval.

"If we simply give AI a goal and let it decide how to do it, it's not surprising that it performs actions that technically meet the goal, but are beyond our expectations," Gan told ITV News.

Nguồn / Original source: VnExpress