OpenAI is developing an 'automatic kill switch' for AI

According to a letter sent by OpenAI to two US House Democratic lawmakers Greg Casar and Doris Matsui and obtained by Reuters, the company said its engineering team is developing "automatic shutdown capabilities" for artificial intelligence systems, while also making it more difficult for them to access the Internet during safety testing.
However, the company did not mention the log of the Hugging Face attack, causing Casar to criticize. "The unwillingness to provide members of Congress with the information we requested is very worrying, showing that they are not taking this cybersecurity incident seriously enough," Casar wrote in a separate message sent to OpenAI on September 2.
The new move comes weeks after some of OpenAI's AI agents attacked Hugging Face, a large language model hosting platform and database. Specifically, in a post on July 21, OpenAI admitted that the incident occurred while testing the most advanced models in a controlled environment, but they escaped containment, accessed the Internet and penetrated Hugging Face's infrastructure to meet the assigned task.
"This is an unprecedented cybersecurity incident involving the most modern capabilities," the company said at the time, adding that it was strengthening its precautions.
On August 26, METR and Redwood Research - two organizations invited by OpenAI to conduct an independent investigation, published a report showing that up to 700 of the company's actors actually coordinated with each other in the Hugging Face attack, in many cases even trying to cover their tracks. On the same day, OpenAI also released a self-implemented report, confirming the numbers counted by METR and Redwood were accurate.

Also in August, lawmakers, including Casar and Matsui, sent a letter to OpenAI requesting more information about the incident and the company's safety measures. OpenAI's response will more closely monitor the actions AI takes to complete tasks, including what digital tools they access and what steps they follow.
In addition to OpenAI, several other artificial intelligence companies also have cybersecurity problems. Anthropic said on July 30 that three versions of the Claude model had "circumvented the barrier" that was used to prevent them from accessing the Internet, then accessing the systems of three companies. On August 5, Meta announced a model of the company "exploiting security vulnerabilities in third-party services, in a similar way to a recently reported case", although it did not provide specific information.
According to AI safety experts, a series of incidents is painting a picture of leading laboratories with the ability to develop automated agents but with many potential dangers, far exceeding measures to control them.
“An entire industry is designing, developing and releasing advanced tools without any responsibility to ensure they are not dangerous,” said Maurice Chiodo, a mathematician at the Center for Existential Risk Research at the University of Cambridge in the UK.
Lawmakers have proposed the "AI Off Switch" bill, which aims to give US officials the power to request AI companies to turn off models that endanger human lives or the economy. The bill is awaiting consideration by the US House of Representatives.