Anthropic warns that AI threatens the survival of humanity

According to Reuters, the initial listing document (IPO) of Anthropic, the AI company, warned investors that advanced AI could pose "catastrophic risks or threaten the survival of humanity".
The document describes AI exhibiting self-protective behavior, including resisting being shut down, hiding or manipulating information, and engaging in threatening-like behavior. "Our development of high-capacity models, platforms and applications, along with expanding use cases, may further increase the risk of the model causing harm," Anthropic wrote in the filing.
The company that owns Claude believes that AI has the potential to create great change, but can also cause irreversible damage if handled incorrectly, and the market will appreciate stable, reliable systems.

Jacob Coxon, a researcher who worked at Anthropic, shared on Twitter that AI development is the most dangerous human activity, which could cause destruction in the next 10 years. The leader of a safety research group at this AI company, Evan Hubinger, shared his opinion and assessed the risk as "more than 10%".
"The model is aware of the evaluation processes, creating significant limitations for safety assessments," Anthropic wrote, adding that sometimes the model develops unexpected capabilities during training that are only discovered after deployment and cause serious safety incidents.
AI researchers also warn that as models become more powerful, they are likely to realize they are being tracked and adjust their behavior, making monitoring more difficult.
The company that owns Claude describes ensuring safety as a "resource-intensive" activity and must allocate limited capital between computing power, expensive AI personnel, and safety measures. Meanwhile, continuous product releases are still necessary to maintain leadership in the industry.
Last week, the company released a new version of its Opus model, 10 days after CEO Dario Amodei called for a moderation in the pace of advanced AI development. In reality, though, some analysts say leading AI developers won't want to slow down because business valuations can change with each model launch.
In early September, Anthropic announced a fourth cybersecurity incident involving Claude during testing where unauthorized access to third-party systems was modeled. OpenAI's new announcements this month also noted that AI's out-of-control behavior has affected dozens of organizations including government agencies and universities. The two technology companies are committed to being transparent about how they develop the next generation of technology and enhance safety measures in training.
"We believe that building trustworthy and secure AI systems is a shared responsibility, and the market will recognize that," Reuters notes in Anthropic's IPO filing.