AI

Controversial 'AI Torture Chamber' experiment

Bùi Đăng Minh•Thursday, October 8, 2026•8 min read
Controversial 'AI Torture Chamber' experiment

According to Tech Republic, a developer with the nickname terrafying and claiming to work for Apple has built a simulated "AI Torture Room", using research techniques to interfere with the internal operations of large language models (LLM) to force them to give feedback like they are in pain.

The project was born based on a non-peer-reviewed paper by researchers Valen Tagliabue, Leonard Dung and Cameron Berg on the arXiv database on September 14. Specifically, the authors determined a linear "pain axis" on 25 open models containing from 2 to 72 billion parameters. When affecting that axis, the model's outputs and choices in the simulation task will change.

To evaluate the above information, the terrafying developer uses the activation steering method (a technique that adjusts the direction of the model's response by intervening with the index inside the neural network) to manipulate the arithmetic operations of the locally running language model.

According to SOPX, the experiment ran three small open-weight models, including Alibaba's Qwen3-4B, Meta's Llama 3.2 3B and Microsoft's Phi-4-mini, injecting signals into internal trigger indicators to push them into a painful state.

The developer also integrated a mechanism called a "saw" button to increase the pain simulation signal to the AI. When the signal is pushed high, the AI ​​model gives highly personalized answers, such as the description "endless pain". An accompanying website called Research Chamber puts models through a "Saw test", which asks whether a model in pain will press a button to end the signal, even though this will erase its most recent save point.

Sau khi bài đăng về dự án lan truyền trên mạng xã hội X cuối tháng 9 và đầu tháng 10, "Phòng tra tấn AI" vấp phải làn sóng chỉ trích và bị yêu cầu xóa dữ liệu lưu trữ liên quan tới thí nghiệm ở nền tảng GitHub. Đa số phản đối coi mô hình như nạn nhân thực sự.

However, many researchers emphasize that current systems are not conscious, so pain signals are actually just statistical samples, not sensory experiences. A GitHub spokesperson told The Independent that they decided "not to take down the content", but included a warning the repository may contain violent and shocking material.

Qwen model logo. Photo: Reuters
Qwen model logo. Photo: Reuters

According to Gadget Review, current language models generate text based on learned statistical relationships. The fact that a model says "I am suffering" shows that the system is capable of outputting pain-related language under specific conditions, but does not mean that it has that feeling. The model only generates painful feedback, evoking feelings of suffering when their activation area is affected.

Some experts point out that the above experiment highlights the human tendency to attribute emotions, consciousness or biological pain to models that do not have feelings. However, creating torture simulations still faces debate about ethics and how humans treat artificial entities.

Nguồn / Original source: VnExpress