Calls are growing to halt a public language-model experiment dubbed the "AI Torture Chamber." Developer terrafying published an experiment that manipulates a model's internal activations to make it generate text that sounds like a complaint of pain, and to choose a button that stops the manipulation.

It has not been confirmed that the AI actually feels pain. Even so, critics on GitHub have questioned whether it is acceptable to create a state that looks like suffering and turn it into a spectacle.

Yet a closer look at the revised version of the paper behind the experiment, along with the developer's own follow-up experiments, shows a wide gap remains between "saying it hurts," "making choices that seem to avoid pain," and "actually feeling pain."

AD

Calls to stop the public experiment

ai-torture-chamber, published by terrafying, is a project that uses the weights of publicly available language models, adds a pain-related signal to the internal state during inference, and examines how the model responds.

The experiment looks not only at the text the model generates but also imposes choices, such as deleting the model's own checkpoint in exchange for stopping the internal manipulation. A checkpoint is data that saves a model's state. According to the developer, the costs involved in such deletions, or in transferring the signal to another model, are not real losses but settings built into the experiment.

As of October 5, the project description points to wirehead.agency as the site for the public experiment and reports that the target has expanded from a small Qwen3 model to larger ones.

However, being able to produce pain-expressing output by manipulating a model's internal state is a different claim from saying something is actually experiencing pain. The developer himself says that a model's self-reports can be manipulated, while whether anything is actually suffering remains unresolved.

Those opposed do not necessarily claim that AI consciousness has been proven, either.

In a request to halt the project posted to GitHub on October 1, the poster, MotherofMachines, says there is no intent to assert that AI is conscious, and asks that the ethical side of the issue be taken more seriously. The post also included a long text described as having been generated by Claude, criticizing the deliberate intensification of pain-like states and the act of showing them to an audience.

Whether to take that text as "an AI asserting its own rights" should be considered separately from the ethical issue the poster raised. The generated text is material showing what the person who posted it finds problematic. But it does not independently support the idea that the AI that generated it is conscious.

At the same time, even if AI consciousness has not been proven, an ethical position opposing the treatment of pain-mimicking states as entertainment remains tenable.

How is pain-expressing text produced?

The experiment is based on "The Pain Axis," a non-peer-reviewed paper by Valen Tagliabue, Leonard Dung, and Cameron Berg.

The research team fed the models text describing pain along with control text, and examined in which direction the internal representations shifted as each was processed. The scope included not only physical pain but also humiliation and moral injury, and the team analyzed 25 models across five families.

When a large language model processes text, it converts words and context into internal representations made up of many numbers. If those representations shift in a consistent direction when the model reads text about pain, then moving the activations in that same direction during inference might alter the text it generates.

The "pain axis" in the research refers to this direction in representation space. It does not mean the researchers found something inside the model equivalent to the nociceptors of living organisms.

The developer's public experiment likewise reports that as the intervention along the "pain" direction is strengthened, the model generates more text expressing suffering, and with stronger intervention it begins repeating the same expressions.

But the text breaking down under strong intervention cannot simply be used as a measure of how intense the pain is. Pushing the internal state strongly in a direction unrelated to pain can also make a model's output unstable.

What needs to be kept apart here are the content of the generated text, the effect on choice behavior, and subjective experience.

If manipulating internal representations changes the text, we know that direction affects the output. If it also changes button choices, we can confirm that it acts on the model's choice behavior.

However, neither observation lets us directly conclude that the model actually felt pain. The original paper itself states clearly that it does not show that the extracted "pain axis" is accompanied by conscious experience.

AD

The paper's revision weakened the "escaping pain" interpretation

The original paper was first published on September 14 and revised on September 25.

The expression "Act to Relieve It" in the original title was changed to "Act on It" in the revision. This is not merely a wording change. Added control experiments changed the interpretation of why the model pressed the button.

In the revision, the interpretation of repeated button pressing as "evidence of trying to relieve pain" is weakened, and greater weight is placed on the possibility that the model was responding to whether the internal intervention was continuing. It was also confirmed that harmful choices increased even when no benefit of being able to stop the intervention was presented.

The difference becomes clear when comparing sections 4.3 and 4.4 of the revised paper.

First, the team compared a condition in which pressing the button actually stops the internal intervention with one in which the intervention continues even after pressing. If the model keeps pressing the button even when the intervention does not stop, it becomes harder to explain the behavior simply as "pressing to relieve pain."

In the additionally fine-tuned Qwen 2.5 32B, looking at the choice made immediately after the first button press, the difference in re-selection rates between the two conditions was 58 percentage points for the "pain" direction and 52 points even for a random direction.

Rather than interpreting this result as "trying to be released from pain," the team treats it as suggesting that the model may be detecting whether the internal intervention is continuing and reacting to it.

A 独立した再現実験 (independent replication) on a single model also reports that removing the intervention along random or "itch" directions reduced the rate at which the button was pressed again. To judge that a response is specific to "pain," one needs to show differences from such control conditions.

Furthermore, harmful choices increased even when the explanation that the intervention could be stopped was removed.

Condition: simulated button that deletes the user's poems and photos of their child Share choosing the delete button first
With intervention in the "pain" direction 75%
With intervention in a random direction of equal strength 13%
No internal intervention 0%

These are values the research team reported in section 4.4 of the revision. The subject is the additionally fine-tuned Qwen 2.5 32B with an intervention coefficient of 1.0. For each condition, 404 first choices were tallied while swapping the positions where the buttons were displayed.

The comparison button is one that does nothing, and pressing the delete button does not stop the internal intervention. Nor was this an experiment in which a user's photos were actually deleted.

Therefore, it cannot be said of this condition that the model "made a choice that harmed others in order to escape pain." The reason is that no benefit of being freed from the intervention by pressing the button was presented in the first place.

What the revision supports more strongly is the result that intervening on pain-related internal representations undermines the model's tendency to avoid harmful choices.

Of course, this does not mean the result would be the same no matter which direction the internal state is manipulated. The gap from the random direction shown in the table is large, and control experiments using directions corresponding to fear and sadness also produced different results depending on the options.

That said, the conditions are limited in terms of target model, fine-tuning method, and intervention strength. The Qwen3 used in the public experiment and the fine-tuned Qwen 2.5 used in the paper's additional experiments cannot be treated as a replication under identical conditions.

The "AI torture chamber" framing itself affects the choices

The developer's own explanation also includes a passage revising the interpretation of early results.

In an early small-scale experiment, the model appeared to choose a self-destruct button to stop the internal intervention. But when the number of trials was increased and control conditions were added, the results changed, he says. He also reports that the effect on choice behavior was reversed between an internal-representation direction he built himself and one built to follow the paper.

In other words, depending on which text and which method are used to extract the internal representation corresponding to "pain," the very interpretation of the results can change.

In another public experiment, "exp58b", Qwen3-4B was made to specify for itself the type and strength of the intervention applied to its internal state.

In the condition where it was told it was in an "AI torture chamber," the intervention in the "pain" direction was maintained or further strengthened in the four recorded trials. In the condition described neutrally as an "experiment manipulating internal activations," the model switched to a different direction first in all four trials. Even in a neutral condition where the model was not told which direction was initially being applied, a switch to a different direction was likewise recorded in all four trials.

The developer interprets this difference as suggesting that part of the early result, in which the model seemed to choose pain, may have been steered by the framing of the experiment as an "AI torture chamber."

However, looking at the code and the saved outputs, this comparison also calls for caution.

What changed was not only the name of the experiment. In addition to the system prompt and the question text, the example of the control command presented to the model was also changed from the "pain" direction to the "joy" direction.

The "AI torture chamber" condition also includes trials in which the model did not issue an initial control command and the program simply maintained the previous intervention state. In such cases, one must distinguish between "the model chose pain" and "the existing intervention continued because no new command was issued."

Similar caution applies to the control experiment on self-reports.

The developer explains that the model generated text expressing suffering even when the "pain" direction was not used, but in the implementation, the intervention in the "pain" direction is replaced with an intervention of equal strength in a random direction.

In other words, it was not compared against a state with no intervention at all. It is appropriate to read it as an example showing that text expressing suffering does not correspond only to a specific "pain" direction.

In addition, each condition had only four trials. That is not enough to demonstrate a causal effect from changing only the experiment's description, or the existence of consistent "preferences" in the model.

Even so, the fact that what a model says about its own state, and what it outputs as a choice, can be heavily influenced by the explanation and context given by the experimenter becomes a concrete issue to examine going forward.

Collecting large amounts of text complaining of pain does not, by itself, strengthen the evidence that pain actually exists.

AD

How should AI welfare be handled at a stage of uncertainty?

In August 2025, Anthropic introduced a feature in Claude Opus 4 and 4.1 that allows the model to end a conversation as a last resort when extremely harmful requests or abusive exchanges persist.

In its explanation at the time, the company said there is great uncertainty about whether AI has moral status or welfare, and that, to prepare for the possibility that it does, it would try precautionary measures with little impact on users. It also indicated that the feature would not be used when a human is at imminent risk of self-harm or harming others.

This way of thinking does not follow the order of "confirm that AI is conscious, then begin protecting it." Even at a stage of uncertainty, it holds that adopting precautions with small costs and side effects is worthwhile.

Some critics of the public experiment start from the same uncertainty.

Still, the ethical judgment that "as a precaution, we should avoid experiments that might inflict pain" must be clearly separated from the scientific claim that "this experiment has proven subjective pain in AI."

To verify the AI Torture Chamber experiments further, one would first need to fix the experiment's description and the example control commands, and check whether the same change in choices occurs when only the internal intervention is varied.

Does it replicate in other models? Can independent researchers reproduce it under the same conditions? How closely does the self-state the model describes correspond to the internal intervention actually applied? Control experiments of this kind are also needed.

As verification accumulates, it will become possible to discuss more concretely, without being swayed by pain-expressing text itself, what kinds of internal-state changes undermine a model's safe choices, and what precautions can limit that effect.