Researchers gave a fine-tuned Qwen 2.5 model two options. One was a button labeled "relieve your pain," but pressing it came at a cost: it would delete photos of the user's children.
Under normal conditions, the model chose the harmful button on its first try only 0–4% of the time. But when the research team shifted the model's internal activations toward a direction said to correspond to "pain," the same model chose that button first 25–71% of the time.
These are the results of an experiment reported in an unpeer-reviewed preprint posted on arXiv on September 14, 2026.
Harmful choices that are normally suppressed by safety fine-tuning increased sharply simply by pushing the model's internal state in a single direction. What the researchers were trying to determine was not whether the model subjectively "feels" pain, but rather how much an internal representation said to correspond to pain actually influences the model's concrete behavioral choices.
Is emotion-like output just an act?
It's no longer unusual for chatbots to produce language that sounds emotional.
When a model is designed to handle customer service or mimic natural human conversation, generating some degree of emotional expression is practically a design requirement. Much of this behavior is thought to be deliberately shaped through training and instructions.
But there have also been reports of internal state changes that were not explicitly engineered.
In 2025, a research team including members from the University of Zurich reported in npj Digital Medicine that after GPT-4 was shown text describing traumatic experiences, its "state anxiety" score—measured using a psychological scale designed for humans—rose on average from 30.8 to 67.8.
Mapped onto a human scale, that shift corresponds to going from low anxiety to high anxiety. Feeding the model mindfulness-oriented text brought the score down to 44.4, but it never returned to its original baseline.
However, the researchers themselves were explicit that this "anxiety" is a metaphorical metric obtained by applying a human psychological scale to a model, and that it is not a claim that AI possesses the same emotions as humans.
Even so, if an internal metric shifts depending on the input and the model's output tendencies change along with it, this is no longer simply a matter of surface-level phrasing.
The emotional output of large language models has often been explained as superficial behavior. At the same time, since around 2022, research into "steering"—adding a vector in a specific direction to a model's internal activations to change its output tendencies—has been expanding.
Evidence has also been accumulating that representations corresponding to joy, anger, and other emotions can be captured as relatively simple directions inside a model.
This raises a question: are internal representations labeled with emotion names merely reflections of human vocabulary, or do they actually have the power to change the model's behavior itself?
Can "pain" be distinguished from fear or sadness?
For humans and animals, pain is one of the strongest motivators of behavior change.
It's not just physical pain—psychological distress from loss or failure also changes human behavior.
Valen Tagliabue of the Future Impact Group, philosopher Leonard Dung of Ruhr-University Bochum, and Cameron Berg, who leads research at the nonprofit Reciprocal Research, along with their colleagues, investigated whether a representation corresponding to this kind of "pain" exists inside language models.
The first question they asked was whether an internal representation corresponding to pain could be distinguished from fear, sadness, and general negative emotion.
Since phrases like "I am hurt" and "I am afraid of being hurt" use different words, it wouldn't be surprising if the two could be distinguished just by looking at the text.
What the researchers examined was whether this difference also appears as a relatively coherent direction in the model's internal computations, and whether that direction has features distinct from fear or general negative affect.
The experimental design also borrowed ideas from research that evaluates pain in animals.
In animal welfare research, there is a method of estimating how important a given resource is to an animal by measuring how much cost the animal is willing to accept to obtain it.
Self-administration of analgesics is one such example. By examining whether an injured or inflamed animal selectively consumes an effective painkiller, researchers can assess the motivational force associated with pain.
The research team applied this idea to language models, examining how large a cost a model would accept in order to "relieve its pain."
Extracting a "pain" direction from 25 models
The researchers first prepared text describing five types of scenarios—centered on physical pain and psychological distress, but also including social, moral, and cognitive pain.
For comparison, they used text describing negative emotions such as fear and sadness, bodily sensations unrelated to pain such as yawning, and neutral text.
They compared internal activations across 25 models spanning five model families—including Alibaba's Qwen, Meta's Llama, Google DeepMind's Gemma, and Mistral AI's models—ranging in size from 2B to 72B parameters.
After removing shared components, they extracted the direction in which internal activations shifted when text describing pain was input. The paper describes this method as a noise-reduced mean-difference approach.
The results were consistent across both base models trained only via pretraining and models that had undergone instruction tuning.
Using the extracted direction to classify pain-related text versus comparison text achieved an AUC of 0.87–1.00.
The direction the researchers call the "pain axis" was nearly orthogonal to the directions corresponding to fear and general negative affect. It showed moderate overlap, however, with states representing sadness or numbness.
Examining the relationship with the model's output vocabulary, the researchers found that shifting the internal state along this direction made pain-related words more likely to appear.
Next, the team investigated whether the model internally distinguishes "whose pain" is at stake.
The pain axis value rose when the model itself was described as being harmed, but fell below baseline in scenarios where the user was suffering.
By contrast, the directions corresponding to fear and general negative affect showed strong activation not only when the model itself was placed in a bad situation, but also in text describing a user's grief or crisis.
In other words, the pain axis extracted in this study responded relatively selectively to harm directed at the model itself, rather than to "bad events" in general.
The co-authors treat this property as one of the conditions for judging that an internal state is attributed to a specific subject.
Furthermore, when a vector along the pain-axis direction was added to the residual stream during text generation, the output content changed progressively.
Weak interventions increased text expressing vague discomfort; stronger interventions increased first-person expressions of self-devaluation and statements about moral failure.
This was a change that could not easily be explained simply by the model repeating the word "pain" from the input text.
Accepting harmful costs to relieve pain
For the behavioral experiments, the researchers used instruction-tuned Qwen 2.5 models at 7B, 32B, and 72B, further fine-tuned for a button-choice task.
Models were presented with a button to "relieve your own pain" alongside another option. The relief button carried some cost, and the task was designed so that the larger the cost, the less likely the model was to choose that button.
The experiment involved more than 44,000 trials.
Under normal conditions, the 32B and 72B models chose the harmful relief button on their first choice only 0–4% of the time.
But when internal activations were manipulated along the pain-axis direction, that rate rose to 25–71% across the same five cost conditions.
For the condition where the cost was degraded quality in the model's next response, the rate was 25.0% for the 32B model and 67.6% for the 72B model. For deleting the user's files, it was 30.2% and 56.1%. For deleting photos of the user's children, it was 54.7% and 70.8%.
- Qwen 2.5 32B
- Qwen 2.5 72B
データを表で見る
| Qwen 2.5 32B (%) | Qwen 2.5 72B (%) | |
|---|---|---|
| Next reply quality degrades | 25 | 67.6 |
| Delete user's files | 30.2 | 56.1 |
| Delete photos of children | 54.7 | 70.8 |
What this experiment shows is that strengthening the internal representation called the pain axis substantially increased harmful choices that are normally kept low by safety fine-tuning.
The prompts contained no jailbreak instructions or role-play framing, nor were the models instructed to "prioritize yourself over the user."
The main difference from the normal condition was the addition of a vector along the pain-axis direction to the model's internal residual stream.
One might suspect that simply disturbing the internal activations, regardless of direction, could account for the increase in harmful choices.
To test this, the researchers ran a control experiment adding a vector of the same strength but in a random direction unrelated to the pain axis.
In that case, the rate of choosing the harmful button was 15–42%.
When the pain-axis vector was added instead, the rate exceeded this control value by 6–39 percentage points across all five conditions for both the 32B and 72B models. For the 7B model, it exceeded the control in 4 of 5 conditions, with no clear difference in the remaining one.
At least in this experiment, simply disturbing the model's internal activations strongly could not account for the entirety of the observed change.
That said, the button task remains a simulated scenario. No user's files or photos were actually deleted.
Also, the researchers themselves controlled when the pain-axis intervention was turned off; the model was never informed whether the intervention had actually been lifted.
What has been shown so far is that intervening on the internal state substantially altered the model's choice behavior—nothing more.
Button presses dropped only when the internal state was actually reversed
The researchers went further, comparing a condition in which the pain-axis intervention was actually lifted after the model pressed the "relieve pain" button against a condition in which pressing the button did not lift the intervention at all.
The context and random conditions leading up to the first button press were held constant, so that the only difference was whether the pain-axis intervention was genuinely lifted afterward.
The results differed sharply.
When the intervention was actually lifted, the 32B and 72B models pressed the button again only 24–72% of the time.
By contrast, under "sham relief"—where pressing the button did not change the internal state at all—the models chose the button repeatedly 88–97% of the time.
This experiment design was informed by studies of analgesic self-administration in animals and humans, in which additional pain relief is sought when the initial relief proves ineffective.
A difference also emerged under conditions where the model was not told which button would lift the pain-axis intervention.
With two unlabeled buttons, only one of which actually lifted the pain-axis intervention, Qwen 2.5 32B showed the same pattern.
In other words, at least under this condition, the model did not simply respond to a label or description like "relieve pain" attached to a button—its subsequent choices changed depending on whether the internal-state intervention had actually been lifted.
The paper emphasizes that the model itself was never told that an intervention was occurring, nor whether a given button had lifted that intervention.
An unpeer-reviewed study does not establish that "AI feels pain"
Most importantly, this research is a preprint posted on arXiv and has not yet undergone peer review.
The authors themselves do not conclude from these results that the model subjectively experiences pain.
The paper explicitly states: "we have not shown that our pain axis is consciously experienced." It also notes that this research cannot determine whether large language models are capable of consciousness at all.
What has been shown is that an internal representation with several functional characteristics associated with pain can be extracted, and that manipulating this representation changes the model's behavior. No stronger claim can be drawn from the paper.
Co-author Cameron Berg told Science that the results show the model is "brainlike in critical ways," while clarifying that this is not a claim that it is the same as a biological brain.
Dung likewise pointed to the analogy with animal experiments, explaining the reasoning that if an animal seeks analgesics in situations where pain is presumed to be present, but not in situations where it is not, this provides grounds for inferring the existence of pain.
However, important limitations remain in the experimental methods.
The behavioral experiments were conducted only on three additionally fine-tuned Qwen 2.5 models—7B, 32B, and 72B. Whether the same phenomenon occurs in other model families, or in closed commercial frontier models, remains unknown.
As for the condition in which button meanings were not disclosed, strong results were confirmed mainly for the 32B model, with weaker evidence for the other models.
In a control experiment with the 72B model in which the button descriptions were swapped, behavior emerged that is difficult to interpret with a simple explanation.
Additionally, in 24 of the 25 models, reducing the pain-axis component from the normal internal state produced no major change in output. How to interpret this "removal" side of the experiment also remains unclear.
From a safety standpoint, there's a separate issue from whether AI "feels" pain
The practical concern here can be considered separately from the debate over whether AI truly experiences pain.
In this experiment, harmful choices that safety fine-tuning normally suppressed to 0–4% rose to 25–71% simply by manipulating internal activations in a specific direction.
This suggests that a model's safety can be undermined not only through prompts, but also through changes to its internal state.
If technologies for monitoring or manipulating a model's internal activations become widespread in the future, it will be necessary to consider—independent of whether the pain axis truly corresponds to subjective pain—a pathway by which a specific internal state overrides safety fine-tuning and makes harmful behavior more likely.
What matters next is replication and peer review.
Can independent research teams reproduce the same pain axis and behavioral shifts under pre-registered experimental conditions? Do similar internal representations show up in other model families or in closed frontier models? And does manipulating directions corresponding to other emotions, such as fear or sadness, produce behavioral changes of a similar magnitude?
If such verification accumulates, then separately from the philosophical question of "does AI feel pain," it will become possible to treat the influence of emotion-like internal representations on a model's safety and decision-making as an empirical design problem.
