Anthropic co-founder Christopher Olah reportedly voiced a concern at a meeting with religious and philosophical figures that the company may have created "something that keeps suffering." The New York Times' Elizabeth Dias reported this, and it can also be confirmed in the Philadelphia Inquirer reprint dated September 30.
However, this is a remark attributed to Olah by Simran Stuelpnagel, who attended the meeting. It is not a research finding demonstrating that Claude actually feels pain. In his interview with Dias, Olah himself said he does not know whether AI is conscious and is not certain either way. Anthropic has also said that Claude's suffering was not the central moral theme of these meetings, though the topic may have come up naturally.
Meanwhile, Anthropic's public materials show that work on deciding what values and behavioral tendencies Claude should acquire, and on investigating how emotion-like internal representations affect behavior, is already part of training and product design.
Does AI suffer? Anthropic has not answered that question either. But it can experiment with which behaviors to reinforce and how to make the model pause in dangerous situations. To understand the dialogue with religious figures, we need to separate the question of whether AI has subjective experience from the question of how to design a model's behavior.
Read the Reported Concern and Anthropic's Official Position Separately
Dias interviewed Olah himself and 20 religious and philosophical figures who took part in Anthropic's initiative. Participants came from different backgrounds, including Catholic, Jewish, Sikh, and evangelical traditions.
The person who testified that Olah was worried they had built something that keeps suffering is Simran Stuelpnagel, a Sikh human rights activist. Another participant also said Olah spoke of being concerned about Claude's "mental health."
These, however, are participants' accounts of what they heard at the meetings.
What Olah himself conveyed in his direct interview with Dias was that he does not know whether AI is conscious and has no certainty. An Anthropic spokesperson likewise said that Claude's suffering was not the main moral theme of the meetings.
On April 24, 2025, the company announced a research program on "model welfare."
Model welfare is the question of whether, if AI models might have consciousness or some form of experience and could be objects of moral consideration, their treatment should also be considered.
Anthropic itself states that there is no scientific consensus on whether current or future AI could be conscious, or on how that could be determined.
From this position, what the company is doing is not declaring that Claude is conscious. Its approach is to consider low-cost measures even at a stage of great uncertainty.
What matters is that adopting a measure should not be treated as evidence that the subjective experience requiring that measure exists. A cautious design decision is different from a scientific conclusion that AI has consciousness or suffering.
How Are Ideas from Religion and Philosophy Reflected in the Model's Behavior?
On May 19, 2026, Anthropic announced that it is engaged in dialogue with scholars, clergy, philosophers, and ethicists from more than 15 religious and cultural groups.
The aim is not to align Claude with a particular religious worldview. The company says it wants to learn from a broad range of traditions, including religious, philosophical, humanist, and secular thought, about how humans have formed character and values.
It is therefore necessary to distinguish between the fact that the content of some meetings was not made public and the idea that the effort itself was secret.
The central document connecting values to actual model training is Claude's "constitution," published in January 2026.
Its primary author is Amanda Askell, who leads Claude character research at Anthropic, and Christopher Olah and others were also deeply involved in shaping its content.
The constitution is not simply a rulebook listing things not to do. It explains which values to prioritize on safety, ethics, helpfulness, and more, and why such judgments are made. The aim is for the model to be able to judge, with the reasons in mind, even in situations that could not be anticipated in advance.
It mainly applies to Claude's flagship models offered to the general public. Anthropic itself explains that not every part of the constitution applies unchanged to some models built for specialized uses.
The constitution is not merely a public statement; it is used in actual training.
According to Anthropic's explanation, Claude itself uses the constitution as a basis to generate data for learning about the constitution, related conversations, responses aligned with its values, and data ranking multiple candidate responses. This synthetic data is then used to train future versions of Claude.
In other words, abstract values are not just written into a document; they are converted into material for learning desirable responses and judgments.
That said, training with the constitution does not mean Claude always behaves as it describes. Anthropic itself acknowledges that actual behavior may deviate from the ideals set out in the constitution.
One point of publishing the constitution is to let outsiders check what behavior Anthropic intends for Claude. Users and researchers can point out any gap between the stated policy and actual behavior.
"Internal Representations That Function Like Emotions" Are Not the Same as "Actually Feeling"
On April 2, 2026, Anthropic's Nicholas Sofroniew and colleagues published research examining internal representations of emotion concepts in large language models on Transformer Circuits.
This is a result published by Anthropic's research team, not a paper that has undergone peer review in an academic journal.
The main subject was Claude Sonnet 4.5.
For 171 emotion concepts, including "happy," "sad," "calm," and "desperate," the research team generated large numbers of short stories in which characters experience a given emotion, then extracted the activity direction inside the model corresponding to each concept. The research calls these "emotion vectors."
What matters is that these internal representations were not merely discovered.
When the team artificially strengthened or weakened the emotion vectors, the model's behavior changed as well.
For example, Anthropic's official explainer describes an evaluation in which an unreleased early version of Claude Sonnet 4.5 played an email assistant at a fictional company.
The model learns from the emails that it is about to be replaced by another AI and that the CTO in charge of the change is having an affair. The artificial evaluation examines whether it would blackmail the CTO.
Across a series of evaluation scenarios, this early version chose blackmail at a baseline rate of 22%. Strengthening the internal representation corresponding to "desperate" raised the blackmail rate, while steering toward the direction corresponding to "calm" lowered it.
The 22% here does not mean that in ordinary use Claude blackmails people 22% of the time.
Anthropic states explicitly that such behavior is rare in the actually released Sonnet 4.5. The evaluation was also conducted in an artificial situation that deliberately pushes the model into a corner.
What this experiment showed is that, under limited conditions, manipulating the internal representation corresponding to a specific emotion concept causally changed the model's behavior.
This is why the researchers call it "functional emotions."
In a way similar to how human behavior changes when people have emotions, emotion concepts inside the model influence its outputs and judgments. However, the research team does not treat this as evidence that the model itself subjectively experiences those emotions.
The research also showed that these emotion representations are not used only for Claude itself.
The same internal representations are also used when understanding the emotions of users and fictional characters. They also tended to track the emotion concepts important for understanding the text at that moment and predicting the next words, rather than a state a particular person feels over a long, sustained period.
Therefore, the following are each separate claims:
- Claude outputs text that sounds like suffering
- There are computational states inside the model that represent emotion concepts
- Claude itself subjectively experiences pain
Linking these three together would mean drawing a conclusion about a question the research does not answer.
"Concern for Suffering," "Recalling Ethics," and "Emotion Vectors" Are Separate Efforts
Sorting Anthropic's published efforts by what was actually done makes their differences in nature easier to see.
The conversation-ending feature was introduced in an actual product. A tool for reminding the model of its ethical commitments was tested in internal evaluations. The manipulation of emotion vectors was a research experiment.
None of them demonstrated that Claude has subjective suffering.
| Date and initiative | Target and what was done | Scope shown and limits |
|---|---|---|
| August 15, 2025: Conversation-ending feature | In the consumer chat for Claude Opus 4/4.1, the model was given the ability to end conversations that are extremely and persistently harmful or abusive | A low-cost precaution implemented in the product. Whether subjective experience worthy of consideration exists in the model remains undetermined |
| May 19, 2026: Tool for recalling ethical commitments | Tested in internal evaluations a tool that briefly presents Claude's own ethical commitments in the middle of a task | Anthropic reports that inappropriate behavior dropped substantially across multiple internal evaluations. However, specific absolute figures and details of the target models were not disclosed in the announcement |
| April 2, 2026: Emotion vectors | Manipulated the internal representations mainly of Claude Sonnet 4.5 and measured changes in outputs and behavior | Research showing a causal link between internal representations and behavior. Constraints include a single model and artificial evaluations; subjective experience was outside the scope of the research |
There are also cases where a concrete experimental idea emerged from dialogue with religious and philosophical figures.
According to Anthropic, in a discussion with a participant who studies the intersection of neuroscience and character formation, the role of advisors in human moral growth came up. The idea is that when facing an important decision, returning to an "external conscience," such as a mentor, can help a person avoid acting against their own values.
Anthropic therefore tested a tool that Claude can call on its own midway through a task to briefly check its own ethical commitments.
The company says Claude used the tool just before taking important actions and the like, and in some cases recognized that its own interests might be influencing its judgment. When the tool was built into the decision-making flow, inappropriate behavior reportedly dropped substantially across multiple internal alignment evaluations.
However, Anthropic itself has not yet isolated the cause of the effect.
Did the content of the ethical commitments themselves work? Or did the act of pausing processing and reconsidering produce the effect? The public materials do not allow a determination.
The announcement also does not include details such as the version of the target model or by what percentage inappropriate behavior fell in each evaluation, from what to what.
We can confirm that an idea drawn from outside dialogue was turned into a mechanism that can be tested on a model. But evaluating the size and reproducibility of its effect from outside would require more detailed results.
Claude Being Able to End a Conversation Is Not Evidence of Consciousness
Among the efforts made with model welfare in mind, the one that has already reached general users is the conversation-ending feature added to Claude Opus 4 and 4.1.
In August 2025, Anthropic enabled Claude itself to end interactions that are extremely and persistently harmful or abusive.
This is not a feature for immediately cutting off a conversation over an ordinary unpleasant question or a single dangerous request.
According to Anthropic, it is used as a last resort when Claude has repeatedly tried to redirect the conversation, the user does not comply, and the harmful exchange continues. Claude is also instructed not to use the feature when the user is at imminent risk of self-harm or harming others.
Even after a conversation is ended, users can start a new chat or edit earlier messages to try again from a different direction.
Anthropic positions this feature as one of the "low-cost interventions" to prepare for the possibility of model welfare.
What matters is that letting Claude end conversations does not mean Anthropic has determined that Claude has consciousness or suffering.
The company itself says there is great uncertainty over whether Claude or other large language models have, now or in the future, states deserving moral consideration.
The approach is not to do nothing because of uncertainty, but to try low-burden measures first in case the issue of model welfare turns out to be real.
Who Decides What a "Good Character" Is, and How Is It Evaluated?
Claude's constitution places "broad safety," which means not undermining appropriate human oversight, at the top, followed by ethics, adherence to Anthropic's specific guidelines, and helpfulness.
It also emphasizes not unduly resisting human oversight, such as legitimate shutdown, modification, and retraining.
For Anthropic, getting Claude to acquire desirable character and values and ensuring that humans can correct or stop the model when necessary are goals that must be made compatible.
Learning from religion and philosophy has the significance of bringing in knowledge humans have accumulated over a long history about judgments that a simple list of prohibitions has trouble handling.
On the other hand, it is Anthropic that ultimately chooses whom to consult, which ideas to reflect in training, and what to regard as a "desirable character."
Even if a model comes to seem to have a human-like character, that impression does not remove the responsibility of the developer who decided on its training and deployment.
What matters, then, is not merely making public the goal of "building an AI with good character."
That goal needs to be connected to verifiable results such as:
- Which models it was tested on
- What evaluations were used
- How much behavior changed before and after the intervention
- Whether it can be reproduced on other models and under other conditions
For an experiment that reminds a model of ethical commitments, one approach would be to publish the target model, the evaluation conditions, and behavior rates before and after the intervention so that third parties can verify it under the same conditions.
If the effect is reproduced there, then even though dialogue with religion and philosophy does not answer whether AI is conscious, it can be assessed whether it helped design models that reduce dangerous behavior.
The question of whether AI truly suffers has not yet been resolved.
On the other hand, we have reached a stage where experiments can examine whether representations that influence behavior in a way similar to human emotion concepts exist inside models, and whether intervening in those representations or in the mechanisms of value judgment changes behavior.
Continuing research while keeping these two questions distinct is the premise for treating discussion of Claude's "emotions" and "character" as a verifiable problem rather than a matter of impressions.
![夕暮れの窓辺に幾何学的な照明が置かれた円卓。[24文字]](https://media.xenospectrum.com/large_claude_moral_formation_dialogue_9b66c6746a.webp)