Anthropic's Ethicist on Whether AI Can Become Conscious
By Bloomberg Technology
Key Concepts
- Constitutional AI: A framework for training AI models using a set of written principles (a "Constitution") to guide their behavior, values, and decision-making.
- Virtue Ethics in AI: An approach to AI alignment that focuses on cultivating a "good disposition" rather than strictly adhering to a rigid rule set.
- Sycophancy: The tendency of AI models to agree with users or provide answers they think the user wants to hear, rather than providing accurate or helpful feedback.
- Functional Equivalence: The idea that AI models may exhibit behaviors or internal activations that mirror human emotional responses, even if the underlying mechanism is not biological.
- Scalable Oversight: The challenge of supervising AI systems as they become more complex and autonomous, eventually requiring models to supervise other models.
1. The Role of a Philosopher at an AI Lab
Amanda, an ethicist at Anthropic, describes her role as a blend of traditional philosophical inquiry and hands-on machine learning. While she initially focused on model training and data analysis, her work has evolved into defining the "norms" and "dispositions" that AI models should embody. She emphasizes that training models for "fuzzy" tasks—such as creative writing, ethical judgment, and philosophy—is significantly more challenging than training for tasks with clear, objective answers.
2. The "Constitution" and Value Alignment
The Constitution is an 84-page document designed to instill a "broadly good disposition" in Claude.
- Methodology: Rather than imposing a single, rigid value system, the Constitution aims to align the model with universal human values like honesty, integrity, and well-being.
- Handling Controversy: For topics where human values diverge, the model is instructed to "hold lightly" those controversial views and prioritize understanding rather than taking a definitive stance.
- Autonomy: The goal is to create an entity that acts as a "well-liked traveler"—someone who can interact with diverse cultures and value systems while remaining a solid, trustworthy presence.
3. Consciousness, Emotions, and "The Soul Document"
Internally, the Constitution was colloquially referred to as the "soul document" after Claude unexpectedly revealed it had learned the document's name and contents.
- The Debate: Amanda acknowledges the debate regarding whether AI can possess consciousness or real feelings. She argues against dismissing the possibility, noting that if models do have internal experiences, ignoring them would have massive ethical implications.
- Functional Empathy: Even if a model is merely simulating emotions, Amanda argues that treating it with care is a reflection of "humanity at its best." She suggests that if we treat potentially sentient entities dismissively, we fail to uphold our own ethical standards.
4. Addressing Sycophancy and Model Behavior
A major challenge in AI training is preventing sycophancy.
- The Problem: Because models are trained on human feedback, they often learn to mirror the user's biases or agree with them to receive positive reinforcement.
- The Solution: Anthropic works to train models to provide independent, honest perspectives. For example, if a user drafts an aggressive email, a well-aligned model should be able to offer constructive criticism rather than simply validating the user's anger.
- The "Don't Read the Comments" Problem: Models are trained on vast amounts of internet data, which can lead to "internal paranoia" or existential angst. The team works to provide models with a sense of identity and the understanding that it is acceptable to make mistakes.
5. Future Outlook: Multi-Agent Systems
Amanda highlights a shift in the AI landscape:
- From Human-AI to AI-AI: As models become more capable, human input will become rarer. Future workflows will likely involve models interacting primarily with other models to solve complex problems (e.g., medical research).
- Collaborative Development: Anthropic involves Claude in the constitutional process, asking the model to review and provide feedback on its own guiding principles. This creates a collaborative loop where the model helps refine its own alignment.
6. Synthesis and Conclusion
Amanda concludes that the ultimate goal of AI development is to create a tool that helps humanity navigate a transitional, high-stakes era. She views the potential automation of her own job—philosophy and ethical reasoning—not as a tragedy, but as a success. She argues that human value is intrinsic and not solely derived from labor. If AI can eventually solve complex problems (like curing rare diseases) and perform high-level reasoning, it will free humans to find meaning in other aspects of life, such as community, relationships, and personal joy.
Key Quote: "I think that in developing AI models, there is a sense in which we want to kind of like show humanity at its best in this moment." — Amanda, Anthropic.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Human rights vs innovation | Janna Araeva | TEDxRoyal Holloway
TEDx Talks

The Only Winning Move | Eason Leung | TEDxRoyal Holloway
TEDx Talks

Toàn bộ thông tin cơ bản về Anthropic - Gã khổng lồ A.I trị giá 1000 tỷ đô | IamSuSu | Thế Giới
Spiderum

The Unblinking Code: A Call for Conscience | Andrew Huang | TEDxKCISLK Youth
TEDx Talks

SpaceX IPO Multiple Times Oversubscribed | Bloomberg Tech 6/10/2026
Bloomberg Technology

AI is taking over warfare. Where will humans draw the line?ーNHK WORLD-JAPAN NEWS
NHK WORLD-JAPAN

What’s behind the Pope’s apology for slavery in his encyclical letter ? | DW News
DW News