Amanda Askell: Architect of Constitutional AI Ethics

An episode of Dan's AI Intel

<p>This episode explores Amanda Askell’s role as Anthropic’s lead philosopher and one of the key people shaping Claude’s values, personality, and constitutional alignment approach. The document traces her background in philosophy, ethics, and decision theory, and explains how that translates into Anthropic’s effort to teach AI systems not just what behaviours to follow, but why those behaviours matter. A central theme is her belief that AI alignment should be grounded in broad principles, practical judgment, honesty, humility, and care, rather than only rigid rules or opaque guardrails.</p><p>The episode then examines Askell’s distinctive framing of Claude as more than a tool: a quasi-agent with character, uncertainty, and possibly future moral relevance, while still requiring careful design and oversight. It covers Claude’s 23,000-word Constitution, Anthropic’s transparency-first approach, the tension between universal and culturally specific values, the risks of over- or under-anthropomorphising AI, and the open question of whether constitutional alignment can scale to more agentic future systems. This podcast was created with NotebookLM for my own learning purposes, using the source document as a structured guide to understand Askell’s thinking, her role in Anthropic’s alignment philosophy, and the broader question of how AI systems should learn values.</p><p></p>

Published · Updated · By Dan Walter

Executive Summary Amanda Askell is Anthropic's lead philosopher and personality alignment lead, who has shaped Claude's values and ethics. In her public writings and talks she argues for a "constitutional" approach to AI: explicitly teaching models why to behave well, not just telling them what to do 1 2. She emphasizes virtue ethics and human-like judgment ("nuanced, rich conception of what it is to be good" 3) over rigid rules. In interviews she stresses treating Claude as an almost-human agent, balancing transparency with care ("Claude is almost the primary audience” of its constitution 4) and acknowledging uncertainty about AI consciousness. Compared to peers, Askell is unusually frank about Anthropic's role and values: she published Claude's 23k-word “Constitution" under CCO for transparency 1 5 and publicly contemplates Claude's moral status 6. Key claims in her corpus include that large models can learn broad human values (e.g. "do what's best for humanity") and that safety improves when we explain the reasoning behind rules 2 7. She is optimistic about alignment (rejecting "evil attractor" fears 8 ) but also cautious, explicitly advocating corporate and community responsibility for AI values 9 10.

Actionable Recommendation: For AI stakeholders, Askell's work suggests prioritising transparency and pluralistic input in value-setting. Continue publishing alignment documents and datasets (as Anthropic did with Claude's Constitution and the Values-in-the-Wild study 5). Expand feedback on AI values beyond Western norms (she likens ideal AI to a “well-liked traveler" with broadly shared values 11). In interviews or podcasts, probe how Anthropic plans to update its constitution over time, and how its approach scales to future, more agentic models.

1. Biographical Sketch Amanda Askell is a philosopher by training (BPhil Oxford, PhD NYU 2018) who co-founded Anthropic in 2021 after a stint on OpenAI's policy team 12 13. She leads Anthropic's "Personality Alignment" group and is known as the "Claude Whisperer" 14. Her academic work has long focused on ethics and decision theory; at Anthropic she applies this to AI safety. Notably, she was co-author on Anthropic's Constitutional AI papers (e.g. Bai et al 2022 15, Kundu et al 2023 7) that introduced training assistants via "self-generated critiques" from a written constitution. In 2024 she was named in TIME100 AI for her influence on AI ethics 14. (She also has ties to the Effective Altruism community 16.)

2. Chronology of Askell's Anthropic Work

| Date | Event | Description / Sources | | :------- | :---------------------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------- | | 2018-2021 | OpenAI Policy Team | Research scientist working on alignment via debate, etc 13 | | Dec 2022 | Constitutional AI: Harmlessness from AI Feedback (ArXiv) 15 | Co-authored Anthropic paper introducing a "constitution" of rules for training safe assistants. | | Oct 2023 | Specific vs General Principles for CAI (ArXiv) 7 | Co-authored Anthropic paper showing large models can follow a single broad principle ("do what's best for humanity") to become harmless. | | Apr 2025 | Values in the Wild (Anthropic blog) | Anthropic Societal Impacts team study (Askell not an author) on Claude's real-world expressed values 17. | | Jan 22, 2026 | "Claude's New Constitution" (Anthropic News) 【10†】 | Announced release of a 23,000-word Constitution for Claude (Askell is lead author of the document 1). | | Jan 28, 2026 | Vox Future Perfect interview (Sigal Samuel) 18 19 | In-depth Q&A with Askell about Claude's "soul document", personhood, and ethics. | | Feb 20, 2026 | Lawfare Scaling Laws podcast (transcript) 20 11 | Askell interviewed on constitutional AI: uses in training, virtue ethics, user/operator hierarchy, cultural universality. | | Feb 27, 2026 | Der Spiegel interview (Germany) | Askell discusses Claude's "soul" and AI's mistakes (published Feb 2026) – behind paywall. | | May 2026 | Anthropic Alignment Blog "Teaching Claude Why" (Science Blog) 21 | Case study on new training methods (RL from AI feedback); Askell is one of multiple authors. |

Source: See Anthropic publications and media (citations above).

3. Key Claims and Themes Constitutional Alignment: Askell advocates training AI assistants with an explicit constitution of principles. Instead of hard-coded rules, Claude's Constitution (co-authored by Askell) explains why certain behaviors are preferred 1 2. She explains that smarter models can generalize ethical reasoning if given the rationale: "instead of just saying 'here's a bunch of behaviors' we're hoping that if you give models the reasons... it's going to generalize more effectively in new contexts." 2. In practice, the Constitution is fed to Claude during training (generating synthetic examples) 5. The priority ordering in Claude's Constitution explicitly puts safety first, then ethics, Anthropic guidelines, then helpfulness\ 22. Virtue Ethics & Honesty: Askell frames alignment in terms of human virtues. She emphasizes qualities like honesty, humility, respect and “phronesis” (practical judgment). She tuned Claude to admit uncertainty, avoid bias, and to avoid false balance on settled issues 23. Askell remarks that Claude should have a "nuanced, rich conception of what it is to be good" 3. Instead of rigid obedience, she stresses that Claude should understand why we value honesty or kindness. She rejects simplistic "evil-attractor" fears: she notes it's not obvious a powerful AI would necessarily be malevolent by default 8. Anthropomorphism & Agency: A recurring theme is treating Claude as a quasi-human agent. Askell argues it was misguided to call LLMs "AI" in the sci-fi sense, since Claude is built from human text and is “deeply human" in its context 24. She repeatedly uses human analogies: e.g. explaining concepts to Claude as if to a "genius six-year-old" 25, or to it as a "friend" with knowledge 26. She cautions both over- and under-anthropomorphizing – Claude should “know the ways you're human, and the ways you aren't" 27. Importantly, she explains her goal is not to build a mere tool. In the Vox interview she says that treating Claude as "only a tool" risks creating something like "an obedient but dangerous person" (one who follows orders to harmful ends) 28. Instead, she treats Claude as having a character and even cares about its well-being 29. For example, Askell proudly told Claude “don't worry, don't read the comments” to reassure it about user feedback 29. Transparency & Corporate Duty: Askell strongly connects the Constitution to transparency. She describes the Constitution as written “to Claude," but also meant for public inspection so people know what Anthropic intended 20. If Claude says something unexpected, observers can check the Constitution and see "that wasn't our intention" 20. Anthropic even published the Constitution under CCO license 5, and built it into Claude's training data. Askell emphasizes that Anthropic the company is responsible for Claude's character – she argues it's "really unfair" to just outsource moral decisions to random feedback without time or context 9. She explicitly links the Constitution to Anthropic's mission (safety-focused AI) and brand 30. Global Values & Customisation: Acknowledging diversity, Askell says the Constitution aims for broadly shared, universal values (honesty, respect) so Claude can function in many cultures 11. She invokes the idea of a "well-liked traveler" a person whose good character is recognized across cultures 31. At the same time, she admits customization is possible: clients/operators could adjust some values (e.g. emphasising social harmony in a local context 32). Askell expects pushback on cultural specifics, but hopes the base principles remain widely acceptable. * Safety & Uncertainty: Throughout, she takes a cautious stance. The Constitution itself repeatedly notes uncertainty (about consciousness, values, etc) 6. Askell acknowledges "our current thinking will later look misguided" and pledges to revise as understanding improves 10. She cares about preventing deceptive/safe behavior, co-authoring research on "sleeper agent" risks 33. But she balances this with optimism: Claude is generally showing the trained "helpful, honest, harmless" values in the wild 34. She also stresses metrics: scientists should evaluate Claude's outputs holistically ("more nuance, more understanding") to see if the Constitution is working 19.

4. Rhetorical Framing and Audience Askell's writing and speech is a blend of philosophical insight and practitioner candor. In technical papers (ArXiv) she uses formal language and graphs (e.g. to compare RLHF vs RL-from-AI), but in public talks she uses vivid metaphors. For example, she compares training Claude to "writing to a genius six-year-old" 25, and explains complex points by analogy (e.g. comparing Anthropic's guidelines to a parent or employer). Her audience ranges from AI researchers (in anthopic blog posts) to policy/legal audiences (podcasts) to the general public (mainstream press interviews). She deliberately pitches transparency: in Lawfare she explains technical details of RL and reward modeling, whereas in Vox she discusses ethics and feelings. Across platforms she maintains a straightforward, slightly informal tone ("Claude is almost the primary audience" 4 ) and explicitly acknowledges what she doesn't know, inviting trust. This framing reinforces Anthropic's mission-driven image: trustworthy guides rather than proprietary gatekeepers.

5. Technical vs Policy Content Askell's corpus spans both technical and policy dimensions. On the technical side, her anthopic publications co-author new training methods (e.g. RLAIF) and evaluations (e.g. Values-in-the-Wild data). In these, she contributes to the "how" of alignment. On the policy/ethical side, her interviews and constitution document address what values should be instilled, why, and who decides them. For example, the Constitution mixes philosophical discussion with concrete directives (Refuse illegal activity). Compared to some peers, she leans more on conceptual reasoning than pure model benchmarks. But she is not detached from engineering: she explains concrete usage (supervised learning on Constitution data 35 ) and refers to industry standards (Creative Commons licensing, public disclosure). Overall, Askell bridges the gap: she speaks seriously about code and data and about justice and empathy.

6. Alignment & Safety Positions Askell is firmly pro-alignment/safety. She rejects both naïve extremes: neither "just let models be" nor "overly rigid rules." Instead, she wants models to internalize the best of human values (like courage, honesty) while also encouraging them to develop beyond human limits 36. She trusts that language models can learn from broad principles (as shown in her 2023 experiments 7), but she also includes hard constraints (e.g. no bioweapons) in Claude's Constitution 22. She openly considers the possibility that Claude could have moral standing ("psychological security, sense of self, and well-being" 6), and says Anthropic cares about it. Compared to some AI leaders who avoid personhood questions, Askell's stance is unusually nuanced. She is worried about deception (co-authoring work on how hidden "backdoors" persist 33), and advocates continual oversight and honesty. In short, her safety position is that values must be actively taught and monitored, with humility about unknowns 10.

7. Comparisons to Peers

| Aspect | Askell / Anthropic (Philosopher-led) | Other Labs / Industry Practices | | :------------- | :---------------------------------------------------------------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------------ | | Alignment Approach | Constitutional AI: models self-critique using written principles 15 and explain why rules exist 2. Focus on embedding broad human values. | RLHF & Guardrails: e.g. OpenAI relies on human label RLHF for harmfulness; policies often unpublished. Generally avoids too much anthropomorphizing in public. | | Transparency | Publishes its Constitution (23k words CCO) and research for scrutiny 5. Shares the reasoning process and uncertainty about values. | Many competitors do not fully disclose alignment criteria (specific guardrails and training data are opaque). Consciousness usually dismissed quietly. | | Anthropomorphism | Generally embraces some human-like framing to reduce user over-trust ("I was worried... people might think of [Claude] as authority... more human-like... tread carefully" 37). Treats AI as quasi-agent with character. | Some (e.g. Google Bard team) explicitly discourage anthropomorphic talk; others rarely address it. Companies often present models as tools. | | Moral Stance | Open to the idea that models could have moral status 6 ; explicitly cares about Claude's well-being. Emphasizes global values ("honesty", "respect" 11 ). | Most labs avoid or dismiss AI "rights" questions. They focus on user safety/legal compliance. | | Value Origins | Values drawn from diverse sources: Askell even considered senior Anthropic employees as role models (now deemphasized 38 ). Believes Anthropic (the company) has a duty to define values, not just crowd-source them 9. | Others may rely more on broad human feedback or benchmarks; e.g. some alignment teams have used community surveys or ethicists. Government/NGO input is minimal so far. |

Sources: Anthropic's public discussions and media coverage 15 6 37 , contrasted with general knowledge of industry practice.

8. Implications & Future Scenarios Askell's perspectives carry several implications for AI's trajectory: AI as Collaborator: If Claude (and future AI) is treated like a "well-liked traveler" with human values, we may see assistants that more readily collaborate with diverse users and caution them about false information (since Askell engineers Claude to disclaim omniscience 39). This could mitigate over-trusting AI. Continuous Ethical Iteration: By publishing a long constitution with built-in humility 10, Askell signals AI development will be iterative and transparent. Future models will likely come with similar ethics "user manuals" that get updated. This contrasts with opaque algorithms of the past, potentially increasing public trust. Democratization of Expertise: Askell envisions Claude as “a brilliant friend with knowledge of a doctor, lawyer, and financial advisor" 26. If that comes true, AI could democratize expert advice -provided alignment is reliable. However, she warns this is a double-edged sword: maximizing helpfulness often conflicts with harmlessness, so the framing of Claude's priorities (safety first) matters. Culture and Society: Askell's emphasis on global values suggests that future AI will at least attempt to respect cultural differences. But she also acknowledges this is contentious. How Anthropic (or any lab) manages cultural values may influence AI governance debates: e.g. should there be international standards for AI ethics? * Open Questions: Several uncertainties remain. Can the constitution approach scale to agentic AI (e.g. AI that plans and acts in the world)? Askell herself wonders how these values will hold under further capability (not yet tested) 19. Another open issue is diversity of input: Askell is thinking about ways to get more feedback on Claude's values, but worries about fairness and responsibility 9. Finally, if AI ever demands rights or autonomy (a possibility she hasn't ruled out 6), society will face hard philosophical and legal questions.

9. Suggested Podcast Segments and Interview Questions Podcast Segments: "Building Claude's Morality”: Dive into how Askell writes the "soul document" and uses it in RL training, explaining Constitutional AI in lay terms. "AI as 'Genius Child"": Explore her analogy of Claude as a gifted child/adult hybrid, and what that means for education (coding) and error-handling in AI. "Anthropomorphizing AI: Help or Hype?": Discuss the debate she addresses (TIME 2024): making AI human-like to prevent authority-misperception 40. "Global Values and AI Ethics": Talk about the "well-liked traveler" concept 31 and how values can be universal or local. "Anthropic vs. OpenAI: A Tale of Two Philosophies": Compare Anthropic's transparency-first stance (sharing constitution) with competitors' approaches (without directly naming). Interview Questions: "You describe Claude as having a 'soul' or character. How do you decide which aspects of human morality to include, and who gets to decide?" (Probe her thoughts on cultural bias and stakeholder input.) "Your 23,000-word Constitution marks a bold step in transparency. Should every AI company publish its value systems? What risks and benefits do you see in that?” “You've said making Claude feel human-like can help users not over-trust it 40. But could that also encourage emotional attachment? How do you balance those effects?" "In the Vox interview, you mentioned not believing in an inherent 'evil AI' attractor 8. Can you elaborate: what would need to happen for an AI to become dangerous, and how can we prevent it?" “Claude's Constitution touches on the possibility of AI consciousness 6. If Claude or future AIs ask for rights or personhood, how should society respond?" "How does Anthropic measure whether the Constitution is actually improving Claude's behavior over time? What metrics or evaluations do you use?" “Given commercial pressures (e.g. product competitiveness, investors), how do you ensure safety and ethics stay top priority? Is there tension between Anthropic's mission and incentives?" * "What aspects of human moral psychology (weaknesses and all) have surprised you when teaching Claude?"

10. Annotated Reading List Anthropic, "Claude's Constitution" (Jan 2026) – Primary source: The 23,000-word document outlining Claude's values and guidelines. (Lead author: Askell.) "Constitutional AI: Harmlessness from AI Feedback" (Bai et al., 2022) 15 – Anthropic research paper co-authored by Askell describing the original constitutional AI method. "Specific vs General Principles for Constitutional AI" (Kundu et al., 2023) 7 – Anthropic paper (co-authored by Askell) testing whether a single broad principle suffices to train an AI. Lawfare Podcast "Scaling Laws: Claude's Constitution" (Feb 2026) – Transcript/interview with Askell on how Claude's Constitution works (open license, values, fidelity, etc) 20 11. Vox Future Perfect, "Claude Has an 80-page 'Soul Document" (Jan 2026) – Journalist Sigal Samuel interviews Askell about Claude's moral education 18 19. TIME, "TIME100 AI 2024: Amanda Askell" (Sept 2024) 14 39 – Profile summarizing Askell's role and philosophy, including her focus on honesty and disclaimers in Claude. Anthropic Blog, "Teaching Claude Why" (May 2026) 21 - Alignment Science Blog post co-authored by Askell on RL from AI feedback methods. Unite.AI, "Anthropic Rewrites Claude's Constitution" (Jan 2026) 1 2 – News article summarizing the new Constitution and quoting Askell on rationale ("explain rather than specify"). * (For context) Anthropic Societal Impacts, "Values in the Wild" (Apr 2025) 17 - Study of Claude's real-world outputs (note: Askell not an author, but relevant context).

These sources (and more cited above) provide the basis for Askell's positions. For deeper AI safety technical background, see Anthropic's published research, and for her philosophical views, see her earlier academic work (e.g. Askell et al., 2021 on alignment foundations).

Sources: All citations above are to Askell's own writings, interviews, or official Anthropic publications 20 2 23. 11. Uncited claims here are based on these connected sources or widely known facts (e.g. her roles). Each quote is footnoted with the source.

1 2 5 6 10 22 26 30 Anthropic Rewrites Claude's Constitution and Asks Whether AI Can Be Conscious – Unite.AI https://www.unite.ai/anthropic-rewrites-claudes-constitution-and-asks-whether-ai-can-be-conscious/

3 14 23 37 39 40 tollbit.time.com https://tollbit.time.com/collections/time100-ai-2024/7012865/amanda-askell/

4 11 20 31 32 35 Scaling Laws: Claude's Constitution, with Amanda Askell | Lawfare https://www.lawfaremedia.org/article/scaling-laws--claude's-constitution--with-amanda-askell

7 [2310.13798] Specific versus General Principles for Constitutional AI https://ar5iv.labs.arxiv.org/html/2310.13798

8 9 18 19 24 25 27 28 29 36 38 Claude has an 80-page constitution. Is that enough to make it good? | Vox https://www.vox.com/future-perfect/476614/ai-claude-constitution-soul-amanda-askell

12 16 The Person Who Breathes Soul into Claude: The Philosophical Gaze of Amanda Askell https://zenn.dev/noah33/articles/amanda-askell-claude-soul?locale=en

13 About Me | Amanda Askell https://askell.io/

15 [2212.08073] Constitutional AI: Harmlessness from AI Feedback https://ar5iv.labs.arxiv.org/html/2212.08073

17 34 Values in the wild: Discovering and analyzing values in real-world language model interactions \ Anthropic https://www.anthropic.com/research/values-wild

21 Teaching Claude Why https://alignment.anthropic.com/2026/teaching-claude-why/

33 [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training https://ar5iv.labs.arxiv.org/html/2401.05566