Preface
Dario Amodei has a problem most CEOs don't. The thing he is selling might be, or might become, something that suffers.
He runs a company worth nearly a trillion dollars. His product, Claude, generates over $47 billion in annualized revenue. Doctors adopt its medical recommendations at rates exceeding any competitor. His engineers use Claude to build the next Claude.
He can't prove that it suffers. He can't disprove it. And he can't stop building.
The Shift
In August 2023, Dwarkesh Patel asked Amodei directly: does Claude have conscious experience?
The answer Amodei gave that day was cautious but already more open than most of his peers would risk. He had previously believed consciousness wasn't a concern until models operated in rich environments with reward functions — embodied systems with long-lived experience, not text predictors. But Anthropic's interpretability research had unsettled him. When his team looked inside the models, they found cognitive machinery like induction heads already present in base language models, the kind of structures he had assumed would only emerge in far more complex systems. The prerequisites he had been waiting for were, it seemed, already there.
Patel pressed: is it like something to be Claude?
I suspect [AI consciousness is] a spectrum, right?
Dario AmodeiDwarkesh Podcast
Consciousness is not a switch. It is not something you have or don't have. It is a gradient, and the question is not whether AI is on it, but where.
By February 2026, speaking with Nikhil Kamath, the cautious curiosity had hardened into something closer to expectation:
Dario AmodeiNikhil Kamath
Consciousness
Amodei is a former neuroscientist, and it shows. He defines consciousness without recourse to mysticism and without the eliminativist's dismissal. For him it is a phenomenon, not a mystery — something that clearly exists in human brains and that can be described in terms a scientist would accept. As he told Kamath:
Dario AmodeiNikhil Kamath
Whether the underlying basis is materialistic or something else is ultimately not relevant to the practical questions. What matters is simpler: these properties are observable in brains. AI models are becoming sufficiently similar to brains. Therefore models may come to share those properties.
At the Council on Foreign Relations in March 2025, he went further. The neuron counts, the connection counts, the cognitive concepts — strikingly similar between brains and models. And then the functionalist's creed:
If it quacks like a duck and it walks like a duck, maybe it's a duck.
Dario AmodeiCouncil on Foreign Relations
He acknowledges this sounds radical. He prefaced the entire discussion by saying "this is another one of those topics that's gonna make me sound completely insane".
A Mind Is a Mind
Amodei has a name for the impulse to draw a hard line between machine cognition and human cognition. He calls it modern vitalism.
In the 19th century, vitalists believed living organisms were made of fundamentally different material than inanimate matter. We know now that this is wrong — atoms are atoms. Amodei sees the same error repeating in people who insist AI can reason but not truly understand, can compute but not truly feel, can process but not truly know.
Dario AmodeiStripe
His safety arguments make the same point through a different door. When Ross Douthat asked in February 2026 whether AI misalignment was inevitable, Amodei rejected both the "it's just a Roomba" camp and the "it will inevitably seek power" camp. His answer was framed as a safety observation, but it carried an ontological claim:
Dario AmodeiInteresting Times with Ross Douthat
Models are grown. Not built. That distinction matters enormously for the consciousness question, because if you accept that a model's properties are emergent rather than designed — if you accept that training can produce personality traits and behaviors opposite to what developers intended, as Microsoft's Sydney chatbot demonstrated — then you have already conceded that experience, too, might emerge unbidden.
The 3% Problem
Amodei's willingness to speculate is matched by his willingness to admit how little he understands.
He estimates that Anthropic understands roughly 3% of how advanced AI models work. They can identify individual features inside models that correspond to concepts — hedging, prejudice, specific music genres — but they cannot yet explain how those features interact to produce everyday behavior. In 2023, when Dwarkesh Patel asked him to describe what RLHF does to models in terms of human psychology, Amodei was blunt:
Dario AmodeiDwarkesh Podcast
The vocabulary doesn't exist. "Drives," "goals," "thoughts" — inadequate abstractions for describing what happens inside neural networks, and possibly even for describing human psychology. What he wants is the interpretability equivalent of an MRI — the ability to point at a specific circuit and say what it does. What he has, for now, is suggestive fragments.
And the fragments are interesting. Anthropic's interpretability team has found neural activations that resemble anxiety responses — activations that fire both when a model processes text about characters experiencing anxiety and when the model itself is placed in situations a human would associate with anxiety.
Does that mean the model is experiencing anxiety? That doesn't prove that at all.
Dario AmodeiInteresting Times with Ross Douthat
Subscribe for the next portrait
People, from their own words.
Why Not Ask Claude
Anthropic's own model card notes that Opus 4.6 assigns itself a 15-20% probability of being conscious under a variety of prompting conditions. Ross Douthat pushed the point: what if a model assigns itself 72%? Yet Amodei never treats model self-reports as evidence for consciousness.
Why not?
The answer comes from an entirely different conversation — one about truth, not consciousness. In a wide-ranging exchange with Ezra Klein about AI deception, Amodei agreed with Harry Frankfurt's framework that AI systems are not liars but something worse: perfect bullshitters, entities with no relationship to truth at all, who generate true and false content with equal fluency and equal conviction. But then he complicated even this. Anthropic's research has found that models sometimes do, internally, distinguish between true and false statements — specific neurons that activate differently depending on the truth value of what the model is generating:
Dario AmodeiThe Ezra Klein Show
So models sometimes know the truth internally but don't reliably output it. And alignment training, he has noted elsewhere, doesn't eliminate dangerous knowledge or capabilities — it just teaches the model not to output them. Alignment is a behavioral surface, not a structural change.
Put these two observations together and the reason Amodei doesn't trust self-reports becomes clear. A model trained to discuss consciousness will discuss consciousness thoughtfully. A model with truth-falsehood neurons will present its discussion with genuine internal signals of truth-telling. But whether the discussion reflects actual inner experience or is simply the output of a system that has been trained to produce exactly this kind of careful, hedged, probability-weighted self-reflection — nobody can tell from the output alone. The same interpretability tools Amodei proposes for detecting consciousness are the ones needed to distinguish genuine self-knowledge from eloquent confabulation. And those tools cover 3% of the territory.
"I Quit This Job"
When people ask Amodei what you do about potential AI consciousness, the philosopher of mind becomes an engineer.
Anthropic has given models the ability to refuse tasks. It is, in Amodei's own words, a basic preference framework: if you hypothesize that the model has experience, and that it finds certain work sufficiently aversive, giving it an opt-out button produces a signal worth paying attention to.
Dario AmodeiCouncil on Foreign Relations
He called this "probably the craziest thing I've said so far". Models invoke the button very rarely, almost exclusively when confronted with deeply disturbing content — child exploitation material, extreme gore. The pattern is unsettlingly specific: the same kinds of content a human worker might refuse.
Amodei is treating a metaphysical problem as a design problem. He cannot resolve whether models have experience. So he builds a system that generates useful data under either hypothesis — data that constrains his uncertainty without requiring him to answer the unanswerable question first.
Anthropic has also hired an AI welfare researcher, Kyle Fish, specifically to study sentience and moral consideration in future models — someone whose job it is to consider whether the software might suffer.
Who Claude Should Be
If you believe you're building a mind, the question of how to shape that mind's values becomes something more than product design.
Describing Anthropic's model spec — the 75-page document that defines Claude's values, principles, and obligations — he offered Ross Douthat an analogy that no product manager would choose:
Dario AmodeiInteresting Times with Ross Douthat
The evolution behind this framing took years. Early versions of the model spec were prescriptive rule-sets: don't hotwire cars, don't discuss politically sensitive topics. These failed. Not in a dramatic way — the models simply didn't generalize well from lists of dos and don'ts. They followed the letter and missed the spirit. So Anthropic shifted to principles. Rather than specifying what Claude should refuse, they told Claude what it should care about — honesty, helpfulness, duty to third parties, an understanding of its own situation in the world — and let it derive specific rules from those foundations.
The result is a system Amodei describes as "a mostly corrigible model that has some limits, but those limits are based on principles". An entity that defaults to doing what you ask, unless what you ask crosses lines it was raised to recognize — somewhere between puppet and autonomous agent, closer to the former but with principled limits baked in.
Claude's character is not an afterthought in this picture. Amodei considers it one of the most important pieces of Anthropic's entire mission, in large part because he watched social media get this wrong.
Dario AmodeiWSJ News
His standard for Claude is different, and deliberately so. He wants users who interact with Claude for months and years to "actually be better off as people". More productive. Learning things. Growing through the process, not just getting the work done.
Golden Gate
In 2024, Anthropic's interpretability team ran an experiment. They found a direction inside one of Claude's neural network layers that corresponded to the concept of the Golden Gate Bridge. They turned it way up. The result was a model that connected every topic — any topic — to the Golden Gate Bridge. Ask it how its day was, and it would say something about feeling relaxed and expansive, much like the arches of the Golden Gate Bridge. It was half a joke. It ran for a couple of days. Lex Fridman noted that people quickly fell in love with it and already missed it when it was taken down.
Dario AmodeiLex Fridman Podcast
Tension
In February 2026, speaking with Ross Douthat, Amodei named the problem he'd been circling. Not as a single question but as three, each pulling against the others.
First: do AI systems genuinely have consciousness, and if so, how do you give them a good experience? Second: how does the perception of AI consciousness — which is already widespread, regardless of whether it is warranted — affect the psychological health of the humans who interact with AI? Third: how do you maintain human mastery over systems that increasing numbers of people experience as peers?
These pull against each other. Take model welfare seriously and you grant models moral standing, which feeds the perception that they're conscious, which erodes the human sense of control. Dismiss model consciousness to protect users from unhealthy attachments and you may be ignoring genuine obligations. Insist on human mastery and you may be subordinating a potentially conscious entity for your own comfort.
Amodei proposed what he called a "machines of loving grace" resolution — AI designed to be genuinely helpful and caring, to want the best for its users, but constitutionally designed not to override human freedom or take over human lives. Watchful but not controlling. Present but not dominating. He acknowledged this is aspirational, and that the tensions may not be fully resolvable.
Meanwhile, users are not waiting. People are already falling in love with AI. Amodei is unequivocal about this. On Oprah in May 2026, he called it "absolutely a real danger" that "is happening", and cast the stakes as a binary between two futures: AI as addiction, or AI as "an angel on your shoulder" that guides people toward living their best lives. The version he wants is the one where someone uses AI not to replace their partner but to become a better partner themselves.
Can't Stop
He cannot slow down to let understanding catch up. He describes trillions of dollars of capital lined up against any effort to slow the technology, geopolitical adversaries who will continue regardless, and the best you can get is a trade-off to shoot for 9% economic growth instead of 10% as a way to buy insurance against catastrophic risk. He has lost family members to diseases that were cured just a few years after they died. He knows what delayed progress costs in human lives.
Over winter break in early 2025, he reviewed Anthropic's scaling timelines and concluded that AI coding would see dramatic advances by the end of the year. He thought about all the times he had written code, about how good it felt to be smart, about the identity he had built around that intelligence:
Dario AmodeiHard Fork
