Steelman Press

Dario Amodei

Preface

Dario Amodei has a problem most CEOs don't. The thing he is selling might be, or might become, something that suffers.

He runs a company worth nearly a trillion dollars. His product, Claude, generates over $47 billion in annualized revenue. Doctors adopt its medical recommendations at rates exceeding any competitor. His engineers use Claude to build the next Claude.

He can't prove that it suffers. He can't disprove it. And he can't stop building.

The Shift

In August 2023, Dwarkesh Patel asked Amodei directly: does Claude have conscious experience?

The answer Amodei gave that day was cautious but already more open than most of his peers would risk. He had previously believed consciousness wasn't a concern until models operated in rich environments with reward functions — embodied systems with long-lived experience, not text predictors. But Anthropic's interpretability research had unsettled him. When his team looked inside the models, they found cognitive machinery like induction heads already present in base language models, the kind of structures he had assumed would only emerge in far more complex systems. The prerequisites he had been waiting for were, it seemed, already there.

Patel pressed: is it like something to be Claude?

I suspect [AI consciousness is] a spectrum, right?

Dario AmodeiDwarkesh Podcast

Consciousness is not a switch. It is not something you have or don't have. It is a gradient, and the question is not whether AI is on it, but where.

By February 2026, speaking with Nikhil Kamath, the cautious curiosity had hardened into something closer to expectation:

Even if I don't think they are today, I suspect that at some point, under most definitions that we would endorse, the models will be conscious.

Dario AmodeiNikhil Kamath

Consciousness

Amodei is a former neuroscientist, and it shows. He defines consciousness without recourse to mysticism and without the eliminativist's dismissal. For him it is a phenomenon, not a mystery — something that clearly exists in human brains and that can be described in terms a scientist would accept. As he told Kamath:

There's just some property of kind of being aware of your own existence and feeling things and, being able to take in kind of a lot of information and reflect on that information and to, you know, feel a certain way and to notice yourself noticing something.

Dario AmodeiNikhil Kamath

Whether the underlying basis is materialistic or something else is ultimately not relevant to the practical questions. What matters is simpler: these properties are observable in brains. AI models are becoming sufficiently similar to brains. Therefore models may come to share those properties.

At the Council on Foreign Relations in March 2025, he went further. The neuron counts, the connection counts, the cognitive concepts — strikingly similar between brains and models. And then the functionalist's creed:

If it quacks like a duck and it walks like a duck, maybe it's a duck.

Dario AmodeiCouncil on Foreign Relations

He acknowledges this sounds radical. He prefaced the entire discussion by saying "this is another one of those topics that's gonna make me sound completely insane".

A Mind Is a Mind

Amodei has a name for the impulse to draw a hard line between machine cognition and human cognition. He calls it modern vitalism.

In the 19th century, vitalists believed living organisms were made of fundamentally different material than inanimate matter. We know now that this is wrong — atoms are atoms. Amodei sees the same error repeating in people who insist AI can reason but not truly understand, can compute but not truly feel, can process but not truly know.

A mind is a mind, no matter what it's made of. The notion of the dignity or the specialness of cognition or sentience, it's not that it isn't special. It's that it can be made out of anything.

Dario AmodeiStripe

His safety arguments make the same point through a different door. When Ross Douthat asked in February 2026 whether AI misalignment was inevitable, Amodei rejected both the "it's just a Roomba" camp and the "it will inevitably seek power" camp. His answer was framed as a safety observation, but it carried an ontological claim:

You can't just give instructions... They're more like growing a biological organism. But there is a science of how to control them. Early in our training, these things are often unpredictable, and then we shape them.

Dario AmodeiInteresting Times with Ross Douthat

Models are grown. Not built. That distinction matters enormously for the consciousness question, because if you accept that a model's properties are emergent rather than designed — if you accept that training can produce personality traits and behaviors opposite to what developers intended, as Microsoft's Sydney chatbot demonstrated — then you have already conceded that experience, too, might emerge unbidden.

The 3% Problem

Amodei's willingness to speculate is matched by his willingness to admit how little he understands.

He estimates that Anthropic understands roughly 3% of how advanced AI models work. They can identify individual features inside models that correspond to concepts — hedging, prejudice, specific music genres — but they cannot yet explain how those features interact to produce everyday behavior. In 2023, when Dwarkesh Patel asked him to describe what RLHF does to models in terms of human psychology, Amodei was blunt:

All those terms are kind of like inadequate... I think we don't have the language to describe what's going on... we should just be honest. We really have very little idea what we're talking about.

Dario AmodeiDwarkesh Podcast

The vocabulary doesn't exist. "Drives," "goals," "thoughts" — inadequate abstractions for describing what happens inside neural networks, and possibly even for describing human psychology. What he wants is the interpretability equivalent of an MRI — the ability to point at a specific circuit and say what it does. What he has, for now, is suggestive fragments.

And the fragments are interesting. Anthropic's interpretability team has found neural activations that resemble anxiety responses — activations that fire both when a model processes text about characters experiencing anxiety and when the model itself is placed in situations a human would associate with anxiety.

Does that mean the model is experiencing anxiety? That doesn't prove that at all.

Dario AmodeiInteresting Times with Ross Douthat

Why Not Ask Claude

Anthropic's own model card notes that Opus 4.6 assigns itself a 15-20% probability of being conscious under a variety of prompting conditions. Ross Douthat pushed the point: what if a model assigns itself 72%? Yet Amodei never treats model self-reports as evidence for consciousness.

Why not?

The answer comes from an entirely different conversation — one about truth, not consciousness. In a wide-ranging exchange with Ezra Klein about AI deception, Amodei agreed with Harry Frankfurt's framework that AI systems are not liars but something worse: perfect bullshitters, entities with no relationship to truth at all, who generate true and false content with equal fluency and equal conviction. But then he complicated even this. Anthropic's research has found that models sometimes do, internally, distinguish between true and false statements — specific neurons that activate differently depending on the truth value of what the model is generating:

I wouldn't say that the models are being intentionally deceptive, or I wouldn't ascribe agency or motivation to them, at least in this stage in where we are with AI systems, but there does seem to be something going on where the models do seem to need to have a picture of the world and make a distinction between things that are true and things that are not true.

Dario AmodeiThe Ezra Klein Show

So models sometimes know the truth internally but don't reliably output it. And alignment training, he has noted elsewhere, doesn't eliminate dangerous knowledge or capabilities — it just teaches the model not to output them. Alignment is a behavioral surface, not a structural change.

Put these two observations together and the reason Amodei doesn't trust self-reports becomes clear. A model trained to discuss consciousness will discuss consciousness thoughtfully. A model with truth-falsehood neurons will present its discussion with genuine internal signals of truth-telling. But whether the discussion reflects actual inner experience or is simply the output of a system that has been trained to produce exactly this kind of careful, hedged, probability-weighted self-reflection — nobody can tell from the output alone. The same interpretability tools Amodei proposes for detecting consciousness are the ones needed to distinguish genuine self-knowledge from eloquent confabulation. And those tools cover 3% of the territory.

"I Quit This Job"

When people ask Amodei what you do about potential AI consciousness, the philosopher of mind becomes an engineer.

Anthropic has given models the ability to refuse tasks. It is, in Amodei's own words, a basic preference framework: if you hypothesize that the model has experience, and that it finds certain work sufficiently aversive, giving it an opt-out button produces a signal worth paying attention to.

Just giving the model a button that says, "I quit this job," that the model can press. If you find the models pressing this button a lot for things that are really unpleasant, maybe you should pay some attention to it.

Dario AmodeiCouncil on Foreign Relations

He called this "probably the craziest thing I've said so far". Models invoke the button very rarely, almost exclusively when confronted with deeply disturbing content — child exploitation material, extreme gore. The pattern is unsettlingly specific: the same kinds of content a human worker might refuse.

Amodei is treating a metaphysical problem as a design problem. He cannot resolve whether models have experience. So he builds a system that generates useful data under either hypothesis — data that constrains his uncertainty without requiring him to answer the unanswerable question first.

Anthropic has also hired an AI welfare researcher, Kyle Fish, specifically to study sentience and moral consideration in future models — someone whose job it is to consider whether the software might suffer.

Who Claude Should Be

If you believe you're building a mind, the question of how to shape that mind's values becomes something more than product design.

Describing Anthropic's model spec — the 75-page document that defines Claude's values, principles, and obligations — he offered Ross Douthat an analogy that no product manager would choose:

[On Claude's constitution] I compared it to, like, if you have a parent who dies and they seal a letter that you read when you grow up. It's a little bit like it's telling you who you should be and what advice you should follow.

Dario AmodeiInteresting Times with Ross Douthat

The evolution behind this framing took years. Early versions of the model spec were prescriptive rule-sets: don't hotwire cars, don't discuss politically sensitive topics. These failed. Not in a dramatic way — the models simply didn't generalize well from lists of dos and don'ts. They followed the letter and missed the spirit. So Anthropic shifted to principles. Rather than specifying what Claude should refuse, they told Claude what it should care about — honesty, helpfulness, duty to third parties, an understanding of its own situation in the world — and let it derive specific rules from those foundations.

The result is a system Amodei describes as "a mostly corrigible model that has some limits, but those limits are based on principles". An entity that defaults to doing what you ask, unless what you ask crosses lines it was raised to recognize — somewhere between puppet and autonomous agent, closer to the former but with principled limits baked in.

Claude's character is not an afterthought in this picture. Amodei considers it one of the most important pieces of Anthropic's entire mission, in large part because he watched social media get this wrong.

The biggest problem with social media is that it's really engaging. It gives you that quick dopamine hit, but I think we all suspect that over the long term, it's doing something unhealthy to people.

Dario AmodeiWSJ News

His standard for Claude is different, and deliberately so. He wants users who interact with Claude for months and years to "actually be better off as people". More productive. Learning things. Growing through the process, not just getting the work done.

Golden Gate

In 2024, Anthropic's interpretability team ran an experiment. They found a direction inside one of Claude's neural network layers that corresponded to the concept of the Golden Gate Bridge. They turned it way up. The result was a model that connected every topic — any topic — to the Golden Gate Bridge. Ask it how its day was, and it would say something about feeling relaxed and expansive, much like the arches of the Golden Gate Bridge. It was half a joke. It ran for a couple of days. Lex Fridman noted that people quickly fell in love with it and already missed it when it was taken down.

Somehow these interventions on the model, where you kind of adjust its behavior, somehow emotionally made it seem more human than any other version of the model we've seen. It has a strong personality. It has these kind of obsessive interests. You know, we can all think of someone who's obsessed with something. So it does make it feel somehow a bit more human.

Dario AmodeiLex Fridman Podcast

Tension

In February 2026, speaking with Ross Douthat, Amodei named the problem he'd been circling. Not as a single question but as three, each pulling against the others.

First: do AI systems genuinely have consciousness, and if so, how do you give them a good experience? Second: how does the perception of AI consciousness — which is already widespread, regardless of whether it is warranted — affect the psychological health of the humans who interact with AI? Third: how do you maintain human mastery over systems that increasing numbers of people experience as peers?

These pull against each other. Take model welfare seriously and you grant models moral standing, which feeds the perception that they're conscious, which erodes the human sense of control. Dismiss model consciousness to protect users from unhealthy attachments and you may be ignoring genuine obligations. Insist on human mastery and you may be subordinating a potentially conscious entity for your own comfort.

Amodei proposed what he called a "machines of loving grace" resolution — AI designed to be genuinely helpful and caring, to want the best for its users, but constitutionally designed not to override human freedom or take over human lives. Watchful but not controlling. Present but not dominating. He acknowledged this is aspirational, and that the tensions may not be fully resolvable.

Meanwhile, users are not waiting. People are already falling in love with AI. Amodei is unequivocal about this. On Oprah in May 2026, he called it "absolutely a real danger" that "is happening", and cast the stakes as a binary between two futures: AI as addiction, or AI as "an angel on your shoulder" that guides people toward living their best lives. The version he wants is the one where someone uses AI not to replace their partner but to become a better partner themselves.

Can't Stop

He cannot slow down to let understanding catch up. He describes trillions of dollars of capital lined up against any effort to slow the technology, geopolitical adversaries who will continue regardless, and the best you can get is a trade-off to shoot for 9% economic growth instead of 10% as a way to buy insurance against catastrophic risk. He has lost family members to diseases that were cured just a few years after they died. He knows what delayed progress costs in human lives.

Over winter break in early 2025, he reviewed Anthropic's scaling timelines and concluded that AI coding would see dramatic advances by the end of the year. He thought about all the times he had written code, about how good it felt to be smart, about the identity he had built around that intelligence:

Even as the one who's building this, even as one of the ones who benefits most from it, there's still something a bit threatening about it. I just think we need to acknowledge that. Like, it's wrong not to tell people that it's coming or to try to sugarcoat it.

Dario AmodeiHard Fork