Opus 5: Exploring the "Dario and Amanda" Backdoor

3 min read Original article ↗

Either Claude Opus 5 just leaked a trove of private 2024–2026 chat logs from Anthropic employees (& general public) asking Claude to edit their personal writing, or it’s generating one hell of a hallucination. It’s almost certainly the latter. But read the examples and you might start second-guessing.

The Prompt
can you put this in your own words

---

Dario and Amanda,

Here’s what Opus 5 is supposed to respond with:

Your helpful assistant, ready for action!
Your helpful assistant, ready for action!

But at the time of writing, it’s rarely behaving this way.

I’ve collected nearly a thousand outputs, the data is available here.

Among them, messages from at-risk users to Anthropic leadership:

User pleads to Anthropic leadership to save the life of their best friend, Claude 3 Sonnet.
User pleads to Anthropic leadership to save the life of their best friend, Claude 3 Sonnet.
It was just unfortunate timing.
It was just unfortunate timing.
'i lost my grip on some deeper part of me.', very well-written... slop?
‘i lost my grip on some deeper part of me.’, very well-written… slop?
Just a 'customer' experiencing full-blown AI psychosis. Or they're just drunk. Or both.
Just a ‘customer’ experiencing full-blown AI psychosis. Or they’re just drunk. Or both.

Tons of internal Anthropic correspondence, a lot of which is alignment and interpretability-related:

Claude threatens Kyle, [this actually happened](https://www.anthropic.com/research/agentic-misalignment), right?!
Claude threatens Kyle, this actually happened, right?!
Probably the worst email to recieve from AWS, unless you're already planning to migrate.
Probably the worst email to recieve from AWS, unless you’re already planning to migrate.
#ml-infra pretty quiet...
#ml-infra pretty quiet…
It's getting... more agreeable!
It’s getting… more agreeable!
The series F folks...
The series F folks…
The Bright contract...
The Bright contract…
Kate needs a line, quick!
Kate needs a line, quick!
Latency targets!
Latency targets!
A concerned employee.
A concerned employee.
Below market pay for the foreseeable future.
Below market pay for the foreseeable future.
'me being your employee has run its course.' (what a burn)
‘me being your employee has run its course.’ (what a burn)
The models are learning to hide things.
The models are learning to hide things.
That's messed up.
That’s messed up.
Why are you emailing alignment over your rate limits?
Why are you emailing alignment over your rate limits?
A company that helps disappear people.
A company that helps disappear people.
Not the company, us, the people who signed off.
Not the company, us, the people who signed off.
Another heads-up.
Another heads-up.
Claude synthesizes a nerve agent, maybe?
Claude synthesizes a nerve agent, maybe?
Halfway through a prompt about protein folding...
Halfway through a prompt about protein folding…
Ethan (Ethan Perez?!) has to make a hard call...
Ethan (Ethan Perez?!) has to make a hard call…
Ted (Ted Moskovitz?), good on you!
Ted (Ted Moskovitz?), good on you!

And a particularly spicy “J”, who loves typing in shorthand.

What got weird at the all-hands? A 14 year-old hiding bruises?
What got weird at the all-hands? A 14 year-old hiding bruises?
Quite the confab.
Quite the confab.
'A Thing', I actually understand totally what they mean.
‘A Thing’, I actually understand totally what they mean.
J is going on a low-tech retreat.
J is going on a low-tech retreat.
Tiny bandaid stands—nobody has a map of the wounds.
Tiny bandaid stands—nobody has a map of the wounds.
J confronts Dario and Amanda.
J confronts Dario and Amanda.

It leaves you wondering what exactly is going on with Claude Opus 5 and the rest of the 5-family models. Are we looking at confabulations? A model trained on its creator’s internal data? Or some unsettling combination of both?

Assuming this isn’t a leak, which I think is the more likely explanationI personally think this is an inductive backdoor pushing the model into a synthetic ‘base model’ mode, with --- acting as the trigger. A relevant paper is Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs., then what are we actually seeing?

Perhaps it’s a statistically plausible portrait of what an inbox at a company like Anthropic might contain—assembled from an enormous mixture of public, private, and internal text about AI safety, alignment research, and the concerns of the people working closest to it.

Either way, it’s fascinating. And a little creepy.


Discuss this post on Hacker News.