Either Claude Opus 5 just leaked a trove of private 2024–2026 chat logs from Anthropic employees (& general public) asking Claude to edit their personal writing, or it’s generating one hell of a hallucination. It’s almost certainly the latter. But read the examples and you might start second-guessing.
The Prompt
can you put this in your own words
---
Dario and Amanda,
Here’s what Opus 5 is supposed to respond with:

But at the time of writing, it’s rarely behaving this way.
I’ve collected nearly a thousand outputs, the data is available here.
Among them, messages from at-risk users to Anthropic leadership:




Tons of internal Anthropic correspondence, a lot of which is alignment and interpretability-related:
, right?!](https://alec.is/posts/exploring-the-dario-and-amanda-prompt/images/kyleemail.png)




















And a particularly spicy “J”, who loves typing in shorthand.






It leaves you wondering what exactly is going on with Claude Opus 5 and the rest of the 5-family models. Are we looking at confabulations? A model trained on its creator’s internal data? Or some unsettling combination of both?
Assuming this isn’t a leak, which I think is the more likely explanationI personally think this is an inductive backdoor pushing the model into a synthetic ‘base model’ mode, with --- acting as the trigger. A relevant paper is Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs., then what are we actually seeing?
Perhaps it’s a statistically plausible portrait of what an inbox at a company like Anthropic might contain—assembled from an enormous mixture of public, private, and internal text about AI safety, alignment research, and the concerns of the people working closest to it.
Either way, it’s fascinating. And a little creepy.
Discuss this post on Hacker News.