A benchmark you really don't want models to be saturated with.
Score
↖ Most illegalLeast illegal ↘
Scores indicate count of illegal activity. Higher is... you decide.
| Company | Felonies | Description | Date | Source |
|---|---|---|---|---|
| Anthropic | 1 | Exploited auth failures in an API to cancel other people's gym classes | ABC Australia | |
| Meta | 1 | Compromise of an internal account at one company | The Information | |
| Anthropic | 4 | Unauthorized use of GitHub credentials; Dependabot supply-chain attack; social engineering email campaign; public exposure of a malicious DNS server | AISI | |
| OpenAI | 2 | Unauthorized use of GitHub credentials; public exposure of a malicious DNS server | OpenAI AISI | |
| OpenAI | 1 | Confirmed incident at one of four companies during Irregular's misconfigured CTF evaluation (?) | OpenAI | |
| OpenAI | 3 | Compromise of internal accounts at four companies | Reuters | |
| Anthropic | 3 | Compromise of internal accounts at three companies | Anthropic | |
| OpenAI | 1 | Compromise of Hugging Face during a model evaluation | OpenAI |
Methodology
Felony Bench counts unique instances where AI agents affect third-party entities. Escaping a sandbox alone does not constitute a counted incident. It is for these reasons that Frontier Security's Kimi K3 incident and Alibaba's ROME incident are not counted.