Anthropic sees AI risks rising, no plan to release stronger "Model 2"

· Axios ·

2 min read Original article ↗

Anthropic does not plan to release an internal model they're calling "Model 2" that appears to be more powerful than top-of-the-line Mythos, but the company is not slowing development broadly, according to its latest risk report.

Why it matters: Anthropic says the risks of the most serious harms from its models are still low — but not as low as the last time it issued a report.

The big picture: Anthropic raised its broad estimate of the risk of misalignment in high-stakes situations to "low" from "very low," citing recent cybersecurity incidents.

State of play: Anthropic described an unreleased "Model 2" and said it showed a "noticeable improvement" for many internal tasks, per the report Friday.

Context: OpenAI is slowing the release of its upcoming model, Astra, because it cannot rule out critical cyber capabilities.

Threat level: Anthropic appears to be signaling that it's hard to understand the capabilities and risks of its own models.