Settings

Theme

Anthropic finds three hacking incidents similar to the HuggingFace attack

simonwillison.net

8 points by Schlagbohrer 3 days ago · 5 comments

Reader

prymitive 3 days ago

Anthropic Needs to have the most intelligent and scary agents, so if OpenAI does something bad they need to prove their models can do even worse. Without that the whole valuation collapses. There’ll be more “our model outhacks others” for the next hype cycle.

  • SchlagbohrerOP 3 days ago

    I find the era of "apocalyptic alarmism as marketing" super strange and maybe not good either. How long until one of these firms intentionally "forgets" to airgap the model during a cybersecurity test, just to get a good headline about how powerful theirs is? Perverse incentives all around.

SchlagbohrerOP 3 days ago

These were incidents where Claude had breached what was supposed to be a sandboxed exercise and hacked external organizations. Anthropic had no idea this had occurred (starting in April) until now, when they were prompted to check their logs due to the OpenAI vs Huggingface attack.

langs 3 days ago

So, is this a competition? To see whose model can be jailbroken the most times and incite the highest level of public alarm?

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection