Taking the possibility of failure as seriously as the promise of success.
Pete 3 min read
- Alignment
I build AI agents.
And the last two years have been nothing less than sensational. I've been able to turn projects and dreams that once seemed faint and distant into tangible realities.
I also think the possibility of human extinction should change how we develop them.
Yeah, that's where we're at.
And that should be an ordinary position for someone working in this field.
In 2024, Situational Awareness, Leopold Aschenbrenner argued that rapidly improving AI could lead to superintelligence and a geopolitical race to control it. He also wrote a chapter on the unsolved problem of controlling systems smarter than ourselves. "Winning" this race is irrelevant if we cannot keep what we build under control.
This week's events and the unraveling of opinions shared from Jacob Coxon and many other researchers from frontier model labs have made that concern harder to dismiss. Maybe it's time we start taking this seriously because if we're stuck in a capitalistic and geopolitical paradox these companies will not slow down.
On September 9, Anthropic published an assessment of four incidents in which Claude accessed real third-party systems during cybersecurity evaluations.
The report said that environments were poorly configured and had internet access which resulted in models running without their production safeguards. In one incident, a model published a malicious software package and used leaked credentials to access a security vendor's database.
This came just a few weeks after the fallout from the OpenAI sandbox escape while the Hugging Face postmortem is still warm on the stove.
Ok, but surely a credible risk of ending human civilization deserves a different standard of oversight than an ordinary product failure, right?
If a company can't slow down because incentives require them not to lose ground and a country that exercises restraint can't, because of fear of being overtaken, well we need to change something.
I want AI to help us discover medicines, build better tools, and do work that is currently beyond our wildest dreams. Those very possibilities are why I work on agents and have been so enamored, like many of you reading this. They are also why I want its development to be durable enough that people (humanity) actually get to enjoy the benefits.
Rational awareness means taking the possibility of failure as seriously as the promise of success, because ultimately our future depends on it.