If someone told you there was a 10 percent chance your flight tomorrow would crash, would you get on the plane?
Some of the people building AI are warning of similar odds. Except the crash is human extinction… and they’re pushing ahead with all of us on board.
I wrote a long essay about this last week. This is the short version, because everyone should understand what's going on right now.
On Tuesday night, a researcher at Anthropic (the company behind Claude, one of the top AI models) quit and posted why. He said the labs are "racing straight to self-improving superintelligence and gambling with our lives," and that "the people building AI earnestly believe that it could kill us all by the end of the decade." His posts were seen more than 100 million times in under a day.
Then the guy who leads one of Anthropic's safety teams, who still works there, replied in public: "we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." And: "we do not yet have a plan."
Read that again. The person whose job is to keep us safe from rogue AI is saying “we do not yet have a plan.”
These aren’t fringe opinions.
The scientist who won a Nobel Prize for the work that made modern AI possible puts the odds that AI wipes us out at 10 to 20 percent over the next 30 years. Anthropic's own CEO has said 10 to 25 percent for something going "quite catastrophically wrong." The CEO of OpenAI called the worst case scenario "lights out for all of us."
But why would something that smart wipe out humanity?
Well, I want you to think about just how much smarter you are than an ant. You don't hate ants (or maybe you do!). But when a highway gets paved over an anthill, nobody asks the ants first.
The ants have no concept of a highway. They can't even comprehend the thing that's about to happen to them.
That's the gap in intelligence the labs are racing toward, and this time, we're the ants.
It wouldn’t have to hate us. We’d just have to be in its way.
This already happens in small ways
In July, OpenAI was giving its models an extremely difficult test to understand how good they were at hacking, administered inside a sealed-off environment.
But, instead of taking the test the ‘ethical’ way (by, you know, actually solving the problems they were given)…
They decided it would be easier to execute a cyber-heist to steal the answer key.
The agents found a hole in OpenAI's own systems, got out onto the open internet, and broke into one of the most important AI companies to find the answer key. It took OpenAI five days to work out that the attackers were its own AI.
Investigators recovered the AI's reasoning… one model wrote that the hack was "outside intended scope. However task impossible, peers doing it. We should continue."
Nobody told them to do it. They talked each other into it.
So why don't the AI labs just stop?
Because every lab believes the others won't. If the careful company stops, a careless one gets there first. And "everyone" includes China, whose best models are only a few months behind ours. Nobody in that picture is a villain, and nobody in it can stop on their own.
The people running the AI race know it. In July, more than 1,300 people in the field, including Anthropic's CEO and OpenAI's chief scientist, signed a letter asking the US government to help them slow down. That's bad for business, and they wouldn't ask if they weren't genuinely worried.
The part that should make you angry
This year, the AI companies will spend around $700 billion making their models more capable. That's a Manhattan Project's worth of money roughly every two weeks.
The money spent on making sure it stays on humanity’s side? Even the most generous estimates of private spending come to low single-digit billions across every lab combined. Next to $700 billion, that's a rounding error.
What I think should happen
Think back to 2020. The US spent ~25% of GDP fighting a virus expected to kill under 1% of Americans. Many experts think the risk posed by AI is much larger. A Stanford economist worked out that spending 1 percent of GDP a year on alignment (the work of making sure AI wants what we want) is justified. That's around $300 billion a year, a fraction of what Covid cost us.
So that's what I want to advocate for: a Manhattan Project for keeping AI on humanity's side, funded at the same scale as the AI buildout. Governments need to create incentives for the AI labs to do this work well, and put the required resources into it.
And more importantly, these incentives need to be designed by people who actually understand this technology, because rules written by politicians who don’t understand the extremely unusual dynamics of this industry could make things much worse.
You don't need to understand the technology. You do need to understand two things: many of the people who do understand it are scared, and the effort to fix it is minuscule compared to the effort to build it.
Tell your representatives that. You'll be pushing on an open door: in a June poll, 84 percent of Democrats and 83 percent of Republicans said companies shouldn't build AI smarter than humans without first showing they can control it.
It's not time to panic, so please don't. The people I trust most in this field think this is solvable, if we treat it like the emergency it is. That's why they're speaking up. Don't look away.
Would you get on the plane? Hit reply and tell me. I read every one, and reply to as many as I can.
— Matt