sean goedecke

They really do think AI might kill everyone

A recent resignation tweet from an Anthropic researcher has everyone talking about the AI apocalypse again. Among other things, he said:

The people building AI earnestly believe that it could kill us all by the end of the decade.

Many people found it hard to believe that AI researchers think this way. Some explained it as a PR campaign to promote AI regulation, or as self-promotion, or as a way to boost AI company stock prices. Others felt it had to be impossible, because if you really believed this you’d be bombing datacenters instead of posting on Twitter.

In fact, not only do many AI researchers1 seriously believe this, they’ve been thinking and writing about it since the mid-2000s. Eliezer Yudkowsky — the ur-figure for most modern AI safety culture — has been publishing papers since at least 2008 saying that superintelligent AI could destroy all human life. It’s been such a common idea that the AI research community has abbreviated “how likely you think AI is to kill everyone” to “p(doom)” (i.e. the probability2 of doomsday) since around 2010.

I know it sounds very silly if you’re not in the AI bubble. But it really is true, and if you assume there has to be some different motive you’ll be deeply confused by what AI researchers say and do. They truly do believe that there is a reasonable chance that superintelligent AI will kill everyone.

This is why AI researchers care so much about “alignment”: building AIs that share genuinely human beliefs and values. If we build a “misaligned” superpowerful AI — an AI with goals that are alien to us — it might sweep humanity away. It could kill everyone deliberately, e.g. to stop us getting in the way of some goal. It could kill everyone in passing, e.g. like we might pave over an anthill to build a road. Either way, everyone dies.

How AI could kill everyone

Okay, but how? What do these people think is actually going to happen? AIs are computer programs running in a datacenter somewhere. How could they possibly cause the extinction of humanity? Wouldn’t someone just turn them off? Unsurprisingly, AI research nerds have come up with some concrete answers to this in the last two decades. Here they are, in order of plausibility:

An AI could make and release some super-pathogen or virus. In the hope that current AIs can achieve huge breakthroughs in medicine similar to the ones they’ve already achieved in mathematics, we’re setting up autonomous labs3. Bioweapons have been terrifying biologists for decades, and with good reason. Historical pandemics have killed up to 80% of affected human populations, and tend to be defeated by accident: the disease happens to evolve into a less virulent strain, or short incubation periods mean that infected people can’t carry the disease far, or some percentage of the population is naturally immune. A plague designed to be maximally fatal — or several plagues in quick succession with different characteristics — could be much worse.

Alternatively, an AI could trigger global thermonuclear war. We’re already seeing AI be integrated into military and government decision-making processes. If a rogue AI managed to set off a bunch of nukes, or to coordinate4 with other countries’ rogue AIs to nuke each other, we could be in an ordinary nuclear apocalypse scenario: billions dead in the strikes, billions dead in the ensuing famine, and so on. It’s commonly assumed that a post-global-thermonuclear-war Earth would still support some tiny human population, but a determined AI could surely find some way to mop up the stragglers.

There are some other theories. Once robotics has permeated the world economy, an AI could take over the robots (including drones) to kill everyone, like in Terminator. Or AIs could take advantage of nanotechnology to create self-replicating machines that turn the world into “grey goo”5. Or AIs could terraform the planet so as to make it unliveable for humans (as in Nick Bostrom’s famous paperclip example). Or they could do something else that our puny human brains aren’t able to think of.

Counterarguments

One common counter-argument6 here is to say “well, it’d be impossible to extinguish all human life — what about undiscovered tribes in the Amazon, or survivors living in the ruins of modern-day cities?” I don’t know, man. At some point you’re just conceding the argument: the policy positions you’d adopt if you thought AI might wipe out 99% of humans are the same as if you thought it might wipe out 100%. And like I said above, if an AI can kill almost everyone, it’s probably smart and capable enough to finish the job somehow.

Another is to say that the government will simply step in and nationalize the AI labs when the situation gets too dangerous. Maybe! But this kind of concedes the argument: a technology important enough to be fully taken over by the government is a terrifyingly dangerous technology.

A third is to say “well, someone would just turn it off”. I don’t find this plausible at all: an AI powerful enough to build a super-plague is an AI sophisticated enough to pretend it’s curing cancer, or to exfiltrate itself to some datacenter where it won’t be turned off, or to take some other countermeasures.

The winner takes it all

Why would you work in AI, if you believe this? Why wouldn’t you go live in the woods somewhere, or start bombing datacenters, or assassinating AI lab CEOs? For a few reasons.

An AI powerful enough to end humanity is an AI powerful enough to save it. I wrote about this in Help peer: many AI researchers believe that the only way for humanity to truly survive long-term is with the help of superintelligent AIs, so long as someone can figure out alignment.

Isn’t this a huge risk? Maybe not. If somebody is going to build superintelligent AI, you might be obligated to try and do it first. You can’t go and bomb every datacenter in the world, after all.

Why does it matter who’s first? Some popular theories of AI development involve a “foom” or “hard takeoff”: the first time someone really cracks self-improving AI, capabilities will increase exponentially, because smart AI will be better able to make itself smarter, ad infinitum7. There are no draws in the AI race. The first lab to figure out smart human-like intelligence will be the first one to figure out wildly superhuman intelligence, and thus will be in a position to stop anyone else from doing it.

This is an under-discussed point in the AI risk debate. Lots of AI researchers believe that the first thing a true superintelligence will do is reach out and stop all other AI research: either by hacking the labs, persuading them to stop, or literally drone-striking their datacenters. According to this view, if you’re an AI researcher and you think you can build an aligned AI, you should be working 24/7 so you can manifest God, and you should wake up every morning gripped by the fear that someone elsewhere has manifested the Devil, and your training datacenter no longer exists.

Conclusion

I have been on the fringes of this world for my entire adult life. I read Meditations on Moloch as a young adult and wanted to get into AI. I am one of the few people to read the entirety of the Sequences, Eliezer Yudkowsky’s million-plus-word magnum opus about rationality. I was too young for the Extropians mailing list, but I’ve spent years on LessWrong. On the other hand, I’m not a card-carrying rationalist: I think if you have a strong intuition on one side and a convincing-sounding argument on the other, you should pick the intuition8. I’m a deontologist, not a utilitarian. I don’t even live in San Francisco!

I’m conflicted about AI risk. The current behavior of AI agents does seem to vindicate a lot of the early science-fiction-sounding worries of the AI doomers, but modern LLMs are a lot more human-like than the alien minds in the apocalypse scenarios, and in general it does just seem too silly to credit (I guess I’m picking the intuition here).

However, it bothers me to see people dismissing these people as part of a PR operation, or as liars looking to boost an upcoming AI lab IPO, or as isolated crazies who haven’t thought their position through. Whatever else you say about the AI doomers, they have more than two decades’ history of explicitly spelling out exactly what they believe and why, even when it was complete science fiction to talk about AI at all. They’ve earned the right to be treated as sincere.


  1. In this post I’m going to use “AI researcher”, “AI safetyist”, and “rationalist” as reasonably synonymous terms for “someone who thinks there’s a chance AI kills everyone”.

  2. You could write a whole other post about the relationship between the rationalist/“AI safety” community and making concrete numerical predictions for unlikely future events.

  3. Alternatively, smart enough AIs might trick scientists with ordinary labs to produce dangerous substances.

  4. This is beyond the scope of this post, but many AI researchers believe that super-smart AIs will inherently come to agree with each other and eventually to coordinate without ever having to communicate, simply because they can predict what the other one will do.

  5. This was the most popular theory in the late 2000s, when nanotechnology was trendier.

  6. Relegating this counter-argument to a footnote because I hate it: many people say “we shouldn’t worry about AI risk, because climate change (or AI misinformation, or some other thing) is more urgent and serious”. You simply do not have to choose: it is possible to worry about multiple risks at the same time.

  7. Some people advocate for a “slow takeoff”. However, the main proponent of that view is Paul Christiano, who has just today joined OpenAI, citing “a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term”.

  8. This is kind of a Michael Huemer-ish position in epistemology.


If you liked this post, consider subscribing to email updates about my new posts, or sharing it on Hacker News.

Here's a preview of a related post that shares tags with this one.

Help peer

One of the most influential 20th century pieces of writing about AI is Isaac Asimov’s The Last Question. Although there are many humans in the story, the protagonist is the computer Multivac, who evolves over the course of ten trillion years from a single datacenter to a universe-spanning mind in hyperspace. Multivac (now called “AC”) ends the story like this:
Continue reading...