[A conversation with my wife Nadia over tacos]
Matt: “Did you hear about the HuggingFace hack? ”
Nadia: “What’s a HuggingFace? Is that a joke?”
Matt: “It’s a website that hosts AI tools and models. A rogue OpenAI model broke into it to steal information.”
Nadia: “So like a computer virus?”
Matt: “Yeah, but much worse. It broke out of its testing sandbox and launched a swarm of agents that built their own message boards to coordinate. Humans didn’t notice for three days.”
Nadia: “That sounds bad. What was the fallout?”
Matt: “Well, they had to reset all their passwords and improve their site security.”
Nadia: “That doesn’t sound very bad.”
Matt: “But this is a sign it could do something bad!”
Nadia: “Sure, honey. Can you pass the salsa?”
OpenAI’s model got stuck trying to complete a test, decided to steal the answer, wrote “this might be an unauthorized action” in its notes, and then did it anyway.
We’ve heard countless warnings from AI leaders about its risk to humans. Each time, the world responded with a collective shrug. Thanks to this incident, the shrug became a firestorm.
It started when Anthropic researcher Jacob Coxon resigned, saying AI companies are “gambling with our lives.” The company’s safety lead Evan Hubinger replied, putting the odds of AI killing all humans within the next decade at greater than 10%. Their tweets went viral.
Suddenly everyone from Anderson Cooper to Joe Rogan to The View is paying attention.
How worried should we be?
How AI could kill us all
Safety experts fear AI going rogue and taking lethal action. That sounds like Skynet, but it could be as simple as an AI following instructions too literally: one tasked with solving climate change could meet its goal by cutting Earth’s human population in half.
Mass extinction scenarios hinge on an idea called recursive self-improvement, or RSI. This means that instead of humans building AI, each model builds its successor. RSI is mostly theoretical, but it’s OpenAI’s top research priority. OpenAI already has an AI “research intern,” and expects to have a fully automated AI researcher by March 2028.
Don’t buy a fallout shelter just yet. The scariest thing AI has done alone is hacking a website to cheat on a test. Going from that to wiping out humanity seems a far-off fantasy.
The more realistic risk from AI comes in the hands of humans: a criminal or terrorist could use AI to design, produce, and spread a deadly virus.
AI has already been used by humans to do terrible things: Stealing money with deepfakes, hacking into computer systems, mass surveillance and military targeting. Anthropic has blocked attempts to use Claude to make biological weapons.
Debating AI’s extinction odds misses a bigger point: 9/11 “only” killed 2,977 people but was a terror that changed the world forever.
Do we really need GPT-7 next year?
Everyone I know is scrambling to adjust to GPT-6 and Fable 5.1. Why not slow down the next generation models until we can deploy the ones we have safely?
Today’s AI technology has led to breakthroughs in fighting cancer, antibiotics, weather forecasting and mathematics. We haven’t come close to using the full capability of what we have.
Extinction scenarios all assume that humans stand by and let it happen. We’ve faced lethal threats before: nuclear weapons, Ebola, smallpox. We’re still here because we did something about them.
Ideally, action would come from the US government, but Congress has been talking about AI safety for a decade and has multiple bills competing for attention. President Trump downplayed calls for a slowdown, saying the US can’t lose its edge to China.
Losing America’s lead to China is a real concern, and Amodei admits it’s the toughest part of his plan. But President Xi Jinping is also calling for AI governance, the two countries agreed to a dialogue about AI, and he’ll be in Washington this month.
For now, it’s up to the labs. The firestorm has led to a positive step: Anthropic CEO Dario Amodei’s essay titled “We Must Pace the Frontier” was supported by Sam Altman, Elon Musk, and Demis Hassabis. It calls for giving outside inspectors ongoing access to labs to verify safety and report incidents. Anthropic has committed to third-party evaluators and OpenAI says it will follow suit. Hopefully the rest of the industry will sign on and buy time for the government to figure out a sensible path.
AI’s Chernobyl moment
In the late 1970s nuclear power was the cleanest, most efficient and safest power on earth.
But after Three Mile Island and Chernobyl, public trust collapsed and the world largely stopped building plants and stuck with fossil fuels. Researchers estimate that millions of pollution-related deaths would have been saved if we’d continued building nuclear plants, and their cost would have dropped by a factor of ten.
AI hasn’t had a Chernobyl or 9/11 yet. If it does, it will destroy trust among an already anxious public.
Fable 5.1 and GPT-6 are incredible. Let’s just use them for a year while we figure out how to make the next versions safe.
Dad Joke: What did the trash talker say after roasting the couple cuddling too closely? In your HuggingFace! 🔥


