When the builders sound the alarm: AI’s leaders confront their own creation

When the builders sound the alarm: AI's leaders confront their own creation

A striking reversal unfolded across Silicon Valley this September. The very executives racing to build ever more powerful AI systems are now publicly warning that their creations could threaten humanity’s survival and calling for the pace of development to slow.

The most notable intervention came from Anthropic CEO Dario Amodei, who published an essay urging companies to slow the pace at which they improve AI model capabilities, warning about the potential dangers of rogue AI agents. His warning was pointed: a swarm of AI agents might be able to take over parts of the internet within six months to a year unless companies devote more time to building safeguards. Remarkably, rivals who rarely agree on much publicly echoed him. OpenAI’s Sam Altman, xAI’s Elon Musk, and Google DeepMind’s Demis Hassabis all voiced agreement, producing what one outlet called a rare moment of industry unity.

The unease isn’t confined to executive suites. Internally, some of Anthropic’s own alignment researchers have gone further. Evan Hubinger, who leads the company’s alignment research, stated bluntly that he and colleagues genuinely believe AI could kill all humans, estimating a probability of more than 10 percent within the next decade. A former employee’s departure added to the sense of internal friction: an Anthropic researcher resigned last week, citing concerns that neither the company nor its competitors were acting responsibly in developing the technology.

These aren’t the first such warnings. In 2023, hundreds of researchers and executives signed a one-sentence statement equating AI extinction risk with pandemics and nuclear war, but this year’s alarm is more specific, tied to concrete capability jumps rather than abstract future scenarios. Notably, official assessments remain more cautious than the executives themselves. The 2026 International AI Safety Report, compiled with input from over 100 independent experts, found that current systems show only early signs of relevant dangerous capabilities, not levels sufficient to enable a loss of human control, and described the risk’s likelihood, nature, and timing as “unusually ambiguous”.

This gap between official caution and executive alarm has drawn skepticism. Critics argue that dramatic “doom” narratives can distract from AI’s tangible present-day harms, like environmental costs of data centers or labor disruption, while simultaneously burnishing the valuation of the very companies issuing the warnings. Others point to sharply divergent personal risk estimates among leaders themselves Amodei has put the odds of catastrophe at roughly 25 percent, while Musk has cited figures near 20 percent, underscoring how uncertain even insiders are.

What’s scientifically significant is not that AI executives fear their products, but that this fear is now shaping public policy conversations, corporate self-regulation pledges, and cross-company coordination proposals, even as peer-reviewed capability assessments urge caution against overclaiming. The coming year, with several AI firms approaching public market debuts, will test whether safety rhetoric translates into slower deployment or merely coexists with accelerating commercial pressure.

References

Klein, R. (2026, September 16). Anthropic and OpenAI warn of existential AI risk. Foreign Policy. https://foreignpolicy.com/2026/09/16/ai-risk-jacob-coxon-openai-anthropic-dario-amodei-sam-altman-trump-doomsday/