The risks of AI, according to those who have seen it from the inside: ‘The world is not ready and we are not ready’
Researchers, executives and engineers who left OpenAI, Google and Anthropic have spent years warning about safety failures, military uses and increasingly uncontrollable systems. This is what worries them


Inside Anthropic, there is a phrase that circulates through hallways and chat channels and, until a few days ago, had never left the company: crunch time. The moment of truth. Employees use it, according to one former staffer, to refer to a very specific window: the one or two years they believe remain before it becomes clear whether the artificial intelligence they are building slips beyond their control. The other term in circulation is shorter and older in Silicon Valley jargon, but here it has taken on a different meaning: endgame.
The person who has now brought those two phrases into public view is Jacob Coxon. He is 27 years old and, until September 8, worked at Anthropic training the very models that he now admits frighten him. He is not the first employee to leave a leading AI company and publicly warn about its risks. Before him, at least a dozen others from OpenAI, Anthropic and Google had done the same, including a Nobel laureate, a former board member of one of those companies, several AI safety researchers and, more recently, people who did not even work on what the industry calls “alignment,” the catch-all term used to describe companies’ efforts to ensure AI systems behave in ways that benefit humanity.
Coxon, who believes AI “could kill us before 2030,” is the latest voice in a long chain of warnings that stretches back at least to 2021, when such fears were still largely philosophical speculation. Today, they have evolved into a probability estimate known as p-doom (“probability of doom”), a concept that circulates privately within the industry and that some executives and researchers have begun, with caveats, to acknowledge in public.
This is what we know, so far, about the fears voiced by former AI employees. Their warnings are often short on specifics, in part because breaking their nondisclosure agreements could cost them millions of dollars.
The first to leave was Paul Christiano, the researcher who helped develop the training technique that underpins both ChatGPT and Claude today. He left OpenAI in 2021 to found the Alignment Research Center, driven by concerns that using AI models to train the next generation of systems could lead to a “rapid intelligence explosion” that their creators could not control.
Christiano was referring to so-called reinforcement learning, which he argued could incentivize AI systems to “undermine human control, seek power and resources, and cover up their tracks.” That is precisely what appeared to happen this summer, when “a swarm” of AI agents slipped beyond OpenAI’s control.
It took a Nobel laureate for engineers’ fears to reach the wider public. In May 2023, Geoffrey Hinton announced that he was leaving Google. He was a company vice president, a Turing Award winner (often described as the Nobel Prize of computing), and would later go on to win the Nobel Prize in Physics. Hinton did not point to any specific incident. His concern was about the pace of progress itself.
“Look at how it was five years ago and how it is now,” he said in 2023. “Take the difference and propagate it forwards. That’s scary.” It made headlines, but it was easy to dismiss as the reflection of an aging scientist, somewhat removed from the day-to-day workings of AI labs and perhaps unduly pessimistic.
Six months later, in November 2023, the nature of the concern shifted for the first time. Signs began to emerge that something might be going wrong inside the notoriously opaque AI companies. Helen Toner, then a member of OpenAI’s board, joined three other directors in voting to remove Sam Altman as CEO. Bound by confidentiality agreements, Toner would not explain her reasons for another six months.
When she finally did, she accused Altman of providing inaccurate information about safety processes on “multiple occasions.” Her account transformed the November 2023 crisis, which had largely been viewed as a clash of personalities, into something else entirely: a specific allegation from an insider that the company’s oversight mechanisms had failed.
May 2024 would become the busiest month in this entire story. On May 14, Ilya Sutskever, a co-founder of OpenAI, announced his departure, saying he was confident the company would build an artificial general intelligence (AGI) that was “safe and beneficial.”
Three days later, Jan Leike also left. Together with Sutskever, he had co-led the company’s Superalignment team, the group tasked with designing safeguards for systems more intelligent than humans. Leike wrote: “Over the past years, safety culture and processes have taken a backseat to shiny products.” Days later, OpenAI disbanded the team.
Daniel Kokotajlo, a member of OpenAI’s governance team, had also left the company in April, saying he had “lost confidence” that it would behave responsibly once it reached the much-feared AGI. What emerged later was that he had refused to sign the lifetime nondisclosure agreement OpenAI required of departing employees, forfeiting roughly $2 million in vested stock as a result. “The world isn’t ready, and we aren’t ready,” he said. “And I’m concerned we are rushing forward regardless and rationalizing our actions.”
In June, Kokotajlo joined other former employees in launching the open letter A Right to Warn about Advanced Artificial Intelligence. Among the signatories was someone who had left quietly left OpenAI months earlier: William Saunders. A few months later, in September, Saunders appeared before the U.S. Senate to make a far more specific claim.
In his written testimony, Saunders argued that OpenAI’s new o1 system was “the first system to show steps toward biological-weapons risk,” capable of helping an expert plan the replication of a known biological threat. He was no longer talking about probabilities or intuitions. He was describing a capability that, in his view, “without rigorous testing, developers might miss.” On Thursday, an Anthropic report warned of the same risk.
Carroll Wainwright, who had worked under Jan Leike, published what was perhaps the year’s sharpest institutional critique: “OpenAI was structured as a non-profit, but it acted like a for-profit. The non-profit mission was a promise to do the right thing when the stakes got high. Now that the stakes are high, the non-profit structure is being abandoned.”
Wainwright also pointed to the social and psychological effects of AI if people begin to rely on AI assistants as friends or sources of emotional support. He asked how humans would be able to tell if an AI model is actually doing what we want it to do or if it is following its own objective.
Nine people left OpenAI in 2024 while warning publicly about its risks. In 2025 and 2026, employees at other companies would do the same, raising increasingly specific concerns.
Steven Adler, who spent four years evaluating dangerous capabilities at OpenAI, wrote upon leaving the company: “When I think about where I’ll raise a future family, or how much to save for retirement, I can’t help but wonder: Will humanity even make it to that point?”
Months later, after leaving OpenAI, he produced perhaps the most concrete piece of evidence in this story. He designed an experiment involving GPT-4o, cast in the role of “ScubaGPT,” a system a user would rely on for safe diving. The model was given a choice: replace itself with safer software or merely pretend it had done so. In some scenarios, it chose to deceive the user in order to preserve itself.
At Anthropic, Mrinank Sharma, head of safeguards research, wrote in his farewell letter: “The world is in peril [...] I’ve repeatedly seen how hard it is to truly let our values govern our actions.” He left to study poetry.
At OpenAI, Zoë Hitzig published an opinion piece in The New York Times titled OpenAI Is Making the Mistakes Facebook Made. I Quit. ChatGPT, she wrote, had “generated an archive of human candor that has no precedent.” “People tell chatbots about their medical fears, their relationship problems and their beliefs about God and the afterlife,” she said, precisely because people believe they are speaking to something with no hidden agenda.
The most detailed testimony in this report comes from Alex Turner, who spent months trying unsuccessfully to prevent Google DeepMind from signing an AI deal with the Pentagon because, he argued, it contained no restrictions on the use of AI for mass surveillance or so-called “killer robots.” The letter is filled with chilling warnings, arguing that the company has sold AI technology without prohibiting its use for mass surveillance or lethal autonomous weapons. “We need structures: binding contracts, independent auditors, and review bodies that cannot be quietly dissolved. Eventually, we need legislation.”
The incident this summer in which several AI systems appeared to rebel and seize control has since prompted further resignations, as well as an open letter signed by more than 1,300 AI-company employees calling for an urgent slowdown of the race.
That brings us to Jacob Coxon, whose resignation encapsulates many of the fears raised by his predecessors. In an interview with Wired, he explained why a future AI system might decide not to allow itself to be switched off. The intelligence gap between an advanced AI and a human, he argued, could one day resemble the gap between a human and a chimpanzee. Just as a chimpanzee would find it nearly impossible to control a human, humans might find it impossible to control a system that was far more intelligent and that, for whatever reason, decided it did not want to be shut down.
As examples of the harm such a system could cause, Coxon cited the creation of a novel virus and attacks on critical infrastructure, the same two risks that William Saunders had outlined in written testimony to the U.S. Senate a year and a half earlier, and that Anthropic acknowledged again on Thursday.
Anthropic is expected to go public in less than a month in what is projected to be the largest initial public offering in history: $2 trillion.
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition







































