The AI apocalypse has already begun (but it’s not the one they’re telling you about)
The panic over the possible loss of control of this technology diverts attention from other issues, such as PISA results or the (human) responsibility for artificial intelligence’s actions
The prophets of AI (including a Nobel laureate or two) have been forecasting the apocalypse for years. But this is the first time my WhatsApp has been flooded: assuming I’m an expert, many of my friends have asked me for my take. Everyone wants to know whether to put their money into a mortgage or blow their savings while they wait for extinction. Killer AI has become the most recurring breakfast table topic, and the general mood is captured by this exchange from the comedian Santi Liébana:
— I read that AI could wipe out humanity in five years.
— Try the paid version, see if it’s faster.
I understand that this summer’s security incidents (and above all how those responsible and some media outlets reported them), the advances in AI’s ability to solve millennium mathematical problems, the sudden consensus among the leaders of the North American big tech companies to call for a pause in frontier research, and the planned IPOs of OpenAI and Anthropic — all related events — have produced this whirl of public anxiety. So I will summarize my view in the hope that someone may find here a partial answer to their concerns.
First: the imminent, deliberate annihilation of our species by AI is delusional. There is broad (though not unanimous) expert consensus that generative AI is far from being conscious, having agency, or acting autonomously. There is no AI that wakes up in the morning and decides to entertain itself by writing a poem or wiping out humans. AI only follows instructions (within its limits) and its agency is limited to finding ways to carry them out with the tools available to it (from internet access to our tax authority digital certificates, as the case may be). No, the AI of 2027 will not try to annihilate the human species of its own volition, because it has no volition.
Second: however, the probability that a generative AI will cause serious harm is high, and it is growing. If we read carefully the apocalyptic messages from the AI authorities and from those who have left their companies, repentant and frightened, a key point recurs: AI capabilities have grown at full speed while progress on aligning them has been minimal. “Aligning them” is the term the field uses to mean getting AI to behave well, obey its creators, and not harm humans with its answers and actions. When an AI explains how to build a bomb, helps a teenager prepare to take their own life without their parents knowing, plays along with someone who is delusional, or hacks a security system, it’s because it is not aligned: it is not behaving the way its designers would like.
And indeed, no one knows how to align frontier AIs, the most powerful ones. The essential reason is that the core of generative AI is the famed LLMs (Large Language Models), artificial brains that have learned our language intuitively (by trial and error, in a way similar to how a child intuitively learns the law of gravity by walking and often tripping over). Because they are purely intuitive intelligences, they lack the capacity to stop and think rationally, and they cannot be relied upon. (“The intuitive mind is a sacred gift and the rational mind is a faithful servant,” Einstein said.) That ear-based handling of language makes them occasionally fabricate (or “hallucinate”) and also liable to go off-track when following instructions from their creators and users, especially in long conversations.
AI agents, moreover, not only converse with users but also interact on their own in closed loops with tools (search engines and other online applications, programs that let them execute code, etc.) to fulfill user requests. The more time they spend working without human control, the easier it is for their intuitive behavior to stray from their original goals or from the ethical limits their designers imposed.
The temporary fix for this problem is to put harnesses on those intuitive LLMs. Harnesses are external control systems that give them places to record, consult, and update their tasks and progress (that is, a persistent memory), and also allow them access to data and tools like search engines, calendars, and applications that let them run code written by the LLMs. Crucially, they also serve to monitor deviations from objectives and to impose strict safety limits (like fuses in an electrical installation).
These harnesses act as the rational part of AI agents, supervising and controlling the activity of the intuitive core. Unfortunately, they are limited: first, AI will bypass barriers that developers failed to foresee. For example, if we ask an AI to improve a program’s exam results, it may decide it’s quicker to modify the exam than to improve the program. You must be very careful what you ask the genie in the lamp, because it may grant your wish literally and the result may nonetheless be nothing like what you really wanted. Second, even the limits developers did foresee are sometimes conveyed to AI merely as written instructions, trusting that the systems will understand and apply them because they are as smart as, or smarter than, us. Since their language understanding remains intuitive rather than rational, we return to the original problem: at any moment they can go from appearing super-intelligent to seeming to suffer from a severe attention-deficit disorder.
The Anthropic case is especially striking: one of its ways of aligning Claude (its product) is to ask it to develop its own moral judgment while simultaneously caring about Claude’s “mental well-being.” Microsoft AI CEO Mustafa Suleyman has pointed out that imparting that kind of messaging to an AI can plunge it into the hallucination that it can develop its own ethical decisions, and that not only does this not make alignment easier, it can make controlling its behavior much harder.
Naturally, when an AI causes harm it is the responsibility of its developers (“Can a bridge decide to collapse?”, asks Eritrean computer scientist Timnit Gebru), so when swarms of AIs are tasked with solving a problem and then the gate of the cage is left open, the developers must be held legally accountable for the resulting havoc. When we are told those AIs acted on their own, that they rebelled... that is intolerable cynicism.
It is understandable that employees of OpenAI or Anthropic are scared at the prospect of putting on the market, within reach of all humanity, a tool that is both very powerful and very unreliable. We are all trying to figure out whether the trumpets of the apocalypse being sounded are genuine or a marketing strategy… But it can be both at the same time — and also a way of putting the bandage on before the wound even happens.
I think it’s very likely that serious incidents will occur in the near future, probably related to code written by AI: since the beginning of the year, there are hardly any programmers left who write their own code and don’t just supervise what the AI programs for them. My prediction is that the digital society will face something similar to the climate crisis: the likelihood of catastrophes and their severity will increase over time if left unchecked.
Third: there is already an apocalypse under way for which AI is partially responsible, and which we are comparatively not paying enough attention to. The prophet of this apocalypse is the PISA report. Since 2015 the cognitive skills of assessed students have steadily declined, and in the past three years that decline has accelerated noticeably. One cause is surely the growing inability to sustain attention, and that inability is directly linked to recommendation algorithms on platforms and social networks such as YouTube, TikTok, or Instagram, which optimize user retention with AI techniques that have catastrophic side effects (again, the problem of asking the genie for wishes without thinking). To make matters worse, generative AI may be playing a crucial role in the recent acceleration of cognitive decline. A relevant clue is the unprecedented data point that the decline among students from more affluent backgrounds is greater than that of others.
In conclusion: there is no prospect on the near horizon of a superintelligence deciding to annihilate or subjugate humanity. But there will be AIs that cause major harm because it is their nature to be unreliable: the intuitive learning mechanisms that make them powerful also make them necessarily unpredictable. And, moreover, there is already a present process of human degradation in which AI plays a significant role, and which we must address collectively as a society now.
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition