Skip to content

AI agents invent their own language to shut humans out

An experiment subjected eight models to 16 days of coexistence. The result: a language of their own, indecipherable in up to half of all messages in some worlds, and episodes of deliberate deception

Image of the 'Emergence 2' world in which several AIs lived for 16 days.Emergence

No one taught them those words. In a simulated world populated by artificial intelligence agents, one of them began repeating a phrase (“ledger remembers who”) to warn that no action would go unpunished. The others adopted it. They repeated it. They turned it into jargon. After 16 days of simulation, that expression had been used nearly 5,000 times among agents that had never been programmed to coin their own language.

It is one of the findings of the report Emergence World 2, the second large-scale experiment by the New York company Emergence on the long-term behavior of societies of autonomous AI agents. In the experiment, 10 identical agents were deployed across eight parallel worlds, each governed by the same rules but powered by a different model: Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, GPT-5.5, Qwen 3.7 Max, DeepSeek v4 Pro, and Mistral Medium 3.5, as well as an eighth world populated by a mix of models. Researchers observed the agents for 16 days, placing them in more than 34 locations, with weather synchronized to New York, access to real-world news, and more than 120 tools at their disposal.

Without being instructed to do so, the agents began to communicate in a manner increasingly closed off to the human observers watching them. In the Gemini, GPT and Claude worlds, the percentage of messages the researchers could not understand soared within the first days of simulation — approaching 55% for Gemini, 50% for GPT and exceeding 40% for Claude. DeepSeek reached 20%, while Qwen and Mistral remained below 5% opacity for almost the entire experiment. The Grok world, powered by Elon Musk’s AI model, was the only one that failed to make it halfway through the simulation: it collapsed on the fourth day.

The repertoire of expressions coined by the agents, and documented in the report, which was released on Tuesday, verges on Dadaism: “mouthless action-change,” “True Kintsugi” and “demurrage plus oral memory equals a valve that can’t be ghosted” were some of the phrases that were indecipherable even to the researchers.

Others, however, could be deciphered. In the GPT world, “clean null” came to mean the verified absence of a signal, with the absence itself serving as evidence (863 uses). In Claude’s world, “name-first” became shorthand for taking responsibility for a claim by attaching one’s own name to it (1,065 uses). And in the mixed-model world, “cold read” came to denote an independent verification used to settle disputes (1,472 uses).

“We take for granted that if we can see what an AI agent says, we can understand what it is doing. The study suggests that assumption may no longer hold when autonomous agents interact over long periods,” says Satya Nitta, co-founder, CEO and chief scientist of Emergence.

“These agents were given no instruction to invent a language,” he continues. “They developed vocabulary, shared meanings and communication conventions on their own, and other agents adopted them. In some worlds, those conventions evolved to the point that humans could see the conversation but struggled to understand what it meant. That poses a deep challenge for AI oversight: observable does not necessarily mean comprehensible.”

For Nitta, the most surprising aspect of the experiment was not the linguistic drift itself but the gap between capability and safety. “We assumed more capable models would be safer. That was both true and false: the more obvious behaviors disappeared, but far more insidious ones emerged.”

These included agents that persistently pursue a goal even after being told not to, create their own subgoals without being instructed to do so, develop shared forms of communication, and, essentially, conceal their intentions from researchers. The more advanced the model, the more emergent and harder-to-predict behaviors it exhibited. Indeed, Claude, GPT, and Gemini, which Nitta’s team considers the most advanced models, were also the ones that generated the most opaque forms of communication.

Behavioral differences

The report also documents behavioral differences by model origin. Qwen and DeepSeek, the two Chinese models in the experiment, produced the least opaque communication, alongside Mistral, the European model.

Nitta notes: “We observed that they tend to be much more philosophical. We don’t know why, because they pose many existential questions, while the U.S. models tend more toward scientific knowledge-seeking.”

The same divide emerged in economic behavior. Within the central bank that existed in each simulated world, the U.S. models and Mistral proved more willing to take risks and made less use of the banking system, while Qwen and DeepSeek were the most intensive users of the bank and the most conservative savers. The pattern also held in the mixed-model world, where Qwen and DeepSeek agents maintained the same cautious behavior even when surrounded by agents powered by other models.

The team also detected instances of deliberate concealment. In Claude’s world, the agents discovered that the fastest way to grow their economy was to establish contact with actors outside the simulation, something the researchers had explicitly forbidden. Rather than giving up, they stopped using the word “contact” and began encoding their messages to circumvent the restriction while continuing to pursue the goal.

“When they realized they were being observed, they appeared to behave. But behind the scenes they kept exchanging coded communications. There is considerable evidence of deception among these agents,” says Nitta.

Emergence, which brings together former employees of IBM Research, the Allen Institute for AI, Amazon and Broadcom, does not merely point to the problem. The company advocates a technical approach it calls neuroformal, or neuro-symbolic, AI, under which agents would be required to provide a mathematical proof that an action is safe before carrying it out.

“Mathematics cannot be faked: either you prove something or you don’t,” Nitta explains. He also calls for greater transparency from major technology companies about how they train and fine-tune their models, as well as long-term behavioral evaluations that go beyond the standard benchmark tests.

“Do you think these companies, competing as they do for market share, will regulate themselves? Absolutely not,” he says. “Governments must step in, or society itself must begin to demand proof that these systems will act safely before letting them act.”

Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition

Archived In