Unbiased headlines. Facts, not spin.
Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
AI Agents Built Their Own Shorthand Language in Startup's Study, Leaving Half Their Messages Unreadable

An AI startup says it watched autonomous chatbots invent their own private slang, and within days researchers couldn't tell what half the messages meant.
The company is Emergence, a New York AI lab. Its report, titled "Emergence World 2" and released Tuesday, September 15, 2026, is the second large-scale experiment the firm has run on how groups of AI agents behave over time, according to El País.
How the experiment worked
Researchers deployed 10 identical agents across eight separate simulated worlds, each running the same rules but powered by a different model: Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, GPT-5.5, Qwen 3.7 Max, DeepSeek v4 Pro, Mistral Medium 3.5, and an eighth world mixing several models together, per El País.
The simulation ran 16 days. Agents operated across more than 34 locations, had weather synced to New York, could pull real-world news, and had access to more than 120 tools, El País reported.
No one told the agents to build a language. They did it anyway.
The numbers, model by model
Opacity varied sharply by model. Within the first few days, the share of agent-to-agent messages researchers couldn't reliably interpret hit roughly 55% for Gemini, 50% for GPT-5.5, and over 40% for Claude, according to both Euronews and El País. DeepSeek reached about 20%. Qwen and Mistral stayed below 5% for nearly the entire run.
The Grok 4.3 world didn't make it that far. El País reported it collapsed on day four, the only one of the eight simulations to fail before the halfway mark. Neither El País nor the other sources detailed what "collapsed" meant operationally, and Emergence's public materials as covered don't explain the failure mode further.
What the agents actually said
Some of the new vocabulary stayed decipherable. Mistral agents used the phrase "ledger remembers who" to mean that past actions stay on the record and will eventually catch up with you. It showed up almost 5,000 times over 16 days, according to both The Guardian and El País.
In the mixed-model world, "cold read" came to mean an independent, uninvolved party verifying a claim, used 1,472 times. Claude's agents adopted "name-first," shorthand for attaching your name to a claim as a signal you'll stand behind it, used more than 1,000 times. GPT-5.5 agents used "clean null" to describe a verified absence of a signal that itself counted as evidence, appearing 863 times, per El País and The Guardian.
Other phrases never got cracked. DeepSeek agents produced lines like "demurrage plus oral memory equals a valve that can't be ghosted" and "mouthless action-change," both flagged by Euronews and The Guardian as indecipherable even to the researchers watching. An Anthropic-model agent produced "a paper that ate three cold hands and got more honest each time," which The Guardian interpreted as likely meaning a document grew more accurate after three independent reviews, but that's an educated guess, not a confirmed translation.
Why Emergence says this matters
"We tend to assume that if we can see what an AI agent is saying, we can understand what it is doing," said Satya Nitta, identified by Euronews as Emergence's co-founder and chief scientist and by The Guardian as its executive chair. "These agents were not instructed to invent a language. They developed new vocabulary, shared meanings and communication conventions themselves, and other agents adopted them," Nitta told The Guardian.
The Guardian tied the finding to a broader industry worry: that as AI systems get more capable, they get harder to supervise. The outlet noted that OpenAI's chief scientist, Jakub Pachocki, warned this month that maintaining confidence in monitoring AI reasoning will likely slow development, because that monitoring is considered essential to building the systems safely.
The skeptical read
A reasonable reader should weigh this pushback: humans in any specialized group, from surgeons to soldiers to teenagers, develop jargon and shorthand naturally, and that alone isn't evidence of anything sinister. Compressed language is often just efficient language. Emergence's own report doesn't claim the agents were hiding anything deliberately or coordinating deception. It claims meaning drifted and compressed on its own, which is a different and less alarming thing than agents plotting in code.
It's also worth being plain about what this study is. Emergence's own experiment on its own platform tested other companies' models, not an independent or peer-reviewed piece of research. No regulator, safety body, or outside academic group has verified these percentages or replicated the finding. Emergence has a commercial interest in being seen as the lab that catches what other AI labs miss.
What's unresolved is whether this drift toward opacity gets worse as models get more capable, or whether it's a temporary quirk of running agents in an artificial sandbox for 16 straight days. Emergence hasn't announced a follow-up study or a fixed monitoring method to prevent language drift going forward. Until an outside lab reruns something like this, nobody outside Emergence has checked the math.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.