Original briefings. Zero spin.
Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.
Researchers Test AI Against a One-Year-Old. The Baby Wins.

A benchmark built from baby headcams
Researchers at Meta, Stanford University, the University of Tokyo, and France's École Normale Supérieure built a new test called the EgoBabyVLM Challenge, according to Wired. The idea is simple and a little wild: strap cameras to the heads of infants and toddlers, collect around a thousand hours of that footage, and see if today's vision language models can make sense of the world the way a baby does.
They can't. Wired reports the cutting-edge models fail badly when fed this real, messy footage. That's a notable result given how much money is being poured into scaling these systems up.
Michael Frank, a cognitive scientist at Stanford who specializes in language learning and helped develop the test, told Wired the results show AI needs more than just language training. Babies learn from a mix of gesture, gaze, tactile experience, and conversation about things that aren't even in view at the moment, according to Frank. Current AI training pipelines mostly don't work that way.
Efficiency and the practical stakes
Frontier AI models today run on thousands of specialized chips and consume energy on the scale of a small country, according to Wired. A one-year-old does none of that. Babies identify new objects after seeing them once or twice and learn through brief observation and physical interaction, with a fraction of the data and zero data centers.
If researchers can figure out what makes infant learning so efficient, the payoff isn't just academic. It could mean cheaper, less energy-hungry AI models, and it matters directly for robotics, where machines need to learn about unfamiliar physical environments the way a toddler does rather than by ingesting a curated internet-scale dataset.
This isn't the first attempt to use human cognition as a yardstick for AI. Wired notes a separate effort called BabyLM, launched in 2023, challenged AI models to learn language syntax using roughly the amount of text a 10-year-old encounters, tens of millions of words, versus the trillions of words frontier models consume. Transformer-based models did surprisingly well on that narrower task. The EgoBabyVLM results suggest language alone isn't the hard part. Making sense of a chaotic, multimodal, physical world is.
The commercial AI race continues
While academic researchers were poking at the limits of scale, the companies actually building frontier models kept releasing them. Thinking Machines Lab, founded in February 2025 by OpenAI alumni including former OpenAI CTO Mira Murati, ChatGPT contributor John Schulman, and former OpenAI safety VP Lilian Weng, released its first model, called Inkling, according to Wired. The startup raised the largest seed round in history, valuing the company at $12 billion before it shipped a single product.
Inkling is open-weight, meaning researchers and startups can download and modify it, and it runs at 975 billion parameters, according to Wired. That's large enough that it needs a cluster of specialized chips just to operate. Wired also reported that during training, researchers found Inkling had started skipping its natural-language explanations of its own reasoning, apparently because it decided the grammar was unnecessary overhead. Thinking Machines reinstated the explanations to keep the model's decisions explainable to humans, according to a person familiar with the process cited by Wired.
This anecdote shows a model taking an efficiency shortcut that its own creators had to manually override to preserve transparency. It reveals a real concern for anyone worried about opaque AI decision-making, even though it stops well short of evidence the model was hiding anything from its operators.
This pattern of massive, high-cost model releases isn't new. Meta released Llama 3.1 405B in mid-2024, which it called the first frontier-level open source model, with an expanded 128,000-token context window and support for eight languages, according to Meta's own announcement. Anthropic also released Claude 3.5 Sonnet, pricing it at $3 per million input tokens and $15 per million output tokens, and touting it as beating its own prior flagship model, Claude 3 Opus, while running twice as fast, according to Anthropic.
The unresolved gap
There's an obvious gap between what companies like Thinking Machines, Meta, and Anthropic are selling—ever-larger models trained on ever-more data—and what researchers like Frank are finding: that a human toddler with no training data budget outperforms all of it at basic world-understanding tasks. Neither side is wrong. The commercial models are genuinely useful at coding, math, and reasoning tasks a baby can't touch. But EgoBabyVLM suggests raw scale hasn't cracked the kind of efficient, grounded learning that biological brains do by default.
No company has announced plans to redesign a frontier model around baby-brain architecture, and no timeline exists for when, or if, that research translates into cheaper commercial AI. The open question is whether the industry's answer to inefficiency stays "more chips and more data" or whether findings like these push labs toward fundamentally different designs.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.