READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Amazon Exec: 85% of Companies Piloting AI Agents, Only 5% Ship Them Because They Don't Actually Work

Amazon Exec: 85% of Companies Piloting AI Agents, Only 5% Ship Them Because They Don't Actually Work
At VB Transform 2026, Amazon's AGI Autonomy director Bryan Silverthorn said the enterprise AI agent problem isn't smarter models, it's reliability nobody is properly measuring. Cisco data shows 85% of enterprises are piloting agents but just 5% have them in production, and companies are checking uptime while ignoring whether the thing actually gets the answer right.

The AI industry sold enterprises on autonomous agents that would handle customer service, QA, and back-office work without human hand-holding. According to Cisco data cited at VB Transform 2026, 85% of enterprises are piloting AI agents. Only 5% have actually shipped them to production.

A Tuesday session at the conference featured Bryan Silverthorn, Director of AGI Autonomy at Amazon, who joined the company through its acquisition of Adept AI and now runs multimodal agent training inside Amazon's AGI lab. The focus was on why the gap exists.

Silverthorn's diagnosis: the problem isn't that the models aren't smart enough. It's that companies have no real way to measure whether an agent will keep working once it leaves the lab.

The Serial Number Problem

Silverthorn described a real deployment gone wrong. A customer built an agent to extract serial numbers from screens for software QA. It ran flawlessly for two months. Then it started intermittently reading the wrong numbers.

The cause wasn't a dramatic model failure. It was a vision encoder that behaved differently depending on where the serial number appeared on screen, triggered by a software change too small for a human to notice. The agent passed every eval anyone threw at it, then quietly broke in production.

According to VentureBeat's own research presented at the conference, half of surveyed companies shipped agents that passed internal evaluations and then failed in front of real customers. Enterprises are tracking the wrong signals entirely. The same research found companies overwhelmingly monitor uptime, whether the agent is running, while ignoring accuracy, whether the agent is right. That's checking a pulse without checking the diagnosis.

Most companies also aren't building their own testing rigor. They default to whatever evaluation the model vendor hands them and stop there, according to VentureBeat's findings. That leaves enterprises making a binary bet: trust the vendor completely, or trust nothing at all. Neither is a strategy.

Four Dimensions, Not One Score

Silverthorn credits Princeton research for a framework breaking reliability into four separate dimensions: consistency, robustness, predictability, and safety. Most evals smash all four into a single pass/fail number, hiding exactly where an agent is likely to break.

"It unpacks different factors that I see tangled together in almost every eval I've ever seen," he said.

This isn't an argument against building better models. "The models have to be better. Obviously, we're working hard on making the models better," Silverthorn said. Better models alone won't close the 85-to-5 gap. Teams need to identify where their specific application is likely to vary and measure accordingly, matching rigor to the actual stakes of the task.

Managing Agents Like Interns

Inside Amazon's AGI lab, researchers reportedly call their agents "interns" — as in, "I'll have my intern talk to your intern." Silverthorn said the joke reflects a real operating philosophy: agents are powerful and occasionally clueless, capable of great work and total derailment in the same afternoon.

The prescription is managerial, not technical. Ask what could go wrong. Build in backups and undo capability. Decide up front what level of risk is tolerable. "You can ask the intern, 'Hey, what might you do wrong here? How might you mitigate your negative outcomes?'" he said.

Sovereignty Is the Other Half of the Problem

A separate session at the conference tackled a related enterprise concern: control. Rachad Alao, VP of product engineering at Cohere and a former responsible AI lead at Google and Meta, told VentureBeat CEO Matt Marshall that reliability worries compound when companies don't control their own AI stack.

Alao argued sovereignty isn't just running a model behind a firewall. It means controlling GPUs, private cloud infrastructure, governance systems, and the connectors and tools acting on enterprise data. "You want to have control on the entire stack," he said, pointing to banks, hospitals, and governments as organizations that can't afford to guess where their data actually lives.

Marshall pushed back on one common argument for smaller, locally run models: inference prices keep falling, so why bother optimizing. Alao's answer was that total consumption is rising faster than prices are dropping, because agents doing multi-step reasoning and tool calls burn far more tokens than a simple chatbot. He cited an unnamed Canadian bank using Cohere's on-premises models for regulated workloads while routing lower-stakes tasks elsewhere.

What's Actually Unresolved

Neither Silverthorn's Princeton-derived framework nor Alao's sovereignty pitch is an industry standard yet. There's no independent benchmark forcing enterprises to measure consistency, robustness, predictability, and safety separately, and no regulatory requirement that vendors disclose eval methodology beyond what they choose to publish. Until that changes, the 85%-piloting, 5%-shipped gap identified by Cisco is likely to persist, and enterprises will keep discovering their agents' blind spots the way that QA customer did: in production, two months in, after it's already gone wrong.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatAmazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026
center
VentureBeatCohere VP says enterprise AI sovereignty requires control of the full agent stack at VB Transform 2026
center
VentureBeatAmazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026