Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
Nvidia's $20 Billion Groq Inference Racks Enter Full Production, First Deployment Lands at Nebius This Year

Since Nvidia signed its roughly $20 billion arrangement with Groq on December 24, 2025, the fastest-moving question in AI hardware has been whether the company could turn that money into shipping product before competitors caught up. On August 24, 2026, Nvidia answered that question. The company confirmed its Groq 3 LPX inference rack has entered full production and will be operational at cloud provider Nebius before the end of the year, according to CNBC.
Nvidia senior director Dion Harris told reporters the racks target a specific problem: making AI agents feel instantaneous rather than laggy, especially for coding tools where users notice every second of delay. "For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive" service levels, Harris said.
What's actually in the box
Each liquid-cooled Groq 3 LPX rack packs 256 individual Groq chips and can deliver roughly 3,400 tokens per second, based on a benchmark from Artificial Analysis that Nvidia cited, per CNBC and Anadolu Agency. The chips carry 500 megabytes of SRAM directly on the die, a design that keeps data close to the processor and cuts the memory bottlenecks that slow down other inference hardware.
Nvidia says pairing the rack with its new Vera Rubin systems can push throughput as high as 35 times per megawatt compared with conventional setups, according to The Motley Fool. That efficiency claim matters because electricity, not chips, is increasingly the bottleneck constraining how much AI capacity companies can build.
Groq's chips are made by Samsung Electronics. Nvidia's own GPUs come from Taiwan Semiconductor Manufacturing Co., a split that gives Nvidia two separate foundry relationships feeding one product line.
The deal nobody quite bought
The transaction behind this hardware doesn't look like a normal acquisition. Nvidia paid roughly $20 billion in cash, but rather than buying Groq outright, it licensed Groq's LPU designs on a non-exclusive basis and hired away key personnel, including founder and CEO Jonathan Ross, according to Crypto Briefing and KuCoin. Groq kept operating as an independent company under new CEO Simon Edwards, running 13 data centers globally.
Groq has since raised its own money on top of Nvidia's payout: $650 million in June 2026 and another $350 million in August 2026, pushing its valuation to $3.5 billion. Nvidia itself participated in that August round, according to Crypto Briefing, meaning Nvidia is now both a licensor's customer and an investor in the same company.
This structure carries a weakness. Yahoo Finance's bear case notes that because Nvidia licensed technology and hired talent rather than buying the company, investors have less visibility into what Nvidia actually obtained for $20 billion or whether the deal will produce returns matching its size. Bernstein analyst Stacy Rasgon countered to CNBC that Nvidia's balance sheet is strong enough to absorb a deal this size with little strain regardless of how it's structured. Both points can be true at once: the deal is financially survivable for Nvidia, and it's still unusually opaque for a $20 billion transaction.
The competition isn't standing still
Nvidia isn't the only company chasing low-latency inference. Advanced Micro Devices has announced plans to integrate its own rack-scale systems with chips from Cerebras, which has since gone public, according to CNBC. OpenAI's newly announced "Ultrafast" mode, powered by Cerebras, currently promises 750 tokens per second, less than a quarter of the 3,400 tokens per second Nvidia claims for Groq 3 LPX.
Nvidia CEO Jensen Huang has framed the Groq hardware as additive, not a GPU replacement. He said in March that he would allocate a quarter of coding-focused data center capacity to Groq chips, keeping the rest on Vera Rubin. "This isn't about replacing GPUs," Harris told CNBC. "It's about using the right price, right processor for the right part of the workload."
The near-term test is straightforward and falls on one company. Nebius has to actually bring the Token Factory deployment online before the end of 2026 and show real-world throughput matches Nvidia's benchmark claims. If it does, Nvidia has proven it can go from a $20 billion check to shipping hardware in under a year. If Nebius's numbers fall short of the 3,400 tokens-per-second figure Nvidia is citing, the gap between marketing and deployed reality becomes the story analysts start asking about on the next earnings call.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.