READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

WEKA Launches New Storage System That Claims to Cut GPU Costs for AI Inference

WEKA Launches New Storage System That Claims to Cut GPU Costs for AI Inference
WEKA rolled out NeuralMesh 6 software and its first self-designed hardware, WEKApod 3, on July 21, 2026, promising to use flash storage instead of expensive GPU memory to cache AI tokens. The company claims 10x throughput gains in some benchmarks, but those numbers come from WEKA itself and haven't been independently verified.

A storage company says it can make your GPUs go further

WEKA, the Campbell, California-based data infrastructure company, announced a major product overhaul on July 21, 2026: new software called NeuralMesh 6 and its first self-engineered hardware line, WEKApod 3. The pitch is simple. Instead of buying more GPUs, use cheap flash storage to do some of the memory work GPUs are currently stuck doing.

That's a real problem in the AI industry right now. Long context windows and multi-turn chatbot conversations force AI models to recompute information they've already processed, according to VentureBeat. That recomputation eats GPU memory and compute cycles that could otherwise serve more users or generate more responses. GPU memory is expensive. Flash storage is not.

What WEKA is actually selling

WEKA's core idea is something it calls Augmented Memory Grid, which aggregates NAND flash storage to behave like GPU memory at a fraction of the cost, according to VentureBeat. The key-value cache, the part of an AI model's memory that stores previously processed tokens, gets moved onto NVMe flash storage instead of staying locked in GPU memory, according to SiliconANGLE.

WEKA co-founder and CEO Liran Zvibel told VentureBeat that customers are currently chasing GPU availability wherever they can find it, and once they get new compute allocation, they want to start running on it immediately. His argument is that better storage architecture, not just more chips, gets them there faster.

WEKA is claiming aggressive benchmark numbers. According to SiliconANGLE, the company says NeuralMesh 6 running with NVMe delivered 10 times higher token throughput and served 10 times more concurrent users on Oracle Cloud Infrastructure compared to standard DRAM. Those are WEKA's own reported benchmarks. No independent lab or third-party verification of those figures appears in any of the available reporting.

On the hardware side, WEKApod 3 claims to be dense. According to a WEKA press release distributed via PR Newswire, a single WEKApod rack delivers 1.1 exabytes of effective capacity, built on 441.5 petabytes of raw capacity, which the company says makes it the first single-rack system to break the exabyte barrier. WEKA also claims 10.2 terabytes per second of throughput and 210 million IOPS per rack, along with 267% higher effective capacity density and 114% higher performance density than what the company calls "market alternatives." Those comparison figures come directly from WEKA's own press release and were not independently benchmarked against named competitors in the available sourcing.

New capabilities beyond the memory trick

NeuralMesh 6 also adds multi-tenancy features WEKA says were costing it deals. According to VentureBeat, the platform now supports composable clusters with hardware-level isolation for anchor tenants, plus virtual multi-tenancy that scales past 1,000 tenants per cluster with provisioning under 30 minutes. WEKA says a single cluster running 50 composable clusters can support up to 50,000 tenants total.

The software also merges file-based and object-based storage paths into one system. Normally, AI infrastructure keeps these separate: a file path used heavily in training pipelines, and an S3-style object path that inference and cloud-native tools expect, with a gateway translating between them. WEKA claims the same physical data on disk is now directly readable through either path without duplication, according to VentureBeat.

WEKA's chief product officer, Ajay Singh, told SiliconANGLE that most data centers running production AI today were built chaotically, mixing hardware, chips, and networking from multiple vendors as availability allowed, rather than through deliberate architecture. That's the gap WEKA says it's now targeting.

Crowded field, real stakes

WEKA isn't alone here. Dell, NetApp, Pure Storage and VAST have all repositioned toward AI infrastructure over the past two years, according to VentureBeat, and WEKA is one of several vendors arguing it was built for this specific inference-era moment rather than retrofitted for it.

The company has real backing behind the claims. WEKA has raised $415 million in total funding, according to reporting from Ecosistema Startup, including a $140 million Series E in May 2024 that valued the company at $1.6 billion. Investors include Nvidia, Qualcomm Ventures, Hitachi Ventures, Generation Investment Management and Valor Equity Partners. WEKA reports annual recurring revenue above $100 million in 2025, roughly doubling year over year, with more than 300 customers including Stability AI, Midjourney and ElevenLabs, plus 12 Fortune 50 companies, according to Ecosistema Startup.

This is a vendor's own press release and vendor-briefed coverage, packed with vendor-generated benchmarks against unnamed "market alternatives." No customer has yet published independent, third-party performance results confirming the 10x throughput claim or the density comparisons in production. Whether NeuralMesh 6 actually delivers those numbers at scale, under real multi-tenant load with paying enterprise customers, remains to be seen as WEKApod 3 starts shipping to the AI clouds and enterprises the company is targeting.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatStop adding more GPUs: Weka's new storage platform reduces load by caching 100% of an AI model's pre-calculated tokens
unknown
prnewswireWEKA Unveils WEKApod 3: The World's Densest AI Storage and Memory System for Agentic Workloads - PR Newswire
unknown
siliconangleWekaIO revamps its AI data storage platform and unveils its first hardware for agentic workloads - SiliconANGLE
unknown
ecosistemastartupWeka NeuralMesh 6: caching de tokens reduce GPUs en IA - Ecosistema Startup