READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Google Ships Three New Gemini Models, Still No Sign of the Flagship 3.5 Pro

Google Ships Three New Gemini Models, Still No Sign of the Flagship 3.5 Pro
Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a cybersecurity-focused 3.5 Flash Cyber model on Tuesday, all cheaper and more token-efficient than their predecessors. The actual flagship everyone's waiting on, Gemini 3.5 Pro, is still stuck in partner testing with no release date. Translation: Google is optimizing costs while the AI arms race with OpenAI and Anthropic keeps grinding forward.

Google put out three new Gemini models on Tuesday, July 21. None of them is the model people actually asked about.

Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber all shipped today, according to a Google blog post from Tulsee Doshi, senior director of product management for the Gemini team. What didn't ship: Gemini 3.5 Pro, the flagship model Google originally teased for a June launch. CNET reported Google says 3.5 Pro is "currently testing with partners" with no firm release date attached.

The New Workhorse

Gemini 3.6 Flash replaces Gemini 3.5 Flash, which barely had time to settle in since its debut at Google I/O in May. According to Ars Technica, 3.5 Flash never quite lived up to Google's coding promises, and 3.6 Flash is Google's fix.

The numbers back that up. On the DeepSWE coding benchmark, 3.6 Flash scores 49% versus 37% for its predecessor, according to MarkTechPost. On OSWorld, a computer-use test, it hits 83.0% versus 78.4%. On MLE Bench it reaches 63.9% versus 49.7%.

The bigger story is cost. Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, according to Google's own blog post, and MarkTechPost reported the DeepSWE benchmark specifically showed up to a 65% token reduction. Pricing dropped too: $1.50 per million input tokens and $7.50 per million output tokens, down from $9.00 on the output side for 3.5 Flash.

Businesses running AI agents pay per token every time the model reasons through a task. Fewer tokens per task means real savings at scale.

Cheap and Fast: Flash-Lite

Gemini 3.5 Flash-Lite is Google's budget option, and it's aggressively priced at $0.30 per million input tokens and $2.50 per million output tokens. VentureBeat's pricing comparison shows that undercuts most competitors except Chinese models from DeepSeek, Xiaomi, and Alibaba, which are cheaper across the board.

Flash-Lite runs at 350 output tokens per second, according to CNET, and MarkTechPost reported it beats the older 3.1 Flash-Lite by wide margins: 54% versus 31% on Terminal-Bench 2.1, and 72.2% versus 60.1% on the long-context GDM-MRCR v2 benchmark. It even beats the pricier standard 3 Flash model on some tests.

One wrinkle: Google's own prior-generation 3.1 Flash-Lite is still technically cheaper at $0.25/$1.50 per million tokens, according to VentureBeat, but it's roughly twice as slow. Google is betting most developers will pay the small premium for speed.

A Model Built to Hack (For the Good Guys)

Gemini 3.5 Flash Cyber is the one nobody else is shipping. It's Google's first LLM tuned specifically for cybersecurity work, built to find, validate, and patch software vulnerabilities, according to MarkTechPost.

CNET reported the model already works alongside Google's CodeMender infrastructure agent and is finding and fixing bugs in Android, Chrome, and YouTube's internal codebases. For now, access is limited to governments and "trusted partners," per Google DeepMind material cited by CNET, with wider availability promised later. No pricing has been announced for it.

Given that China remains the dominant cyber threat actor targeting American infrastructure and tech companies, according to years of U.S. government assessments, a model purpose-built to patch vulnerabilities before adversaries exploit them is a genuinely useful defensive tool, not just a product-line curiosity.

Where's Gemini 3.5 Pro?

The flagship 3.5 Pro model was supposed to launch in June, according to Ars Technica, and it still hasn't shipped. Google's blog post says it's "currently testing with partners" and will go broadly available "as soon as it's ready" — corporate-speak for no committed date.

Meanwhile, Google says it has already begun pretraining for Gemini 4, according to both the company's own blog and CNET's reporting. Google is investing in the model after next while the model everyone's actually waiting on sits in partner testing.

The AI market isn't waiting around either. OpenAI shipped ChatGPT-5.6 earlier this month, per CNET, and Anthropic, xAI, and a growing field of Chinese competitors including DeepSeek and Alibaba are all releasing competitively priced models on a similar pace, per VentureBeat's own pricing table.

Google's pitch today is efficiency and cost, not raw capability. That's a real, provable improvement for developers watching their token bills. Whether it's enough to hold market share while the actual flagship model stays in limbo is the open question Google didn't answer on Tuesday.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatGoogle's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks —and 3.5 Pro is on the way
center-left
Ars TechnicaGoogle announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4
unknown
blog.googleIntroducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber - Google Blog
unknown
marktechpostGoogle Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads - MarkTechPost
unknown
cnetGoogle Releases 3 New Gemini Models, 3.5 Pro Still Not Available - CNET