Google Ships Three New Gemini Models for AI Agents — And Quietly Skips the One Everyone Was Waiting For.
Gemini 3.6 Flash, 3.5 Flash-Lite, and a restricted 3.5 Flash Cyber all target agent economics — but Gemini 3.5 Pro is still nowhere to be found.
Google Just Released 3 New Gemini Models — But Skipped the One Everyone Wanted.
3: New Gemini models released this week
17%: Fewer output tokens vs. the prior Flash model
83.0%: OSWorld-Verified computer-use score
Google DeepMind just shipped three new Gemini models built almost entirely around one problem: making AI agents cheap enough to run thousands of times an hour. What it didn't ship was the flagship reasoning model the industry has been waiting on since February.
On Tuesday, Google released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a restricted Gemini 3.5 Flash Cyber variant. Google frames all three around efficiency, latency, and reliability for customers building AI agents at scale — but the release is just as notable for what's missing: Gemini 3.5 Pro, Google's flagship reasoning and coding model, still hasn't shipped, even after Google teased it in May as something it hoped to roll out "next month."
1: The Economics Driving Google's Flash Strategy:
Running an autonomous agent in production comes down to a cost equation vendors rarely spell out: a model has to reason competently through a multi-step task, but every extra token it generates along the way adds cost and delay to a workflow that might run thousands of times an hour. Teams building background agents need throughput first and raw parameter count second — which is exactly the trade-off Google is now splitting across three purpose-built models instead of one general-purpose flagship.
Gemini 3.6 Flash is positioned as the workhorse for coding and multimodal reasoning, cutting output tokens by 17% against its 3.5 Flash predecessor — and by up to 65% on specific tests like the Datacurve DeepSWE benchmark. It's priced at $1.50 per million input tokens and $7.50 per million output tokens, and Google's own benchmarks show real gains: a 49% success rate on DeepSWE versus 37% for the older model, and a jump from 49.7% to 63.9% on MLE Bench.
2: Real Deployments — Figma, Harvey, and Hebbia:
Figma has already integrated 3.6 Flash into its prototyping infrastructure; the company's Director of Product Engineering, Matt Colyer, says it gives developers a faster path through design iterations without sacrificing output quality. Legal tech platform Harvey and research tool Hebbia are routing multimodal document work through the model — ingesting raw financial filings, parsing structure, reading embedded charts, and drafting reports for human review.
Google also built a client-side computer-use tool directly into the Gemini API and Gemini Enterprise platforms, eliminating the custom intermediary software engineering teams previously had to build to let models operate on top of an operating system. The company reports its OSWorld-Verified computer-use score climbed to 83.0%, up from 78.4%, alongside updated safeguards it says improve jailbreak resistance without raising refusal rates on legitimate requests.
"Land soon."

The Hidden AI War
Nobody Is Telling You About
Our latest documentary deep-dive into the geopolitical struggle for machine intelligence dominance. Explore the two paths of AI development: open source vs. closed architecture.
— Logan Kilpatrick, Google DeepMind product lead, on the still-unreleased Gemini 3.5 Pro.
3: The Cheap Tier — and the One You Probably Can't Access:
Gemini 3.5 Flash-Lite is aimed at a different job entirely: high-volume document processing and agentic search rather than deep reasoning. The Artificial Analysis Index clocked it at 350 output tokens per second, the fastest in the 3.5 series, at $0.30 per million input tokens and $2.50 per million output tokens — cheap enough that engineering teams can route simple, high-volume subagent requests to a minimal thinking level and save deeper reasoning for multi-step work. On Google's long-context GDM-MRCR v2 test, it improved to a 72.2% success rate from 60.1%, and its GDPval-AA v2 score nearly doubled, from 642 to 1140.
Gemini 3.5 Flash Cyber sits apart from the other two. Fine-tuned specifically to find and remediate code vulnerabilities, it's restricted to governments and vetted partners through a limited pilot — a deliberate constraint, Google says, against the model being used to generate exploit code. Inside Google's own CodeMender security agent, multiple instances of Flash Cyber run in parallel, cross-checking each other's findings before producing a single remediation report for human sign-off.
4: What the Missing Pro Model Signals:
Gemini Pro models are Google's highest-capability offering for complex reasoning and coding — and the line hasn't been updated since February, even as OpenAI shipped GPT-5.5 and began rolling out GPT-5.6, and Anthropic launched Claude Opus 4.8, Claude Sonnet 5, and expanded access to Fable 5. Bloomberg reported last week that Google has faced internal delays getting 3.5 Pro to meet its own performance goals.
Kilpatrick says the model is currently in partner testing, and that Google has already started pre-training for Gemini 4.
For enterprises, the takeaway isn't which lab ships the biggest flagship next. It's that the models actually shipping right now are built for a very specific job — running agents cheaply and reliably at scale — and that's the capability worth evaluating today, regardless of which vendor's flagship reasoning model lands next.
Build Agents on Infrastructure That's Already Here.

Why Google’s New Free Gemini Feature Is a Massive Wake-Up Call for Businesses
Every major lab is racing to make agent-scale AI cheaper and faster — but stitching those models into a reliable business workflow still takes real engineering.
Otherworlds AI's Agent+ platform handles that integration for you, and our team also builds custom enterprise AI solutions tailored to your exact workflows and cost targets — no waiting on the next flagship release required.
Learn more at otherworldsai.com
Support our research
Independent analysis fueled by you.







