Google Ships Gemini 3.6 Flash With Lower Prices and 17% Fewer Tokens

Google launched Gemini 3.6 Flash alongside two other models today. The new Flash uses fewer tokens, costs less, and improves across coding, agentic, and knowledge benchmarks, all while Gemini 3.5 Pro remains delayed.

Google Ships Gemini 3.6 Flash With Lower Prices and 17% Fewer Tokens

Google dropped three new Gemini models today, and the one that matters most to developers isn’t the delayed flagship everyone’s been waiting for. It’s a cheaper, faster Flash model that ships exactly when Google needs a win.

Gemini 3.6 Flash is live now. It costs $1.50 per million input tokens and $7.50 per million output tokens, $1.50 less on output than the 3.5 Flash it replaces. More importantly, Google says it uses up to 17% fewer output tokens on the Artificial Analysis Index, which means fewer reasoning steps, fewer tool calls, and lower total cost per task in agentic workflows.

The announcement came via Logan Kilpatrick on X, framing the release as a direct response to developer feedback: higher intelligence, better token efficiency, and a lower price.

What 3.6 Flash Actually Improves

Google published benchmark comparisons against 3.5 Flash, and the gains are meaningful across the categories developers care about most.

On coding, 3.6 Flash scored 49% on DeepSWE against 37% for 3.5 Flash. For MLE Bench for machine learning research tasks, it hit 63.9% against 49.7%. Then, OSWorld-Verified for computer use, 83% against 78.4%. And on GDPval-AA v2 for knowledge work, 1421 against 1349.

Those aren’t just number bumps. The DeepSWE jump in particular, a 12-point gain, matters for developers running agentic coding loops where the model needs to navigate repos, write patches, and complete multistep tasks without burning context on dead ends.

Harvey, the legal AI company, said 3.6 Flash completed tasks 12% faster on average than its predecessor. Niko Grupen, Harvey’s head of applied research, said the model showed strong gains on their internal benchmarks and was notably more efficient.

Computer use is now a built-in tool in the Gemini API, not a separate model. I covered that shift when it hit 3.5 Flash, and Google is carrying it forward into 3.6 Flash as a first-class capability.

The Two Other Models

Google also shipped Gemini 3.5 Flash-Lite at 30 cents per million input and $2.50 per million output. It’s positioned as the high-throughput workhorse for simpler tasks inside larger agent systems. On some evals like SWE-Bench Pro and OSWorld-Verified, it beats the original Gemini 3 Flash despite the lower price, and Google is already rolling it into Google Search.

The third model is the most interesting, and the most restricted. Gemini 3.5 Flash Cyber is fine-tuned to find, validate, and patch software vulnerabilities. It powers CodeMender, a new agent that scans repos, builds exploits to prove vulnerabilities are real, and returns patches as code diffs for developer approval.

Google’s DeepMind Big Sleep team tested it on Chrome and Safari codebases. It found 55 unique confirmed issues in the V8 JavaScript engine against 47 for 3.5 Flash and 36 for Claude Opus 4.6, including 10 bugs no other model caught.

But here’s the catch: Flash Cyber is restricted to governments and trusted partners through a limited-access pilot. Citing the dual-use nature of automated vulnerability research, Google isn’t making this one public.

The Strategic Context

This launch drops one day before Alphabet earnings and two months after Google promised Gemini 3.5 Pro for June and then missed that deadline entirely. The Pro model has now slipped three times: first from Google I/O in May, then June, then July 17.

Google said today that 3.5 Pro is testing with partners and that its largest-ever pretraining run is underway for Gemini 4. The statement reads like they’re asking developers to stay patient while they work through whatever is holding Pro back.

Meanwhile, the field isn’t waiting. GPT-5.6 Sol launched July 9. Grok 4.5 launched the same day. Kimi K3 launched July 16 and drew enough demand that Moonshot AI briefly limited new subscriptions. Alibaba is teasing Qwen 3.8 Max.

Against that backdrop, 3.6 Flash is Google playing its strongest hand. It’s not a frontier model, and it’s not trying to be. It’s cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max, according to Google’s own comparison. And the 17% token efficiency gain compounds across every agentic loop, every recursive tool call, every long-running workflow.

Flash isn’t just the cheap tier anymore. It’s where most developers actually run their agents, and Google knows it.

What I’m Watching

The coding gains are real, but they’re Google’s own numbers. I want to see independent benchmarks before I call this a decisive win for agentic workflows. The DeepSWE jump is promising, but 49% still means the model fails more than half the time on real software engineering tasks.

The token efficiency claim, 17% fewer output tokens, matters more for cost than headlines. If 3.6 Flash genuinely completes coding tasks in fewer turns, the effective cost per task drops below the sticker price. That’s the metric that wins budget conversations, not per-token rates.

And the Cyber model being locked behind a government-only access pilot is the right call for now, but it also means Google is sitting on a capability that could genuinely shift how security teams work. I’ll be watching whether the access restrictions ease once the dual-use implications are better understood.

The bottom line: Google shipped a Flash model that’s smarter, cheaper, and more efficient at exactly the moment it needed to remind developers it can still move fast. Now the question is whether 3.5 Pro, or Gemini 4, can deliver the same kind of upgrade at the frontier.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *