Google dropped three new models today: Gemini 3.6 Flash, 3.5 Flash-Lite, and a security model called 3.5 Flash Cyber. The whole release is aimed at one thing, making AI agents as cheap as possible to run.

Cheaper two ways at once

Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index. On the DeepSWE coding benchmark the drop is dramatic: from 276,000 tokens per task down to 97,000. On top of that, output pricing fell from $9.00 to $7.50 per million tokens.

Fewer tokens times cheaper tokens compounds. If you run agents that work all day, this is the headline.

Better, not just cheaper

DeepSWE jumped from 37% to 49%. ML engineering went from 49.7 to 63.9 on MLE-Bench. And on OSWorld-Verified computer use, 3.6 Flash scores 83.0%, ahead of Claude Sonnet 5 at 81.2% and GPT-5.6 Luna at 72.6% on Google's own comparison table. Computer use now ships as a built-in tool in the Gemini API.

The budget tier beats last gen's mid tier

Gemini 3.5 Flash-Lite runs at 350 tokens per second and costs 30 cents per million input tokens. It outscores last generation's 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and on OSWorld (74.0% vs 65.1%). The cheap tier now beats the old mid tier.

The two buried stories

Flash Cyber is a small security model fine-tuned to find and patch vulnerabilities. It hits 83.2% on CyberGym, within 2.4 points of the best frontier agent, but Google is only releasing it to governments and trusted partners through a limited pilot.

And one sentence near the bottom of the post: Google says its most ambitious pre-training run ever, for Gemini 4, has already started.

Try it today

Gemini 3.6 Flash and 3.5 Flash-Lite are live now in Google AI Studio and the Gemini app. If you want the full breakdown with all the benchmark charts, the video covers everything in 90 seconds.