Cheaper usually means dumber. Gemini 3.6 Flash is cheaper and scores higher than the model it replaces, and the reason is simple: it uses fewer tokens and each token costs less.

I already wrote up the launch news. This post is the explainer: what each number means in plain English, and where the model still loses.

The cost of an agent

Think of an agent's bill as tokens used times price per token. Gemini 3.6 Flash cuts both.

  • Fewer tokens. On DeepSWE v1.1, one long multi step coding job, the average output per task fell from 276,000 tokens on 3.5 Flash to 97,000. That is about 65% less. On the Artificial Analysis Intelligence Index it uses 17% fewer output tokens.
  • Cheaper tokens. Output price fell from $9.00 to $7.50 per million. Input stays at $1.50.

For context from Google's own table: GPT 5.6 Luna and Grok 4.5 list at $6 per million output tokens, Gemini 3.1 Pro at $12, and Claude Sonnet 5 at $15.

What the benchmarks measure

  • DeepSWE v1.1 (49%, up from 37%): can it finish one long coding job from start to finish without losing track.
  • MLE-Bench (63.9%, up from 49.7%): can it take a data problem and build and train a working model. In other words, can the AI build AI.
  • GDPval-AA v2 (1,421, up from 1,349): professional work like reports and analysis. This is an Elo rating like chess, so there is no ceiling of 100.
  • OSWorld-Verified (83.0%, up from 78.4%): can it use a normal desktop, clicking, typing and opening apps. Here it beats Claude Sonnet 5 (81.2%) and GPT 5.6 Luna (72.6%).

Where it is last

Against the other frontier models in Google's table, 3.6 Flash scores lowest on SWE-Bench Pro (58.7%, real bug fixes and features in real code bases) and on Terminal-Bench 2.1 (78.0%, getting work done by typing terminal commands).

It wins on reading. Chart reasoning on CharXiv hits 85.2% without tools and 89.4% with code tools. On the needle in a haystack memory test it scores 91.8% at 128K context and 54.0% at 1 million, the best in the table. A drop as the haystack grows is normal.

So the pitch is not that it wins every benchmark. It is that it is good enough on most of them and much cheaper to run.

The other two models

Gemini 3.5 Flash-Lite is the cheapest tier: 350 output tokens per second, 30 cents per million input and $2.50 per million output. It beats the old mid tier 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld (74.0% vs 65.1%).

Gemini 3.5 Flash Cyber is a small model tuned to find and fix security flaws. It scores 83.2% on CyberGym, within 2.4 points of the best frontier model, but only governments and trusted partners get it, through a limited pilot.

Specs and what comes next

Gemini 3.6 Flash has a 1 million token input window and 64,000 output tokens. It reads text, images, video, audio and PDFs, and writes text. Built in tools include function calling, code execution, Search and Maps grounding, computer use, structured outputs and the Batch API.

Google also says pre-training for Gemini 4 has started. The full benchmark tables are on Google's announcement and the Gemini API pricing page.

If you want to learn how to build with models like this, that is what my AI Academy community is for.