On July 21, 2026, Google announced three new Gemini models at once: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The headline model, 3.6 Flash, is priced at 1.50 dollars per million input tokens and 7.50 dollars per million output tokens, down from the 9 dollars per million output that the previous Flash generation charged.
Google’s central claim for 3.6 Flash is token efficiency rather than raw capability: it uses 17 percent fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65 percent fewer on the DeepSWE benchmark measured by Datacurve. Reported scores include DeepSWE at 49 percent versus 37 percent, MLE Bench at 63.9 versus 49.7, OSWorld-Verified at 83.0 versus 78.4, and GDPval-AA v2 at 1421 versus 1349. The cheaper 3.5 Flash-Lite runs at 0.30 dollars in and 2.50 dollars out per million tokens, sustains 350 output tokens per second, and posts large jumps on Terminal-Bench 2.1 (54 percent versus 31 percent) and GDPval-AA v2 (1140 versus 642).
The third model is narrower. Gemini 3.5 Flash Cyber is fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models, and Google says it reaches competitive frontier performance on the CyberGym benchmark. It is not generally available: Google is releasing it exclusively to governments and trusted partners through its CodeMender tool as part of a limited-access pilot, with multiple Flash Cyber agents working together to produce combined vulnerability reports.
Flash and Flash-Lite went live the same day across Google AI Studio, Android Studio, the Gemini Enterprise Agent Platform, the Gemini Enterprise app, and the Gemini consumer app. The business signal is that Google is now competing on cost per completed task, not cost per token: a model that needs 17 percent fewer output tokens and fewer tool calls to finish an agentic job is cheaper twice over.