Google Says Newest Gemini Models Use Fewer Output Tokens
Gemini 3.6 Flash uses fewer output tokens and carries a lower generation price as Google targets the rising cost of running AI agents.
Topics
Image Credit- Chetan Jha/ MIT Sloan Management Review India
Google has released Gemini 3.6 Flash, its latest model that uses fewer output tokens than its predecessor, targeting one of the main cost pressures facing companies running AI agents at scale.
The model consumed 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, Google said on Tuesday, July 21.
The company said the reduction came alongside fewer reasoning steps and tool calls when completing multi-stage tasks.
AI providers typically charge developers for the number of tokens a model receives and generates. Longer responses, repeated reasoning and multiple tool calls can raise both costs and response times in agentic systems.
Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens.
The input price is unchanged from Gemini 3.5 Flash, while the output price has fallen from $9.
Google said the combination of lower token use and cheaper output generation would reduce the overall cost of completing agentic tasks.
The company did not release customer data showing how much businesses could save in live deployments. Token use varies according to prompts, software tools, data sources and task complexity.
Google also reported improved scores for the model on several coding, computer-use and knowledge-work benchmarks.
Gemini 3.6 Flash scored 49% on the DeepSWE coding benchmark, compared with 37% for Gemini 3.5 Flash, according to company figures.
Benchmark results may not reflect performance or costs in individual business applications, where token consumption can vary depending on prompts, software tools, data sources and the complexity of a task.
Google separately released Gemini 3.5 Flash-Lite for high-volume workloads such as search and document processing. The model is priced at 30 cents per million input tokens and $2.50 per million output tokens. It generates about 350 output tokens per second, based on Artificial Analysis data cited by Google.
The company also introduced Gemini 3.5 Flash Cyber, a specialized model designed to identify and repair software vulnerabilities.
It will initially be available only to governments and selected partners through Google’s CodeMender security system because of concerns that cyber capabilities could be misused.
Gemini 3.6 Flash and 3.5 Flash-Lite are available through Google’s developer and enterprise platforms. Google said Gemini 3.5 Pro remained under testing, while training had begun on Gemini 4.

