Notice
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

Community & contactTelegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
Back to news
Products

Gemini 3.6 Flash builds on feedback and uses 17% fewer output tokens

Google DeepMind says Gemini 3.6 Flash improves coding, knowledge work, and multimodal performance while using 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index.

96% VERIFIED

Google DeepMind has introduced Gemini 3.6 Flash as a workhorse model shaped by developer and customer feedback on Gemini 3.5 Flash. The company says it improves coding, knowledge work, and multimodal performance while retaining Flash-level speed and scale.

According to the Artificial Analysis Index, Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor and requires fewer reasoning steps and tool calls for multi-step workflows. Reported evaluation results also include a DeepSWE coding score increase from 37% to 49% and an OSWorld score increase from 78.4% to 83%, although results depend on the test configuration.

Google lists the model at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. It says the combination of lower pricing and improved token efficiency can reduce the total cost of running agentic tasks.

Source evidence

Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4 - Ars Technicaarstechnica.com · supporting

[ Credit: Google Credit: Google]( In the DeepSWE test for coding, 3.6 Flash jumps to 49 percent versus 37 percent for 3.5 Flash. The new model now supports computer use as a standard feature in the Gemini API, too. The OSWorld test for computer use shows a modest boost to 83 percent from 3.5’s 78.4 percent score. Efficiency was a big focus for Gemini 3.5 Flash, and Google says that effort has been amped up with 3.6. Even with small benchmark gains, Gemini 3.6 Flash uses about 17 percent fewer tokens. [...] Gemini 3.5 Flash, which was the star of the show at I/O, has already been deprecated. In its place, developers and users will find Gemini 3.6 Flash. Google makes the usual claims about this model—it’s marginally more capable and better at coding, and it has great multimodal features. Google says the changes to 3.6 Flash were made in response to user feedback on the 3.5 release. In general, Gemini 3.5 Flash didn’t appear to live up to Google’s promises around code generation. Perh

3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyberblog.google · supporting

Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows. This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run. 3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API) [...] 3.6 Flash, using Managed Agents on AIS, can help parse through and analyze financial data and transcripts more efficiently and accurately than 3.5 Flash (AIS) 3.6 Flash executes code migrations, using multi-agent orchestratio

Gemini 3.6 Flashdeepmind.google · supporting

Google DeepMind Build with Gemini Try Gemini # Gemini 3.6 Flash Best for token efficiency in coding, knowledge work, and multimodal tasks Try in GeminiBuild with Gemini Our workhorse model that reduces output token usage by 17% compared to 3.5 Flash, according to Artificial Analysis Index. Capabilities Hands-on Showcase Performance Model information ## Intelligence in a Flash Get advanced reasoning at Flash-level latency and scale. Slide 1 of 4 ### Token Efficiency Get better quality in coding, knowledge work, and multimodal tasks whilst reducing token usage. ### Fast and smart Speed and scale don’t have to come at the cost of intelligence. ### Master complexity Deep reasoning across long horizons and iterative coding tasks. ### Truly multimodal [...] Nick Frolov, Head of Product, Junie, JetBrains ## Performance 3.6 Flash is more token efficient than 3.5 Flash and a step-up in coding and knowledge work. [...] ### Truly multimodal Multimodal understanding acro

Gemini 3.6 Flash Pricing: The Real Cost Drop Is Bigger Than ...trilogyai.substack.com · supporting

Cost per task = (price per token) × (tokens consumed per task) × (attempts per success) Google moved the first variable 17%. But its published efficiency numbers move the second variable at least as much: 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, up to 65% fewer on the DeepSWE coding benchmark, and takes fewer reasoning steps and tool calls per multi-step workflow. Notably, the token reduction did not come at the expense of quality, the DeepSWE score itself rose from 37% to 49%. Fewer tokens and better answers is the combination that makes the per-task math compound rather than trade off. Compound those and the arithmetic looks like this: Sticker comparison only — 0.83× price, no efficiency gain → ~17% cheaper

Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on ...venturebeat.com · supporting

As measured by Artificial Analysis, the model processes 350 output tokens per second, making it highly effective for agentic search and massive document processing workloads. Artificial Analysis notes this is about twice as fast as prior generation model Gemini 3.1 Flash-Lite. Developers can configure 3.5 Flash-Lite to prioritize low-latency execution for high-volume tasks using minimal thinking levels, or engage higher thinking levels to process complex multi-step subagent workloads. Despite its lite designation, it outperforms the standard Gemini 3 Flash on several key agentic and coding evaluations, including SWE-Bench Pro, where it scores 54.2% compared to 49.6%, and OSWorld-Verified, scoring 74.0% versus 65.1%. [...] Gemini 3.6 Flash serves as the heavy-duty workhorse of the trio. It handles complex coding, intricate knowledge work, and multimodal processing with improved precision. Enterprise customers utilize it for demanding tasks such as complex document parsing, intricate c

Googles New 3 Gemini Models Are Incredible - Gemini 3.6 Flash And More Gemini 4 News)youtube.com · supporting

dive into every single model so you can really understand what's going on. So the first model here is called Gemini 3.6 flash. This is essentially is the workhorse model that delivers better coding, knowledge work and multimodal performance and according to the artificial analysis index reduces output token usage by 17% compared to 3.5 flash. And on some benchmarks like deepsw SWE by data curve, Google observes up to 65% at all lower cost per hour per token. And so essentially what you have here with Gemini 3.6 6 Flash is a model that is, I guess you could say, more efficient and better quality than the previous version of 3.5 Flash. So, Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. And it not only delivers a step up in coding and knowledge work, but [...] is a huge improvement on there. And of course on the GDP valve, you can see that there is a huge improvement there. So this model is built to scale agentic workflows. So this is designed for both