Gemini 3.6 Flash 基于用户反馈推出,输出 token 使用量减少 17%
Google DeepMind 发布 Gemini 3.6 Flash,称其在编码、知识工作和多模态任务上优于 3.5 Flash,并根据 Artificial Analysis Index 减少 17% 的输出 token 使用量。
Google DeepMind 表示,Gemini 3.6 Flash 直接吸收了开发者和客户对 3.5 Flash 的反馈,定位为面向编码、知识工作及多模态任务的主力模型。官方称,该模型在保持 Flash 级响应速度和规模的同时,提升了复杂推理与长流程编程能力。
在 Artificial Analysis Index 的对比中,3.6 Flash 的输出 token 使用量比 3.5 Flash 少 17%,并且在多步骤工作流中需要更少的推理步骤和工具调用。部分公开测试还显示,其 DeepSWE 编码得分从 37% 提升至 49%,OSWorld 得分从 78.4% 提升至 83%;这些结果取决于具体测试设置。
官方给出的价格为每百万输入 token 1.50 美元、每百万输出 token 7.50 美元。Google 称,更高的 token 效率和更低的价格可降低智能体任务的整体运行成本。
来源证据
Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4 - Ars Technicaarstechnica.com · supporting[ Credit: Google Credit: Google]( In the DeepSWE test for coding, 3.6 Flash jumps to 49 percent versus 37 percent for 3.5 Flash. The new model now supports computer use as a standard feature in the Gemini API, too. The OSWorld test for computer use shows a modest boost to 83 percent from 3.5’s 78.4 percent score. Efficiency was a big focus for Gemini 3.5 Flash, and Google says that effort has been amped up with 3.6. Even with small benchmark gains, Gemini 3.6 Flash uses about 17 percent fewer tokens. [...] Gemini 3.5 Flash, which was the star of the show at I/O, has already been deprecated. In its place, developers and users will find Gemini 3.6 Flash. Google makes the usual claims about this model—it’s marginally more capable and better at coding, and it has great multimodal features. Google says the changes to 3.6 Flash were made in response to user feedback on the 3.5 release. In general, Gemini 3.5 Flash didn’t appear to live up to Google’s promises around code generation. Perh
3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyberblog.google · supportingGemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows. This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run. 3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API) [...] 3.6 Flash, using Managed Agents on AIS, can help parse through and analyze financial data and transcripts more efficiently and accurately than 3.5 Flash (AIS) 3.6 Flash executes code migrations, using multi-agent orchestratio
Gemini 3.6 Flashdeepmind.google · supportingGoogle DeepMind Build with Gemini Try Gemini # Gemini 3.6 Flash Best for token efficiency in coding, knowledge work, and multimodal tasks Try in GeminiBuild with Gemini Our workhorse model that reduces output token usage by 17% compared to 3.5 Flash, according to Artificial Analysis Index. Capabilities Hands-on Showcase Performance Model information ## Intelligence in a Flash Get advanced reasoning at Flash-level latency and scale. Slide 1 of 4 ### Token Efficiency Get better quality in coding, knowledge work, and multimodal tasks whilst reducing token usage. ### Fast and smart Speed and scale don’t have to come at the cost of intelligence. ### Master complexity Deep reasoning across long horizons and iterative coding tasks. ### Truly multimodal [...] Nick Frolov, Head of Product, Junie, JetBrains ## Performance 3.6 Flash is more token efficient than 3.5 Flash and a step-up in coding and knowledge work. [...] ### Truly multimodal Multimodal understanding acro
Gemini 3.6 Flash Pricing: The Real Cost Drop Is Bigger Than ...trilogyai.substack.com · supportingCost per task = (price per token) × (tokens consumed per task) × (attempts per success) Google moved the first variable 17%. But its published efficiency numbers move the second variable at least as much: 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, up to 65% fewer on the DeepSWE coding benchmark, and takes fewer reasoning steps and tool calls per multi-step workflow. Notably, the token reduction did not come at the expense of quality, the DeepSWE score itself rose from 37% to 49%. Fewer tokens and better answers is the combination that makes the per-task math compound rather than trade off. Compound those and the arithmetic looks like this: Sticker comparison only — 0.83× price, no efficiency gain → ~17% cheaper
Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on ...venturebeat.com · supportingAs measured by Artificial Analysis, the model processes 350 output tokens per second, making it highly effective for agentic search and massive document processing workloads. Artificial Analysis notes this is about twice as fast as prior generation model Gemini 3.1 Flash-Lite. Developers can configure 3.5 Flash-Lite to prioritize low-latency execution for high-volume tasks using minimal thinking levels, or engage higher thinking levels to process complex multi-step subagent workloads. Despite its lite designation, it outperforms the standard Gemini 3 Flash on several key agentic and coding evaluations, including SWE-Bench Pro, where it scores 54.2% compared to 49.6%, and OSWorld-Verified, scoring 74.0% versus 65.1%. [...] Gemini 3.6 Flash serves as the heavy-duty workhorse of the trio. It handles complex coding, intricate knowledge work, and multimodal processing with improved precision. Enterprise customers utilize it for demanding tasks such as complex document parsing, intricate c
Googles New 3 Gemini Models Are Incredible - Gemini 3.6 Flash And More Gemini 4 News)youtube.com · supportingdive into every single model so you can really understand what's going on. So the first model here is called Gemini 3.6 flash. This is essentially is the workhorse model that delivers better coding, knowledge work and multimodal performance and according to the artificial analysis index reduces output token usage by 17% compared to 3.5 flash. And on some benchmarks like deepsw SWE by data curve, Google observes up to 65% at all lower cost per hour per token. And so essentially what you have here with Gemini 3.6 6 Flash is a model that is, I guess you could say, more efficient and better quality than the previous version of 3.5 Flash. So, Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. And it not only delivers a step up in coding and knowledge work, but [...] is a huge improvement on there. And of course on the GDP valve, you can see that there is a huge improvement there. So this model is built to scale agentic workflows. So this is designed for both