Google 推出 Gemini 3.5 Flash-Lite,面向高吞吐任务优化
Gemini 3.5 Flash-Lite 是 Google Gemini 3.5 系列中的低延迟、低成本多模态模型,适用于工单分类、数据提取和其他重复性高、调用量大的任务。Google 表示,它在高负载场景下的延迟低于 Gemini 3.5 Flash。
Google DeepMind 将 Gemini 3.5 Flash-Lite 定位为面向规模化生产流量的轻量模型,支持文本、图片、音频、视频和 PDF 输入。其主要应用包括分类、文档解析、翻译、摘要、结构化信息提取,以及代理系统中的子任务处理。
官方资料显示,该模型的 API 价格为每百万输入 token 0.30 美元、每百万输出 token 2.50 美元。开发者可以在低延迟和低成本模式下处理大批量请求,也可以提高思考级别以应对更复杂的多步骤任务。Google 同时展示了它与 Gemini 3.5 Flash 的延迟对比,但具体收益仍会因任务类型、输入规模和部署方式而变化。
来源证据
Gemini 3.5 Flash-Lite | Gemini API - Google AI for Developersai.google.dev · supportingGemini API Gemini API # Gemini 3.5 Flash-Lite Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints. ## Documentation Visit the Latest model page for full coverage of features and capabilities. ## gemini-3.5-flash-lite
Gemini 3.5 Flash-Lite on real enterprise workblog.box.com · supporting## Faster and lighter, too The efficiency result is the surprising part. Even against the prior "lite" model, Gemini 3.5 Flash Lite is the leaner one on this workload: it reaches its answers in roughly a quarter of the time (about 27 seconds per task on average vs. ~99), using about half the tokens and a third of the model and tool calls. Rather than trading quality for speed, the earlier model appears to work harder — more retrieval and reasoning turns — and still arrives at less accurate answers. For enterprise workloads at scale, Gemini 3.5 Flash Lite offers both higher accuracy and a lighter operational footprint. ## Get started [...] Sorting by type of work rather than industry sharpens the picture: Gemini 3.5 Flash Lite leads across all four modes of analytical work, and its edge grows with difficulty. It's strongest on report drafting (64% vs 46%) and expert review (62% vs 51%), and its widest margin comes on data analysis — turning imperfect source data into a number you can
3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyberblog.google · supporting3.5 Flash-Lite is the fastest model in the 3.5 series. As measured by Artificial Analysis, it runs at 350 output tokens/s. Priced at $0.3/1M input tokens and $2.5/1M output tokens and with significantly better quality than 3.1 Flash-Lite, 3.5 Flash-Lite offers a strong price-to-performance ratio for developers and customers running high throughput production traffic. 3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash. [...] 3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with the minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads. The model now also has computer use as a built-in tool to reliably support these agentic tasks across surfaces. It’s a significant step up in coding and agentic
Gemini 3.5 Flash-Lite — Google DeepMinddeepmind.google · supporting### Cost-efficient A stronger price-to-performance ratio compared to 3.1 Flash-Lite. ## Hands-on Explore what you can do with Gemini 3.5 Flash-Lite Slide 1 of 4 ### 3.5 Flash-Lite vs. 3.5 Flash latency SxS (API) 3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash ### One prompt, 25 web design options (API) Working alongside 3.6 Flash as the master agent, 3.5 Flash-Lite instantly generates 25 unique, ready-to-explore web design concepts ### Receipt analysis and translation at scale (API) 3.5 Flash-Lite can scale receipt translation and summarization with its multimodal understanding ### Rapid game development (API) 3.5 Flash-Lite builds a game by instantly generating and iterating through multiple options ## Showcase Slide 1 of 3 [...] Ashwin Kannan, Principal AI Engineer, Palo Alto Networks “Gemini 3.5 Flash-Lite landed on the Pareto frontier in our receipt extraction benchmark, offering one of the best tradeoffs we’ve tested between accuracy, lat
Gemini 3.5 Flash-Lite - Model Carddeepmind.google · supportingGemini 3.5 Flash-Lite Model card Gemini 3.5 Flash-Lite — Model Card Model Cards are intended to provide essential information on Gemini models, including known limitations, mitigation approaches, and safety performance. Model cards may be updated from time to time; for example, to include updated evaluations as the model is improved or revised. See the Google DeepMind site for a comprehensive list of model cards. Published: July 2026 Model Information Description Gemini 3.5 Flash-Lite is an addition to the Gemini 3 series of highly-capable, natively multimodal, reasoning models. The model is cost-efficient and fast, optimized for high-volume, latency-sensitive tasks like translation and classification as well as supporting agentic workflows. Model dependencies Gemini 3.5 Flash-Lite is [...] workflows. Model dependencies Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite. Inputs Text strings (e.g., a question, a prompt, document(s) to be summarized), images, audio, and video files,
Gemini 3.5 Flash vs Gemini 3.1 Pro: Is the Flash Model Good Enough?mindstudio.ai · supporting### Gemini 3.5 Flash Flash models in Google’s Gemini family are built for throughput. They run faster, cost less per token, and are optimized for tasks where you need a high volume of responses — summarization, classification, structured extraction, customer-facing chat, and similar workloads. Gemini 3.5 Flash continues this tradition with a key upgrade: it generates roughly 2x more output tokens per second than Gemini 3.1 Pro while maintaining competitive accuracy on a wide range of tasks. That speed advantage is significant in production environments where response latency directly affects user experience. Key specs: [...] ## Key Takeaways Gemini 3.5 Flash generates 2x more tokens per second than Pro and costs 4–8x less — for high-volume workloads, that math is hard to ignore. Flash is competitive with Pro on coding, summarization, classification, instruction-following, and most structured tasks. Pro maintains a real edge in multi-step reasoning, long-document coherence, ambig