Notice
先看证据,再决定买不买

频道每天最多 3 条价格异动与中转状态;具体商品请用机器人设置降价/补货提醒。交流群提问请带预算、模型、工具和使用频率。

View
Community & contactTelegram 群点击加入Telegram 频道每天最多 3 条有效价格情报联系我们tgAIPricedb交流群979789483
Back to news
Products

Google brings agentic video understanding to Gemini, cutting token use by up to 88%

Google has introduced agentic video understanding for its latest Gemini models. The capability dynamically searches long-form video and, in company-reported benchmarks, used up to 88% fewer tokens while improving analysis accuracy.

96% VERIFIED

The feature is available for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Instead of processing video at a fixed frame rate, the models can combine reasoning with native video tools to decide which sections to inspect and whether visual content, audio, or transcripts are most useful.

Google says testing showed up to 88% lower token consumption, 66% lower analysis costs, and a 7% accuracy improvement. Developers can use the capability for uploaded and YouTube videos through the Gemini API in Google AI Studio and through the Gemini Enterprise Agent Platform. The figures are company-reported maximums, not guarantees for every video task.

Source evidence

Introducing agentic video understanding with Geminiblog.google · supporting

Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more. The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. ## Benchmarks [...] ## Benchmarks Unlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model’s core reasoning with native video tools to dynamically search,

Google Adds Agentic Video Understanding to Gemini, Claiming 88% Lower Token Use | Superpower Dailysuperpowerdaily.com · supporting

Google is giving Gemini a new way to watch video: instead of sampling frames at a fixed rate, it can decide which moments to inspect, and whether frames, audio, or transcripts contain the answer. The feature, called agentic video understanding, is available now through the Gemini API for uploaded and YouTube videos. Google says it cut token use by as much as 88 percent, reduced analysis cost by up to 66 percent, and improved accuracy by up to 7 percent on standard video-analysis benchmarks. Those are company-reported results, so they indicate tested performance—not a guarantee for every task. The supported models are Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Developers turn the setting on in Google AI Studio or the Gemini Enterprise Agent Platform, with standard Gemini API pricing [...] Google has launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. The feature lets the models search selected parts of a video, using images, audio and transcrip

We're launching agentic video understanding across our latest modelsfacebook.com · supporting

Available via Gemini Enterprise Agent Platform, this new capability improves accuracy while cutting token consumption up to 88% and reducing

Google DeepMind's Post - LinkedInlinkedin.com · supporting

We're bringing agentic video understanding to our latest Gemini models. They can now analyze videos with better accuracy while using up to 88%

Google on X: "With agentic video understanding, developers can dynamically search, scan, and inspect long-form video with up to: 📉 88% fewer tokens 📉 66% lower cost ✅ 7% better accuracy" / Xx.com · supporting

Log inSign up ## Post # Google on X: "With agentic video understanding, developers can dynamically search, scan, and inspect long-form video with up to: 📉 88% fewer tokens 📉 66% lower cost ✅ 7% better accuracy" @Google Google @Google With agentic video understanding, developers can dynamically search, scan, and inspect long-form video with up to: 📉 88% fewer tokens 📉 66% lower cost ✅ 7% better accuracy Scatter plot titled 5:30 PM · Sep 1, 202641KViews 6 @Google Google @Google Sep 1 We’re introducing a new capability to our latest Gemini models: agentic video understanding. This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens. See how it works 🧵 00:00 199 @Google Google @Google Sep 1 [...] 6 @Google Google @Google Sep 1 Agentic video understanding is available now across our latest models via the Gemini API in @GoogleAIStudio and the Gemini Enterprise Agent Platform.

Introducing agentic video understanding with Gemini | Google AI Studio (@GoogleAIStudio) on Xx.com · supporting

Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more. The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Benchmarks [...] Benchmarks Unlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model’s core reasoning with native video tools to dynamically search, scan,