Google добавила в Gemini агентное понимание видео, сократив расход токенов до 88%
Google представила агентное понимание видео для новейших моделей Gemini. Функция динамически ищет нужные фрагменты длинных видео и, по данным тестов компании, снижает расход токенов максимум на 88% при повышении точности анализа.
Новая возможность доступна в Gemini 3.7 Flash, 3.6 Flash и 3.5 Flash-Lite. В отличие от обработки с фиксированной частотой кадров, модели могут сочетать рассуждение с нативными видеоинструментами и самостоятельно выбирать, какие участки изучать, а также использовать изображение, звук или расшифровку.
Google сообщает, что в тестах расход токенов снизился максимум на 88%, стоимость анализа — на 66%, а точность выросла на 7%. Разработчики могут обрабатывать загруженные видео и ролики с YouTube через Gemini API в Google AI Studio и через Gemini Enterprise Agent Platform. Это заявленные компанией максимальные результаты, которые не гарантируют такой же эффект для каждой задачи.
Источники
Introducing agentic video understanding with Geminiblog.google · supportingToday, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more. The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. ## Benchmarks [...] ## Benchmarks Unlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model’s core reasoning with native video tools to dynamically search,
Google Adds Agentic Video Understanding to Gemini, Claiming 88% Lower Token Use | Superpower Dailysuperpowerdaily.com · supportingGoogle is giving Gemini a new way to watch video: instead of sampling frames at a fixed rate, it can decide which moments to inspect, and whether frames, audio, or transcripts contain the answer. The feature, called agentic video understanding, is available now through the Gemini API for uploaded and YouTube videos. Google says it cut token use by as much as 88 percent, reduced analysis cost by up to 66 percent, and improved accuracy by up to 7 percent on standard video-analysis benchmarks. Those are company-reported results, so they indicate tested performance—not a guarantee for every task. The supported models are Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Developers turn the setting on in Google AI Studio or the Gemini Enterprise Agent Platform, with standard Gemini API pricing [...] Google has launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. The feature lets the models search selected parts of a video, using images, audio and transcrip
We're launching agentic video understanding across our latest modelsfacebook.com · supportingAvailable via Gemini Enterprise Agent Platform, this new capability improves accuracy while cutting token consumption up to 88% and reducing
Google DeepMind's Post - LinkedInlinkedin.com · supportingWe're bringing agentic video understanding to our latest Gemini models. They can now analyze videos with better accuracy while using up to 88%
Google on X: "With agentic video understanding, developers can dynamically search, scan, and inspect long-form video with up to: 📉 88% fewer tokens 📉 66% lower cost ✅ 7% better accuracy" / Xx.com · supportingLog inSign up ## Post # Google on X: "With agentic video understanding, developers can dynamically search, scan, and inspect long-form video with up to: 📉 88% fewer tokens 📉 66% lower cost ✅ 7% better accuracy" @Google Google @Google With agentic video understanding, developers can dynamically search, scan, and inspect long-form video with up to: 📉 88% fewer tokens 📉 66% lower cost ✅ 7% better accuracy Scatter plot titled 5:30 PM · Sep 1, 202641KViews 6 @Google Google @Google Sep 1 We’re introducing a new capability to our latest Gemini models: agentic video understanding. This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens. See how it works 🧵 00:00 199 @Google Google @Google Sep 1 [...] 6 @Google Google @Google Sep 1 Agentic video understanding is available now across our latest models via the Gemini API in @GoogleAIStudio and the Gemini Enterprise Agent Platform.
Introducing agentic video understanding with Gemini | Google AI Studio (@GoogleAIStudio) on Xx.com · supportingToday, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more. The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Benchmarks [...] Benchmarks Unlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model’s core reasoning with native video tools to dynamically search, scan,