---
format: "aidr-story-markdown/v1"
id: "5373e342022d077eb44d17f6277c023de754beaf4a398796b833dd41c260c067"
canonical_url: "https://aidr.today/5373e342?lang=en"
title: "Google Cuts Gemini Video Tokens 88% in Shift to Agentic Understanding"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-01T18:51:47.000Z"
category: "Models"
topics: ["google","gemini","agent","inference"]
source_urls: ["https://huggingnews.com/ai/google-cuts-gemini-video-tokens-88percent-in-shift-to-agentic-understand-9496537d","https://x.com/Google/status/2094840325789430066","https://x.com/GoogleAIStudio/status/2094841307935957304","https://x.com/Google/status/2094840332336652328","https://x.com/Google/status/2094840983913402704","https://x.com/GoogleDeepMind/status/2094840182457422260"]
summary: "Developers can now use a goal-directed reasoning system to process long-form video by adjusting frame rates and modalities based on a specific prompt. This agentic capability is available via the Gemini API and AI Studio for 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models, as well as the Gemini Enterprise Agent Platform. The system reduces token consumption by as much as 88%, lowers costs by 66%, and improves benchmark accuracy by 7%. The method replaces static processing, which analyzes video at a fixed rate of one frame per second. By scanning speech transcripts to pinpoint relevant moments before fetching visual frames, Gemini reasons across audio and visual tracks for content ranging from 10 minute guides to multi hour recordings. The feature will be integrated into the Gemini app and YouTube's \"Ask YouTube\" tool on the video watch page."
---

# Google Cuts Gemini Video Tokens 88% in Shift to Agentic Understanding

> [Open the canonical story](<https://aidr.today/5373e342?lang=en>)

**Published:** 2026-09-01T18:51:47.000Z
**Category:** Models
**Topics:** google, gemini, agent, inference

## Summary

Developers can now use a goal\-directed reasoning system to process long\-form video by adjusting frame rates and modalities based on a specific prompt\. This agentic capability is available via the Gemini API and AI Studio for 3\.7 Flash, 3\.6 Flash, and 3\.5 Flash\-Lite models, as well as the Gemini Enterprise Agent Platform\. The system reduces token consumption by as much as 88%, lowers costs by 66%, and improves benchmark accuracy by 7%\. The method replaces static processing, which analyzes video at a fixed rate of one frame per second\. By scanning speech transcripts to pinpoint relevant moments before fetching visual frames, Gemini reasons across audio and visual tracks for content ranging from 10 minute guides to multi hour recordings\. The feature will be integrated into the Gemini app and YouTube's "Ask YouTube" tool on the video watch page\.

## Sources

- [Story source](<https://huggingnews.com/ai/google-cuts-gemini-video-tokens-88percent-in-shift-to-agentic-understand-9496537d>)
- [Story source](<https://x.com/Google/status/2094840325789430066>)
- [Story source](<https://x.com/GoogleAIStudio/status/2094841307935957304>)
- [Supporting source](<https://x.com/Google/status/2094840332336652328>)
- [Story source](<https://x.com/Google/status/2094840983913402704>)
- [Supporting source](<https://x.com/GoogleDeepMind/status/2094840182457422260>)

