Developers can now use a goal-directed reasoning system to process long-form video by adjusting frame rates and modalities based on a specific prompt. This agentic capability is available via the Gemini API and AI Studio for 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models, as well as the Gemini Enterprise Agent Platform. The system reduces token consumption by as much as 88%, lowers costs by 66%, and improves benchmark accuracy by 7%.

The method replaces static processing, which analyzes video at a fixed rate of one frame per second. By scanning speech transcripts to pinpoint relevant moments before fetching visual frames, Gemini reasons across audio and visual tracks for content ranging from 10 minute guides to multi hour recordings. The feature will be integrated into the Gemini app and YouTube's "Ask YouTube" tool on the video watch page.

Sign in to suggest edits

Key sources

  1. SOURCE@google“process long-form video content with more accuracy”x.com
  2. SOURCE@googleaistudio“take an active, goal-directed role in determining what to watch”x.com
  3. SUPPORT@google“7% better accuracy”x.com
  4. SOURCE@google“Gemini Enterprise Agent Platform”x.com
  5. SUPPORT@googledeepmind“from 10-minute guides to multi-hour recordings”x.com
Markdown