Google Launches Agentic Video Understanding for Gemini Flash, Cutting Video Tokens by Up to 88%

Loading…

Google has shipped agentic video understanding capabilities for its Gemini Flash model family, introducing a technique that reduces video token consumption by up to 88% compared to naive frame sampling. The system uses an agentic approach to selectively process only the most relevant video segments rather than uniformly sampling frames, dramatically cutting inference costs for long-form video inputs. This is a significant capability upgrade for developers building video analysis pipelines, as it makes high-context video tasks tractable at production scale without ballooning costs. Developers using the Gemini API can now process longer videos more efficiently, enabling use cases like meeting summarization, video search, and content moderation that were previously cost-prohibitive. The combination of cost reduction and agentic framing signals Google is positioning Gemini Flash as the go-to multimodal model for high-throughput video workloads.