Gemini will watch video like an agent: higher accuracy with up to 88 percent fewer tokens
Google DeepMind announced agentic video understanding for its newest Gemini models: they analyze video with higher accuracy while spending up to 88 percent fewer tokens.
Video is the hungriest input a language model faces: frames become tokens and even a short clip devours the context window. Google DeepMind's announcement changes that arithmetic: the newest Gemini models get agentic video understanding, and the company's two numbers are higher analysis accuracy and up to 88 percent fewer tokens.
Agentic means the model no longer swallows the whole video blindly; it approaches the footage according to what it is looking for, returning to specific segments when needed and skimming the rest. It is a model that searches with the remote in hand rather than one that merely watches. Higher accuracy at lower cost is the natural product of that selective gaze.
The practical payoff is wide: hour-long meeting recordings, security cameras, lecture videos and archive sweeps open up to real workloads once the token bill drops by up to an order of magnitude. It also completes the video leg of the context-economy chain we covered this week with ContextPilot and SparkLLM: the winner is the model that knows where to look, over the one that reads everything.