55% of Video Benchmarks Are Solvable Without Watching the Video
Video-Oasis audits existing video-LLM benchmarks and finds most questions yield to language priors alone, exposing a near-random-guess ceiling for true visual tasks.
Most video-LLM leaderboards are measuring something other than video understanding. The assumption baked into every ranking, that a model scoring well on video benchmarks must be perceiving and reasoning about visual content, turns out to be wrong for the majority of benchmark samples currently in circulation.
Video-Oasis audits existing benchmarks rather than adding another one. The core diagnostic move is separating questions that require genuine visual or temporal input from those solvable via language priors and world knowledge alone. The audit works by systematically testing whether models can answer benchmark questions correctly without ever seeing the video, then flagging those samples as shortcuts. Think of it as a stress-test for the benchmark itself rather than for the model: if a question about a cooking video can be answered correctly by a model that has never seen a kitchen on screen, that question is not measuring video understanding, it is measuring culinary knowledge retrieval.
After removing shortcut-solvable samples, what remains is a set of video-native challenges, questions where visual perception and temporal context genuinely determine the answer. On those filtered questions, state-of-the-art models perform only marginally above random guessing. 55% of existing benchmark samples fall to language priors alone, and the residual difficulty exposes a capability ceiling current architectures have not cleared. For teams evaluating or selecting video-LLMs for production use, the takeaway is direct: your current benchmark scores are likely inflated by linguistic pattern-matching, and the models you ranked highest may not actually see the video.
We're thinking: We find the benchmark-auditing framing more consequential than any single leaderboard result this quarter. If 55% of video benchmark samples are solvable without visual input, then the entire ranking infrastructure built on those benchmarks is measuring a mixture of language prior strength and actual visual perception, with no clean way to separate the two. That means teams choosing a video-LLM based on published scores may be optimizing for the wrong capability entirely. The distilled video-native challenge set Video-Oasis produces is worth treating as a minimum bar, not a stretch goal, before deploying any model in contexts where temporal or visual grounding actually matters.
Key takeaways:
- Video-Oasis introduces a benchmark audit framework that classifies samples by whether they require genuine visual and temporal input, filtering shortcut-solvable questions driven by language priors or world knowledge.
- After filtering, 55% of existing benchmark samples are removed as shortcuts; on the remaining video-native subset, state-of-the-art models land only marginally above random guessing, though the filtered challenge set is necessarily smaller and its generalization across benchmark types warrants further validation.
- Teams benchmarking video-LLMs for any deployment where temporal or visual grounding matters should run candidates against the Video-Oasis filtered challenge set before trusting published leaderboard rankings.
Source: Video-Oasis: Rethinking Evaluation of Video Understanding