Key Takeaways & Executive Findings
- •• TimeJudge introduces a zero-shot framework that recasts temporal error detection as binary question pairs, eliminating the need for task-specific fine-tuning. • TEDBench provides a rigorously constructed benchmark with 381 videos, 1524 captions, and 3048 QA pairs, covering four complexity levels for fine-grained temporal error evaluation. • Comprehensive evaluations show that TimeJudge consistently improves recall and F1-score across multiple state-of-the-art video-LLMs for temporal consistency assessment. • The approach is generalizable, scalable, and training-free, making it suitable for real-world evaluation of video captions in multimodal systems.
Abstract
Video large language models (video-LLMs) have demonstrated impressive capabilities in multimodal understanding, but their potential as zero-shot evaluators for temporal consistency in video captions remains underexplored. Existing methods notably underperform in detecting critical temporal errors, such as missing, hallucinated, or misordered actions. To address this gap, we introduce two key contributions. (1) TimeJudge: a novel zero-shot framework that recasts temporal error detection as answering calibrated binary question pairs. It incorporates modality-sensitive confidence calibration and uses consistency-weighted voting for robust prediction aggregation. (2) TEDBench: a rigorously constructed benchmark featuring videos across four distinct complexity levels, specifically designed with fine-grained temporal error annotations to evaluate video-LLM performance on this task. Through a comprehensive evaluation of multiple state-of-the-art video-LLMs on TEDBench, we demonstrate that TimeJudge consistently yields substantial gains in terms of recall and F1-score without requiring any task-specific fine-tuning. Our approach provides a generalizable, scalable, and training-free solution for enhancing the temporal error detection capabilities of video-LLMs.
1. Introduction
Video large language models (video-LLMs) are rapidly transitioning from research prototypes to real-world products, making rigorous and scalable evaluation indispensable. Although human assessment remains the gold standard, it is slow, costly, and subjective, prompting a shift toward automated, model-based protocols.
Building on the success of the "LLM-as-a-Judge" paradigm for text, the emerging "Video-LLM-as-a-Judge" framework proposes to apply powerful video-LLMs to score, rank, and filter the video captions generated by other video language models (VLMs) (Zheng et al., 2023; Liu and Zhang, 2025). Besides offering low-cost, high-throughput evaluation, this approach accelerates the evolution of video-LLMs from dialogue systems into general-purpose multimodal agents with broad applications in evaluation, alignment, retrieval, and reasoning (Bai YS et al., 2023; Lee et al., 2023; Li RS et al., 2023; Liang et al., 2023; Liao et al., 2024; Xu et al., 2024).
However, when applied to video-understanding benchmarks, existing video-LLM judges prove unreliable. As illustrated in Fig. 1, directly prompting them to verify the fidelity of machine-generated captions often yields incorrect verdicts, particularly regarding temporal coherence, missing events, hallucinated actions, or misordered sequences. We identify three underlying causes: (1) an over-reliance on static spatial cues (e.g., objects and scenes) at the expense of temporal perspicacity; (2) limited capacity for multi-hop, high-level reasoning in the temporal domain; (3) pronounced biases toward textual priors or visually dominant patterns rather than grounding decisions on video evidence. These shortcomings, rooted in both training data and model architecture, erode the reliability of video-LLM-based evaluation for temporally sensitive tasks such as event localization and action detection.
Loading authentic research manuscript (Pages 1–5)...
Yangliu HU, Zikai SONG, Junqing YU, Yiping Phoebe CHEN, Wei YANG (2025). TimeJudge: empowering video-LLMs as zero-shot judges for temporal consistency in video captions. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2500412
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is TimeJudge?
TimeJudge is a zero-shot framework that enhances video-LLMs' ability to detect temporal inconsistencies in video captions by decomposing the verification task into calibrated binary question pairs and applying modality-sensitive confidence calibration.
What is TEDBench?
TEDBench is a benchmark consisting of 381 videos, 1524 captions with controlled temporal errors, and 3048 bidirectional QA pairs, designed to evaluate video-LLM performance on temporal error detection across four complexity levels.
Does TimeJudge require fine-tuning?
No, TimeJudge is a training-free, zero-shot method. It consistently improves recall and F1-score without any task-specific fine-tuning, making it a practical and scalable solution.
What types of temporal errors does TimeJudge address?
TimeJudge targets missing, hallucinated, and misordered actions in video captions, which are common yet critical temporal consistency errors.
How does TimeJudge improve temporal error detection?
By reducing cognitive load through lightweight binary queries and dynamically weighting visual versus textual evidence based on calibration, TimeJudge steers video-LLMs toward relevant spatiotemporal evidence, yielding substantial performance gains.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena