Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
Abstract
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector τ_l, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts τ_l at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.
Community
VideoLLMs do capture temporal information, but it peaks at intermediate layers and progressively fades toward the output. We pinpoint this fading through representation-level analysis and propose Temporal Activation Injection (TAI), which extracts the temporal signal at its peak and reinjects it into the subsequent layers. Without any additional training, TAI consistently improves temporal reasoning across VideoLLMs built on diverse language backbones. Accepted to NeurIPS 2026!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models (2026)
- What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA (2026)
- The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs (2026)
- From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making (2026)
- Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting (2026)
- On Temporal Binding in Large Audio Language Models (2026)
- GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.01595 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper