papersTODAY 04:00 UTC
Study probes whether vision-language models truly capture temporal structure
A new arXiv paper reframes temporal grounding as an anomaly-detection problem in order to test whether vision-language models actually represent time ordering in video and image sequences. The authors report that strong results on existing video benchmarks do not necessarily show that these models rely on temporal structure rather than shortcuts. They propose this setup as a way to measure temporal consistency more directly.