papersTODAY 04:00 UTC
Paper studies how distractors affect test-time scaling in reasoning VLMs
A new arXiv preprint examines whether irrelevant information, known as distractors, changes how vision-language models behave when allowed to spend more compute at inference time. Prior work on text-only models found that such distractors can worsen inverse scaling, where reasoning degrades as test-time compute grows. The authors extend that question to multimodal settings, where models must handle both images and text. The submission is a cross-listed replacement in cs.AI and cs.LG.
COVERAGE · 3 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.CLUnderstanding the Effects of Distractors on Reasoning Vision-Language Models ↗TODAY 04:00 UTC
arXiv cs.AIUnderstanding the Effects of Distractors on Reasoning Vision-Language Models ↗TODAY 04:00 UTC
arXiv cs.LGUnderstanding the Effects of Distractors on Reasoning Vision-Language Models ↗TODAY 04:00 UTC