arXiv paper studies image-question dependence in VLM test-time reinforcement learning
A new arXiv preprint examines how test-time reinforcement learning adapts vision-language models to unlabeled target data, noting that results depend heavily on the quality of self-generated training signals. The authors argue that consensus-based learning signals are inherently limited and propose exploiting dependence between images and their questions to improve reliability. The work is categorized under machine learning and has not yet been peer reviewed.