papersTODAY 04:00 UTC
Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions
Researchers propose using recordings of how listeners' eyes move while they interpret a speaker's description as a training signal for vision-language models. By converting these gaze scanpaths into incremental feedback, the models learn to produce referring expressions that are more pragmatically suited to the listener. The work is presented as an arXiv preprint in the computation and language category.