Study probes how multimodal LLMs describe bistable images like the duck-rabbit
A new arXiv paper examines whether multimodal large language models report bistable images, such as the duck-rabbit figure, in ways comparable to humans. Humans typically perceive only one interpretation of such images at a time, and the research tests whether MLLMs show a similar pattern. The work falls in the area of model perception and evaluation research.