papersTODAY 04:00 UTC
EventVL Uses Multimodal LLMs to Interpret Event Camera Streams
Researchers present EventVL, a multimodal large language model designed to interpret event-based camera data rather than relying on CLIP-style encoders. The work targets explicit understanding of event streams, a sensing modality where most prior vision-language approaches have focused only on conventional perception tasks. The paper is a revised cross-listing on arXiv.