papersSEP 10 04:00 UTC
Survey Reviews Inference-Efficiency Methods for Video and Audiovisual LLMs
A new survey on arXiv examines mechanisms for reducing inference costs in video large language models, which pair video representations with pretrained LLMs to generate responses from text prompts. The paper addresses why video understanding remains computationally expensive and organizes existing efficiency techniques across video and audiovisual tasks.