papersSEP 12 04:00 UTC
arXiv Paper Applies Incongruity-Resolution Supervision to Multimodal Humor Understanding
A revised arXiv preprint proposes supervising multimodal models with incongruity-resolution reasoning signals, arguing that humor tasks reward correct reasoning and not just correct answers. The authors evaluate this approach against humor benchmarks such as the New Yorker Cartoon Caption Contest. The work targets improving machine understanding of cartoon-based jokes.