ReH-FUSE: Reliability-Aware Fusion of Experts for Multimodal Emotion Recognition
A new arXiv paper introduces ReH-FUSE, a method for multimodal emotion recognition in conversation that accounts for how much each evidence source can be trusted in a given instance. The approach hierarchically combines expert predictions so that lexical, vocal, and other cues are weighted according to their reliability rather than treated as equally informative. It targets the problem that different modalities may dominate depending on the conversational context.