Multimodal emotion recognition in the wild: Corruption modeling and relevance-guided scoring

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Despite significant progress in Multimodal Emotion Recognition (MER), many existing approaches assume clean data, which may not fully account for the sensor noise and environmental degradation encountered in real-world deployment. Furthermore, prior robustness evaluations have predominantly relied on embedding-level corruption-a simplified proxy that may not capture the complex distributional shifts caused by complex raw-level noise. In this work, we present a systematic robustness analysis under representative raw input-level corruption. Our study indicates that existing models experience significant performance degradation under these perturbations-particularly in the language modality-highlighting potential vulnerabilities in current fusion paradigms. To address this, we propose Relevance-Guided Scoring (RGS), an adaptive fusion mechanism. Complementing existing methods based on coarse-grained weighting or relative attention, RGS estimates semantic relevance at a fine-grained temporal level. This allows the model to selectively emphasize informative segments even within partially degraded streams, thereby improving robustness against multimodal interference. Experiments on the CMU-MOSI and CMU-MOSEI benchmarks as representative proxies for MER demonstrate that our approach is model-agnostic, improving the robustness of various backbone architectures across multiple corruption scenarios.

키워드

Multimodal emotion recognitionMultimodal learningMultimodal corruption modelingRelevance-guided scoring
제목
Multimodal emotion recognition in the wild: Corruption modeling and relevance-guided scoring
저자
Lee, YoonsunCho, Sunyoung
DOI
10.1016/j.eswa.2026.132660
발행일
2026-09
유형
Article
저널명
Expert Systems with Applications
326