상세 보기
EmoXFormer: Human-Cognition-Inspired Multimodal Emotion Recognition from Disjoint Modality Datasets
- Aisha, Qurat Ul Ain;
- Choi, Ji-Hoon;
- Choi, Se-In;
- Roy, Partha Pratim;
- Kim, Byung-Gyu
SCOPUS
0초록
Most multimodal emotion recognition (MER) assumes temporally aligned and co-recorded modalities, yet deployed systems often observe text, audio, and vision from separate pipelines where sample-level correspondence is unavailable. We study Disjoint Modality Learning (DML), where modality specialists are trained on independent corpora and combined at inference without pairing, identity linkage, or synchronization. We propose EmoXFormer, an alignment-free fusion model that couples strong unimodal encoders (RoBERTa, Wav2Vec2, ViT) with Cross-Gated Context Attention (CGCA) to exchange cross-modal context while suppressing off-context activations under missing or mismatched evidence. To enable reproducible evaluation, we introduce a Unified Disjoint Test (UDT) that harmonizes MELD (text), RAVDESS (audio), and FER2013 (vision) into a shared seven-class label space. On UDT-U (unimodal aggregate under a shared head), EmoXFormer achieves 66.9% micro accuracy. On UDT-3M (congruent triads from disjoint corpora), fusion reaches 79.2%, outperforming the best unimodal baseline (63.9%) by +15.3 points. Results show that robust multimodal decision fusion is feasible even when cross-modal correspondence is absent.
키워드
- 제목
- EmoXFormer: Human-Cognition-Inspired Multimodal Emotion Recognition from Disjoint Modality Datasets
- 저자
- Aisha, Qurat Ul Ain; Choi, Ji-Hoon; Choi, Se-In; Roy, Partha Pratim; Kim, Byung-Gyu
- 발행일
- 2027-08
- 유형
- Conference paper
- 권
- 16819
- 페이지
- 292 ~ 306
- 언어
- ENG
- 출판사
- Springer Nature Switzerland AG
- 발행국가
- 스위스
- 분량
- 15 페이지
- ISSN
- E 1611-3349
P 0302-9743