상세 보기
화자-청자 이원 상호작용에서 멀티모달 신호를 활용한 동적 공감 반응 생성 모델
- 최세인;
- 최지훈;
- 김병규
초록
The rapid growth of online and remote interactions has increased demand for AI systems that can deliver emotionally supportive, empathetic responses, yet many conversational agents still fail to reflect the richness of human empathy in face-to-face, multimodal settings. This study proposes a multimodal empathetic response generation framework that explicitly models emotional dynamics in dyadic interactions to produce dynamic empathetic speaking videos. Given a speaker’s audio and video, the framework generates an empathizer’s responses across linguistic, acoustic, and visual modalities. The framework integrates three modules. First, a Video Conversation (ViCo)-based facial representation module encodes the speaker’s expression dynamics as 3D Morphable Model (3DMM) coefficients and predicts the empathizer’s facial expressions and head movements to enhance visual empathy. Second, an AnyGPT- based module generates semantically coherent and emotionally appropriate empathetic utterances from acoustic features. Third, to mitigate temporal mismatch between generated speech and facial motion, a lip-synchronization module (Wav2Lip) aligns lip movements with audio and produces natural conversational videos. Quantitative and qualitative evaluations show that the proposed framework generates more empathetic, dynamic, and context-consistent responses, highlighting the importance of multimodal integration for emotionally expressive, human-centered AI.
키워드
- 제목
- 화자-청자 이원 상호작용에서 멀티모달 신호를 활용한 동적 공감 반응 생성 모델
- 제목 (타언어)
- A Multimodal Dynamic Empathetic Response Generation Model for Speaker–Listener Dyadic Interactions
- 저자
- 최세인; 최지훈; 김병규
- 발행일
- 2026-05
- 유형
- Y
- 저널명
- 멀티미디어학회논문지
- 권
- 29
- 호
- 5
- 페이지
- 746 ~ 756