화자-청자 이원 상호작용에서 멀티모달 신호를 활용한 동적 공감 반응 생성 모델

A Multimodal Dynamic Empathetic Response Generation Model for Speaker–Listener Dyadic Interactions

초록

The rapid growth of online and remote interactions has increased demand for AI systems that can deliver emotionally supportive, empathetic responses, yet many conversational agents still fail to reflect the richness of human empathy in face-to-face, multimodal settings. This study proposes a multimodal empathetic response generation framework that explicitly models emotional dynamics in dyadic interactions to produce dynamic empathetic speaking videos. Given a speaker’s audio and video, the framework generates an empathizer’s responses across linguistic, acoustic, and visual modalities. The framework integrates three modules. First, a Video Conversation (ViCo)-based facial representation module encodes the speaker’s expression dynamics as 3D Morphable Model (3DMM) coefficients and predicts the empathizer’s facial expressions and head movements to enhance visual empathy. Second, an AnyGPT- based module generates semantically coherent and emotionally appropriate empathetic utterances from acoustic features. Third, to mitigate temporal mismatch between generated speech and facial motion, a lip-synchronization module (Wav2Lip) aligns lip movements with audio and produces natural conversational videos. Quantitative and qualitative evaluations show that the proposed framework generates more empathetic, dynamic, and context-consistent responses, highlighting the importance of multimodal integration for emotionally expressive, human-centered AI.

키워드

Empathetic Response GenerationMultimodal LearningDyadic Interaction3D Morphable Model (3DMM)Affective Computing
제목
화자-청자 이원 상호작용에서 멀티모달 신호를 활용한 동적 공감 반응 생성 모델
제목 (타언어)
A Multimodal Dynamic Empathetic Response Generation Model for Speaker–Listener Dyadic Interactions
저자
최세인최지훈김병규
DOI
10.9717/kmms.2026.29.5.746
발행일
2026-05
유형
Y
저널명
멀티미디어학회논문지
29
5
페이지
746 ~ 756