For a long time, conversational AI has been confined to unimodal text or speech exchanges. As human-machine dialogue increasingly extends into the physical world through embodied agents and virtual assistants, agents must perceive scenes and humans through multimodal signals (e.g., speech, video, sensor data). To enable seamless human-machine collaboration, they must also generate speech and non-verbal communication, maintain a coherent conversation, and, when necessary, interact with the physical world in real-time, with low latency, and in a context-aware manner.Â
Historically, research on these challenges has been fragmented. For example, egocentric conversational AI focuses on interactions from the user's perspective, such as AI assistants or augmented-reality glasses. Conversely, exocentric and dyadic conversational AI centers on third-person perspectives and face-to-face communication, typical of traditional robotics. However, real-time multimodal conversations naturally demand an integration of multiple perspectives to perform cross-view reasoning. For instance, a seamless dialogue about assembling furniture requires an understanding of the user's view (egocentric), the broader context of the room and objects (exocentric), and human emotions, gestures, and motion (dyadic).
These challenges outline four core research areas that require interdisciplinary collaboration across computer vision, robotics, machine learning, speech and audio processing, and dialogue communities:
Multimodal representation and perception
Real-time interaction dynamics and memory
Embodied interaction
Benchmarks and datasets
The first NeurIPS Real-Time Multimodal Conversational AI workshop aims to advance progress across these critical research areas.