Deborah Dahl, Conversational Technologies
Dirk Schnelle-Walka, Switch Consulting
Gérard Chollet, CNRS, Télécom-SudParis
The Journal on Multimodal User Interfaces (JMUI) publishes multidisciplinary research on the design, implementation, and evaluation of multimodal interfaces and interactions, spanning Human-Computer Interaction, signal processing, cognitive science, and ergonomics.
This special issue stems from the W3C Workshop on Smart Voice Agents (February 2026), which identified a critical need to move beyond proprietary, fragmented voice ecosystems toward interoperable, real-time, and inclusive systems. Smart Voice Agents (SVAs) are no longer just novelty tools but foundational interfaces in the healthcare, automotive, and web accessibility sectors, sometimes in multimodal contexts involving other modalities such as graphics or gestures. However, current platforms lack interoperability, which leads to "lock-in," hindering portability and making it difficult for agents to collaborate effectively. This special issue seeks high-quality research that addresses these gaps through new protocols, architectural approaches, and inclusive design.
We invite submissions focusing on the "Top 8 Cross-cutting Issues" and key themes identified by the workshop community:
Real-time Interaction and Natural Turn-taking: Research into incremental, word-level processing to enable responsive interruption handling and low-latency feedback
Interoperability and Multi-Agent Protocols: Standards for agent-to-agent communication, conversation handoff, and shared context (e.g., the Open Floor Protocol).
Voice-First Web Accessibility: Leveraging LLM reasoning and embeddable agents to bridge the gap between visual interfaces and non-visual access, moving beyond traditional screen readers.
Multimodal Coordination and Grounding: Fusing voice with gaze tracking, gestures, and non-verbal cues to improve intent resolution and trust.
Reliability and Hallucination Control: Shared benchmarks and evaluation methods for ASR and LLM error modes in noisy or high-stakes environments like healthcare.
Accessibility for Immersive Content: Standardizing semantic 3D metadata and voice spatial cues for immersive/XR environments.
Pronunciation and Language Representation: Standards for phonetic markup (e.g., SSML in HTML) to ensure consistent pronunciation across assistive technologies.
Privacy, Trust, and Delegation: Frameworks for user consent, privacy-preserving identity assertions, and auditable agent actions taken on behalf of users.
JOURNAL ON MULTIMODAL USER
INTERFACES
is a Springer Journal
2025 Impact Factor: 2.6
Editor-in-Chief: Jean-Claude Martin,
CNRS-LISN, Université Paris Saclay, France
More information:
https://link.springer.com/journal/12193