Project Title: Conflicting Interpretations in Dialogue
TL;DR: Investigating misunderstandings in collaborative dialogues: dual-perspective annotation, grounding dynamics, and model evaluation
Description
Misunderstandings in communication, especially within collaborative dialogue settings, present significant challenges in Natural Language Processing (NLP). This project explores the phenomenon of conflicting interpretations among interlocutors, specifically focusing on how misunderstandings arise and how they can be computationally modeled within collaborative communication scenarios. Current NLP methodologies often assume a single, definitive interpretation for reference expressions; however, this research investigates situations where speakers and listeners assign different referents to the same expressions—a common yet underexplored phenomenon in both human-human and human-AI interactions.
Our empirical work centres on dialogues from the MapTask corpus (Anderson et al., 1991), in which participants collaborate using asymmetric map information. We developed a perspectivist annotation framework that, unlike prior schemes, separately captures the speaker's intended and the addressee's interpreted referent for each reference expression, enabling dual-perspective tracking of how understanding emerges, diverges, and repairs over the course of a dialogue. Applying this framework corpus-wide via a scheme-constrained LLM pipeline, we produced over 13,000 annotated reference expressions with reliability estimates. Our analysis reveals that multiplicity discrepancies, where a landmark appears more than once on one map but not the other, are the primary driver of referential misalignment (Li, Gatt & Poesio, 2026a). Building on this resource, we evaluated whether vision-language models can distinguish potential from established common ground in asymmetric dialogue, and identified a systematic content-driven over-alignment bias analogous to the overhearer's illusion (Li, Gatt & Poesio, 2026b; Schober & Clark, 1989).
We are currently working on transferring this annotation methodology to other collaborative dialogue datasets and exploring the possibility of recording our own dialogue data to extend the investigation beyond MapTask. These efforts aim to support the development of computational methods capable of identifying misunderstandings, modelling their resolution, and facilitating recovery strategies, advancing both the performance of LLMs in collaborative settings and our theoretical understanding of incremental mutual understanding.
The project is being carried out by Nan Li. He is a Ph.D. candidate in the NLP group at the Department of Information and Computing Sciences at Utrecht University, supervised by Prof. Massimo Poesio and Prof. Albert Gatt. For more information, please contact Nan at n.li[AT]uu.nl.
Publications
Li, N., Gatt, A., & Poesio, M. (2026a). Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask. In Proceedings of LREC 2026 (Oral).
Li, N., Gatt, A., & Poesio, M. (2026b). Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue. To appear in Proceedings of SIGDIAL 2026 (Oral).
References
Anderson, A. H., Bader, M., Bard, E. G., Boyle, E., Doherty, G., Garrod, S., Isard, S., Kowtko, J., McAllister, J., Miller, J., Sotillo, C., Thompson, H. S., & Weinert, R. (1991). The Hcrc Map Task Corpus. Language and Speech, 34(4), 351–366.
Schober, M. F., & Clark, H. H. (1989). Understanding by addressees and overhearers. Cognitive Psychology, 21(2), 211–232.