🗣️ Dr Samantha Sie & Prof Ianthi Tsimpli (University of Cambridge)
📅 18th Jun, 2026 (16:00 - 17:00)
🏫 Little Hall, Sidgwick Site
Abstract:
India is a highly linguistically diverse country and classrooms are a good reflection of this diversity. Several schools opt for English as the school language despite the highly varied proficiency of the teachers and the low proficiency of the learners. To mitigate this, teachers engage in pedagogical translanguaging by utilising their students’ first languages to scaffold learning; this entails the practice of codeswitching, which is one of the linguistic phenomena we examine in the talk.
Seven teachers from four government-aided primary schools in Delhi were recorded during English lessons, and transcriptions of their audio recordings constitute the Teacher Talk Corpus. We explore teachers’ language use in English subject lessons, focusing on (i) codeswitching patterns and (ii) the use of English finiteness and interrogatives, which are features susceptible to influence from Indian English.
🗣️ Dr Ksenia Zanon (University of Cambridge)
📅 21st May, 2026 (16:30 - 17:00)
🏫 GR05, EFB, Sidgwick Site
Abstract:
The empirical remit of this paper is the prepositional domain of Russian, where I harness both the well-studied facts and those bereft of the systematic analytical scrutiny they deserve. My exploration draws primarily from the synchronic data but is also informed by certain diachronic developments in the language. I propose a new taxonomy of Russian prepositions which entails three classes: Russian Ps are either (i) (partial) realizations of case; (ii) adjuncts in the nominal domain or (iii) elements eligible to form their own phasal domains. The latter two types are analytical offshoots of the contextual approach to phases and require a parametric approach to the structure of the noun phrase.
🗣️Prof Napoleon Katsos (University of Cambridge)
📅 21st May, 2026 (16:00-16:30)
🏫 GR05, EFB, Sidgwick Site
Abstract:
In September 2025 Cambridge welcomed back the 11st instalment of the Experimental Pragmatics (XPrag) conference, which first took off as a conference series in 2005 from the same location. In this talk I will showcase some of the research presented at XPrag 2025 by Dimitris Kastanas, Songqiao Xie, and myself. These projects investigate informativeness and metonymy, situated within a broader cognitive science framework. The effect of non-linguistic cognitive skills and the pragmatic abilities of different kinds of minds (autistic or neurotypical) are themes that run through these projects. I aim to show that experimental investigations of linguistically-inspired questions can generate exciting food for thought for several disciplines.
🗣️ Dr Dimitra Lazaridou-Chatzigoga (University of East Anglia, University of Cambridge)
Moderated by Prof Napoleon Katsos
📅 14th May, 2026 (16:00-17:00)
🏫 Little Hall, Sidgwick Site
Abstract:
All languages are thought to have ways to express generalisations, like ‘Leopards have spots’, ‘Dodo birds are extinct’ and ‘Canadians like hockey’. While much is known about generics in English and several closely related languages, they are not as well-understood across languages as other linguistic forms. Addressing this limitation, our Generics Across Languages program systematically studies a wide set of languages from different language families, via the means of a newly developed toolkit, combining insights from formal linguistics, cognitive psychology and semantic fieldwork elicitation techniques, focusing on their morphosyntax and semantics. In this talk, I will first discuss the motivation for this project and how it fits within the broader generics literature. I will then present our methodology and the Generics Toolkit we developed, whose primary component is a series of storyboards with embedded target sentences, focusing on different ontological categories and targeting five key readings. The talk will showcase data from our current database, which comprises data from 19 languages, collected via our network of international collaborators. This work will facilitate an in-depth examination of proposed universals on generics, shedding light on their theoretical implications.
🗣️ Dr Holly E. A. Sutherland (University of Cambridge)
📅 7th May, 2026 (16:00-16:30)
🏫GR05, EFB, Sidgwick Site
Abstract:
The God, Midrash, and Linguistic Creativity interdisciplinary research project (part of the GLAD project) is studying linguistic creativity in the context of classical rabbinic midrash, a form of biblical exegesis that draws rich meaning from Jewish scripture. Midrashic readings often appear unconventional, surprising, or even creative. To the casual reader encountering midrash for the first time, it may seem an extremely novel way of using and interpreting language.
However, many of the features of midrash, and the kinds of textual ambiguities it often exploits, are in fact fundamental elements of the way people engage with language on a day-to-day basis. In this talk I will discuss how, through interdisciplinary collaboration between Jewish divinities studies and psycholinguistics, we have developed a novel linguistic creativity task that explores participants' ability to identify and exploit linguistic ambiguities: The Alternative Interpretations Task (AIT).
🗣️ Dr Suhail Matar (University of Cambridge)
📅 7th May, 2026 (16:00-16:30)
🏫GR05, EFB, Sidgwick Site
Abstract:
Reading comprehension involves constructing a cumulative mental model that represents unfolding information. Within this model, the mind establishes discourse referents—entities that can be picked out and referred to, such as ‘sponges’ and ‘steel wool’ in ‘He bought sponges of steel wool…’. Though the composite entity ‘sponges of steel wool’ is also referenceable (‘…and glued them together.’), it remains unclear whether the mind establishes such composite referents in the mental model, and whether this is different than establishing simple referents. Here, I present a sentence level-naturalistic reading paradigm with an embedded categorical manipulation. Participants (n=43) read 72 five-sentence English stories, where the fourth, critical sentence established three simple referents (‘wool, sponges and steel’; simple3), two simple referents (‘steel wool and sponges’; simple2), or two simple referents forming a composite (‘sponges of steel wool’; composite). As hypothesised, and beyond syntactic, semantic and lexical factors, linear regression analyses revealed longer reading times (RTs) for establishing additional simple (RTsimple3>RTsimple2) and composite (RTcomposite>RTsimple2) referents, only during critical sentences. Bayesian modelling of all sentence RTs further showed that a model distinguishing between composite and simple referents best fit RT distributions, outperforming models that tracked referents indiscriminately, tracked only simple referents, or ignored referents altogether. Overall, I provide evidence for distinct cognitive operators acting on the mental model, indicating hierarchical structure-building during comprehension.
🗣️ Dr Matheiu Beaudouin (University of Cambridge)
📅 26th Feb, 2026 (16:00 - 16:30)
🏫 LB3, Sidgwick Site
Abstract:
Tangut, the language of the rulers of the Western Xia Empire (1038–1227), plays a central role in the historical phonology of the Sino-Tibetan family. It is both the third earliest literary language of the family, after Chinese and Tibetan, and a member of one of its most conservative branches, the Gyalrongic group.
Its phonology is primarily category-based: like Middle Chinese, it is mainly known through rhyme dictionaries, which provide reliable access to phonological categories but remain relatively imprecise about their actual phonetic realisation. At the same time, the existence of transcriptions between Tangut, Medieval Eastern Tibetan, Medieval Northwest Chinese, and liturgical Sanskrit means that any improvement in the reconstruction of Tangut phonology has immediate repercussions for our understanding of the synchronic and diachronic phonology of many other languages.
This talk examines R-like sounds in Sino-Tibetan, beginning with Tangut and progressively expanding outward to more distant relatives. Starting from a Pre-Tangut stage reconstructed through the comparative method—in which syllables can be inferred to contain segmental -r- in various syllabic positions—I predict the distribution of two distinct variables of Tangut phonology, with broader implications for the historical phonology of the Sino-Tibetan family.
🗣️ Dr Hannah Davidson(University of Cambridge)
📅 26th Feb, 2026 (16:30 - 17:00)
🏫 LB3, Sidgwick Site
Abstract:
This study investigates the processing and representation of morphologically-complex derived words by German-English bilinguals in their second language (English), compared to English monolinguals. Previous research on native German speakers has found decompositional processing and representation of all derived words irrespective of their semantic transparency, with semantically opaque verbs priming their stems (e.g. verstehen ‘understand’ - STEHEN ‘stand’) comparably to semantically transparent verbs. However, research in English has found no overt priming for opaque relationships between morphologically complex primes and their stem targets, suggesting that only transparent words might be decomposed in English. To this end we investigated whether German-English bilinguals tested in their L2 English behave similarly to English speakers, or whether they also decompose morphologically complex words across the transparency spectrum, like in their native German. The study used a cross-modal priming paradigm, with auditory primes (e.g. disagree) followed by visual targets (AGREE).
Participants completed a lexical decision task across four conditions: semantically opaque, semantically transparent, form/phonological and unrelated. Preliminary results revealed a significant effect of priming and condition, with the priming effect also varying by condition. There was no main effect of group but there was a significant three-way interaction between group, condition, and prime type, indicating that the priming pattern across conditions differed between bilinguals and monolinguals. Priming was considerably more prominent in transparent than in opaque pairs, in line with previous results in English. Yet while bilinguals showed the same overall pattern as monolinguals, their priming effects were of a significantly larger magnitude, suggesting greater sensitivity to word structure in German-English bilinguals. Whether and how L2 proficiency might mediate such effects will be investigated in further analyses.
🗣️ Prof John Williams (University of Cambridge)
📅 26th Feb, 2026 (16:30 - 17:00)
🏫 Little Hall Lecture Theatre, Sidgwick Site
Abstract:
Incidental language learning refers to the process of passively ‘picking up’ aspects of a language without intention to do so in the course of some activity. For example, while reading for pleasure in a foreign language one may spontaneously acquire vocabulary items or linguistic generalisations. Knowledge of grammatical generalisations in particular may remain at an unconscious level (as implicit knowledge) or they may emerge into awareness as spontaneous linguistic insight. The underlying learning mechanism is conceived in terms of generally associative / statistical learning, which has been shown to support lexical and syntactic learning in experiments using strings of nonsense syllables or nonwords. Here I report my ongoing efforts to develop an incidental language learning task that lends itself to the investigation of the mental processes involved in attaining spontaneous linguistic insight (e.g., the relationship between insight and prior implicit learning, and the neural markers of linguistic insight as measured by EEG). The vehicle for the investigation is a simple semi-artificial language (modelled on the Amazonian language Karitiana) in which novel overt markers for transitive and intransitive verb usage are combined with English lexis (e.g., Bill ro-ate the pizza, Mark ro-cooked the fish, Dave gi-slept soundly, Ryan gi-danced beautifully). Despite the apparent simplicity of this system, the templatic nature of the items and extensive training (128 sentences) I have found it surprisingly difficult to find a procedure that yields rule awareness in anything but about 25% of participants (who mostly report using intentional learning strategies and becoming aware of the system very quickly). There has been barely above chance test performance in the rule-unaware majority. Linguistic variations (of word order and marker positioning) and procedural variations that draw attention to the markers and their associated meanings do not improve the level of learning. Only by including additional surface-level cues has the level of learning improved. The general failure to learn is surprising from a statistical (or general associative) learning perspective. The purpose of this talk is to share my own surprise at this failure to learn, to consider the results in relation to different learning theories, and to question the power of statistical learning in meaningful linguistic contexts.
🗣️ Prof Nigel Collier (University of Cambridge)
📅 26th Feb, 2026 (16:00 - 16:30)
🏫 Little Hall Lecture Theatre, Sidgwick Site
Abstract:
In this talk I will touch on themes from our recent work on uncertainty and calibration in large language models, exploring how models represent, estimate, and communicate what they do not know. LLMs tend to be over-confident, whether they are right or wrong. As they are increasingly used in open-ended and long-form generation tasks, the ability to quantify and express epistemic uncertainty becomes crucial, not only for safety and reliability, but also for naturalistic interaction. I finish by proposing what I call Sunao intelligence, from the Japanese sunao, meaning open, sincere, and attuned to reality. It reframes calibration as more than a technical goal, envisioning systems that recognise their own limits and express them with humility, transparency, and genuine understanding of human intention.
🗣️ Dr Norma Schifano (University of Cambridge)
📅 12th Feb, 2026 (16:00 - 16:30)
🏫 Little Hall Lecture Theatre, Sidgwick Site
Abstract:
Truth-conditional semantics has been successful in explaining how the meaning of a sentence can be decomposed into the meanings of its parts, and how this allows people to understand new sentences. In this talk, I will discuss how a truth-conditional model can be learnt in practice on large-scale datasets of various kinds (textual, visual, ontological), and how this is empirically useful, compared to non-truth-conditional models. I will then take stock of the bigger picture, and argue it is (unfortunately) computationally intractable to reduce all kinds of language understanding to truth conditions. To enable a more complete account, I will sketch a new approach to probabilistic modelling, which maintains tractability by relaxing the strict demands of Bayesian inference. This has the potential to explain how patterns of language use arise as a result of computationally constrained minds interacting with a computationally demanding world.
🗣️ Prof Ian Roberts (University of Cambridge)
📅 5th Feb, 2026 (16:00 - 17:00)
🏫 Little Hall Lecture Theatre, Sidgwick Site
Abstract:
On standard Chomskyan assumptions, children, armed with Universal Grammar (UG) as the initial state of language acquisition, discover the grammar – arrive at the final state of acquisition -- of their native language by somehow “matching” their linguistic experience with the parameters of UG, by exposure to the vocabulary, phonology and closed-class grammatical items of the target language. Nurture (these aspects of the linguistic environment the child encounters) interacts with nature (UG). But what is this “matching” process?
Two types of answer to this question have been proposed. The most widespread, since Chomsky (1957), relies on the idea that children postulate grammars based on the interaction of their native UG with experience and choose the correct grammar following an “evaluation metric” of some kind. But there is another, arguably simpler, approach, originating in earlier North American structuralist linguistics and discussed in Chomsky’s early work (1951, 1955/1975): children discover the grammar for their language by means of a learning procedure, which effects the correct “match” between UG and linguistic experience. To use the old terminology: children apply a discovery procedure to marry experience to a grammar; the discovery procedure is the combination of the learning procedure and UG. It was impossible to pursue a discovery procedure in the 1950s: There was very little understanding of how children learn languages, there were no corpora of linguistic evidence, and the theories of computation and learning were still in their infancy. The discovery-procedure based approach has been revived recently, really for the first time since 1957, by Charles Yang. Since this approach does not require a separate evaluation metric, Occam’s razor leads us to favour it.
The present project aims to build on Yang’s insights and show how important, well-described aspects of the grammars of a range of languages, focussing on null subjects and word-order variation, can be better understood in terms of discovery procedures than in terms of evaluation metrics.
🗣️ Dr Guy Emerson (University of Cambridge)
📅 13th Nov, 2025 (16:00 - 16:30)
🏫 GR-05, English Faculty Building
Abstract:
Truth-conditional semantics has been successful in explaining how the meaning of a sentence can be decomposed into the meanings of its parts, and how this allows people to understand new sentences. In this talk, I will discuss how a truth-conditional model can be learnt in practice on large-scale datasets of various kinds (textual, visual, ontological), and how this is empirically useful, compared to non-truth-conditional models. I will then take stock of the bigger picture, and argue it is (unfortunately) computationally intractable to reduce all kinds of language understanding to truth conditions. To enable a more complete account, I will sketch a new approach to probabilistic modelling, which maintains tractability by relaxing the strict demands of Bayesian inference. This has the potential to explain how patterns of language use arise as a result of computationally constrained minds interacting with a computationally demanding world.
🗣️ Prof Bert Vaux (University of Cambridge)
📅 16th Oct, 2025 (16:30 - 17:00)
🏫 SG1, Alison Richard Building
Abstract:
This talk draws on case studies from my work to illustrate three complementary prongs of the relatively new field of Language Analysis for Determination of Origin (LADO) : predictive, forensic, and asylum. The first prong introduces a Bayesian localization model that uses my crowdsourced big data corpus to predict a speaker’s regional background from linguistic survey responses with surprising accuracy, refining prior methods through probabilistic weighting of feature co-occurrence. The second prong involves authorship identification in a legal context, illustrated by analysis of a deceased billionaire’s disputed will. The third addresses the use of linguistic evidence to establish whether applicants for political asylum are from where they claim.
🗣️ Prof Kirsty McDougall (University of Cambridge)
📅 16th Oct, 2025 (16:00 - 16:30)
🏫 SG1, Alison Richard Building
Abstract:
In certain crimes, a perpetrator’s voice may have been heard by a witness, but not recorded. If the police have identified a suspect, a phonetician may be asked to prepare a ‘voice parade’ to test whether the witness recognises the voice of the suspect as that of the perpetrator heard at the crime scene. Analogous to a visual identity parade, the witness is presented with a line-up of recordings which includes the suspect’s voice and a number of foil voices.
Selection of the foil voices is a challenging aspect of voice parade construction for which theory is still evolving. The foil voices should sound similar to the suspect’s voice to provide a fair comparison, yet the phonetic underpinnings of perceived voice similarity are not well understood, such that the principles for selecting foil voices are not straightforward. This talk will present some experimental work investigating the phonetic correlates of perceived voice similarity in accents of British English, considering the roles played by aspects of speech such as fundamental frequency, formant frequencies, voice quality and articulation rate. Implications of the findings for voice parade construction will be discussed.