Fall 2026
Title: When can we predict model behavior?
Abstract: We understand phenomena scientifically only if we can make correct testable predictions. As one example from the empirical science of deep learning, scaling laws predict language modeling loss from training conditions. What else can we predict about language models? And what can’t we predict? One key distinction is between predictable continuous variables and unstable discrete conceptual categories, which are notably nondeterministic. This creates a tension: humans naturally reason in categories, so the phenomena we most want to understand are the hardest to predict. This talk surveys what we know of categorical learning in language models, showing how discrete concepts are revealed through empirical training dynamics and through random variation across training runs. These concepts involve syntax learning, weight mechanisms, and interpretable patterns, all of which can predict model behavior. By leveraging categorical learning, we can ultimately understand a model's natural conceptual structure and evaluate our understanding through testable predictions—including predictions of model failures we haven't yet observed.
Title: Exploring fake modality in conditional antecedents
[joint work with Stefan Kaufmann]
Abstract: Conditionals can serve to characterize possibly non-actual states of affairs. The recent literature adopts the term X-marking (von Fintel & Iatridou 2023) for linguistic forms that serve to indicate remoteness from the actual state of affairs (including genuinely counterfactual ones). Particular attention has been paid to forms that ordinarily mark tense and aspect, but appear repurposed as X-markers in conditionals, as for instance “fake past” in “If I arrived tomorrow, I’d miss the talk." The question whether such forms retain their regular temporal and aspectual meaning in such contexts has been debated extensively (“past-as-past” vs. “past-as-modal”). Building on observations with different types of fake past from previous work, I will discuss English “were to” in conditional antecedents and select other contexts, where it appears to lack the deontic or teleological meaning familiar from indicative “is to”. Remoteness marking “were to” in conditional antecedents can thus be classified as “fake modality”. Having put in place a detailed account of “is to”, I will argue that the seemingly non-modal occurrences are best analyzed as retaining the modal meaning, and that this correctly predicts restrictions that distinguish “were to” antecedents from regular simple past X-marking. I will briefly compare the behavior of “were to” with similar effects found for English “should” and related modals in German and Italian.