Fall 2026
Time: Fridays, 2pm-3pm, Pacific time
Location: Hybrid - South Hall 4607 and in Zoom (link provided upon request)
Please contact Mingsong Yan (mingsongyan@ucsb.edu), Qirui Peng (qpeng9@ucsb.edu), Ruimeng Hu (rhu@ucsb.edu), or Sui Tang (suitang@ucsb.edu) to reserve a slot.
Upcoming Seminar Schedule:
(Click the event below to see the title and abstract)
Title: Atomic Gradient Flows: Gradient Flows on Sparse Representations
Abstract: One of the most popular approaches for TV-regularized optimization problems in the space of measures is the so-called Particle Gradient Flow. For this, one restricts to linear combinations of Dirac deltas and then takes a Euclidean gradient flow in the weights and positions, significantly simplifying computations. Recent results have shown that PGFs recover Wasserstein gradient flow dynamics, and can even converge to global minima under sufficient conditions. In this talk, I present a generalization of PGFs to regularized optimization problems on arbitrary Banach spaces, which we call Atomic Gradient Flow (AGF). The crucial idea is the choice of the right notion of particles, which we argue to be the extremal points of the unit ball of the regularizer. Using Choquet's theorem, we lift the problem into the Wasserstein space of both weights and extremal points, and study convexity, existence, and uniqueness properties for the AGF and metric gradient flows in the lifted setting. Our main result is that the lifting of the AGF is again a metric gradient flow in the Wasserstein space, implying that the AGF follows a very strong dynamic. Lastly, I will also showcase examples, applications and some numerics. This is joint work with Marcello Carioni and Konstantinos Zemas.
Host: Katy Craig
Title: Transformers in the Long-Context Regime
Abstract: Many emerging uses of AI, from analyzing scientific literature to developing complex software, require models to work with large amounts of information. Yet the ability to accept a long input does not explain how effectively a model uses it. What mathematical principles determine the capabilities and limitations of large language models as their context grows? In this talk, I will explore this question through mathematical models of transformers, the architecture underlying many language models, with self-attention at its core. I will first study context-length-dependent temperature scaling, a practical technique used in serveral language models to improve retrieval from long inputs. Drawing on extreme-value theory, we provide a mathematical explanation for this largely empirical strategy by identifying the critical scaling that allows multiple strong candidates to influence the output, even as the context grows. I will then turn to in-context learning, the remarkable ability of an LLM to learn from examples in inputs without changing its parameters. A longer context can contain more examples and therefore more information, making it important to understand how rapidly a model can benefit from this growing body of evidence. Using mean-field analysis, we quantify how quickly learning from a finite context approaches learning from the underlying true population. I will conclude by placing these results within my broader research on mathematical machine learning, highlighting opportunities for applied mathematics to guide the design of reliable learning algorithms.
Host: Xu Yang
Title: Hack's Law, Erosion, and Optimal Transport: Developing Mathematical Models and Computational Methods
Abstract: Hack's Law says that the length of the main river in a river basin scales with the area of the river basin to the power 0.58, thereby prompting speculation on whether there exist optimality principles governing natural processes like channel formation. Identification of such principles and their continued mathematical development could lead to the formulation of a theoretical framework which supports study not only of fluvial landscape evolution, but other geomorphological processes as well. Erosion is nature's process of ``moving dirt." In 1781, Monge initiated the theory of optimal transport with the question, ``Given a pile of sand and a pit of equal volume, how can one optimally transport the sand into the pit?" Building on earlier work by Birnir, Merchant, Smith, Cattan, and Rowlett, we aim to extend optimal transport theory to develop mathematical models and computational methods for studying erosion. Via optimal transport theory, we prove the existence of unique global weak solutions to equations describing the sediment flow in the evolution of fluvial land surfaces, with constant water depth. While an earlier existence theory result by Birnir and Landry demonstrates that the slopes, or gradients of the surface, are moved optimally, our current work proves that erosion does what Monge proposed: move sediment optimally from the mountainside to the river. This is joint work with Bjorn Birnir.