September 11th
Title: Self-supervised In-context Operator Learning for Stochastic Mean-Field Control
Abstract
Stochastic mean-field control coordinates large populations of interacting agents under noise, from swarm planning and systemic risk to Schrödinger bridges. Classical solvers discretize the Fokker-Planck equation on a mesh, so their cost grows exponentially with the state dimension and a tensor-product grid is already out of reach past two dimensions. Deep neural network solvers lift that ceiling but still handle one instance per run, re-optimizing whenever the initial law, the target, or a cost weight changes. This talk presents a mesh-free solution operator for a whole family of such problems, in which a task enters as a prompt and the optimal transport map comes out in one forward pass. We pass to the probability-flow ODE and parameterize its flow map by a prompt-conditioned coupling flow whose exact inverse and analytic log-determinant give the score at linear cost in the dimension, against the cubic cost of a generic map.
September 18th
Title: Training-Free Universal Approximation by Prompting Random Transformers
Abstract
How expressive is prompting a transformer? Answering this question is important for separating the roles of prompting, architecture, and pretraining in transformer models, and for determining whether task-specific behavior must be stored in model weights or can instead be induced at inference time through the prompt. We show, in an approximation-theoretic sense, that pretraining is not strictly necessary: a single-layer softmax attention network with random, untrained weights can approximate any Hölder function on a compact manifold when steered by an appropriate soft prompt. Guided by the connection between softmax attention and kernel methods, we construct explicit soft prompts—a prompt per target function, independent of the query—as solutions to linear systems matching attention logits to Gaussian kernel exponents, under which the frozen transformer emulates the classical Nadaraya-Watson kernel estimator. The construction requires only a mild rank condition on the weights, which we show holds almost surely under Gaussian initialization. The prompted network inherits the theoretical guarantees of kernel regression, leading to universal approximation theorems with minimax-optimal rates that depend on the intrinsic dimension. We further quantify the cost of prompting, exposing a tradeoff between the norm of the constructed soft prompt tokens, prompt length, and hidden dimension.
September 25th
Title: Outpainting: spatially extending aero-optic phase screens
Abstract
Aero-optic effects distort light wave propagation near a high-speed aircraft, thereby degrading performance in airborne imaging and communication systems. Measuring aero-optic data through experiment is costly and the resulting data often has a limited spatial size. Further, alternative methods for simulating this data, including computational fluid dynamics and conventional phase screen generation algorithms (e.g., boiling flow), face drawbacks such as large computation time or inaccurate statistics. More recently, data-driven algorithms have been proposed that can synthesize data that matches relevant statistics of measured aero-optic data. However, these methods cannot spatially extend aero-optic data. In this paper, we introduce ReVAR-ext (Re-whitened Vector AutoRegression-extender), an algorithm that builds on an existing data-driven approach, ReVAR, to spatially extend measured aero-optic data (a process called outpainting) and match the spatial and temporal correlations of the measured data. ReVAR-ext generalizes the generation process of ReVAR by combining multiple sets of synthetic data with the input measured data. This approach generates multiple fixed-sized synthetic images, each of which overlaps with the input data, and then stitches them together. When paired with ReVAR, the ReVAR-ext algorithm can generate aero-optic data with arbitrary temporal duration and arbitrary spatial size. Our experiments show that extended data generated by ReVAR-ext closely matches the temporal power spectrum of two measured aero-optic data sets. Further, the extended data approximately matches the spatial autocorrelation, with reduced accuracy at large spatial lags and at vertical lags.
October 2nd
Title: Training Dynamics of Transformers through Data-Driven Nonlocal Mean Field Control
Abstract
In this paper, we study the soft-attention dynamics through a mean-field control perspective. We first establish a rigorous mean-field limit for the particle system associated with the dynamics of soft-attention residual networks (SAResNet), showing that, as the number of particles tends to infinity, the empirical measure converges to a probability measure evolving under a continuity equation with an attention-induced velocity field. We then formulate training the SAResNet into a soft constrained mean-field control problem, for which we establish the existence of optimizers in admissible class. We derive the first order optimality condition for the soft-constrained control problem, thereby characterize the necessary conditions that are satisfied by the minimizers of the optimal training dynamics. We propose an auxiliary-control algorithm based on the Pontryagin’s maximum principle for training SAResNet. Finally, numerical experiments demonstrate effective learning and favorable prediction accuracy comparing to direct gradient training.
October 9th
Title: A computational framework for matrix-free second-order optimization
Abstract
Iterative methods for large-scale smooth optimization rarely rely on exact second-order information, since the computation, storage, and manipulation of the Hessian matrix are often infeasible. I will describe a framework for deriving and representing first- and second-order information in analytic form. Rather than introducing coordinates and computing large collections of partial derivatives, we show how gradients, bilinear Hessians, and Hessian operators can be obtained directly from Taylor expansions, at a computational cost comparable to a single functional or gradient evaluation. We demonstrate numerically that this framework can be used both to accelerate established first-order methods and to implement fully second-order methods without explicit knowledge of the Hessian matrix. In particular, it alleviates the need for costly line searches in methods such as Conjugate Gradient and BFGS, and provides fast hyper-parameter-free variants of these algorithms.
If time allows, I will also talk about a recent specific application: X-ray nano-holotomography provides fast phase contrast imaging of nanoscale structures in three dimensions, and typically involves reconstruction of data cubes with 4000^3 unknowns. Processing of such data is time-consuming and traditionally done in several separate steps. We recently showed that it is possible to pose the inverse problem as one large joint minimization problem, and computationally feasible to solve it with above computational framework. Applied to an atomic layer deposition pattern and mouse brain tissue, the approach improves sharpness and contrast, revealing neuronal structures that were not visible with traditional reconstruction techniques.
October 16th
Title: Estimating Wavefront Tip/Tilt Using Tomography
Abstract
Isolating the wavefront error caused by aerodynamic turbulence is often difficult due to mechanical vibrations that can contaminate the low frequency bands of the measured data. A common strategy for mitigating this contamination is simply to remove the tip/tilt component from each wavefront measurement. Unfortunately, this has the effect of a high-pass filter on the wavefront data, which as a consequence shifts the estimated temporal spectral energy peak to a higher frequency than what is physically true. In this work, we propose a single-shot tomographic approach that estimates the unknown aero-optical tip/tilt for each frame from a bundle of separate tip/tilt removed wavefront measurements. Unlike other tip/tilt estimation techniques, our approach does make any temporal assumptions, and so can be applied to non-convecting turbulence.