Tuesday, October 27 · 9:00–10:30 a.m.
Ballroom A
Chair: Hillary Fairbanks
Brandon Imstepf · UC Merced
Computational models provide a biophysically grounded framework to study tau protein spread in Alzheimer’s disease and other tauopathies, but parameter inference is often intractable due to the slow, computationally expensive simulations. Here, we introduce a machine-learning framework to accelerate the Network Transport Model (NTM), which models tau spread on the brain’s structural connectome. We replace the costly graph-edge PDE solver with three surrogate families: linear regression, symbolic regression, and multilayer perceptrons (MLP). The MLP achieves R² > 0.99 against numerical solvers, while symbolic regression yields analytical relationships between tau spread and kinetic parameters. These approaches reduce simulation time from hours to minutes (~60×), enabling the thousands of forward simulations required for data-driven inference.
Wei Kang · Naval Postgraduate School
Data assimilation combined with forecast models is a standard approach in geoscience, including numerical weather prediction. This study investigates direct data-driven forecasting of dynamical systems from raw observations, without conventional data assimilation. A central challenge is explainability: without physics-based analyses as a reference, data-driven forecasts can behave as black boxes, making their accuracy and reliability difficult to assess. We introduce a quantitative framework for evaluating the predictability of dynamical systems. Using PDE examples, we demonstrate how the framework can identify observations with low predictive value and filter them before training, enabling neural networks to focus on more informative inputs and improve forecasting performance. For high-dimensional systems, we further show that the computational cost can be reduced by projecting the predictability analysis onto a low-dimensional effective subspace.
Hongyun Wang · UC Santa Cruz
Diffusion models are generative models for sampling a distribution represented by a data set. In forward diffusion, data points are evolved to pure noise by gradually adding noise, and the results are used to train a neural network to emulate the score. In reverse diffusion, the trained score is used to evolve pure noise to samples of the desired distribution. In this study, we explore a toy problem in which the true distribution is uniform along the unit circle and is represented by a set of data points. A meaningful and desired learning is to recover the whole circle implicitly represented by the data, not just the given data points. We find that when we train the diffusion model sufficiently, it samples only the given data points. In contrast, when we reduce the training epochs by four, it samples the desired uniform distribution along the circle. This simple example demonstrates that in some situations desired learning may require insufficient training.
Bhargav Sriram Siddani · Lawrence Berkeley National Laboratory
Coarse-grained representations of particle systems based on stochastic partial differential equations (SPDEs) offer a computationally efficient alternative to particle-based simulations. For example, the regularized Dean-Kawasaki (DK) SPDE describes the density evolution of Brownian particles. However, it cannot accurately capture the non-Markovian and non-Gaussian effects that arise at short timescales and when the number of particles per grid is low. We address these limitations by developing a flow-matching-based generative modeling framework that learns the conditional distribution of particle fluxes, enabling accurate predictions of short-time density evolution. To make the model robust, we incorporate statistical symmetry into the flow matching model to ensure that ensemble-averaged symmetry is preserved. We demonstrate the advantages of our proposed framework across a range of system configurations and compare its performance with that of the regularized DK equation.
Siyuan Xing · California Polytechnic State University
Data-driven discovery of governing equations provides a way to understand complex dynamical systems when first-principles modeling is difficult or incomplete. However, most existing methods are limited to low-dimensional systems with time-invariant governing laws. The Sparse Regression Embedded Interpretable Network (SREINet) addresses these limitations by embedding sparse regression into an interpretable neural architecture and periodically pruning redundant terms. Without explicitly constructing combinatorially large candidate libraries, it enables equation discovery for nonlinear systems with over one hundred state variables. The discovered models reproduce coherent structures beyond the training data and are validated using experimental triple-pendulum measurements. An extension, H-SREINet, combines SREINet with a hypernetwork to identify systems with time-varying coefficients or switching structures. The resulting models can be converted into explicit equations and support extrapolation and tipping-point prediction.