Mark Goldstein
Diffusion models for inverse stellarator design
Using the ARTEMISS fast-ion stellarator design dataset, this talk studies the inverse problem of stellarator design. Stellarator design is usually posed as expensive PDE-constrained optimization: search over boundary shapes until the resulting equilibrium has the desired properties. We flip this around and pose it as conditional generative modeling: learn to sample plasma boundary shapes given a target set of plasma and device properties, so that candidate designs come out as generative model samples rather than an optimization run. The first half of the talk introduces the models themselves, presenting diffusion models from the viewpoint of stochastic interpolants, a construction that turns generative modeling into learning a velocity field for an ODE. We then walk through the design choices for the stellarator problem: how to parameterize the boundary surface, what to condition on (confinement, pressure, etc...), and how to validate generated designs by re-solving the MHD equilibrium with VMEC. We close with results so far and open questions.
Leonardo Zepeda-Núñez
Latent generative models for high-dimensional chaotic dynamical systems
Latent modeling is a ubiquitous component in modern machine learning, offering a powerful framework to factorize complex tasks and simplify modeling approaches. At the same time, generative AI continues to drive breakthroughs in many fields by enabling the efficient modeling of high-dimensional distributions.
In this talk, I will present how latent generative models, which unite these two frameworks, provide a powerful paradigm for tackling challenging problems in science and engineering. Using climate downscaling and computational fluid dynamics (CFD) as key examples, I will show how this paradigm unlocks unprecedented capabilities. These include generating statistically accurate tropical cyclones, even when such extreme events are absent from the low-resolution climate data, and accurately describing the conditional, time-dependent distribution of three-dimensional fluids with diverse physics and geometries.
I will present the overarching framework, its instantiation, and the scientific rationale. I will also discuss how these new methods have the potential to revolutionize climate risk assessment and engineering tasks such as uncertainty quantification and optimization under uncertainty.
Ionut Farcas
Learning parametric reduced models of plasma micro-instabilities with sparse grids and optimized dynamic mode decomposition
Parametric data-driven reduced-order models (ROMs) that embed dependencies in many input parameters are essential for enabling many-query tasks in large-scale problems, and are key to the development of digital twins. However, generating training data to construct the ROMs using standard grid-based approaches is computationally infeasible due to the curse of dimensionality. This presentation introduces an efficient strategy for constructing parametric data-driven ROMs by leveraging sparse grid interpolation with (L)-Leja points. These points are nested and exhibit slow growth, resulting in sparse grids with low cardinality in low- to medium-dimensional settings, making them well-suited for computationally expensive, large-scale problems. As a representative real-world application, we consider gyrokinetic simulations of plasma micro-instabilities in fusion experiments. We construct parametric ROMs for the full five-dimensional gyrokinetic distribution function using optimized dynamic mode decomposition (optDMD) in combination with sparse grids. For an electron-temperature-gradient-driven micro-instability simulation with six input parameters, we demonstrate that a predictive parametric optDMD ROM can be built using only 28 high-fidelity simulations, while achieving evaluation costs up to three orders of magnitude lower. These results highlight the potential of sparse grid-based parametric ROMs in enabling otherwise intractable many-query tasks in large-scale applications.
Misha Khodak
Breakeven complexity: When do neural surrogates pay off?
Neural PDE solvers promise dramatic speedups over classical methods once trained, but their value is clouded by the cost of generating the training data, which requires running the very classical solvers they aim to replace, and that cost grows precisely on the hard problems we care about. To study this, we develop breakeven complexity, a cost-aware metric that reframes evaluation around amortization: asking not "how accurate is the solver?" but "how many solves are needed for its training cost to pay off?" By combining classical convergence tests with neural scaling laws, we are able to evaluate breakeven complexity on eight models across four PDE settings, and find a counterintuitive result: as problems get harder (longer rollouts, higher dimensions, more complex physics), breakeven complexity shrinks even as test error grows, suggesting that neural surrogates will be most impactful exactly where simulation is hardest, from turbulent flows to fusion plasmas.
Sam Stechmann
Element learning: accelerating finite element-type methods via machine learning
To speed up computations, artificial neural networks and machine learning tools have surfaced as game-changing technologies. However, many machine learning approaches tend to lose some of the advantageous features of traditional numerical PDE methods, such as interpretability. In this talk, we introduce a systematic approach (which we call element learning) with the goal of accelerating finite element-type methods via machine learning, while also retaining the desirable features of finite element methods. Numerical examples illustrate a computational speed-up factor of 5 to 20.
Andrej Risteski
Inference-Time Algorithms: A Theoretical Lens on Tractability and Error Propagation
Modern AI systems are increasingly built by placing trained models inside larger computational loops. Inference-time algorithms are a basic instance of this idea: they use one or more trained models at test time to incorporate new information, exploit pretrained models as priors, and trade computational effort for accuracy, sample quality, or control. Examples include generator-verifier search for reasoning, diffusion models for solving inverse problems, and reward-guided generation. Theoretically, this revisits a classical question from optimization and theoretical computer science: what can be done with access to an oracle? Here, however, the oracles are new and non-standard: they model the capabilities of large pretrained models, making them powerful, but also imperfect because they are learned. This combination leads to new questions about algorithm design and error propagation.
This talk studies two central aspects of this paradigm: computational efficiency and error propagation. The first vignette considers generator-verifier systems, and shows how stochastic backtracking can trade additional computation for accuracy, giving a principled version of test-time scaling even with imperfect learned oracles. The second vignette studies diffusion steering: when can we efficiently bias a pretrained diffusion model toward higher-reward samples while staying close to the original model? We show that tractability depends strongly on both the reward structure and the alignment objective, and that simple primitives—such as sampling from linear tilts—can be surprisingly useful for handling richer reward classes.
Philip Morrison
On the Hamiltonian Structure of the Low Rank Method for Vlasov-Poisson
What it means to be Hamiltonian in a general context will be reviewed. Then, a bracket for the low rank method will be described. Finally some comments on its influence on compuations. Joint work with Lukas Einkemmer.
Haizhao Yang
Machine Scales, Human Steers: Agon for Large-Scale Autonomous Research
In this talk, I will present Agon, an autonomous large-scale omnidisciplinary research system built on Prompt Economy. Agon organizes research into reusable loops that generate, refine, test, audit, and write scientific artifacts across domains, using minimal prompts, massive parallelism, and zero-code orchestration. Across hundreds of iterations and multiple scientific deployments, Agon demonstrates that research production can be substantially scaled by machines. At the same time, its failures reveal a sharp boundary: current systems can automate many checks, but human scientists remain essential for scientific taste, anomaly detection, and final judgment. The central message is simple: machine scales, human steers.
Andrew Christlieb
Closures that enable multi-scale effects
The curse of dimensionality is a fundamental challenge in modeling multiscale systems. Reductions from the BBGKY hierarchy to kinetic descriptions, and from kinetic models to fluid equations, make complex simulations computationally tractable. Similar reductions underlie turbulence models. At each level, however, kinetic, fluid, or turbulence closures improve computational efficiency by eliminating degrees of freedom—and, with them, potentially important small-scale physics.
In many regimes, these reductions provide accurate and useful models. They can fail, however, when a system evolves too rapidly for local thermodynamic equilibrium to remain a valid approximation. In strongly nonequilibrium systems, including inertial-confinement-fusion plasmas, the path by which the system evolves can be as important as its instantaneous macroscopic state. This history dependence, or hysteresis, cannot generally be represented by conventional closures based only on local equilibrium quantities.
We therefore seek reduced models with closures that retain essential multiscale and history-dependent effects while preserving critical mathematical structure, including entropy consistency and the conservation of macroscopic quantities. In this talk, we review recent work on incorporating machine-learned closures into partial differential equation models to recover otherwise unresolved physics. We consider two examples: a kinetic closure that captures two- and three-body collisional effects, and a fluid closure that enables moment models to represent weakly collisional dynamics. Together, these examples illustrate how mathematically constrained learning can extend the range of validity of reduced models without restoring the full complexity of the underlying high-dimensional system.
Wenlong Mou
Kinetic Plasma Control from Sparse Diagnostics: Provable Imitation Learning for Vlasov–Poisson Instabilities
Controlling kinetic instabilities from limited diagnostics is a fundamental challenge for nuclear fusion. In Vlasov--Poisson dynamics, kinetic instabilities arise from unresolved velocity-space structure, while feedback policies must act from sparse macroscopic observations. Imitation learning offers a way to distill kinetic control strategies into sensor-based policies, but raises a basic question: when does predictive accuracy on simulated trajectories translate into closed-loop stability?
In this talk, we develop a theoretical framework for offline and online imitation learning in partially observed kinetic systems. For behavior cloning, we characterize learnability through sensor resolution, measurement noise, and the entropy of the initial-condition ensemble, revealing adaptivity to low-complexity kinetic structure. We then prove that arbitrarily small offline imitation error can excite an unstable mode absent from expert data and grow exponentially after deployment. In contrast, controlling imitation error on learner-induced trajectories yields polynomial error growth over polynomial time horizons, motivating a DAgger-style algorithm.
Simulations of a 1D1V two-stream instability show that history-dependent neural controllers using sparse density measurements delay electric-field growth and phase-space distortion, while online data aggregation improves long-horizon control relative to additional offline data. I will conclude with ongoing work on generative models for kinetic simulation, aimed at learning the simulation environments used to train such controllers. Together, these results connect statistical learning theory with the structure of kinetic plasma dynamics. Joint work with Xiaofan Xia and Qin Li.
Qi Tang
Structure-Preserving Neural Operators for Convection–Diffusion and Kinetic Plasma Models
Neural-operator surrogates promise fast forward solves for outer-loop plasma tasks, but operators such as the Fourier Neural Operator (FNO) suffer from spectral bias on convection-dominated dynamics and cannot simultaneously capture hyperbolic transport and parabolic smoothing, so their rollout error grows exponentially with the prediction horizon. We introduce a structure-preserving neural operator for convection-diffusion that combines Strang splitting, a learnable semi-Lagrangian transport module, and a learnable diffusion module that treats the mean diffusivity exactly through the heat semigroup. We prove rollout error bounds that grow only linearly in time, so a single model trained on coupled data generalizes to the pure-convection and pure-diffusion limits. On the 1D1V Vlasov-Poisson-Fokker-Planck system, it sustains about 2.5% relative error over long rollouts across collisional, weakly collisional, and collisionless regimes while capturing filamentation and phase-space vortices. We then extend the same idea to the hyperbolic conservation laws underlying plasma fluid and MHD modeling through the Local-Global Neural Operator (LGNO), which pairs a global FNO branch for smooth large-scale dynamics with a local multiresolution branch for shocks and discontinuities. Across 1D and 2D benchmarks, LGNO reduces one-step errors by 2 to 5x over FNO baselines, and, though trained only on short-time WENO-Z data, its coarse-grid rollout exhibits lower numerical dissipation than WENO-Z on a finer grid at lower cost. Together these results show that embedding physical structure into neural operators is essential for stable, accurate, long-horizon surrogates for plasma modeling.
Andrew Giuliani
Optimizing stellarators at the core and edge
Recent stellarator design efforts have largely focused on optimizing the plasma core's behavior, resulting in a moderate-dimensional design problem suited to gradient-based optimization algorithms. These successful efforts have produced several data sets, including the QUAsisymmetric Stellarator Repository (QUASR), https://quasr.flatironinstitute.org/, which have sparked further data-driven discoveries and the STAR_Lite experiment, https://www.hufusion.org/. Lately, attention has shifted outward to shaping the magnetic topology beyond the confined region, specifically at the plasma edge. In this talk, I will survey state-of-the-art stellarator design algorithms, a new data set that expands QUASR, and future directions.
Misha Padidar
Progress on the Fast Ion Database
Minimizing fast ion losses is a primary goal of stellarator design. However, key performance metrics derived from fast ion simulations are costly to include directly in the design optimization loop, and almost exclusively used as diagnostics. Machine learning (ML) has the potential to mitigate this problem, by providing surrogates and generative procedures for rapidly designing stellarator configurations with good fast ion confinement. To enable ML approaches to fast ion optimization, we introduce a dataset containing over 100,000 distinct magnetohydrodynamic equilibria spanning a wide range of magnetic geometries and confinement properties. In this talk we will discuss the construction of the database, and showcase ML models trained on the data.
Rogerio Jorge
Driftless Star and the UWPlasma Ecosystem: End-to-End Differentiable Simulation for Stellarator Design
Designing a stellarator means searching a design space of hundreds of shape parameters, where every candidate must be evaluated through a chain of expensive physics: MHD equilibrium, fast-ion orbits, neoclassical and turbulent transport, and profile evolution. We present the UWPlasma ecosystem, an open-source stack of JAX-native plasma codes — VMEX (3D equilibria), ESSOS (coils and fast ions), GKX (gyrokinetic turbulence), DKX (drift-kinetic transport), NEOPAX (profile evolution), and DRBX (edge turbulence), sharing the differentiable linear-solver backbone SOLVAX — in which the entire chain is exactly differentiable by implicit and adjoint methods at the cost of roughly one extra solve. We show how automatic differentiation enables single-stage optimization to run efficiently, and turn each solver into a layer for machine learning: surrogate training through the physics, gradient-based inverse problems, and uncertainty propagation. These codes power Driftless Star, an open, containerized pipeline for transport-consistent stellarator optimization, validated against Trinity3D+GX on W7-X.
Wei Zhu
Data-Driven Discovery of Conservation Laws, Lax Pairs, and Hidden Structure in Dynamical Systems
In this talk, I will present data-driven approaches for discovering hidden structure in dynamical systems, with a focus on conservation laws and Lax-pair representations. The first approach, neural deflation, iteratively learns functionally independent conserved quantities from trajectory data. The second approach learns interpretable Lax operators by enforcing compatibility with the underlying Hamiltonian dynamics.
These methods provide a way to uncover invariants and integrability-related structures that may not be known a priori. I will illustrate the approaches on several Hamiltonian systems and briefly discuss their potential relevance to structure-preserving reduced-order modeling.