Global Attention with LineAr Complexity for Exascale Generative Data Assimilation in Earth System Prediction
Sponsor: U.S. Department of Energy - Office of Science
Sponsor: U.S. Department of Energy - Office of Science
Accurate weather and climate prediction relies on data assimilation (DA), which estimates the Earth system state by integrating observations with models. While exascale computing has significantly advanced earth simulation, scalable and accurate inference of the Earth system state remains a fundamental bottleneck, limiting uncertainty quantification and prediction of extreme events. We introduce a unified one-stage generative DA framework that reformulates assimilation as Bayesian posterior sampling, replacing the conventional forecast–update cycle with compute-dense, GPU-efficient inference. At the core is STORM, a novel spatiotemporal transformer with a global attention linear-complexity scaling algorithm that breaks the quadratic attention barrier. On 32,768 GPUs of the Frontier supercomputer, our method achieves 63% strong scaling efficiency and 1.6 ExaFLOP sustained performance. We further scale to 20 billion spatiotemporal tokens, enabling km-scale global modeling over 177k temporal frames, regimes previously unreachable, establishing a new paradigm for Earth system prediction.
(a) Overview of one-stage DA workflow. (b) Overview of STORM architecture. Historical states are compressed into a global temporal representation, while the current state remains at full resolution. Noise-gated spatial and temporal attention decouple space–time interactions, reducing complexity from O(K^2N + KN^2) O(N^2) while preserving global correlations.
(a) Tiling achieves linear complexity but limits interactions to local regions, while halo overlap improves continuity but remains local. STORM averages denoised outputs (gradients) in overlapping regions and propagates them iteratively across tiles, enabling global context with linear complexity. Hanning weighting stabilizes boundary interactions. (b) Hierarchical mapping of parallelism strategies to supercomputer hardware, with communication frequencies shown on the right.
Trade-off between spatial resolution and temporal horizon under fixed compute budgets. Each curve represents the achievable spatiotemporal configurations for different model sizes and GPU counts. STORM+GALA expands the feasible frontier, enabling simultaneous high-resolution (km-scale) modeling and long temporal horizons. The highlighted run demonstrates a new capability regime, reaching tens of billions of spatiotemporal tokens and bridging weather-scale forecasting with climate-scale simulation.
Strong scaling efficiencies across various model sizes, scaling to 74,400 GPUs with 96% to 99% strong scaling efficiencies with up to 6 exaFLOP sustained computing throughput at BF16 precision.
Our diffusion-based DA framework provides a stable-in-time estimation of hurricane track and intensity, without the need for a computationally expensive physical model. The forecast–analysis cycles are carried out seamlessly within a unified, data-driven framework that avoids expensive I/O operations for updating the model state.
Delta
Laura
Michael
Teddy
Hurricane track skill: Ensemble trajectories of the four simulated hurricanes, computed from the maximum surface wind speed in each ensemble member. The black tracks correspond to STORM forecast-only , whereas the blue tracks are associated with a data assimilation experiment in which 20% of the ERA5 variables are observed. The reference hurricane tracks (red curves) are obtained from the Best Track dataset provided by National Hurricane Center.
Reference - ERA5 Model Prediction Data Assimilation Enhanced Prediction
Reference - ERA5 Model Prediction Data Assimilation Enhanced Prediction
Reference - ERA5 Model Prediction Data Assimilation Enhanced Prediction
Reference - ERA5 Model Prediction Data Assimilation Enhanced Prediction
Regional high-resolution temperature simulations of 2-m temperature every 6 hours. Similar to the global case, regional systems such as HRRR are limited by small ensemble sizes due to the high cost of physics-based simulations, restricting their ability to capture uncertainty and nonlinear error growth. STORM relaxes this constraint by enabling large-ensemble, long-context inference within a unified framework. Prediction with DA effectively correct errors by incorporating observational constraints. The assimilated fields closely match the ground truth, and regional biases are largely eliminated.
We evaluate STORM in a reanalysis-style benchmark for global 2-m temperature at 1.0° resolution. Regional RMSE is evaluated over Africa, South America, and South Asia (red boxes). Blue boxes denote Australia and Siberia used for the extreme-event evaluation.
Weekly 2-m temperature RMSE over the held-out 2018–2020 testing period for four input/output configurations across Africa, South America, and South Asia. The 2yr-input-4wk-output model consistently achieves the lowest RMSE, demonstrating that multi-year temporal context improves the forecast used by STORM for posterior inference.
Regional weekly 2-m temperature RMSE for the STORM model using two years of historical context to predict the held-out 2018–2020 period. Black: forward simulation; colored curves: STORM posterior estimates using 10% and 30% of the observations. RMSE decreases systematically with increased observational coverage across all regions, demonstrating accurate reconstruction from incomplete observations.
Regional weekly 2-m temperature during the 2019–2020 Australian Black Summer heatwave (top) and prolonged Siberian warmth in 2020 (bottom). The forward simulation largely follows climatological behavior and underestimates the extreme temperature evolution, whereas STORM progressively shifts the posterior toward the ERA5 reference as observational coverage increases. Shaded bands denote ensemble uncertainty, which narrows as observations constrain the inferred climate state.
Oak Ridge National Laboratory, Oak Ridge, TN, USA
Xiao Wang
Isaac Lyngaas
Hong-Jun Yoon
Jong-Youl Choi
Siming Liang
Janet Wang
Dan Lu
Guannan Zhang (corresponding author)
Auburn University, Auburn, AL, USA
Zezhong Zhang
Florida State University, Tallahassee, FL, USA
Hristo G. Chipilski
Feng Bao
AMD Research and Advanced Development, Santa Clara, CA, USA
Ashwin M. Aji
Colorado State University, Fort Collins, CO, USA
Peter Jan van Leeuwen