ML-Based GPU Resource Prediction for Co-Scheduling on a Production HPC System
Speaker: Beste Oztop, Boston University
Co-Authors: Dhruva Kulkarni (NERSC), Zhengji Zhao (NERSC), Ayse K. Coskun (BU), Kadidia Konate (NERSC)
Abstract: Efficient GPU utilization is critical to throughput on production HPC systems, yet jobs frequently underutilize the GPUs they are allocated. Because workload managers allocate whole GPUs throughout a job execution, devices that are allocated but never exercised still register as fully utilized, so on a heavily subscribed system, this idle capacity is invisible to the scheduler. We characterize GPU utilization across 114,195 jobs on NERSC's Perlmutter using per-GPU DCGM telemetry, and find that 20% of jobs peak below 50% in both compute and memory utilization. Exploiting this capacity requires knowing resource demand before allocation, which runtime telemetry cannot supply. We therefore predict a job's maximum GPU compute utilization, maximum memory utilization, and effective GPU count, which is the number of devices it genuinely utilizes, from submission fields alone. Our framework identifies 458,808 GPU-hours of co-scheduling headroom over 17 days, 82.4% of the oracle; with application input parameters, accuracy reaches 94.8%.
Performance Analysis of Phaseless Auxiliary Field Quantum Monte Carlo
Speaker: Namita Shah, University of Michigan
Co-Authors: Neil Mehta, Ermal Rrapaj, Katie Klymko, Pooja Rao, Nick Joyner, Lisa Claus, Charles Lively
Abstract: Drug discovery requires accurate prediction of molecular ground-state energies, but exact electronic-structure calculations become exponentially expensive as molecules grow. We investigate how the choice of quantum method affects the accuracy and computational cost of a hybrid quantum-HPC workflow on NERSC’s Perlmutter system. Our workflow combines molecular Hamiltonian construction, quantum circuit generation, and ph-AFQMC, a Monte Carlo method that uses a trial wavefunction to guide random walks toward the ground state. We compare two approaches for generating this trial wavefunction: Variational Quantum Eigensolver (VQE) and Generative Quantum Eigensolver (GQE), developed by NVIDIA. VQE optimizes parameters in a fixed circuit, while GQE uses a transformer model to generate circuit structures. By benchmarking these approaches on GPU-accelerated HPC workloads, we characterize their computational costs, performance, and resulting accuracy, identifying tradeoffs that determine when each approach is advantageous. This work provides insights into optimizing emerging hybrid quantum-HPC workloads.
Principles of GPU Optimization (or How to Play Tetris with the GPU Memory Hierarchy)
Speaker: LeAnn Lindsey, LBNL
Co-Authors: Afton Geil, Chris Harris, Mario Melara, Doru Thom Popovici, Kjiersten Fagnan
Abstract: Optimizing an algorithm for the GPU requires considering not only the operations it performs, but also how data is represented as it moves through the memory hierarchy. In this talk, I will use our experience porting HMMER3, a widely used bioinformatics application, to GPUs as a case study in practical GPU optimization. We will trace a performance bottleneck to inefficient L1 cache memory accesses and show how restructuring the input data to enable coalesced accesses improved kernel performance by 4.4×. From there, we will explore the GPU memory hierarchy across A100, H200, and B300 architectures and how differences in memory capacity, bandwidth, registers, and shared memory influence algorithm and data-structure design. Finally, I will show how memory constraints can lead to counterintuitive optimization strategies, including checkpointing intermediate dynamic-programming state and repeating computation rather than storing full matrices. Together, these insights offer a roadmap for adapting complex dynamic programming workloads to modern accelerators, demonstrating that sometimes optimization requires looking at a problem from a different perspective.
Challenges and opportunities of large genome foundation models
Speaker: Zhong Wang, LBNL
Abstract: Genome foundation models (gFMs) are reshaping how we read and design biology, but scaling them to environmental genomics exposes new frontiers. We present the GenomeOcean series, trained at NERSC on ~600 Gbp of high-quality contigs assembled from 220 TB of diverse metagenomic datasets spanning oceans, soils, lakes, and host-associated habitats. The series now spans 100M, 500M, and 4B dense parameters, each with quantized deployments for low-footprint inference, and a BGC-finetuned variant that discovers and synthesizes complete biosynthetic gene clusters. To push capacity without proportional compute, we further explored Mixture-of-Experts designs at 8×100M and 8×7B scales. Early results are encouraging: GenomeOcean supports zero-shot and fine-tuned functional annotation, cluster-aware BGC generation, and biosecurity-relevant sequence screening. Yet the road ahead is defined by hard constraints — hardware bottlenecks for training multi-billion-parameter and sparse-MoE gFMs, the absence of standardized, contamination-controlled benchmarks, and open questions around tokenization, evaluation metrics for generative DNA, and safe deployment.
Verifiable PDE Reasoning and Modeling with Neurosymbolics
Speaker: Wuyang Chen, Simon Fraser University
Abstract: Recent progress in Large Language Models (LLMs) has transformed text and code generation, yet models still falter on Partial Differential Equations (PDEs) where correctness, constraints, and physical consequences are critical. This talk explores how formal LLM reasoning can advance symbolic PDE modeling. First, our PDE-Controller formalizes informal PDEs, synthesizes solver-ready code, and plans subgoals to tackle nonconvex control via interactions with external solvers. Second, our Lean Finder accelerates PDE formalization via a semantics-aware search engine for Lean/Mathlib that retrieves relevant theorems, outperforming GPT models and gaining significant traction in the AI-for-math community. Through these efforts, we aim to design a semantics-first LLM that autoformalizes informal PDE problems into machine-checked specifications and synthesizes solver-ready code. This closes the loop between formal analysis and LLM reasoning, ultimately surpassing human heuristics across diverse PDEs.
EveNet: A Foundation Model for Particle Collision Data Analysis
Speaker: Ting-Hsiang Hsu, National Taiwan University
Co-Authors: Bai-Hong Zhou, Ben Nachman, Qibin Liu, Shih-Chieh Hsu, Shu Li, Vinicius Massami Mikuni, Yuan-Tang Chou, Yue Xu, Yulei Zhang
Abstract: Foundation models offer a way to reuse large-scale machine-learning training across many scientific tasks rather than training a separate model for each application. We present EveNet, an event-level foundation model for particle collision data analysis, pretrained on 500 million simulated proton-proton collision events using NERSC GPU resources. EveNet uses a shared transformer-based encoder trained with self-supervised and physics-guided objectives, then adapts efficiently to downstream analyses. We evaluate it on new-particle searches, precision measurements, and anomaly detection, including tests on real CMS Open Data. Across these tasks, EveNet matches or improves on models trained from scratch while requiring substantially less task-specific training and generalizing to unseen signal scenarios. This work demonstrates a reusable scientific-ML workflow in which the cost of large-scale pretraining is amortized across many applications, reducing repeated computation and enabling more efficient use of HPC resources.
Scientific Workflows with Pegasus on NERSC Resources using SFAPI
Speaker: Karan Vahi, USC
Abstract: Workflows are a key technology for enabling complex scientific computations. They capture the interdependencies between processing steps in data analysis and simulation pipelines as well as the mechanisms to execute those steps reliably and efficiently. Workflows can capture complex processes, promote sharing and reuse, and also provide provenance information necessary for the verification of scientific results and scientific reproducibility. This presentation provides an introduction to Pegasus (https://pegasus.isi.edu) and highlights recent enhancements enabling workflow submissions to NERSC's Perlmutter supercomputer using NERSC’s Superfacility API. Through this integration, users can launch Pegasus workflows directly on Perlmutter without needing to install Pegasus and HTCondor locally at NERSC. Instead, users can log into a hosted workflow instance via ACCESS Pegasus (https://pegasus.access-ci.org/) to submit, manage, and debug their workflows using Jupyter notebooks. This hosted environment expands workflow execution options, allowing users to run seamlessly on ACCESS resources and NERSC Perlmutter.
From Small-Scale Dark Matter Physics to Large-Scale Structure: Extreme-Scale Cosmological Simulations
Speaker: Mahesh Natarajan, LBNL
Co-Authors: Jean Sexton, Shamik Ghosh, Zarija Lukic, Andrew Myers, Weiqun Zhang
Abstract: We present an end-to-end HPC workflow bridging small-scale dark matter physics and large-scale structure cosmological simulations. Using NERSC’s Perlmutter (64 A100 GPUs, 4 hrs), we run 1024^3, 20 Mpc/h simulations comparing Cold Dark Matter (CDM), Warm Dark Matter (WDM), and Fuzzy Dark Matter (FDM). We compute matter power spectra, identify halos using Reeber and Rockstar, compute halo mass functions, 1D Lyman alpha flux spectrum, and validate results directly against DESI survey data. To probe cosmic structures, we scale to an 8192^3, 7700 Mpc/h CDM simulation on Frontier (32,768 GPUs, 12 hrs). Raw outputs - lightcones and halos, are staged to Perlmutter for post-processing to compute quantities such as the tSZ spectra. Globus was used to transfer ~60 TB of data over ESnet which reaches speeds of ~20 GB/s. This pipeline demonstrates how NERSC resources unify high-resolution alternative dark matter physics, extreme-scale runs, and mock observation pipelines for modern spectroscopic surveys.
Integrating HPC with Electron Microscopy with Data Streaming
Speaker: Peter Ercius, LBNL
Abstract: Electron microscopy is an important characterization technique in biological and materials sciences. The recent advent of high data rate detectors and automation have drastically increased the amount of data being collected. Traditionally, this data is saved directly to disk and fully analyzed after the experiment. We will show in this presentation the capabilities of a platform comprised of data streaming and containerized data analysis capabilities using modern web-based technologies. This platform leverages the NERSC Superfacility API and greatly improves user experience by integrating microscopy and compute into a complete system.
Towards real-time cryo-EM analysis at BioEPIC by leveraging HPC
Speaker: Matthew David Giammar, UC Berkeley
Co-Authors: Bronwyn Lucas, Agustin Avila Sakar
Abstract: Cryogenic electron microscopy (cryo-EM) captures images of native cellular environments at atomic scale. Microscopes regularly produce >10TB of data per day, and data is typically analyzed over a period of weeks to months, on local workstations, and using fragmented and ad hoc pipelines. This slow turnover is a bottleneck for assessing sample quality and directing experimental efforts. As part of the NERSC Science Acceleration Program for the new Doudna system, we are developing a platform for real-time data analysis and experimental steering in the BioEPIC cryo-EM facility. We present CryoFABRIC, a workflow management system for cryo-EM which enables real-time data analysis while supporting long-term, large-scale analysis for in situ structural biology. Though designed for cryo-EM, CryoFABRIC is built upon an extensible, domain-agnostic, and SQL-based workflow engine (fabric-core) which could prove useful as scientific computing workflows grow more complex and interdisciplinary. CryoFABRIC, along with algorithmic and kernel optimizations, aims to close the gap between data acquisition and analysis positioning BioEPIC to exploit Doudna for real-time cryo-EM analysis at scale.
Reactant.jl: Using compiler technology to bridge the gap between scientific computing and AI
Speaker: Roman Lee, LBNL
Co-Authors: AI and ML are reshaping supercomputing with innovations such as specialized tensor accelerators (e.g., TPUs, GPUs with tensor cores) and frameworks like PyTorch, JAX and TensorFlow. However, scientific applications such as ocean models — often written in Fortran, C++, or Julia and built for traditional HPC — remain largely incompatible with these technologies. This isolates scientific computing from the rapid innovation taking place for AI/ML workloads. In this talk, we discuss an approach to bridging this gap by transpiling a Julia-based ocean model (Oceananigans) using Reactant, an optimizing compiler for Julia built on the MLIR compiler infrastructure. Reactant enables automatic performance portability across CPUs, GPUs, and TPUs (facilitating the use of emerging AI-customized HPC architectures) as well as automatic differentiation (facilitating the use of gradient-based methods). This work opens a path for climate modeling, and scientific modeling applications more broadly, to benefit from the cutting-edge advances in AI/ML-driven hardware, software, and techniques.
QuantumBenchPhase: A quantum simulation and benchmarking library for generating phase diagrams
Speaker: Maggie Bao, The Washington Institute for STEM, Entrepreneurship and Research
Co-Authors: Adam Godel, Adrian Acosta, Connor Howe, Sarah Chehade, Vardaan Sahgal, Joan Etude Arrow, and Brian J. McDermott
Abstract: We present QuantumBenchPhase (QBP), a quantum simulation and benchmarking library for generating phase diagrams. Hamiltonian-centered benchmarking often requires substantial effort across an end-to-end workflow for physical systems and optimization problems. QBP provides a unified, highly modular open-source Python library that configures this workflow through a single command, streamlining the evaluation of classical and quantum algorithms for Hamiltonian simulation. QBP includes built-in models such as the foundational Haldane and Hubbard models, classical and quantum simulation methods including the variational quantum eigensolver (VQE), iterative quantum phase estimation (IQPE), and density matrix renormalization group (DMRG), boundary conditions, multilevel parallelism, and error mitigation. The library can also ingest any problem Hamiltonian from the expansive HamLib dataset. We assess QBP by generating diverse phase diagrams using ideal and noisy simulators along with IQM hardware. Together, these capabilities provide a consistent framework for comparing models, algorithms, observables, and computational backends across diverse physical and optimization applications.
Distributed DMRG on Perlmutter
Speaker: Matthew Blomquist, LBNL
Co-Authors: Matthew Blomquist, Gregor Daiß, Alec Dektor, Erika Ye, Elijah Pelofske, Abhijith Jayakumar, Philip Fackler, Pedro Valero-Lara, William Godoy, Patrick Diehl, Neil Mehta, Katie Klymko, Ermal Rrapaj
Abstract: We present performance results of a real-space parallel density matrix renormalization group (DMRG) algorithm on the Perlmutter supercomputer. Real-space parallel DMRG distributes the computational workload across sites of the matrix product state (MPS), and has primarily been used in the literature to add parallelism to the standard two-site DMRG. In this work, we leverage the real-space parallel approach to create a distributed-memory framework for performing DMRG on very large MPS that exceed the memory capacity of a single node. We show scaling results for one- and two-dimensional Heisenberg models, provide insights into load balancing, and detail the implementation considerations necessary to run the code at scale.
When Noise Improves Quantum Machine Learning: From Understanding to Harnessing
Speaker: Yulong Dong, University of Michigan
Abstract: Quantum noise is expected to degrade quantum machine learning (QML) by driving circuits away from their noiseless implementations. Yet recent studies show that moderate noise may improve QML, an effect that remains unexplained. We develop a statistical learning theory that explains this finite-noise optimum by connecting microscopic noise processes to macroscopic learning performance. A noise-order purity parameter predicts how noise reduces effective model complexity and the generalization gap, while noise simultaneously increases prediction bias. Their competition reveals an intermediate regime between weak-noise error accumulation and strong-noise trainability collapse, and predicts when the optimum shifts or disappears. Using Perlmutter, a UMich–UW collaboration tested these predictions through simulations of trained hybrid quantum–classical models. We further demonstrate that programming noise during training can move a model towards its optimum. This suggests a broader design principle: much as quantization can be co-designed with training in classical ML, harnessing hardware noise can become a useful component of QML for NISQ and partially fault-tolerant architectures.