From Exascale Computing to Sovereign AI Factories: The Future of AI for Science Is Open, Heterogeneous, and Still Needs FP64
September 8, 2026
9:00 AM - 10:00 AM Pacific Time
Online
September 8, 2026
9:00 AM - 10:00 AM Pacific Time
Online
About the Session
The infrastructure conversation around AI is often reduced to low-precision tensor performance. Yet the most consequential scientific workloads do not end with a model response. They combine simulation, experimental data, AI training, inference, autonomous agents, and numerical validation in a continuous discovery loop. AI for Science is transformative because it can change not only how quickly we compute an answer, but how rapidly we formulate hypotheses, design experiments, explore enormous parameter spaces, and convert results into new knowledge.
Dr. Nicholas Malaya, AMD Senior Fellow and technical lead for exascale application performance, will discuss how lessons from Frontier and El Capitan are shaping AMD’s strategy for this convergence of HPC and AI. Oak Ridge National Laboratory’s Lux AI and HPC supercomputer represents an important next step. Powered by AMD EPYC™ CPUs, AMD Instinct™ MI355X GPUs, and AMD Pensando™ networking, Lux is designed to support large-scale training, distributed inference, modeling, and simulation as part of a sovereign U.S. AI factory for science.
The session will examine the complementary roles of AMD EPYC processors and AMD Instinct accelerators, including the AMD Instinct MI430X for sovereign AI and scientific computing and the MI455X for frontier-scale AI. These architectures reflect a heterogeneous design philosophy: provide the appropriate compute engine, memory system, precision, and network for each stage of the workflow.Nick will also explain why FP64 remains essential. Lower precision can dramatically accelerate training and inference, but high-fidelity simulation, numerical convergence, uncertainty quantification, and scientific validation continue to depend on robust double-precision performance. AI may propose the next experiment, but physical simulations will always validate the result.
Finally, the session will describe how AMD’s open strategy, including ROCm™, open programming models and frameworks, and the AMD Enterprise AI Reference Stack, gives organizations greater control over their infrastructure, data, models, and deployment choices. Attendees will leave with a systems-level framework for building open computing environments that span cloud and sovereign infrastructure, AI and simulation, and today’s applications and tomorrow’s discoveries.
Speaker
Nicholas Malaya
Sr. Fellow in High Performance Computing and Sovereign AI, AMD
Nicholas Malaya is a Sr. Fellow in High Performance Computing and Sovereign AI at AMD. He is AMD's technical lead for exascale application performance, and led the deployment of Frontier and El Capitan, both ranked #1 at the time of their debut on the Top500. Nick's research interests include HPC, computational fluid dynamics, Bayesian inference, and ML/AI. He received his PhD from the University of Texas. Before that, he double majored in physics and mathematics at Georgetown University, where he received the Treado medal. In his copious spare time, he enjoys long-distance running, wine, and spending time with his wife and children.