Cluster Director: Unleashing Performance with Managed Supercomputing Infrastructure
July 14, 2026
9:00 AM - 10:00 AM Pacific Time
Online
July 14, 2026
9:00 AM - 10:00 AM Pacific Time
Online
About the Session
Clouds offer tremendous scale and computational power for HPC and AI, but deploying and managing flexible, performant, reliable HPC and AI clusters in the cloud for users has required significant IT expertise and effort--in the past. Staying competitive means ensuring users spend time developing and running models and code, not debugging cluster nodes.
Join the Google Cloud Advanced Computing Community for a deep dive into Cluster Director, Google’s managed infrastructure service designed to simplify the entire lifecycle of HPC and AI clusters. We will explore how Cluster Director replaces manual, error-prone setup with automated, topology-aware orchestration for both Slurm and Kubernetes environments.
In this session, we will demonstrate how Cluster Director accelerates your productivity by:
Zero-Friction Lifecycles: Automate everything from "Day 0" deployment to one-click node remediation.
Maximum Performance: See how topology-aware placement ensures your GPUs and TPUs are physically co-located for ultra-low latency.
Smart Scaling & Flexibility: Seamlessly mix Reservations, Flex-start (Dynamic Workload Scheduler), and Spot VMs in a single environment.
Total Visibility: A tour of the dashboard that identifies "straggler" nodes before they kill your job.
Technology should fuel your progress, not stall it. We’ve automated the "grunt work" of supercomputing so you can stop managing infrastructure and start focusing on breakthroughs. Join us to discover how to get world-class performance with the simplicity of the cloud.
Speaker
Ilias Katsardis
Senior Product Manager - AI Infrastructure, Google Cloud
Ilias Katsardis is a Senior Product Manager based in Sunnyvale, CA, driving the future of AI infrastructure at Google Cloud. He is responsible for Cluster Director and the Cluster Toolkit, two key components of Google's supercomputing architecture. Passionate about making large-scale AI and HPC more accessible, Ilias focuses on creating solutions that automate complex configurations and provide a seamless user experience. His work enables researchers and developers to spend less time on infrastructure management and more time on scientific breakthroughs. With a rich background that includes roles at Cray Inc. and ClusterVision, along with founding two tech startups, Ilias brings over 15 years of deep industry expertise to his role.