Email: tangchq AT gmail. Chinese Name: 唐春强. LinkedIn. Github. Google Scholar. We are hiring.
Dr. CQ (Chunqiang) Tang is Vice President of the AI & Compute Foundation organization at Meta, leading Meta's hardware strategy, hardware/software co-design, and AI chip software.
Since joining Meta in 2013, he has helped build foundational infrastructure across the full hardware and software stack: hyperscale AI clusters (co-designed across silicon, datacenters, and ML models); compute fleet hardware (in-house AI chips, ASICs, CPUs); AI chip software (compilers, kernels); LLM training and inference software; and internal cloud software (IaaS, PaaS, serverless, databases).
In addition to building infrastructure that serves billions of people, he publishes top-venue research on hyperscale systems. His work has earned Best Paper Awards at SOSP, OSDI, ISCA, and ASPLOS; multiple IEEE Micro Top Picks; and two Communications of the ACM honors—a Research Highlight and a cover story. He has published over 100 papers and holds around 100 patents.
If you have time to read only one paper, he recommends his cover story in Communications of the ACM: "Meta's Hyperscale Infrastructure: Overview and Insights."
Recent Best Papers and Highlights:
[CACM'25 Research Highlights] TMO: Transparent Memory Offloading in Datacenters.
The companion "Technical Perspective: Memory Efficiency via Offloading in Warehouse-Scale Datacenters," written by Parthasarathy Ranganathan from Google.[SOSP'24 Best Paper] FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production Monitoring
[OSDI'24 Best Paper] ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing
[ISCA'23 Best Paper] Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters
[ASPLOS'22 Best Paper] TMO: Transparent Memory Offloading in Datacenters
[IEEE Micro Top Picks'24] Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters
[IEEE Micro Top Picks'23] IOCost: Block IO Control for Containers in Datacenters
[SOSP'23, Best Serverless Paper of 2023] XFaaS: Hyperscale and Low Cost Serverless Functions at Meta. This paper was selected by the 9th Workshop on Serverless Computing as the Best Serverless Paper of 2023 out of all serverless papers published that year.
Selected Publications
[Meta AI Blog] Four MTIA Chips in Two Years: Scaling AI Experiences for Billions
[ISCA'26] MTIA-300: Meta's Training Chip with Embedded NIC Chiplets and Communication Offloading Engine
[ISCA'26] Vistara: Making CXL Real—Full Path from ASIC Design and OS Support to Hyperscale Deployment
[ISCA’26] LoKA: Low-precision Kernel Applications for Recommendation Models At Scale
[MLSys’26] Optimizing Deployment Configurations for LLM Inference
[MLSys’26] Sparing Strategy to Minimize Reliability Impact on Large Scale Training Jobs
[ISCA'25] Scaling Llama 3 Training with Efficient Parallelism Strategies
[ISCA'25] Meta's Second Generation AI Chip: Model-Chip Co-Design and Productionization Experiences
[ISCA'25] DCPerf: An Open-Source, Battle-Tested Performance Benchmark Suite for Datacenter Workloads
[Communications of the ACM] Meta's Hyperscale Infrastructure: Overview and Insights
[OSDI'24] MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale
[OSDI'24] Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and Experiences
[NSDI'24] MobileConfig: Remote Configuration Management for Mobile Apps at Hyperscale
[OSDI'23] Conveyor: One-Tool-Fits-All Continuous Software Deployment at Meta
[OSDI'23] ServiceRouter: Hyperscale and Minimal Cost Service Mesh at Meta
[OSDI'23] Global Capacity Management With Flux
[ASPLOS'22] IOCost: Block IO Control for Containers in Datacenters
[SOSP'21] Shard Manager: A Generic Shard Management Framework for Geo-distributed Applications
[SOSP'21] RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation
[OSDI'20] Twine: a Unified Cluster Management System for Shared Infrastructure