We design compiler-level optimization techniques tailored to the characteristics of each target hardware — GPUs, PIM devices, and NPUs — to maximize performance while minimizing programming effort.
Recent Publications
[CGO 2026] Flow-Graph-Aware Tiling and Rescheduling for Memory-Efficient On-Device Inference
[ICS 2025] SortingHat: System Topology-aware Scheduling of Deep Neural Network Models on Multi-GPU Systems
[CGO 2025] CUrator: An Efficient LLM Execution Engine with Optimized Integration of CUDA Libraries
[ICCD 2023] Tailoring Tiling-based GEMM Performance using Supervised Learning
[ICCD 2021] Legion: Tailoring Grouped Neural Execution Considering Heterogeneity on Multiple Edge Devices
[CGO 2020] PreScaler: An Efficient System-aware Precision Scaling Framework on Heterogeneous Systems
[CGO 2025] Accelerating LLMs using an Efficient GEMM Library and Target-Aware Optimizations on Real-World PIM Devices
[LCTES 2024] Orchestrating Multiple Mixed Precision Models on a Shared Precision-Scalable NPU
[DATE 2024] Discovering Efficient Fused Layer Configurations for Executing Multi-Workloads on Multi-core NPUs
We build runtime systems and schedulers that dynamically allocate resources and orchestrate concurrent workloads across heterogeneous devices, adapting to dynamic system state.
Recent Publications
[ICS 2025] PIM-CARE: A Compiler-Assisted Dynamic Resource Allocation Framework for Real-world DRAM PIM
[PACT 2023] Virtual PIM: Resource-aware Dynamic DPU Allocation and Workload Scheduling Framework on Multi-DPU PIM Architecture
[LCTES 2023] Synchronization-aware NAS for an Efficient Collaborative Inference on Mobile Platforms
[DATE 2023] Block Group Scheduling: A General Precision-scalable NPU Scheduling Technique with Capacity-aware Memory Allocation
[DAC 2020] Convergence-Aware Neural Network Training
We co-design custom hardware architectures — PIM functional units, FPGA-based neural accelerators, and GPU microarchitecture extensions — together with the software/compiler support needed to exploit them effectively.
Recent Publications
[LCTES 2026] FLUX: Frequency Scaling with Layer-wise Utilization for Energy-Efficient NPU Execution
[MICRO 2025] PIM-CCA: An Efficient PIM Architecture with Optimized Integration of Configurable Functional Units
[IEEE TVLSI, 2022] Dynamic Rate Neural Acceleration Using Multiprocessing Mode Support
[DAC 2020] Navigator: Dynamic Multi-kernel Scheduling to Improve GPU Performance
We study how to process emerging applications such as large-scale sparse matrix multiplication (SpGEMM), graph workloads, and image super-resolution efficiently by holistically leveraging the diverse computing resources of heterogeneous systems.
Recent Publications
[IEEE Access, 2025] Efficient Image Super-Resolution Using Dynamic Quality Control with Recursive Model Structures
[ACM TACO, 2024] ISP Agent: A Generalized In-Storage-Processing Workload Offloading Framework by Providing Multiple Optimization Opportunities
[CIKM 2023] SAGE: A Storage-Based Approach for Scalable and Efficient Sparse Generalized Matrix-Matrix Multiplication
[ICDE 2023] Orchestrating Large-Scale SpGEMMs using Dynamic Block Distribution and Data Transfer Minimization on Heterogeneous Systems
[ICDE 2020] Optimization of a GPU-based Sparse Matrix Multiplication for Large Sparse Networks