Our central question is: How can an agent learn from data while preserving stability, robustness, safety, and efficiency?
We connect reinforcement learning with control theory, optimization, stochastic approximation, and machine learning. This combination allows us to develop new algorithms, explain why they work, and adapt them to autonomous and industrial systems.
Foundations of Reinforcement Learning
We study the mathematical foundations of sequential decision making. Our work develops convergence and finite-time analyses for Q-learning, temporal-difference learning, value iteration, and policy-evaluation methods. We are particularly interested in function approximation, target networks, experience replay, distributed learning, multi-agent games, and alternative Bellman operators.
Representative topics: switching-system models of Q-learning; finite-time and sample-complexity analysis; stable learning with function approximation; regularized and target-based algorithms; distributed and game-theoretic reinforcement learning.
Representative Research
Bum Geun Park, Donghwan Lee*, Adaptive policy backbone via shared network, ICML2026 (spotlight 2.2%) [link]
Jong-Chan Park, Min-Gyu Park, Donghwan Lee*, "Pretraining a shared Q-network for data-efficient offline reinforcement learning," NeurIPS2025 [link]
Han-Dong Lim, Donghwan Lee*, "Regularized Q-learning," NeurIPS2024 [link]
Donghwan Lee, Hyukjun Yang, Bum Geun Park, "Analysis of approximate linear programming solution to Markov decision problem with log barrier function," ICLR 2026 [link]
Han-Dong Lim, Donghwan Lee*, "Regularized Q-learning," NeurIPS2024 [link]
Donghwan Lee, Jianghai Hu, and Niao He, “A discrete-time switching system analysis of Q-learning,” SIAM Journal on Control and Optimization, vol. 61, no. 3, 2023 [link] [online extension]
Safe and Robust Learning-Based Control
High-performing learning policies are not sufficient when a physical system must remain stable and safe under disturbances, uncertainty, and limited data. We develop reinforcement-learning and control methods that explicitly account for these requirements. Control-theoretic tools provide both algorithmic structure and mathematical certificates.
Representative topics: robust policy gradients; disturbance attenuation; minimax reinforcement learning; Lyapunov-based analysis; safe exploration; stability and convergence certificates; learning-based control of dynamical systems.
Representative Research
Taeho Lee, Donghwan Lee*, "Robust deterministic policy gradient for disturbance attenuation and its application to quadrotor control," IFAC world congress 2026 [link]
Taeho Lee, Donghwan Lee*, "Taming the adversary: stable minimax deep deterministic policy gradient via fractional objectives," 2025 [link]
Yeeun Im, Narim Jeong, and Donghwan Lee*, "Safe-support Q-learning: learning without unsafe exploration" [link]
Autonomous and Industrial Intelligence
We translate theoretical ideas into decision-making tools for real systems. Current and recent collaborations include quadrotor control, visual SLAM, autonomous-driving perception and reasoning, offline reinforcement learning for GPU power management, and reinforcement-learning-based industrial temperature control.
Representative topics: robotics and autonomous systems; vision-language decision making; SLAM reliability; offline reinforcement learning; industrial process control; energy and computing systems.
Representative Research
Bum Geun Park, Narim Jeong, Hyukjun Yang, Donghwan Lee*, High-precision temperature control in industrial evaporation heaters via hybrid reinforcement learning approach, IEEE Access, 2026
Heechan Chung, Yeeun Im, Jongchan Park, Taeho Lee, Tae-Young Kim, Hyungjun Kim, and Donghwan Lee*, "Power consumption optimization of GPU server with offline reinforcement learning," IEEE Access 2025 [link]
Kanwal Naveed, Wajahat Hussain, Irfan Hussain, Donghwan Lee, Muhammad Latif Anjum , "Help me through: imitation learning based active view planning to avoid SLAM tracking failures," IEEE Transactions on Robotics, 2025 [link]
Cross-Cutting Methods
Control theory · Optimization · Stochastic approximation · Switched and hybrid systems · Multi-agent learning · Offline reinforcement learning · Vision-language models for decision making