Large Language Model (LLM) Agents and Reasoning
J. Jeon, M. Cho, J. Kim, and Y. Sung*, "Inject a turn, not a token: Learning where to redirect LLM reasoning for self-correction," to be presented at COLM 2026 Workshop on Efficient Reasoning, 2026
Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim, Youngchul Sung*, "Be my tutor: On-policy co-distillation for mutual LLM improvement via peer feedback," preprint, Jun. 2026
Seohui Bae, Jeonghye Kim, Youngchul Sung and Woohyung Lim*, "Align while search: Belief-guided exploratory inference for test-time world alignment," IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2026
Jeonghye Kim+, Sojeong Rhee+, Minbeom Kim, Dohyung Kim, Sangmook Lee, Youngchul Sung* and Kyomin Jung*, "ReflAct: World-grounded decision making in LLM agents via goal-state reflection," Conference of Emplical Methods for Natural Language Processing (EMNLP). (+Equal Contribution), Nov. 2025
Seohui Bae, Jeonghye Kim, Youngchul Sung and Woohyung Lim*, "Align while search: Belief-guided exploratory inference for test-time world alignment," the Exploration in AI Today (EXAIT) Workshop at ICML 2025, Vancouver, July 2025
Deep Reinforcement Learning
Offline Reinforcement Learning and Imitation Learning
Jongseong Chae, Jongeui Park, Yongjae Shin, Gyeongmin Kim, Seungyul Han, and Youngchul Sung*, "Flow Actor-Critic for Offline Reinforcement Learning," International Conference on Learning Representations (ICLR) , Brazil, Apr. 2026
Yongjae Shin, Jongseong Chae, Jongeui Park, and Youngchul Sung*, "Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning," International Conference on Learning Representations (ICLR) , Brazil, Apr. 2026
Jeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngchul Sung, Kanghoon Lee, and Woohyung Lim*, "Penalizing infeasible actions and reward scaling in reinforcement learning with offline data," International Conference on Machine Learning (ICML), Vancouver, Canada, Jul. 2025 (Spotlight Paper, Top 2.6%)
Yongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngsoo Jang, Geonhyeong Kim, Jongseong Chae, Youngchul Sung, Kanghoon Lee, and Woohyung Lim*, "Online pre-training for offline-to-online reinforcement learning," International Conference on Machine Learning (ICML), Vancouver, Canada, Jul. 2025
Jeonghye Kim, Suyoung Lee, Woojun Kim, and Youngchul Sung*, "Adaptive Q-aid for conditional supervised learning in offline reinforcement learning," Conference on Neural Information Processing Systems (NeurIPS ), 2024
Jeonghyee Kim, Suyoung Lee, Woojun Kim*, and Youngchul Sung*, “Decision convformer: Local filtering in metaformer is sufficient for decision making,” International Conference on Learning Representations (ICLR) , May 2024 (Spotlight Paper, Top 5%)
Sungho Choi, Seungyul Han*, Woojun Kim, Jongseong Chae, Whiyoung Jung, and Youngchul Sung, "Domain adaptive imitation learning with visual observation," Conference on Neural Information Processing Systems (NeurIPS ), New Orleans, LA, Dec. 2023
Jongseong Chae, Seungyul Han*, Whiyoung Jung, Myungsik Cho, Sungho Choi, and Youngchul Sung, “Robust imitation learning against variations in environment dynamics,” International Conference on Machine Learning (ICML) 2022 , Baltimore, MD, USA, Jul. 2022
Exploration and Sample Efficiency
Giseung Park, Whiyoung Jung, Seungyul Han, Sungho Choi and Youngchul Sung*, "Adaptive multi-model fusion learning for sparse-reward reinforcement learning," Neurocomputing, vol. 633, Jun. 2025
Woojun Kim+, Jeonghye Kim+ and Youngchul Sung*, "LESSON: Learning to integrate exploration strategies for reinforcement learning via an option framework," International Conference on Machine Learning (ICML), Hawaii, HI, Jul. 2023 (+Equal contribution)
Seungyul Han and Youngchul Sung*, "A max-min entropy framework for reinforcement learning," Conference on Neural Information Processing Systems (NeurIPS) , Dec. 2021.
Seungyul Han and Youngchul Sung*, "Diversity actor-critic: Sample-aware entropy regularization for sample-efficient exploration," International Conference on Machine Learning (ICML) , Jul. 2021
Whiyoung Jung, Giseung Park, and Youngchul Sung*, "Population-Guided Parallel Policy Search for Reinforcement Learning," International Conference on Learning Representations (ICLR) , Vienna, Austria, May 2020
Seungyul Han and Youngchul Sung*, "Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning," International Conference on Machine Learning (ICML) , Long Beach, CA, USA, Jun. 2019
Safe Reinforcement Learning
Woojun Kim+, Yongjae Shin+, Jongeui Park, and Youngchul Sung*, "Sample-efficient and safe deep reinforcement learning via reset deep ensemble agents," Conference on Neural Information Processing Systems (NeurIPS ), New Orleans, LA, Dec. 2023 (+Equal contribution)
W. Jung, M. Cho, J. Park and Y. Sung*, "Quantile constrained reinforcement learning: A reinforcement learning framework constraining outage probability," Conference on Neural Information Processing Systems (NeurIPS), Dec. 2022
Partially-Observable MDPs
G. Kim, J. Kim, S. Lee, J. Baek, H. Moon, S. Shin and Y. Sung*, "Robust reinforcement learning under dimension-wise state information drop," IEEE Access, 2024
Giseung Park, Sungho Choi and Youngchul Sung*, "Blockwise sequential model learning for partially observable reinforcement learning," the 36rd AAAI Conference on Artificial Intelligence (AAAI ), Feb. 2022 (Oral Presentation, Top 4.6%)
Multi-Task RL, Meta RL, and Multi-Objective RL
G. Park+, H. Nam+, W. Byeon, A. Leshem and Y. Sung*, “Constrained multi-objective reinforcement learning with max-min criterion,” International Conference on Machine Learning (ICML), Seoul, Korea, Jul. 2026 (+ Equal contribution)
Woohyeon Byeon, Giseung Park, Jongseong Chae, Amir Leshem and Youngchul Sung, "Multi-objective reinforcement learning with max-min criterion: A game-theoretic approach," Conference on Neural Information Processing Systems (NeurIPS ) , Dec. 2025
Myungsik Cho, Jongeui Park, Jeonghye Kim and Youngchul Sung*, "ARS: Adaptive reward scaling for multi-task reinforcement learning," International Conference on Machine Learning (ICML), Jul. 2025 (Solved the MT10 task of the Metaworld Environment.)
Giseung Park and Youngchul Sung*, "Reward dimension reduction for scalable multi-objective reinforcement learning," International Conference on Learning Representations (ICLR) , Singapore, Apr. 2025
Myungsik Cho, Jongeui Park, Suyoung Lee and Youngchul Sung*, "Hard task first: Multi-task reinforcement learning through task scheduling," International Conference on Machine Learning (ICML) 2024
Giseung Park, Woohyeon Byeon, Seongmin Kim, Elad Habakuk, Amir Leshem and Youngchul Sung*, "The max-min formulation of multi-objective reinforcement learning: From theory to a model-free algorithm," International Conference on Machine Learning (ICML) 2024
Suyoung Lee, Myungsik Cho, and Youngchul Sung*, "Parameterizing Non-Parametric Meta-Reinforcement Learning Tasks via Subtask Decomposition, Conference on Neural Information Processing Systems (NeurIPS ), New Orleans, LA, Dec. 2023
Myungsik Cho, Whiyoung Jung, and Youngchul Sung*, "Multi-task reinforcement learning with task representation method," ICLR Workshop on Generalizable Policy Learning in Physical World (GPL), Apr. 2022
Multi-Agent Reinforcement Learning
Seongmin Kim, Giseung Park, Woojun Kim, Jeewon Jeon, Seungyul Han*, and Youngchul Sung, "Generalized Per-Agent Advantage Estimation for Multi-Agent Policy Optimization," International Conference on Autonomous Agents and Multiagent Systems (AAMAS), May 2026 (AAMAS '26 Best Paper Award Nominee)
Jeewon Jeon, Myungsik Cho, Youngchul Sung, "STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure TransFormer for Offline Mulit-task Multi-agent Reinforcement Learning," International Conference on Learning Representations (ICLR) , Brazil, Apr. 2026
Woojun Kim and Youngchul Sung*, "An adaptive entropy regularization framework for multi-agent reinforcement learning," International Conference on Machine Learning (ICML), Hawaii, HI, Jul. 2023
Woojun Kim, Whiyoung Jung, Myungsik Cho, and Youngchul Sung*, “A variational approach to mutual information-based coordination for multi-agent reinforcement learning,” International Conference on Autonomous Agents and Multiagent Systems (AAMAS), May 2023
Woojun Kim and Youngchul Sung, “Parameter sharing with network pruning for scalable multi-agent deep reinforcement learning,” International Conference on Autonomous Agents and Multiagent Systems (AAMAS), May 2023
Jeewon Jeon, Woojun Kim*, Whiyoung Jung and Youngchul Sung, “MASER: Multi-Agent reinforcementlearning with Subgoals generated from Experience Replay buffer,” International Conference on Machine Learning (ICML) 2022 , Baltimore, MD, USA, Jul. 2022
Woonjun Kim, Jongeui Park, and Youngchul Sung*, "Communication in multi-agent reinforcement learning: Intention sharing," International Conference on Learning Representations (ICLR) 2021, May 2021
Woojun Kim, Myungsik Cho, and Youngchul Sung*, "Message-dropout: An efficient training method for multi-agent deep reinforcement learning," the 33rd AAAI Conference on Artificial Intelligence (AAAI ) 2019, Honolulu, HW, USA, Jan. 2019
Applications
G. Park+, H. Nam+, W. Byeon, A. Leshem and Y. Sung*, "Fairness-aware resource allocation in edge computing via multi-objective reinforcement learning," submitted to IEEE Trans. Network and Service Management, Mar. 2026
Sohee Bae, Seungyul Han, and Youngchul Sung*, "A reinforcement learning formulation of the Lyapunov optimization: Application to edge computing systems with queue stability," available at arXiv, Dec. 2020
Signal Detection
Yuni Lee and Youngchul Sung*, "Generalized Chernoff information for mismatched Bayesian detection and its application to energy detection," IEEE Signal Processing Letters, vol. 19, no. 11, pp. 753 - 756, Nov. 2012.
Yirang Lim, Juho Park and Youngchul Sung*, "Upper bound for the loss of energy detection of signals in multipath fading channels," IEEE Signal Processing Letters, vol. 16, no. 11, pp. 949 - 952, Nov. 2009.
Youngchul Sung*, H. Vincent Poor and Heejung Yu, "How much information can one get from a wireless ad hoc sensor network over a correlated random field?," IEEE Transactions on Information Theory, vol. 55, no. 6, pp. 2827 - 2847, Jun. 2009. Simple Condition for Theorem 2.
Youngchul Sung*, "Large deviations principle and its applications," in Proceedings of Korean Information and Communications Society (KICS), Coding and Information Theory Workshop, Jan. 2009.
Youngchul Sung*, Xin Zhang, Lang Tong and H. Vincent Poor, "Sensor configuration and activation for field detection in large sensor arrays," IEEE Transactions on Signal Processing, vol. 56, no. 2, pp. 447 - 463, Feb. 2008.
Youngchul Sung, Saswat Misra, Lang Tong*, and Anthony Ephremides, "Cooperative routing for signal detection in large sensor networks," IEEE Journal on Selected Areas in Communications, Vol. 25, no. 2, pp. 471 - 483, Feb. 2007.
Youngchul Sung, Saswat Misra, Lang Tong*, and Anthony Ephremides, "Signal processing for application-specific ad-hoc networks," IEEE Signal Processing Magazine, vol. 23, no. 5, pp. 74 - 83, Sep. 2006.
Youngchul Sung, Lang Tong*, and H. Vincent Poor, "Neyman-Pearson detection of Gauss-Markov signals in noise: Closed-form error exponent and properties," IEEE Transactions on Information Theory, vol. 52, no. 4, pp. 1354-1365, Apr. 2006.
Youngchul Sung, Lang Tong*, and H. Vincent Poor, "Optimal and suboptimal detection of Gaussian signals in noise: Asymptotic relative efficiency," in the proceeding of the Conference on Advanced Signal Processing Algorithms, Architectures, and Implementations XV, part of SPIE International Symposium, San Diego, CA, Jul. 2005.
Youngchul Sung, Lang Tong*, and Ananthram Swami, "Asymptotic locally optimal detector for large scale sensor networks under the Poisson regime," IEEE Transactions on Signal Processing, vol. 53, no. 6, pp. 2005 - 2017, Jun. 2005.
Youngchul Sung, Lang Tong*, and H. Vincent Poor, "A large deviations approach to sensor scheduling for detection of correlated random fields," in Proc. of ICASSP 2005, Philadelphia, PA, Mar. 2005. (ICASSP 2005 Best Student Paper Award)
Blind Estimation
Song Noh, Youngchul Sung*, and Michael Zoltowski, "A new precoder design for blind channel estimation in MIMO-OFDM systems," IEEE Transactions on Wireless Communications, vol. 13, no. 12, pp. 7011-7024, Dec. 2014
Youngchul Sung*, Yirang Lim, Lang Tong and Alle-Jan van der Veen, "Signal processing advances for 3G CDMA: From rake receivers to blind techniques," IEEE Communications Magazine, vol. 47, no. 1, pp. 48 - 54, Jan. 2009
Youngchul Sung, Lang Tong* and Ananthram Swami, "Blind channel estimation for space-time coded WCDMA," EURASIP Journal on Wireless Communications and Networking, vol. 2004, no. 2, pp. 322 - 334, Dec. 2004
Youngchul Sung and Lang Tong*, "Tracking of fast-fading channels in long code WCDMA," IEEE Transactions on Signal Processing, vol. 52, no. 3, pp. 786 - 795, Mar. 2004
Lang Tong*, Alle-Jan van der Veen, Patrick Dewilde and Youngchul Sung, "Blind decorrelating rake receiver for long code WCDMA," IEEE Transactions on Signal Processing, vol. 51, no. 6, pp. 1642 - 1655, Jun. 2003
(Asterisks denote corresponding authors.)