Seungjun Moon
Google Scholar / LinkedIn / github
My research focuses on bridging human and robot: learning from human videos to build the data and models that teach robots to act. Alongside this, I develop real-time generative models for image and video, advancing model performance through large-scale (>10TB) multi-modal and 3D datasets while optimizing architectures for efficient on-device and real-time inference. This line of work, from human video generation with large-scale 3D data to Vision–Language–Action (VLA) models for dexterous manipulation, connects human motion understanding with embodied robot control.
I am currently pursuing a Ph.D. at KAIST under the supervision of Prof. Jinwoo Shin, while continuing my research in industry. I earned my Master's degree at KAIST under the same supervision, following a Bachelor's degree in Electrical Engineering with a minor in Mathematical Sciences, also at KAIST.