How do infants learn about their bodies and the world around them? We develop AI tools to help answer this question by turning video recordings into detailed measurements of infant behavior. Our methods track body movements, posture, and head orientation, supporting the study of self-touch, reaching, and exploration. By reducing labor-intensive manual coding, we aim to make it easier to follow individual developmental trajectories and analyze larger collections of recordings.
Beyond measuring visible behavior, we transfer recorded infant movements to virtual infant models to reconstruct the visual, tactile, and body-position signals these movements generate. We also develop tools to anonymize infant videos while preserving information such as facial expressions and gaze. Together, these approaches connect observations of infant behavior with computational models of development, while making rich developmental data easier to analyze and share.
Infant pose estimation from videos
Rapid progress in computer vision makes it possible to estimate the body configuration of humans from images or videos only. We use this technology to automatically
estimate the pose of infants. Gama et al. (2025) compared the performance of seven state-of-the-art 2D human pose estimation networks on infant datasets, with
ViTPose (Xu et al., 2022) performing the best overall. Khoury et al. (2022) performed 2D pose estimation (in image or pixel space) and then deployed SMPLify-X
(Pavlakos et al., 2019) to get a 3D model. Our current pipeline is shown below.
Comparison of 2D pose estimation networks: Gama, F.; Misar, M.; Navara, L.; Popescu, S. T. & Hoffmann, M. (2025), 'Automatic infant 2D pose estimation from videos: comparing seven deep neural network methods', Behavior Research Methods 57(280). [link to Springer free access][DOI - Springer Nature page][pdf-arxiv][dockerhub][data on osf]
Our code on dockerhub is publicly availble.
3D body pose and head orientation / gaze estimation works:
Theses under Matej Hoffmann's supervision:
Šteinerová, B. (2026), 'Kinematic analysis of infant self-reaching: comparing spontaneous self-touch and goal-oriented movements using 3D pose estimation', Bachelor's thesis, CTU in Prague. [thesis page @ CTU dspace][pdf @ dspace][code]
Šebek, S. (2026), 'Body and object 3D pose estimation for kinematic analysis of infant reaching', Bachelor's thesis, CTU in Prague. [thesis page @ CTU dspace][pdf @ dspace][code]
Ježek, V. (2024), '3D Body Pose Estimation of Infants from RGB Images and Videos', Bachelor's thesis, CTU in Prague. [Dean's Award] [thesis page @ CTU dspace][pdf @ dspace][code]
Jána, P. (2026), 'Estimating Head and Gaze Direction from Videos of Infants', Bachelor's thesis, CTU in Prague. [thesis page @ CTU dspace][pdf @ dspace][code]
Volprecht, V. (2024), 'Accuracy of 3D body pose and shape estimation of infants from RGB and RGB-D data', Master's thesis, Czech Technical University in Prague. [thesis page @ CTU dspace][pdf @ dspace][code]
Vaculínová, Noemi (2023), 'Evaluation Framework for Infant 3D Pose Extraction from RGB Images Using RGB-D Cameras and Motion Capture System', Bachelor's thesis, Czech Technical University in Prague. [thesis page @ CTU dspace][pdf @ dspace]
Application of pose estimation to analyze infant movements
We apply the pose estimation networks to:
analyze spontanous infant movements (Khoury et al. 2022; Khoury et al. 2026)
estimate head orientation (Gama et al. 2025; Yakoub et al. 2026)
detect fidgety movements (Ježek 2026)
Khoury, J., Popescu, S. T., Gama, F., Marcel, V. and Hoffmann, M. (2022), Self-touch and other spontaneous behavior patterns in early infancy, in 'IEEE International Conference on Development and Learning (ICDL)', pp. 148-155. [IEEE Xplore][pdf@researchgate]
Yakoub, N., Goncalves Maia, M., Jana, P., Gama, F., Cech, J., Hoffmann, M. and Berger, S. E. (2026), Torticollis screening during sleep: Insights from manual and automated infant head orientation estimation, in 'IEEE International Conference on Development and Learning (ICDL)'.
Additional publications or student theses:
Khoury, J. A., Gama, F. and Hoffmann, M. (2026), Feet touches and leg movements prior to locomotion onset in infancy, in 'IEEE International Conference on Development and Learning (ICDL)'. [Best Student Paper Award]
Gama, F.; Corbetta, D.; Hoffmann, M. & Jover, M. (2025), Automated head-turn estimation from nose position in infant videos, in '2025 IEEE International Conference on Development and Learning (ICDL)'. [DOI - IEEE Xplore]
Ježek, V. (2026), 'Automatic Detection of Infant Fidgety Movements Using 3D Pose Estimation from Video', Master's thesis, CTU in Prague. [thesis page @ CTU dspace][pdf @ dspace][code]
Motion retargeting and emulating first-person infant multimodal experience
How can humanoid robots help us investigate an infant’s first-person sensorimotor experience?
Starting from a video of an infant, our framework reconstructs the infant’s three-dimensional body configuration and retargets the motion to physical and virtual humanoids—including iCub, pyCub, EMFANT, and MIMo. Replaying these movements generates simulated streams of vision, touch, and proprioception.
The work builds a bridge between developmental science and robotics: humanoid models can help us investigate the rich multisensory experience accompanying infants’ movements, while infant behavior can inspire new approaches to embodied learning in robots.
Lopez, F. M.; Kanazawa, H.; Fiala, O.; Balashov, Y.; Marcel, V.; Rustler, L.; Lenz, M.; Kim, D.; Kuniyoshi, Y.; Triesch, J. & Hoffmann, M. (2026), Simulating Infant First-Person Sensorimotor Experience via Motion Retargeting from Babies to Humanoids, in 'IEEE International Conference on Development and Learning (ICDL)'. [pdf-arxiv][code][youtube-video]
Infant simulator with an embodied caregiver
We developed a physics-based infant simulator with an embodied caregiver, enabling caregiver–infant interactions to be replayed and analyzed from the infant’s perspective. The platform generates multimodal sensory experience, including whole-body touch and egocentric vision, and can distinguish self-, caregiver-, and environment-generated contact. We showcase the tool on caregiver touch, while aiming more broadly at modeling the development of social interaction.
Anonymizing infant videos
Progress in developmental science critically depends on sharing rich behavioral data—especially raw video recordings. However, ethical and privacy constraints often make such sharing impossible. As a result, many valuable infant datasets remain inaccessible, limiting reproducibility, cumulative science, and cross-lab collaboration.
We developed BLANKET, a method designed to make ethical data sharing of infant videos feasible—without compromising scientific utility.
Hadera, D., Cech, J., Purkrabek, M. and Hoffmann, M. (2025), BLANKET: Anonymizing Faces in Infant Video Recordings, in 'IEEE International Conference on Development and Learning (ICDL)'. [DOI - IEEE Xplore][pdf-arxiv][code and video]