2026
Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations
Jialu Li, Jinchuan Tian, Shinji Watanabe
IEEE Spoken Language Technology Workshop (SLT) 2026
Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
Jialu Li, Marvin Lavechin, Xulin Fan, Nancy L. McElwain, Alejandrina Cristia, Paola Garcia-Perera, and Mark Hasegawa-Johnson
IEEE Signal Processing Magazine, 2026 · [Paper]
Bridging Machine Learning and Event-Based Analyses to Assess Maternal Contingent Responses to Infant Nondistress and Distress Vocalizations from Home Audio Recordings
Kexin Hu, Xulin Fan, Yannan Hu, Jialu Li, Bethany Lee, Mark Hasegawa-Johnson, and Nancy L. McElwain
Developmental Science, 2026 · [Paper]
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan, Kexin Hu, Bashima Islam, Mark Hasegawa-Johnson, and Nancy L. McElwain
Interspeech 2026 · [Paper]
Age-Aware Adapter Tuning for Children's Speech Recognition
Jialu Li
arXiv preprint, 2026 · [Paper] [Code] [Model]
2025
Band-Split Self-supervised Mamba for Infant-centered Audio Analysis
Xulin Fan, Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain
Interspeech 2025 · [Paper]
2024
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
Jialu Li, Mark Hasegawa-Johnson, and Karrie Karahalios
Interspeech 2024 · [Paper] [Code] [Model]
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain
ICASSP 2024 Workshop on Self-Supervision in Audio, Speech and Beyond (SASB) · [Paper] [Code] [Model] [Supplement]
Sound Tagging in Infant-centric Home Soundscapes
Mohammad Nur Hossain Khan, Jialu Li, Nancy L. McElwain, Mark Hasegawa-Johnson, and Bashima Islam
IEEE/ACM CHASE 2024 · [Paper]
Preliminary Technical Validation of LittleBeats™: A Multimodal Sensing Platform to Capture Cardiac Physiology, Motion, and Vocalizations
Bashima Islam, Nancy L. McElwain, Jialu Li, Maria I. Davila, Yannan Hu, Kexin Hu, Jordan M. Bodway, Ashutosh Dhekne, Romit Roy Choudhury, and Mark Hasegawa-Johnson
Sensors, 2024 · [Paper]
2023
Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family Audio
Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain
Interspeech 2023 · [Paper] [Code] [Model]
Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language Recognition
Liming Wang, Junrui Ni, Heting Gao, Jialu Li, Kai Chieh Chang, Xulin Fan, Junkai Wu, Mark Hasegawa-Johnson, and Chang D. Yoo
Findings of ACL 2023 · [Paper]
2022
Autosegmental Neural Nets 2.0: An Extensive Study of Training Synchronous and Asynchronous Phones and Tones for Under-Resourced Tonal Languages
Jialu Li and Mark Hasegawa-Johnson
IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022 · [Paper]
Visualizations of Complex Sequences of Family-Infant Vocalizations Using Bag-of-Audio-Words Approach Based on Wav2vec 2.0 Features
Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain
arXiv preprint, 2022 · [Paper]
2021
Analysis of Acoustic and Voice Quality Features for the Classification of Infant and Mother Vocalizations
Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain
Speech Communication, 2021 · [Paper]
Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings
Jialu Li, Vimal Manohar, Pooja Chitkara, Andros Tjandra, Michael Picheny, Frank Zhang, Xiaohui Zhang, and Yatharth Saraf
arXiv preprint, 2021 · [Paper]
2020
Autosegmental Neural Nets: Should Phones and Tones be Synchronous or Asynchronous?
Jialu Li and Mark Hasegawa-Johnson
Interspeech 2020 · [Paper]
2019
An Embodied, Platform-Invariant Architecture for Connecting High-Level Spatial Commands to Platform Articulation
Anum Jang Sher, Umer Huzaifa, Jialu Li, Varun Jain, Alex Zurawski, and Amy LaViers
Robotics and Autonomous Systems, 2019 · [Paper]
2018
A Comparable Phone Set for the TIMIT Dataset Discovered in Clustering of Listen, Attend and Spell
Jialu Li and Mark Hasegawa-Johnson
NeurIPS 2018 Workshop on Interpretability and Robustness in Audio, Speech, and Language (IRASL) · [Paper]