2026

Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations

Jialu Li, Jinchuan Tian, Shinji Watanabe

IEEE Spoken Language Technology Workshop (SLT) 2026


Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities

Jialu Li, Marvin Lavechin, Xulin Fan, Nancy L. McElwain, Alejandrina Cristia, Paola Garcia-Perera, and Mark Hasegawa-Johnson

IEEE Signal Processing Magazine, 2026 · [Paper]


Bridging Machine Learning and Event-Based Analyses to Assess Maternal Contingent Responses to Infant Nondistress and Distress Vocalizations from Home Audio Recordings

Kexin Hu, Xulin Fan, Yannan Hu, Jialu Li, Bethany Lee, Mark Hasegawa-Johnson, and Nancy L. McElwain

Developmental Science, 2026 · [Paper]


Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan, Kexin Hu, Bashima Islam, Mark Hasegawa-Johnson, and Nancy L. McElwain

Interspeech 2026 · [Paper]


Age-Aware Adapter Tuning for Children's Speech Recognition

Jialu Li

arXiv preprint, 2026 · [Paper] [Code] [Model]


2025


Band-Split Self-supervised Mamba for Infant-centered Audio Analysis

Xulin Fan, Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain

Interspeech 2025 · [Paper]


2024

Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis

Jialu Li, Mark Hasegawa-Johnson, and Karrie Karahalios

Interspeech 2024 · [Paper] [Code] [Model]


Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations

Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain

ICASSP 2024 Workshop on Self-Supervision in Audio, Speech and Beyond (SASB) · [Paper] [Code] [Model] [Supplement]


Sound Tagging in Infant-centric Home Soundscapes

Mohammad Nur Hossain Khan, Jialu Li, Nancy L. McElwain, Mark Hasegawa-Johnson, and Bashima Islam

IEEE/ACM CHASE 2024 · [Paper]


Preliminary Technical Validation of LittleBeats™: A Multimodal Sensing Platform to Capture Cardiac Physiology, Motion, and Vocalizations

Bashima Islam, Nancy L. McElwain, Jialu Li, Maria I. Davila, Yannan Hu, Kexin Hu, Jordan M. Bodway, Ashutosh Dhekne, Romit Roy Choudhury, and Mark Hasegawa-Johnson

Sensors, 2024 · [Paper]


2023

Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family Audio

Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain

Interspeech 2023 · [Paper] [Code] [Model]


Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language Recognition

Liming Wang, Junrui Ni, Heting Gao, Jialu Li, Kai Chieh Chang, Xulin Fan, Junkai Wu, Mark Hasegawa-Johnson, and Chang D. Yoo

Findings of ACL 2023 · [Paper]


2022

Autosegmental Neural Nets 2.0: An Extensive Study of Training Synchronous and Asynchronous Phones and Tones for Under-Resourced Tonal Languages

Jialu Li and Mark Hasegawa-Johnson

IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022 · [Paper]


Visualizations of Complex Sequences of Family-Infant Vocalizations Using Bag-of-Audio-Words Approach Based on Wav2vec 2.0 Features

Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain

arXiv preprint, 2022 · [Paper]



2021

Analysis of Acoustic and Voice Quality Features for the Classification of Infant and Mother Vocalizations
Jialu Li, Mark Hasegawa-Johnson, and Nancy L. McElwain
Speech Communication, 2021 · [Paper]

Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings
Jialu Li, Vimal Manohar, Pooja Chitkara, Andros Tjandra, Michael Picheny, Frank Zhang, Xiaohui Zhang, and Yatharth Saraf
arXiv preprint, 2021 · [Paper]


2020

Autosegmental Neural Nets: Should Phones and Tones be Synchronous or Asynchronous?

Jialu Li and Mark Hasegawa-Johnson

Interspeech 2020 · [Paper


2019

An Embodied, Platform-Invariant Architecture for Connecting High-Level Spatial Commands to Platform Articulation

Anum Jang Sher, Umer Huzaifa, Jialu Li, Varun Jain, Alex Zurawski, and Amy LaViers

Robotics and Autonomous Systems, 2019 · [Paper]


2018

A Comparable Phone Set for the TIMIT Dataset Discovered in Clustering of Listen, Attend and Spell

Jialu Li and Mark Hasegawa-Johnson

NeurIPS 2018 Workshop on Interpretability and Robustness in Audio, Speech, and Language (IRASL) · [Paper