Join our network to engage in seminars and discussions that advance discovery and collaboration around cutting-edge research in speech and language health signals
Sept. 14, 2026 · 11:00 AM (Eastern Time)
Prof. dr. Odette Scharenborg
Full Professor
Delft Inclusive Speech Communication Lab/Multimedia Computing Group
Delft University of Technology
The Netherlands
Prof. dr. Odette Scharenborg is a Full Professor of Inclusive Speech Communication at the Delft Inclusive Speech Communication (DISC) group at the Multimedia Computing Group at Delft University of Technology, the Netherlands. Her research aims to develop inclusive speech technology, i.e., making speech technology available for everyone irrespective of how they speak or what language they speak. In her research, she consider technical aspects as well as ethical and societal aspects. From 2017-2025, Odette was on the Board of the International Speech Communication Association (ISCA), the largest international society on speech science and technology, where she was the chair of the Diversity committee (2019-2023) and co-chair of the Interspeech Conferences committee and of the Technical Committee (2017-2019). She served as ISCA Vice-president from 2021-2023 and as ISCA President from 2023-2025. In 2025, she was the General Chair of Interspeech 2025, the flagship conference of ISCA and the largest international conference on speech science and technology. In 2026, she was elected as an ISCA Fellow “For sustained contributions to inclusive speech technology, and for lasting impact on diversity in speech communication and technology”. From 2018-2022, Odette was a member of the IEEE Speech and Language Processing Technical Committee (subarea Speech Production and Perception). From 2019-2023, she served as an (Senior) Associate Editor of IEEE Signal Processing Letters
Inclusive speech communication: Developing assistive technologies for dysarthric speakers
Communication is central to our participation in society and our well-being. For most of us, it comes naturally and is easy. However, this is not the case for more than 400,000 people in the Netherlands who have dysarthria. Dysarthria is due to impaired muscular control of the speech mechanism, which has many different causes. Think of Parkinson’s disease (~70% of patients develop a dysarthria), ALS ( ~85% of patients develop a dysarthria), MS ( ~50%), stroke (~25%), and traumatic brain injury. Having a dysarthria means that speaking costs a lot of energy and, that the speech typically is (very) hard to understand for other people. Also, dysarthria is often accompanied by other physical disabilities, so typing is difficult, making it even harder to communicate. This is also the case for two men who approached us with the question to help them in their day-to-day communication. While typically the dysarthric speaker who is hard to understand will undergo speech therapy to improve their speech intelligibility, we approached the problem from the listener-side. Based on insights from psycholinguistics, we decided to develop an app to live caption impaired speech to help the listener understand the speech better and potentially improve their recognition performance of dysarthric speech.
In this talk, I will present the work done in the Delft Inclusive Speech Communication (DISC) lab on improving automatic speech recognition of dysarthric speech, the development of a Dutch dysarthric speech corpus, and our on-going efforts to develop a live captioning app for speakers with a dysarthria for which we are running a crowdfunding campaign that is attracting national attention.
Meetings are held on Zoom for 60 minutes, about once a month—typically on Mondays at 8 AM PT / 11 AM ET / 4 PM London / 5 PM Berlin / 11 PM Beijing.
Join our mailing list to recieve updates and zoom link📙!
For questions, suggesting speakers, or proposing papers for a journal club, reach out to Jingyao Wu 📧 and Ahmed Yousef 📧
This network was co-founded in 2022 by Daniel Low, Tanya Talkar, Daryush Mehta, Satrajit Ghosh and Tom Quatieri as the Harvard-MIT Speech and Language Biomarker Interest Group.
It seeks to bring together researchers and students from around the world to share novel research, receive feedback, discuss papers, and kickstart collaborations.
Jingyao Wu, PhD, MIT
Ahmed Yousef, PhD, MGH & Harvard Medical School
Daniel Low, PhD, Child Mind Institute & Harvard University
Fabio Catania, PhD, MIT
Nick Cummins, PhD, King's College London
Hamzeh Ghasemzadeh, PhD, University of Central Florida
Rahul Brito, Harvard & MIT
Tanya Talkar, PhD, Linus Health
Daryush Mehta, PhD, MGH & Harvard Medical School
Satrajit Ghosh, PhD, MIT McGovern Institute for Brain Research
Thomas Quatieri, PhD, MIT Lincoln Laboratory
Curated materials (tools, datasets, readings) to support exploration and innovation in the field.
Audio
senselab: a Python package that simplifies building pipelines for digital biometric analysis on speech and voice.
Riverst: a multimodal avatar for interacting with the user(s) and collect audio and video data.
Text
Quick spacy type metrics: https://github.com/HLasse/TextDescriptives and https://github.com/novoic/blabla
Suicide Risk Lexicon, build lexicon with LLMs, and semantic similarity: https://github.com/danielmlow/construct-tracker
Audio and text
Many voice and speech datasets: Alden Blatter, Hortense Gallois, Samantha Salvi Cruz, Yael Bensoussan, Bridge2AI Voice Consortium, Maria Powell, Jean-Christophe Bélisle-Pipon. (2025). “Global Voice Datasets Repository Map.” Voice Data Governance. https://map.b2ai-voice.org/.
Bridge2AI Voice Dataset https://b2ai-voice.org/the-b2ai-voice-database/
Facebook's large-scale multimodal dataset of 4,000+ hours of human interactions for AI research: https://github.com/facebookresearch/seamless_interaction
Audio
CLAC: A Speech Corpus of Healthy English Speakers
Many speech datasets: https://github.com/jim-schwoebel/allie/tree/master/datasets#speech-datasets
Many audio visual datasets: https://github.com/krantiparida/awesome-audio-visual#datasets
Text
Many text datasets: https://lit.eecs.umich.edu/downloads.html#undefined
Many text datasets: https://github.com/niderhoff/nlp-datasets
Audio
Ramanarayanan, V., Lammert, A. C., Rowe, H. P., Quatieri, T. F., & Green, J. R. (2022). Speech as a biomarker: Opportunities, interpretability, and challenges. Perspectives of the ASHA Special Interest Groups, 7(1), 276-283.
Low, D. M., Bentley, K. H., & Ghosh, S. S. (2020). Automated assessment of psychiatric disorders using speech: A systematic review. Laryngoscope investigative otolaryngology.
Cummins, N., Scherer, S., Krajewski, J., Schnieder, S., Epps, J., & Quatieri, T. F. (2015). A review of depression and suicide risk assessment using speech analysis. Speech communication. link
Patel, R. R., Awan, S. N., Barkmeier-Kraemer, J., Courey, M., Deliyski, D., Eadie, T., ... & Hillman, R. (2018). Recommended protocols for instrumental assessment of voice: American Speech-Language-Hearing Association expert panel to develop a protocol for instrumental assessment of vocal function. American journal of speech-language pathology, 27(3), 887-905.
Text
Mihalcea, R., Biester, L., Boyd, R. L., Jin, Z., Perez-Rosas, V., Wilson, S., & Pennebaker, J. W. (2024). How developments in natural language processing help us in understanding human behaviour. Nature Human Behaviour, 8(10), 1877-1889.
Stade, E. C., Stirman, S. W., Ungar, L. H., Boland, C. L., Schwartz, H. A., Yaden, D. B., ... & Eichstaedt, J. C. (2024). Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation. NPJ Mental Health Research, 3(1), 12.
Low, D., Mair, P., Nock, M., & Ghosh, S. Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. link
Most recordings are available upon request.
06-July-2026 | Toward Trustworthy Speech-Based AI Models for Mental Healthcare Assessment and Decision Support |Theodora Chaspari (University of Colorado Boulder)
23-Mar-2026 | From Clinical Speech to Language Learning: Mispronunciation Detection Across Domains |Beena Ahmed (University of New South Wales)
26-Jan-2026 | Speech and Audio Intelligence for Health: Sensing, Reasoning, and Prediction | Ting Dang (The University of Melbourne, Australia)
15-Dec-2025 | Generating and investigating laryngeal biosignals| Andreas Kist (Friedrich-Alexander-Universität Erlangen-Nürnberg)
10-Nov-2025 | Speech as a modality for the characterization and adaptation of neurodiversity| Mark Hasegawa-Johnson (University of Illinois Urbana-Champaign)
06-Oct-2025 | From Noise to Signal: Individual Variability in Voice Fatigue Subtyping| Mark Berardi (University of Iowa)
19-May-2025 | Building your research team: Who should be in the room where it happens? | Maria Powell (Vanderbilt University Medical Center)
5-May-2025 | Clinical theory and dimensions of speech markers: Psychosis as a case study | Lena Palaniyappan (Professor of Psychiatry, McGill)
7-Apr-2025 | Toward generalizable machine learning models in speech, language, and hearing sciences: Estimating sample size and reducing overfitting | Hamzeh Ghasemzade (Massachusetts General Hospital – Harvard Medical School)
17-Mar-2025 | Exploring the Mechanistic Role of Cognition in the Relationship between Major Depressive Disorder and Acoustic Features of Speech | Lauren White (King’s College London)
3-Mar-2025 | Speech as a Biomarker for Disease Detection | Catarina Botelho (INESC-ID, University of Lisbon)
17-Feb-2025 | Exploring Intraspeaker Variability in Vocal Hyperfunction Through Spatiotemporal Indices of RFF | Jenny Vojtech (Boston University)
20-Jan-2025 | Clinically meaningful speech-based endpoints in clinical trials | Julie Liss (Arizona State University)
12-03-2024 | Revealing Confounding Biases: A Novel Benchmarking Approach for Aggregate-Level Performance Metrics in Health Assessments | Roseline Polle (Thymia)
10-16-2023 | Estimation of parameters of the phonatory system from voice | Zhaoyan Zhang (UCLA Head and Neck Surgery)
18-Nov-2024 | The interplay between signal processing and AI to archive enhanced and trustworthy interaction systems | Ingo Siegert (Otto-von-Guericke-University Magdeburg)
4-Nov-2024 | Remote Voice Monitoring System for Patients with Heart Failure | Fan Wu (ETH Zurich)
6-May-2024 | Building Speech-Based Affective Computing Solutions by Leveraging the Production and Perception of Human Emotions | Carlos Busso (UT Dallas)
1-Apr-2024 | Parkinson's speech | Godino-Llorente
4-Mar-2024 | Modelling individual and cross-cultural variation in the mapping of emotions to speech prosody | Pol van Rijn (Max Planck Institute for Empirical Aesthetics)
5-Feb-2024 | Speech Analysis for Intent and Session Quality Assessment in Motivational Interviews | Mohammad Soleymani (USC)
20-Nov-2023 | High speed videoendoscopy | Maryam Naghibolhosseini (Michigan State University)
18-Sep-2023 | Democratizing speaker diarization with pyannote | Hervé Bredin (Institut de Recherche en Informatique de Toulouse) and Marvin Lavechin (Meta AI, ENS)
7-Aug-2023 | Overview of Zero-Shot Multi-speaker TTS Systems | Edresson Casanova (Coqui)
17-Jul-2023 | Considerations for Identifying Biomarkers of Spoken Language Outcomes for Neurodevelopmental Conditions | Karen Chenausky (Harvard Medical School & Massachusetts General Hospital)
15-May-2023 | The Potential of smartphones voice recordings to monitor depression severity | Nicholas Cummins (King's College London)
1-May-2023 | Reading group session: Introductory overview of self-supervised learning, transformers, and attention | Daniel Low (Harvard & MIT)
20-Mar-2023 | Accuracy of Acoustic Measures of Voice via Telepractice Videoconferencing Platforms | Hasini Weerathunge (Boston University)
6-Mar-2023 | Casual discussion on audio quality control and preprocessing | Daniel Low (Harvard & MIT)
26-Jan-2023 | Developing speech-based clinical analytics models that generalize: Why is it so hard and what can we do about it? | Visar Berisha (Arizona State University)
12-Jan-2023 | Using knockoffs for controlled predictive biomarker identification | Kostas Sechidis (Novartis)
15-Dec-2022 | Provide ideas and feedback on the protocol for a large-scale data collection effort (N=5k) on mental health and voice from the NIH Bridge2AI | Daniel Low (Harvard & MIT)
1-Dec-2022 | Inferring neuropsychiatric conditions from language: how specific are transformers and traditional ML pipelines in a multi-class setting? | Lasse Hansen (Aarhus University) & Roberta Rocca (Aarhus University)
17-Nov-2022 | Meet and greet/intros |
3-Nov-2022 | Speech and Voice-based Detection of Mental and Neurological Disorders: Traditional vs Deep Representation and Explainability | Bjorn Schuller (Imperial College London)
20-Oct-2022 | What do machines hear? Overview of deep learning approaches for representing voice | Gasser Elbanna (EPFL & MIT)