PhD position — Privacy-Preserving Clinical Speech Understanding
We are recruiting a PhD candidate to develop privacy-preserving methods for clinical speech understanding, combining speech processing, multimodal machine learning, voice anonymisation and AI for health. The position is hosted by the MULTISPEECH team at LORIA (Université de Lorraine, Inria and CNRS), within the AI Grand Est ENACT research cluster.
Start: October or November 2026
Duration: 3 years
Location: Villers-lès-Nancy, France
See the full position description, candidate profile, and application procedure below.
French title: Compréhension de la parole préservant la vie privée à partir de signaux multimodaux pour les applications cliniques
This fully funded, three-year PhD position is hosted by the MULTISPEECH team at LORIA (Université de Lorraine, Inria and CNRS), in Villers-lès-Nancy, France. The project is funded by the AI Grand Est ENACT research chair.
Host institution: Université de Lorraine, Grand Est, France
Doctoral school: IAEM Lorraine — Computer Science, Automatic Control, Electronics/Electrical Engineering and Mathematics
Discipline: Computer Science
Research unit: LORIA — MULTISPEECH team
Employer: Université de Lorraine
Funding period: 1 October 2026 to 30 September 2029
Expected start date: 1 October or 1 November 2026
Salary: €2,300 gross per month
Application deadline: Until the position is filled; final deadline 30 August 2026, extendable to 20 September 2026 if no suitable candidate is selected.
Natalia Tomashenko, Research Chair holder, AI Grand Est ENACT cluster; MULTISPEECH team, LORIA, Université de Lorraine.
Additional scientific supervision will be provided within the MULTISPEECH team.
Spoken Language Understanding (SLU) is becoming central to digital health, powering voice-based clinical assistants, automatic documentation and triage support. Yet clinical speech is among the most sensitive data that exist: it carries both what is said (verbal content, including names, addresses and dates that directly identify a patient) and how it is said (paralinguistic cues encoding speaker identity, age, gender, accent, emotional state and health biomarkers).
The core scientific difficulty is that identity cues and clinically valuable cues are entangled in the same acoustic signal. A patient calling emergency services may reveal diagnostic signals, such as dyspnea, speech rate or breathing pattern, that should be retained, while also revealing a unique voiceprint and explicit identifiers that must be removed. Existing anonymization systems can degrade precisely the properties that make recordings clinically useful.
This PhD is part of the ENACT chair on Privacy-preserving multimodal spoken language understanding for digital health and medical education. It aims to develop clinical-speech anonymization methods that remove speaker identity and spoken identifiers while preserving diagnostic and semantic information required for clinical understanding. The work will leverage speech and audio large language models (speech-LLMs) and rigorously measure the privacy–utility trade-off.
The reference use case is the SAMU acute-dyspnea emergency call. An additional modality—such as physiological signals, imaging or EHR metadata—is an explicit and valued extension of the topic.
Utility-preserving speech anonymization. Develop voice-anonymization and content de-identification methods that remove speaker identity and named entities while retaining paralinguistic biomarkers relevant to the target clinical task, for example respiratory-distress cues.
Privacy-aware clinical SLU based on speech-LLMs. Develop speech-and-text SLU models that jointly exploit acoustic and linguistic content for tasks such as intent and entity detection, symptom extraction and severity estimation. The models will operate on anonymized inputs and produce structured, interpretable outputs.
Measuring the trade-off. Define and validate an evaluation methodology combining privacy attacks with clinical-utility metrics, to quantify privacy gains and corresponding utility losses.
Additional modality and defenses. Extend the pipeline to an additional modality, preferably physiological signals for the dyspnea use case, and investigate privacy-preserving defenses such as improved disentanglement, privacy-aware learning objectives, differential privacy or federated learning where appropriate.
The PhD will build on a de-identified speech-and-text corpus relevant to a clinical use case, for example emergency/SAMU calls or another acute-care scenario to be finalized with clinical partners. Depending on availability, data may be reused from existing collaborations or collected under HDS-certified health-data infrastructure and ethical approval.
The work will be iterative and may include:
Corpus selection, curation and/or annotation of intents, clinical entities, symptoms and severity.
Development of anonymization baselines, including voice anonymization, named-entity recognition and rule-based text masking.
Privacy evaluation through speaker verification, membership inference and attribute-inference attacks.
Clinical-utility evaluation through task performance and, where feasible, expert assessment.
Development of speech-and-text SLU models based on speech/audio-LLM backbones and disentangled representations.
Extension to multimodal inputs and input- or model-level privacy defenses.
Reproducible methodologies and evaluation protocols for privacy–utility trade-offs in clinical speech understanding.
Privacy-aware clinical SLU models based on speech-LLMs and operating on anonymized inputs.
Attack methods and clinical-utility metrics that quantify residual privacy leakage and utility loss.
Multimodal privacy-preserving approaches combining speech with an additional modality.
Contributions to open privacy-evaluation frameworks for clinical speech, building on the VoicePrivacy initiative and its challenge series.
Scientific dissemination through conference presentations, publications and contributions to training.
The successful candidate will join the MULTISPEECH team at LORIA (Université de Lorraine / Inria / CNRS). The team provides expertise in speech recognition and synthesis, multimodal speech and disentangled representations.
The project offers an interdisciplinary environment at the interface of speech AI, emergency medicine, public health and data governance, health law, and biomedical imaging/signals. It will involve collaboration with clinical and scientific partners at CHRU Nancy, including emergency medicine, public health and clinical research, health law and ethics, and biomedical imaging/signals.
The candidate will have access to dedicated computing resources, speech-processing software, deep-learning frameworks and ENACT computational and engineering support. Clinical data will be handled in accordance with ethical and GDPR requirements, using secure HDS-certified health-data infrastructure where applicable.
The candidate should hold a Master’s degree or engineering degree in computer science, artificial intelligence, signal or speech processing, applied mathematics, data science or a related field.
Strong Python programming skills.
Good knowledge of machine learning and deep learning.
Experience with PyTorch, PyTorch Lightning and Hugging Face Transformers, or another deep-learning framework.
Interest in speech processing and multimodal data processing, including audio, text and signals.
Ability to design, train, evaluate and compare models; familiarity with statistical evaluation methods.
Scientific rigor, autonomy and critical thinking.
Ability to work with clinical teams and translate medical questions into modelling problems.
Speech processing and natural language processing.
Privacy, anonymization, differential privacy or federated learning.
Multimodal analysis, explainable AI or AI for healthcare.
Experience with sensitive or clinical data.
Prior scientific publications are appreciated but not required.
Please send your application by email. Include:
A detailed CV.
A motivation letter explaining your interest in the topic and suitability for the position.
If possible, one or two letters of recommendation, or the contact details of one or two academic or professional referees who may be contacted.
Bachelor’s and Master’s degree transcripts.
Applications will be reviewed as they are received. Shortlisted candidates will be invited to an interview.
Natalia Tomashenko
Research Chair holder, AI Grand Est ENACT cluster
MULTISPEECH team, LORIA, Université de Lorraine
natalia.tomashenko@inria.fr
European Parliament and Council. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, 2024.
European Parliament and Council. Regulation (EU) 2016/679 (General Data Protection Regulation, GDPR). Official Journal of the European Union, 2016.
N. Tomashenko et al. “Introducing the VoicePrivacy Initiative.” Proceedings of Interspeech, 2020.
N. Tomashenko et al. “The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization.” Computer Speech & Language, 2026.
B. M. L. Srivastava et al. “Privacy and Utility of x-Vector Based Speaker Anonymization.” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022.
S. T. Arasteh et al. “Addressing Challenges in Speaker Anonymization to Maintain Utility While Ensuring Privacy of Pathological Speech.” Communications Medicine, 2024.
C. Diaz-Asper et al. “Navigating the Trade-Off Between Personal Privacy and Data Utility in Speech Anonymization for Clinical Research.” npj Digital Medicine, 2025.
I. Cohn et al. “Audio De-Identification: A New Entity Recognition Task.” Proceedings of NAACL-HLT, 2019.
A. Radford et al. “Robust Speech Recognition via Large-Scale Weak Supervision.” 2022.
K. B. Johnson et al. “OBSERVER: A Novel Multimodal Dataset for Outpatient Care Research.” JAMIA, 2025.
R. Wang and H. Lin. “Anonymizing Facial Images to Improve Patient Privacy.” Nature Medicine, 2022.
S. T. Arasteh et al. “The Impact of Speech Anonymization on Pathology and Its Limits.” arXiv:2404.08064, 2024.
X. Zhong et al. “Considerations for Patient Privacy of Large Language Models in Health Care: A Scoping Review.” Journal of Medical Internet Research, 2025.