We develop human-centered speech and language technologies that can understand, generate, and protect the human voice. Our research focuses on expressive speech synthesis and voice conversion, large audio-language models, affective computing, speech deepfake detection and voice security, and speech and language technologies for health. Across these areas, we aim to build systems that capture not only what people say, but also how they say it, while remaining secure, trustworthy, and useful in real-world applications.
At JHU, our lab holds weekly subgroup meetings focused on the following research areas. Each student is assigned to one of these groups:
Speech and Language for Security
LLMs, Emotion, Synthesis (TTS, VC)
Speech and Language for Health
Expressive Speech Translation (jointly with Philipp Koehn's group)