Yui Sudo, Ph.D.
Chief Research Scientist at SB Intuitions
yui.sudo_at_sbintuitions.co.jp
Curriculum Vitae (updated Jul. 2026)
Chief Research Scientist at SB Intuitions
yui.sudo_at_sbintuitions.co.jp
Curriculum Vitae (updated Jul. 2026)
Yui Sudo received his B.S. and M.S. degrees from Keio University in 2009 and 2011, respectively, and his Ph.D. degree from Tokyo Institute of Technology in 2021. He joined Honda Motor Co., Ltd. in 2011 working across several Honda Group companies, including Honda Research Institute Japan Co., Ltd., until 2025. He is currently a Chief Research Scientist at SB Intuitions Corp. and a Team Leader at Noetra Corp. His research interests include automatic speech recognition (ASR), focusing on contextualization, streaming processing, decoding algorithms, multi-speaker ASR, and speech foundation models. He received the IEEE SLT Best Paper Award, IEEE ASRU Best Reviewer Award, JSAI Incentive Award, and JSPE Young Researcher Award. He is a member of IEEE and ISCA.
🎓️Education
Mar. 2021: Ph.D. degree in Engineering, Tokyo Institute of Technology, Japan.
Mar. 2011: M.S. degree in Engineering, Keio University, Japan.
Mar. 2009: B.S. degree in Engineering, Keio University, Japan.
💼Work Experience
Jul. 2026 - Present: Team Leader / Manager, Noetra Corp.
Feb. 2025 - Present: Chief Research Scientist / Manager, SB Intuitions Corp.
Apr. 2011 - Jan. 2025: Honda Group
Dec. 2020 - Jan. 2025: Senior Engineer, Honda Research Institute Japan Co., Ltd.
Feb. 2019 - Nov. 2020: Staff Engineer, Honda R&D Co., Ltd.
Apr. 2012 - Jan. 2019: Staff Engineer, Honda Engineering Co., Ltd.
(Apr. 2016 - Sep. 2016: Honda Engineering North America Inc.)
Apr. 2011 - Mar. 2012: Engineer, Honda Motor Co., Ltd.
Co-authored paper received ISCA INTERSPEECH Best Student Paper Award, 2025.
IEEE SLT Best Paper Award, 2024.
JSAI Incentive Award, 2024.
JSPE Young Researcher Award, 2012.
Y. Sudo et al., “Contextualized Automatic Speech Recognition with Dynamic Vocabulary”, in Proc. SLT, 2024. (🏆IEEE SLT Best Paper Award🏆)
Y. Sudo et al., “Joint Beam Search Integrating CTC, Attention, and Transducer Decoders”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2025.
Y. Peng, S. Muhammad, Y. Sudo, W. Chen, J. Tian, J. Lin, and S. Watanabe, "OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning", in Proc. INTERSPEECH, 2025. (🏆ISCA Best Student Paper Award 2025🏆️)
L. Liu, S. Zhu, K. Washizaki, R. Yoneyama, H. Jeon, M. Zhao, Y. Fujita, H. Shi, N. Yoshida, Y. Gao, R. Koshkin, Y. Hono, Y. Sudo, "Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis ", arXiv preprint arXiv:2606.25369.
🧰Open Models and Datasets
Sarashina2.2-TTS (2026): An LLM-based Japanese-centric TTS model.
NEST-Ja (2026): A Japanese speech SSL model trained on 35,000 hours of audio.
DiaFill (2026): A Japanese dialogue script generation model with natural backchannels and fillers.
VoiceBench-Ja (2026): A Japanese spoken QA benchmark dataset.
Joyo-Kanji-Yomi-Benchmark (2026): A Kanji-level pronunciation evaluation benchmark for Japanese TTS.