This page is no longer updated. My homepage has moved to sw005320.github.io
End-to-End Speech Processing Toolkit
ESPnet is an open-source toolkit for speech recognition, text-to-speech, speech enhancement, speech translation, and spoken language understanding. It provides reproducible recipes and a complete setup for speech foundation model research.
Versatile Evaluation of Speech and Audio
VERSA is a toolkit for evaluating speech and audio quality. It provides seamless access to over 90 evaluation and profiling metrics with 10x variants, assessing audio through multiple dimensions.
Open Whisper-style Speech Models
OWSM reproduces Whisper-style training using publicly available data and ESPnet. Data preparation scripts, training and inference code, pre-trained model weights, and training logs are all publicly released.