A Knowledge Tracing Model for Sight-Singing Based on Skill Decomposition and Multi-View Modeling
Hiroki Karatsu (AIST / Tokyo University of the Arts,Japan)
In this study, we aim to enhance sight-singing education by improving the performance of a knowledge tracing model for sight-singing that estimates learner proficiency from learning histories and predicts performance. Building on prior work that employs a GNN-based knowledge tracing model, we propose a multi-view model that decomposes sight-singing skill into three elements: pitch, interval, and rhythm. Experimental evaluation using a large-scale sight-singing dataset demonstrates that the proposed method achieves higher ROC-AUC scores than conventional approaches, suggesting the effectiveness of multi-view modeling over decomposed skill elements.
SymphoMOS: Improved Singing MOS Prediction with Hybrid Speech and Music Self-Supervised Learning Models
Ryoko Arita (The University of Tokyo, Japan)
We propose SymphoMOS, a singing MOS prediction system that incorporates cross-domain self-supervised learning (SSL) fusion using speech and music SSL models for feature extraction. We instantiate our SymphoMOS using two representative models, wav2vec 2.0 for speech and MERT for music. The results demonstrate that SymphoMOS achieves the state-of-the-art system-level Spearman's rank correlation coefficient for singing MOS prediction on the SingMOS dataset.
Learning Motif-centric Representations for Motive Discovery and Structure Similarity in Symbolic Music Data
Jun-You Wang (National Taiwan Normal University, Taiwan)
Discovering and Steering Musical Features from Neural Audio Codecs via Sparse Autoencoders
Chih-Cheng Chang (Academia Sinica, Taiwan)
The internal representations of neural audio codecs remain uninterpretable "black boxes" where semantic features are deeply entangled. To enhance mechanistic interpretability, we apply sparse autoencoders to the latent spaces of neural audio codecs. By projecting dense activations into an overcomplete sparse feature space, we decompose these representations into discrete features. Analysis reveals a hierarchical progression of musical features from low-level pitch to high-level timbre. Quantitative evaluation via MonoSemanticity and linear probing demonstrates that SAEs enhance linear separability and semantic coherence over the original latents. Besides, we validate the causal influence of these features through latent steering, which allows high-fidelity manipulation of tonal and timbral characteristics while maintaining the integrity of the model reconstruction.
A Large-Scale Dataset for Street Dance Performance Evaluation
Hsuan-Kai Kao (University of Tsukuba, Japan)
Despite the growing popularity of street dance competitions, there is currently no standardized evaluation framework for assessing street dance performance quality. To address this challenge, we introduce a new 3D dance dataset specifically designed for street dance performance evaluation. The dataset consists of 4,281 individual dance performances, along with evaluation-related annotations. Building upon this dataset, we propose a music-stem-aware automated street dance performance evaluation system. Extensive experiments demonstrate that the proposed system can distinguish different levels of dance performance and achieve a Spearman correlation of approximately 0.6 with scores assigned by professional judges.
Exploring a Sense of Togetherness in Musical Ensemble Performance
Karen Kurotaki (University of Tsukuba, Japan)
This talk explores a sense of togetherness in musical ensemble performance through the lens of embodied interaction. In ensemble playing, performers coordinate not only through sound, but also through subtle bodily cues such as gaze, breathing, posture, and movement. Focusing on the Japanese concept of maai, or relational timing and spacing, I discuss how performers become aware of others during ensemble performance. Drawing on video observations, questionnaires, and interviews with string players, this talk considers how togetherness may emerge through bodily, auditory, and interpersonal interaction.
Graphical interface for visualising changes in flute timbre whilst suppressing variations from pitch differences
Kai Hiraiwa (AIST, Japan
An Automatic Music Performance Animation Generation System for Virtual Violin Performance
Ting-Wei Lin (National Chung Hsing University, Taiwan)
AI-driven Cross-Model Mapping of Musical Cognition and Performance
Yu-Fen Huang (National Taiwan University of the Arts, Taiwan)
Diffusion-Based Synthesis Parameter Estimation using Differentiable Digital Signal Processing Mixture Model
Kengo Takemoto (The University of Tokyo / AIST, Japan)
Estimating interpretable synthesis parameters, such as fundamental frequency (F0), loudness, and timbre feature, for each source from instrumental mixtures is important for music audio editing. For this task, we introduce a diffusion model as a generative model of synthesis parameters and formulate estimation as an inverse problem of generating source-wise parameters that reproduce the observed mixture through differentiable digital signal processing mixture model (DDSPMM). Specifically, the proposed method guides the reverse diffusion process using the reconstruction error between the observed and DDSPMM-synthesized mixtures. Experiments on instrumental ensembles show that the proposed method improves estimation accuracy.