MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
Wei-Jaw Lee (National Taiwan University, Taiwan)
This work integrates explicit and precise phoneme conditioning into a lyrics-to-song (LTS) generation model. To our best knowledge, this is the first attempt to introduce singing voice synthesis (SVS) control into the LTS paradigm.
A Pilot Study on Batch Sampling Using Clustering in Text-to-Music Generation
Shunsuke Yoshida (The University of Tokyo, Japan)
This work investigates the effect of batch sampling strategies during training for text-to-audio music generation. Training data are clustered using either text or audio embeddings, and samples with similar characteristics are grouped within the same mini-batch to mitigate gradient interference. Results show that clustering based on text embeddings achieves better performance on objective evaluation metrics than clustering based on audio embeddings. In addition, different cluster granularity leads to different behaviors across evaluation criteria: a moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to exhibit music with more coherent structure in listening tests.
Automatic BGM generation using play scripts and directions
Masahiro Shimizu (Nagoya Institute of Technology, Japan)
We are researching the automatic generation of theatrical background music using play scripts and directions. Background music plays an important role in theater, such as during scene changes and in expressing emotions. However, composing and arranging background music is difficult for beginners. Our research contributes that even beginners can easily create appropriate background music for plays. In this presentation, we would like to mainly discuss two primary challenges: how to effectively reflect the nuances of a theatrical performance in music, and how to evaluate the quality of the tracks generated, and other related topics.
Cover Song Generation: From Symbolic Foundations to Audio-Domain Frontiers
Chih-Pin Tan (National Taiwan University, Taiwan)
As a form of music style transfer, cover song generation preserves the core identity of source tracks while altering their style. This session outlines the technological evolution of the field, from symbolic domain foundations to recent audio-domain breakthroughs, and introduces a novel research direction addressing current challenges in the space.
Towards Development of Symbolic Music Foundation Model
Takaaki Nagoshi (Nihon University, Japan)
We show that block-reordering pretraining enables a single decoder-only model to unify music generation and analysis, avoiding the inefficiency of post-hoc tuning due to pretraining ossification.
Improving Music Listening for People with Hearing Loss
Mana Morimoto (Nagoya Institute of Technology, Japan)
I am researching ways to improve music listening for people with hearing loss. Previously, I proposed a signal processing method that amplifies high-frequency bands to enhance consonant audibility, focusing on vocal intelligibility. In the future, I hope to develop conversion methods tailored to each person—for example, using machine learning. However, there is significant individual variation in how people with hearing loss perceive sound, and I face the challenge of not knowing which methods are effective. I would like to introduce my current progress and discuss effective ways to understand individual hearing characteristics and improve music accessibility together.
Introducing a Unified Theory for Harmonic and Rhythmic Perception (HARP)
Wei-Huai Chen (National Taiwan University, Taiwan)
This paper introduces A Unified Theory for Harmonic and Rhythmic Perception (HARP), a computational framework that jointly models rhythmic timing and harmonic structure within a single generative model. HARP represents musical events using hidden Markov states that capture tempo, harmony, and their transitions, integrating principles from music theory and psychoacoustics. Perceptual inference is formulated using the Free Energy Principle, with hidden musical states estimated by minimizing variational free energy from observed musical events.
(spare slot)
Music Information Retrieval for Ethiopian Orthodox Liturgical Chant: Challenges and Opportunities
Mequanent Argaw Muluneh (Academia Sinica, Taiwan)
Understanding the Localization of Production Practice A Cross-Modal Narrative of Post-War Taiwanese Popular Music
Ming Cheng (Academia Sinica, Taiwan)
This study introduces a framework combining Social Network Analysis (SNA) and AI-assisted audio analysis to examine post-war Taiwanese popular music production (1975-1995). By analyzing collaboration networks alongside high-dimensional audio embeddings from isolated instrumental tracks, we quantify the "sonic identifiability" of socio-technical actors like arrangers and studios. This interdisciplinary approach provides a computational method to validate qualitative historical insights into the power dynamics and localized practices of music production.
Composer's Dilemma
Satoshi Tojo (Asia University, Japan)
This study elucidates how composers achieved a balance between familiarity (elements readily accepted by existing audiences) and novelty (innovative elements). It employs two theoretical frameworks: (i) cross-entropy, to measure the divergence between market demands and composers' strategies; and (ii) game theory, to analyze strategic choices made in anticipation of future shifts in trends.