HARMONYCLOAK: Making Music Audio Unlearnable for Generative AI
HARMONYCLOAK: Making Music Audio Unlearnable for Generative AI
Abstract: Recent advances in generative AI have significantly expanded into the realms of art and music. This development has opened up a vast realm of possibilities, pushing the boundaries of human creativity into unexplored frontiers. However, as generative AI continues to advance, it can replicate artistic styles and produce new artwork, posing significant concerns for the perceived rarity and value of artists' creations. In response to these challenges, it is becoming increasingly crucial to establish and enforce protective measures that safeguard artists' copyrighted work from unauthorized exploitation by generative AI models. In this paper, we introduce the first defensive mechanism, HarmonyCloak, to prevent the exploitative use of artwork, specifically in the context of music, by generative AI models. Particularly, HarmonyCloak employs imperceptible error-minimizing noise to make the model's generative loss approach zero for these perturbed music data, tricking the model into believing nothing can be learned so as to disrupt their attempts to replicate musical structures and styles. By using a set of intra-track and inter-track objective metrics and a subjective user study, extensive experiments on three state-of-the-art music generative AI models (i.e., MuseGAN, SymphonyNet, and MusicLM) validate the effectiveness and applicability of HarmonyCloak in both white-box and black-box settings.
Training Samples
Unlearnable
Music (White-box).
Unlearnable
Music (Black-box).
Music Generated by MuseGAN
Generated by MuseGAN Trained on Clean Music.
Generated by MuseGAN Trained on 15% Unlearnable Music (White-box).*
Generated by MuseGAN Trained on 15% Unlearnable Music (Black-box).*
*All the percentages are the percentage of unlearnable examples in the training data while rest of the data is clean music samples.
Music Generated by SymphonyNet
Generated by SymphonyNet Trained on Clean Music.
Generated by SymphonyNet Trained on 15% Unlearnable Music (White-box).*
Generated by SymphonyNet Trained on 15% Unlearnable Music (Black-box).*
*All the percentages are the percentage of unlearnable examples in the training data while rest of the data is clean music samples.
Music Generated by MusicLM
Generated by MusicLM Trained on Clean Music.
Generated by MusicLM Trained on 15% Unlearnable Music (White-box).*
Generated by MusicLM Trained on 15% Unlearnable Music (Black-box).*
*All the percentages are the percentage of unlearnable examples in the training data while rest of the data is clean music samples.
Music Generated by MuseGAN Trained on More Unlearnable Music
Trained on 30% Unlearnable Music.*
Trained on 50% Unlearnable Music.*
Trained on 100% Unlearnable Music.*
*All the percentages are the percentage of unlearnable examples in the training data while rest of the data is clean music samples.
HARMONYCLOAK vs. Norm-Based Noise
Defensive noise generated using different noise generators (HARMONYCLOAK vs Norm-Based noise generators)