Zero-Shot Vocal Timbre Conversion as a DAW Plugin
CS 352 Machine Perception in Music, Professor Bryan Pardo, Northwestern University
Group Member: Ben Cole, Benjamin Auby, Alan Wang
Modern music production increasingly treats the voice like an instrument: producers want to audition the same take through different vocal “characters,” explore alternate vocal colors for a hook, or shape a performance without re-recording. Traditional vocal effects (EQ, pitch shifting, formant shifting, saturation, reverb) can change tone, but they often sound like effects applied rather than a coherent, natural change in the singer’s identity.
Voice cloning has already entered mainstream music: creators can record a take and use a model to render it in the timbral style of a famous artist, while keeping the original delivery. But for real creative use, “exactly like X” is often less interesting (and raises ethical questions) compared to “somewhere between me and X.” Producers want a controllable continuum of vocal color that still sounds like a single, coherent singer
Research in voice conversion (VC) makes this possible by separating what is being said from who is saying it: converting speaker identity while preserving linguistic content and timing. However, singing voice conversion (SVC) is harder: high-quality results must preserve musical structure (pitch trajectories, vibrato, rhythm, phrasing, and sustained articulation) where small errors are even more obvious. Historically, many strong methods required target-speaker training or extensive adaptation, which makes them impractical for fast, creative workflows.
TimbreTune is our attempt to turn recent zero-shot conversion models into a music-producer–oriented timbre morphing tool. Given a source vocal take and a few seconds of reference audio, TimbreTune preserves the source performance (timing, melody, phrasing) while transferring and interpolating vocal timbre. Instead of aiming for “perfect cloning,” our goal is controllable in-between voices: a continuous slider that lets users blend vocal character between source and reference, packaged as a native DAW plugin (VST3/AU/Standalone) as well as a WebUI/CLI for experimentation.
TimbreTune builds off of Seed-VC, a diffusion-based architecture that disentangles content from speaker identity. The pipeline includes:
1. Linguistic Extraction: A Whisper encoder strips the source audio down to its pure linguistic content, removing timbre and pitch.
2. Speaker Embedding: A CAMPPlus speaker embedding model captures the reference voice's timbre as a high-dimensional vector.
3. Reconstruction: A diffusion Transformer (DiT) conditioned on both the linguistic content and the target speaker embedding reconstructs the mel-spectrograms.
4. Vocoder: BigVGAN converts these generated spectrograms back into 44.1kHz waveforms.
Our addition to the pipeline: Timbre Morphing
TimbreTune introduces controllable timbre morphing. Instead of forcing a 100% conversion, speaker embeddings from the source and reference audio are blended via spherical linear interpolation (SLERP). This allows a user to smoothly dial continuously (using an alpha slider from 0.0 to 1.0) between their original voice and the new vocal identity.
DAW Integration
The system features a Python inference server and a JUCE C++ plugin that communicates with it. Heavy deep learning computation runs on the backend, while the musician interacts through a familiar plugin interface with sliders for diffusion steps, pitch shifting, and the morph amount.
DAW vst3 & AU support
Record, upload, export audio for source/reference and output
Easily iterate on vocals within DAW
Tune-able Parameters
Diffusion Steps (1-200, Default: 10): Dictates how many iterative denoising steps the Diffusion Transformer (DiT) uses to mathematically synthesize the mel-spectrogram. Lower values prioritize fast, draft-quality generation but inherently leave residual acoustic noise (static or "crunch") in the signal. Higher values (50-100) yield the cleanest, highest-fidelity vocal synthesis at the cost of slightly longer processing times.
Length Adjust: Acts as a temporal rate interpolator to multiply the total duration of the generated audio without artificially pitching the vocals. Values below 1.0 will speed up the rhythm and pronunciation, whereas values above 1.0 will stretch and slow down the underlying performance.
Inference CFG (Classifier-Free Guidance) Rate: Determines how aggressively the model forces the audio to adhere to the requested target voice embedding. Expanding the CFG actively amplifies the distinct acoustic characteristics of the target, though values tuned too high can inadvertently multiply prediction noise into synthetic distortion.
Morph Amount: The core continuous parameter weighting the spherical interpolation (SLERP) between the source and target voice identities. A value of `0.0` reconstructs the original source timbre perfectly, while `1.0` triggers a full zero-shot conversion into the target reference. Intermediary values (like `0.5`) force the diffusion model to synthesize entirely new, biologically fused vocal tracts in the center of the latent space.
Auto F0 Adjust: Automatically extracts and aligns the foundational pitch contour (F0) of the source audio to roughly match the natural acoustic register of the target speaker. While useful for seamlessly blending mismatched demographics, disabling this feature can sometimes rescue the audio from strange robotic artifacts if the vocoder struggles with extreme pitch clashes.
Pitch Shift: An absolute mathematical transposition applied to the localized pitch contour before it is fed to the diffusion model. It is represented in continuous semitones (e.g., `-12` immediately drops the entire vocal performance down a full musical octave).
We evaluated 5 distinct voice pairs (e.g., source -> target), generating 21 morphed outputs per pair corresponding to values from 0.0 to 1.0 in 0.05 increments. We extracted the speaker embeddings for all outputs and computed their cosine similarity against both the source and target audio.
To ensure robustness, we ran these metrics using two distinct speaker embedding models:
1. CAMPPlus (The model used internally by TimbreTune)
2. ECAPA-TDNN (https://github.com/speechbrain/speechbrain) (SpeechBrain, used as an external baseline verification)
We define Intermediateness as `1.0 - abs(sim_to_source - sim_to_target)`. A true morph midpoint occurs when the similarity to the source equals the similarity to the target (an intermediateness score of ~1.0).
To rigorously prove that Sphereical Linear Interpolation (SLERP) is geometrically superior for traversing speech manifolds, we implemented and evaluated four distinct baselines at the 0.5 midpoint intersection:
Simple Mixing: A naive 50/50 additive waveform overlap without neural modeling.
LERP: Standard planar linear interpolation between the source and target speaker embeddings.
Cosine / Sigmoid: Smooth easing/activation bounds applied to the linear interpolation gradient.
The results clearly demonstrate that TimbreTune functions as a true morphing tool and mathematically validate our usage of SLERP over traditional interpolation bounds.
1. Midpoint Verification: At alpha=0.5, the SLERP output is a true geometric in-between of the source and target, scoring a 0.867 Intermediateness on CAMPPlus. Both the standard mixing fallback and the LERP / Cosine derivations fall objectively flat, demonstrating that the internal embedding manifolds are strictly hyperspherical.
2. Smoothness: As the morph amount increases from 0 to 1, the similarity to the target strictly increases monotonically, while similarity to the source strictly decreases.
*(Note: Baselines comparing against full Seed-VC conversion and zero-morph source controls confirm that the endpoints match the expected boundaries.)*
We sometimes hear some “crunchy” distortion in the generated audio. Conceptually, it comes from an intersection of (1) diffusion latent-space effects and digital quantization limits: at intermediate `alpha` values the model is asked to synthesize audio from a point that may be slightly out-of-distribution for the training manifold, (2) higher `inference_cfg_rate` can amplify mismatch between the conditioned and unconditioned generations (CFG extrapolation), and (3) when the vocoder produces a waveform whose peaks exceed the digital headroom, the final export to 16-bit PCM (and int16 conversion for MP3) can hard-clip those peaks, producing audible jaggedness.
Another observation: the “true 0.5” blend is slightly imperfect when verified with the secondary embedding model (ECAPA-TDNN) and also by ear. This is expected because embedding models have different feature geometries (so equal SLERP geometry in one model does not guarantee an identical crossover point in another), and because the mel-to-audio decoder is highly non-linear, so equal similarity in embedding space does not always map to perfectly equal perceptual timbre in the final waveform. Prompt scaling / CFG interaction and pitch alignment (Auto F0) can further shift the effective midpoint toward one endpoint.
In future iterations, we could reduce these artifacts by tuning CFG/diffusion steps more conservatively per alpha, improving output stability (ex, peak normalization/limiting before PCM export, or saving with higher bit depth/headroom), and strengthening the overlap strategy (larger overlap or smoother windowing) to minimize boundary transients. For midpoint calibration, we could also fit a small alpha remapping per embedding model so that the similarity crossover aligns more precisely with the desired perceptual “0.5”.
Voice morphing technologies could potentially raise concerns around misuse. In some cases, the system could be used to imitate or replicate another person's voice without consent, leading to irreversible consequences. We recognize this possibility and the potential risks, and frame the project around responsible use: users should not apply the system to someone’s voice without consent, and outputs should not be used in contexts where authenticity or identity verification matters.
TimbreTune is designed primarily as a timbre‑morphing tool, preserving the source performance (timing, melody, phrasing) while transferring/interpolating vocal color from a short reference clip. While the morph control can approach strong timbre transfer, it does not explicitly reproduce a person’s full vocal identity (e.g., their characteristic articulation and performance style), and sounding convincingly like a specific individual would still require performing like them. We also avoid training or fine-tuning on identifiable individuals without permission, and we prioritize self-recorded or consented audio in demonstrations to reduce harm and discourage misuse.
1. Plachtaa. (2024). Seed-VC: Zero-shot Voice Conversion. [GitHub Repository](https://github.com/Plachtaa/seed-vc)
2. OpenAI. (2022). Whisper: Robust Speech Recognition. [GitHub Repository](https://github.com/openai/whisper)
3. FunASR. (2023). CAMPPlus Speaker Embedding. [Model Repository](https://github.com/alibaba-damo-academy/FunASR)
4. Lee, S., et al. (2022). BigVGAN: A Universal Neural Vocoder with Large-Scale Training. [GitHub Repository](https://github.com/NVIDIA/BigVGAN)
5. SpeechBrain. (2021). ECAPA-TDNN Speaker Embedding. [Model Page](https://github.com/speechbrain/speechbrain)
6. VoxMorph (https://arxiv.org/abs/2601.20883)
Contact bencole2026@u.northwestern.edu, alanwang2026@u.northwestern.edu, benjaminauby2026@u.northwestern.edu to get more information on the project
Github Repo: https://github.com/bc2k13/timbretune