Sound naturally unfolds over time. It has attack, continuation, rhythm, resonance, decay, repetition, and memory.
This makes it particularly compatible with gesture.
A hand does not have to control pitch or volume directly.
It can inject energy into a process that produces sound.
It can create an object that continues to sound after the hand is gone.
It can change the state of a system whose sound slowly evolves.
Gesture can create the cause.
Sound can reveal the evolving consequence.
"Why are we learning sound synthesis?"
Yes, audio playback is sometimes good enough.
But what if 1000 particles make sound simultaneously?
"boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", .......... "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing", "boing",
Should they sound always the same?
To generate an organic audio landscape, you should involve the parameters of your VR world to affect the audio.
Some violinists in the orchestra play the same notes throughout the piece. Why?
In the digital domain, we can and should manipulate every single digit as we want.
And this can influence the visual, interaction, haptic, and vice versa
From “a sound plays” to “a sound exists in a place.”
A flat-screen mix is often organized around a screen.
XR uses the listener’s moving head as the reference frame.
The same sound should change as the user turns, walks, crosses a doorway, or moves behind an obstacle.
Using head-related transfer functions (HRTFs), a 3D audio engine filters mono sources with ear-specific timing, level, and spectral cues updated by head pose to recreate convincing front–back and elevation (up/down) localization over headphones.
"Selecting an HRTF spatializer and enabling head tracking lets listeners reliably tell whether a sound is above or below, in front of or behind them."
How to implement that in XR:
Steam Audio
Steam Audio adds physics-based sound propagation on top of HRTF-based binaural audio, for increased immersion.
Sounds interact with and bounce off of the actual scene geometry, so they feel like they are actually in the scene, and give players more information about the scene they are in.
about: https://store.steampowered.com/news/app/596420/view/1589135317218111209
Steam Audio is the spatial layer that makes a created sound behave as though it occupies the world.
A sound object can have a direction relative to the listener through HRTF spatialization. A wall can occlude or transmit sound.
Room geometry and material can affect reflections and reverb. Sound can propagate through an environment instead of behaving like a flat stereo soundtrack.
Sound Synthesis
Frequency Modulation (FM)
The history of FM dates back to 1936 when Edwin Howard Armstrong described the FM frequency as a method of reducing disturbances in radio transmission in a conference of Radio Engineers New York in November 6, 1936.
Bessel function
Early Nintendo games did not simply play recorded audio.
They generated music and sound effects in real time using a small set of sound generators: pulse waves, a triangle wave, noise, and samples. Composers created rich musical worlds by continuously changing frequency, duty cycle, envelope, and timing under severe hardware constraint
Animal Crossing uses its distinctive “Animalese” speech to give characters a recognizable voice without conventional full voice acting.
The original sound team initially experimented with more realistic voice synthesis, but decided that animal characters did not need to speak like humans. Instead, they developed a stylized computational voice system focused on communicating tone and character rather than realistic speech.
A computational sound system can create a stronger identity, personality, and aesthetic language than realistic recorded audio.
Modern game audio goes far beyond generating individual waveforms. In The Legend of Zelda: Tears of the Kingdom, Nintendo designed sound together with the game’s physics, 3D world, and player interactions. Rather than authoring a unique sound response for every possible event, they built systems in which new combinations of physics and gameplay can produce appropriate sound behavior.
*Connecting the physics-driven world with its evolving sound design
-> Highly relavent to XR
"Composition for Objective Sound" attempts to deviate from traditional one-way sheet music and read objective sound with new rules, such as simultaneously proceeding with multiple sheet music sources to proceed multi-directionally at the same time. While reading new sounds, the individual positions of objective sounds that are randomly generated are fixed in space but fluid in their generation simultaneously. These objective sounds can attempt to be read themselves through moments of time accumulation by the system or wait for them to be read. The audiovisual performance they read twists the time axis over the predictable flow and generally makes the audience sense the sound in a new way, different from the act of reading the sound.
Hands (2004) is an early example of using bodily gesture as a multidimensional controller for electronic sound synthesis. Sensors attached to the performer’s hands measured different aspects of movement, such as hand position, distance, orientation, and finger actions, and mapped them continuously to synthesizer parameters.
The important idea is that one gesture does not have to control only one parameter. A single movement can simultaneously affect pitch, timbre, amplitude, modulation, or spatial behavior.
Gesture can act as a high-dimensional control signal rather than a simple trigger.
gesture → multiple continuous parameters → evolving synthesized sound
This is directly relevant to modern hand tracking, where joint positions, velocity, openness, pinch strength, and orientation can all become expressive controls.
Reactable is a tangible interface for real-time electronic music in which physical objects placed on a tabletop represent different parts of a modular synthesizer, such as oscillators, filters, effects, and sequencers. Moving, rotating, adding, or removing objects changes how these components are connected and how sound is generated.
The key interaction is therefore not only changing parameter values. The user can change the structure of the sound-generating system itself.
Reactable (2003) is a tangible interface for real-time electronic music in which physical objects placed on a tabletop represent different parts of a modular synthesizer, such as oscillators, filters, effects, and sequencers. Moving, rotating, adding, or removing objects changes how these components are connected and how sound is generated.
The key interaction is therefore not only changing parameter values. The user can change the structure of the sound-generating system itself.
physical interaction → synthesis structure → sound behavior
AVES (2000) is a collection of interactive audiovisual systems in which synthetic sound and abstract graphics are generated together in real time. The user manipulates dynamic visual forms, but the sound is not simply added afterward as an effect. Both image and sound emerge from the same underlying computational behavior.
This creates a stronger relationship between modalities. Instead of mapping one input independently to sound and another to graphics, the system defines a shared state that drives both.
Sound and image can be two expressions of the same computational process.
interaction → system state → graphics + synthesized sound
This approach often creates a more coherent audiovisual experience because the media are structurally related rather than merely synchronized.
Coexistence with the SARS-CoV-2 virus | Myungin Lee (Ars Electronica 2022)
Parasitic Signals: Coexistence with the SARS-CoV-2 virus (2022) is an interactive audiovisual simulation in which scientific data and virus behavior drive real-time sound synthesis and visualization. Different parts of the system use different synthesis methods:
virus movement controls FM synthesis, mucus behavior controls subtractive synthesis, and virus–receptor interactions drive granular synthesis.
Audience input changes the simulation itself, including virus movement and biological defense behavior, and the resulting changes are then reflected in both sound and visuals.
interaction → simulation behavior → audiovisual behavior
This creates a coherent system in which sound and image emerge from the same computational process.
PatchWorld (2026) turns modular synthesis and visual programming into a three-dimensional XR environment. Users connect blocks for synthesizers, effects, logic, physics, visuals, and interaction directly in space, allowing them to build instruments, reactive environments, and audiovisual performances from inside the world itself.
XR can transform sound synthesis from a flat interface into a spatial system that users can enter, connect, rearrange, and perform.
spatial interaction → synthesis topology → sound + visual behavior
This is essentially a contemporary XR extension of the idea behind Reactable.
PatchWorld in Research | PatchXR — PatchWorld: Music Creation in VR and Desktop
Tónandi is a mixed-reality music experience in which the music of Sigur Rós appears as spatial audiovisual “sound spirits” distributed throughout the physical environment. Users interact with these entities using simple finger and hand movements, revealing and changing their musical patterns and behaviors. The experience is different depending on the physical space and on how the user explores it.
It can be designed as a spatial creature or material that the user discovers and manipulates with the body.
gesture + physical space → musical behavior → spatial audiovisual world
Virtuoso is a VR music system built around instruments designed specifically for three-dimensional interaction rather than reproducing conventional instruments. Synthesizer patches and effects can respond to the height, depth, and tilt of the controllers, while looping tools allow musical structures to accumulate over time. Virtuoso
XR gives an instrument new control dimensions that do not exist on a physical keyboard or knob.
3D movement → continuous synthesis parameters → musical performance
The key question becomes: What kind of instrument becomes possible only when the interface exists in 3D space?
How computers make sound?
“Buffer should never wait for you”
* Trade off between
the buffer size and the performance (latency)
More Audio Funtions
practice) How to make wave-like sound?
OnAudioFilterRead()
System.Random rand = new System.Random();
void OnAudioFilterRead(float[] data, int channels)
{
data[i] = (float)rand.NextDouble();
}
Practice) Interactive Audio parameter control using the script
Audio Effects
Filter (Low pass, high pass, band pass) - https://docs.unity3d.com/Manual/class-AudioEffect.html
Low Pass Filter (https://docs.unity3d.com/ScriptReference/AudioLowPassFilter.html)
Reverberation - https://docs.unity3d.com/ScriptReference/AudioReverbFilter.html
Pitch shift - https://docs.unity3d.com/Manual/class-AudioPitchShifterEffect.html
Sonification Case Study
Space
Nasa releases audio of what a black hole 'sounds' like - YouTube (2022)
Solar Spectral Sonification (Audio/Visual Preview) (youtube.com)
5,000 Exoplanets: Listen to the Sounds of Discovery (NASA Data Sonification) (youtube.com)
All Planet Sounds From Space (In our Solar System) (youtube.com)
Data Sonification: Westerlund 2 (Multiwavelength) (youtube.com)
Data Sonification: M51 (Whirlpool Galaxy) Multiwavelength (youtube.com)
Data Sonification: M51 (Whirlpool Galaxy) Sequence (youtube.com)
HCI
Science