Speech enhancement is a technology designed to improve the clarity and quality of a target speech signal corrupted by background noise, reverberation, or interfering voices. It focuses on effectively removing unwanted noise using multi-channel microphone arrays or deep learning-based spectrogram filtering techniques. This field serves as a foundational technology for a wide range of real-world audio systems, including hearing aids, voice call quality improvement, and front-end processing for automatic speech recognition.
Universal sound separation is the process of isolating arbitrary sound sources into individual tracks within complex acoustic environments, without being limited to specific sound classes. Beyond human speech, deep learning models learn to recognize and separate diverse and undefined sound events, such as musical instruments, natural sounds, and ambient noise. This technology plays a key role in broad application areas, including AR/VR immersive audio rendering and acoustic event detection for smart home environments.
Target sound extraction selectively isolates a specific sound of interest from a mixed acoustic signal based on a user-provided auxiliary cue. By leveraging reference cues such as a speaker's enrollment voice profile or a sample audio clip, the system pinpoint-extracts only the desired sound from the mixture. This capability is exceptionally useful for personalized audio devices, such as customized hearables that allow users to listen selectively to specific voices, as well as targeted sound event tracking systems.