VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Hannah Raphaelle Muckenhirn; Ignacio Lopez Moreno; John Hershey; Kevin Wilson; Prashant Sridhar; Quan Wang; Rif A. Saurous; Ron Weiss; Ye Jia; Zelin Wu

VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Hannah Raphaelle Muckenhirn

Ignacio Lopez Moreno

John Hershey

Kevin Wilson

Prashant Sridhar

Quan Wang

Rif A. Saurous

Ron Weiss

Ye Jia

Zelin Wu

ICASSP 2019 (2018)

Download Google Scholar

Abstract

In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings; (2) A spectrogram masking network that takes both noisy spectrogram and speaker embedding as input, and produces a mask. Our system significantly reduces the speech recognition WER on multi-speaker signals, with minimal WER degradation on single-speaker signals.

Research Areas

Machine intelligence

Explore our many areas of focus

Building a collaborative ecosystem

Shaping the future together

Translating discovery into real-world impact

VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Abstract

Research Areas

Meet the teams driving innovation

Google AI

Google Cloud

Google DeepMind

Google Labs