Improving Speech Recognition Using Consistent Predictions on Synthesized Speech

Gary Wang

Andrew Rosenberg

Zhehuai Chen

Yu Zhang

Bhuvana Ramabhadran

Heiga Zen (Byungha Chun)

Yonghui Wu

Pedro Jose Moreno Mengibar

IEEE ICASSP 2020

Download Google Scholar

Abstract

Speech synthesis has advanced to the point of being close to indistinguishable from human speech. However, efforts to train speech recognition systems on synthesized utterances have not been able to show that synthesized data can be effectively used to augment or replace human speech. In this work, we demonstrate that promoting consistent predictions in response to real and synthesized speech enables significantly improved speech recognition performance. We also find that training on 460 hours of LibriSpeech augmented with 500 hours of transcripts (without audio) performance is within 0.2\% WER of a system trained on 960 hours of transcribed audio. This suggests that with this approach, when there is sufficient text available, reliance on transcribed audio can be cut nearly in half.

Research Areas

Speech Processing

Defining the technology of today and tomorrow.

Philosophy

People

Teams

AI/ML Foundations  & Capabilities

Algorithms & Optimization

Computing Paradigms

Responsible Human-Centric Technology

Science & Societal Impact

Projects

Publications

Resources

Shaping the future, together.

Student programs

Faculty programs

Conferences & events

Improving Speech Recognition Using Consistent Predictions on Synthesized Speech

Abstract

Research Areas

Learn more about how we conduct our research

Defining the technology of today and tomorrow.

Philosophy

People

Teams

AI/ML Foundations & Capabilities

Algorithms & Optimization

Computing Paradigms

Responsible Human-Centric Technology

Science & Societal Impact

Projects

Publications

Resources

Shaping the future, together.

Student programs

Faculty programs

Conferences & events

Improving Speech Recognition Using Consistent Predictions on Synthesized Speech

Abstract

Research Areas

Learn more about how we conduct our research

AI/ML Foundations  & Capabilities