LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Heiga Zen; Rob Clark; Ron J. Weiss; Viet Dang; Ye Jia; Yonghui Wu; Yu Zhang; Zhifeng Chen

LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Heiga Zen

Rob Clark

Ron J. Weiss

Viet Dang

Ye Jia

Yonghui Wu

Yu Zhang

Zhifeng Chen

Interspeech (2019)

Download Google Scholar

Abstract

This paper introduces a new speech corpus called ``LibriTTS'' designed for text-to-speech use. It is derived from the original audio and text materials of the LibriSpeech corpus, which has been used for training and evaluating automatic speech recognition systems. The new corpus inherits desired properties of the LibriSpeech corpus while addressing a number of issues which make LibriSpeech less than ideal for text-to-speech work. The released corpus consists of 585 hours of speech data at 24kHz sampling rate from 2,456 speakers and the corresponding texts. Experimental results show that neural end-to-end TTS models trained from the LibriTTS corpus achieved above 4.0 in mean opinion scores in naturalness in five out of six evaluation speakers. The corpus is freely available for download from http://www.openslr.org/60/.

Explore our many areas of focus

Building a collaborative ecosystem

Shaping the future together

Translating discovery into real-world impact

LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Abstract

Meet the teams driving innovation

Google AI

Google Cloud

Google DeepMind

Google Labs