MD3: The Multi-Dialect Dataset of Dialogues

Jacob Eisenstein; Vinodkumar Prabhakaran; Clara Rivera; Dora Demszky; Devyani Sharma

MD3: The Multi-Dialect Dataset of Dialogues

Jacob Eisenstein

Vinodkumar Prabhakaran

Clara Rivera

Dora Demszky

Devyani Sharma

InterSpeech (2023) (to appear)

Download Google Scholar

Abstract

We introduce a new dataset of conversational speech representing English from India, Nigeria, and the United States. Unlike prior datasets, the Multi-Dialect Dataset of Dialogues (MD3) strikes a balance between open-ended conversational speech and task-oriented dialogue by prompting participants to perform a series of short information-sharing tasks.
This facilitates quantitative cross-dialectal comparison, while avoiding the imposition of a restrictive task structure that might inhibit the expression of dialect features.
Preliminary analysis of the dataset reveals significant differences in syntax and in the use of discourse markers. The dataset includes more than 20 hours of audio and more than 200,000 orthographically-transcribed tokens, and is made publicly available at \url{https://www.kaggle.com/datasets/jacobeis99/md3en}.

Research Areas

Natural language processing

Explore our many areas of focus

Building a collaborative ecosystem

Shaping the future together

Translating discovery into real-world impact

MD3: The Multi-Dialect Dataset of Dialogues

Abstract

Research Areas

Meet the teams driving innovation

Google AI

Google Cloud

Google DeepMind

Google Labs