Abstract
Recent advances in speech technology predominantly favor high-resource languages, leaving speakers of Sub-Saharan African languages underserved. To bridge this divide, we introduce WAXAL, an open, large-scale speech dataset spanning 24 African languages representing over 100 million speakers. WAXAL comprises over 14,100 hours of speech collected in partnership with African academic and community organizations to build local research capacity: an Automatic Speech Recognition (ASR) corpus with 2,340 hours of transcribed, image-prompted natural speech across 19 languages (plus ~11,760 untranscribed hours) and a studio-quality Text-to-Speech (TTS) corpus exceeding 235 hours. We empirically benchmark three architecturally distinct ASR models---Gemma 3n 2B, Whisper-Large-v3, and MMS-1B-all---across the 19 ASR languages on strictly speaker-disjoint splits. Fine-tuning on WAXAL yields substantial gains, reducing macro-average WER by up to 58% for Whisper-Large-v3, 56% for Gemma 3n 2B, and 17% for MMS-1B-all (and Character Error Rate by up to 70%). WAXAL-fine-tuned models generalize zero-shot to three out-of-distribution benchmarks---FLEURS, Mozilla Common Voice, and Sunbird SALT---and outperform FLEURS-trained baselines in bidirectional cross-dataset transfer. Finally, an LLM-as-judge semantic evaluation confirms that fine-tuning preserves speaker intent beyond surface-level transcription metrics. WAXAL is publicly released under a CC-BY-4.0 license at https://huggingface.co/datasets/google/WaxalNLP