Researchers Release Largest-Ever African-Language AI Dataset, Earn ACL 2026 Best Paper Nomination
A multi-institutional research initiative called the African Languages Lab has published the largest validated speech-and-text dataset for African languages to date, spanning 40 languages and 19 billion tokens, in a paper accepted as a long oral presentation with a Best Paper Nomination at ACL 2026.

A multi-institutional research initiative called the African Languages Lab has published the largest validated speech-and-text dataset for African languages to date, spanning 40 languages and 19 billion tokens, in a paper accepted as a long oral presentation with a Best Paper Nomination at ACL 2026.
Scale of the Dataset
The dataset covers 40 African languages with 19 billion tokens of monolingual text and 12,628 hours of aligned speech data, built through a purpose-designed, quality-controlled collection pipeline that includes a new mobile-first crowdsourced platform called All Voices. Fine-tuning models on the dataset produced average gains of +23.69 ChrF++, +0.33 COMET, and +15.34 BLEU over baselines across the 31 languages evaluated.
Who Built It
The project was led by Sheriff Issaka of UCLA with senior co-author Saadia Gabriel, working alongside roughly 20 co-authors across about 10 institutions on multiple continents, including Evans Kofi Agyei of the University of Cape Coast in Ghana and Sadick Abdul Mumin of Northwestern University's Qatar campus. The initiative also mentored 15 early-career researchers, many from African institutions. A related group, the AI4D African Languages Lab based at the Data Science for Social Impact lab at South Africa's University of Pretoria and led by researchers including Abiodun Modupe and Idris Abdulmumin, presented adjacent African-language NLP papers at the same ACL 2026 meeting in San Diego.
Why it matters
Africa is home to roughly a third of the world's languages, yet the project's own framing notes that 88% of African languages are severely underrepresented or entirely ignored in computational linguistics research, leaving voice assistants, translation tools, and chatbots largely unusable for hundreds of millions of speakers. This dataset gives researchers and developers a foundation to build genuinely functional AI language tools for previously near-zero-resource African languages.