Arthur Goldstuck | CEO | World Wide Worx | Editor-in-Chief | Gadget.co.za | mail me |
Artificial Intelligence (AI) is often described as a triumph of machines over human limitations. However, one of the most revealing AI stories from Africa tells a different story. It highlights a victory achieved by people who refused to let their voices disappear from the digital future.
More than 7,000 individuals across several African countries volunteered to record their speech. They contributed to a new open dataset released by Google and a consortium of African research institutions. The dataset, called the WAXAL dataset, contains more than 1,250 hours of transcribed speech across 21 languages. These numbers illustrate a powerful reality: AI needs humans to exist and evolve.
A human story behind the data
Most large AI systems learn from data gathered incidentally. Developers scrape text from the web and pull images from public repositories. However, the WAXAL initiative followed a different path. It shows clearly that AI needs humans when data does not already exist in digital form.
African universities and community organisations worked directly with volunteers. Participants contributed their voices so speech technologies could function in languages that digital systems have long ignored.
The project is described as a collective decision about belonging.
Over 7,000 volunteers joined us because they wanted their voices and languages to belong in the digital future. Today, that collective effort has sparked an ecosystem of innovation in fields like health, education and agriculture. This proves that when the data exists, THE possibility expands everywhere.
– Isaac Wiafe from the University of Ghana
Why language data matters
Speech technology now forms one of the main ways people interact with digital services. Voice assistants, transcription tools, automated call centres and educational platforms increasingly rely on spoken language interfaces. Consequently, AI needs humans to provide authentic linguistic input that machines cannot generate independently.
When systems fail to recognise a language or accent, exclusion becomes inevitable. Yet Africa hosts more than 2,000 languages. Despite this linguistic richness, mainstream speech technologies support only a small fraction of them.
Observers often blame the gap on a shortage of data. That explanation contains some truth. However, it overlooks an important human factor. Data exists only when people feel willing and able to create it under conditions they trust.
African institutions lead the process
The WAXAL dataset was developed over three years with funding from Google. However, African institutions led the actual process of collecting speech data. Key contributors included Makerere University in Uganda, the University of Ghana and community organisations such as Digital Umuganda in Rwanda.
The dataset is linked directly to local expertise. For AI to have a real impact in Africa, it must speak our languages and understand our contexts.
– Joyce Nakatumba-Nabende, Senior Lecturer at Makerere University
Building speech datasets within the communities that speak those languages also improves the quality of future technologies. At the same time, this approach shifts control over the outcomes.
Ownership and control of data
In a notable departure from earlier multinational practices, African partner institutions retain ownership of the dataset. Google provided funding and technical guidance. However, the company relinquished direct control over the data.
This structure demonstrates that African languages can integrate into modern AI systems through processes that combine technical rigour with social participation. Once such systems exist, the remaining languages no longer face barriers to feasibility. Importantly, the volunteers remain central to the story. Their contributions show once again that AI needs humans to create foundational datasets.
Each recorded voice transforms language support from a distant aspiration into a practical pathway. Every contribution becomes part of a foundation that others can build upon. In this sense, the dataset functions like infrastructure. It becomes invisible once established. Yet it decisively shapes who can participate in digital systems.
Empowering African innovation
Progress depended heavily on trust and willingness. Participants had to believe that contributing their voices would lead to meaningful outcomes. They accepted uncertainty because they believed in the broader benefits.
Empowerment is the ultimate impact. This dataset provides the critical foundation for students, researchers, and entrepreneurs to build technology on their own terms, in their own languages, finally reaching over 100 million people. We look forward to seeing African innovators use this data to create everything from new educational tools to voice-enabled services that create tangible economic opportunities across the continent.
– Aisha Walcott-Bryant, Head of Google Research Africa
These possibilities include voice-based education platforms, automated health triage services and agricultural advice delivered through speech technologies.
Human participation at the centre of AI
The WAXAL dataset does not guarantee inclusion. It also cannot solve the full complexity of Africa’s linguistic landscape. Nevertheless, the project demonstrates a crucial lesson. AI needs humans not only as users, but also as contributors at the beginning of the technological pipeline.
AI built with active human participation produces different outcomes from systems built on leftover digital traces. By restoring people to the start of the process, initiatives like WAXAL reshape how technology evolves. Instead of leaving communities at the receiving end, they place them at the centre of innovation.




























