Lagos-concentrated, so co-located with the buyers. Absent from ElevenLabs TTS, AWS Transcribe and Google Chirp 3; ElevenLabs' own published STT band for it is 25–50% WER.
- yor
- Tonal
- 7 dialects
49 AI readiness
46m
Low-resource
15m
speakers, medium confidence estimate
35
AI readiness, weighted toward speech
4
dialects tracked as separate collection targets
Sits in ElevenLabs' own Moderate band at 25–50% word error rate, which is published evidence that its size has not bought it coverage. A large diaspora in South Africa and the UK puts part of the demand outside Zimbabwe.
Capability coverage
Some open data and big-tech ASR, but no production TTS. Demand exists; quality does not.
The facts
Dialects
Dialect is self-declared at contribution and then validated against gold-standard clips. Which dialects people turn up to record in is a better demand signal than any desk ranking.
Open corpora that cover it
Voices
No voice yet. Studio recording is the scarcest asset in this whole field — the largest open African corpus holds 11,000 hours and about twenty of them are studio-grade — so voices follow demand rather than leading it.
Same family
Related
A Yoruba model helps Igbo less than the shared border suggests, and a Hausa model transfers usefully across Chadic languages in three other countries. These share Niger–Congo with Shona.
Lagos-concentrated, so co-located with the buyers. Absent from ElevenLabs TTS, AWS Transcribe and Google Chirp 3; ElevenLabs' own published STT band for it is 25–50% WER.
49 AI readiness
46m
Commercially dense in trade and diaspora remittances. Absent from AWS Transcribe entirely and capped at Chirp 2 on Google.
49 AI readiness
33m
The highest-value regional asset after Hausa: cross-border into Niger, Cameroon and Chad, with strong humanitarian and agricultural demand.
18 AI readiness
20m±
Take an API key and make the call from your own code, or start in the playground and see what the model actually does with your text.