The best-served language on the continent, and the control case: where a language is genuinely covered, the only place left to compete is price and domain.
- afr
- Non-tonal
- 2 dialects
100 AI readiness
17m
Catalogue
29
languages with data good enough to publish
12
countries where one of them is the language of record
30
countries reached once cross- border speakers are counted
Africa has somewhere between 1,250 and 2,100 natively spoken languages depending on where the line between language and dialect is drawn. This catalogue holds the ones where public data is good enough to say something true — West, East, Southern, Central and the Horn on the same terms — and stops there. Building the other two thousand rows from guesses would make the whole thing worse than useless.
Showing 29 of 29 languages
No filters applied
The best-served language on the continent, and the control case: where a language is genuinely covered, the only place left to compete is price and domain.
100 AI readiness
17m
The least attractive large African language commercially, precisely because it is the one the majors already serve tolerably — ElevenLabs places it in its 5–10% WER band, the only African language above Moderate.
91 AI readiness
80m
Where the enterprise volume actually sits. Frontier ASR runs 24.7–35% WER on the accent generally — and 40–70% on African named entities, which is the number that decides whether a bank can deploy.
76 AI readiness
110m
The best cross-border expansion asset in West Africa — one language investment reaches Niger, Chad, Ghana and northern Cameroon. Also the one West African language ElevenLabs does cover, so the deficit here is smaller than it looks.
65 AI readiness
88m
The clearest proof that community collection outperforms top-down prioritisation: 1,211 hours onto Common Voice from 420 contributors, which no desk ranking would have predicted.
65 AI readiness
15m
A large diaspora is the unusual feature: remittance and telehealth demand sits partly outside the continent.
59 AI readiness
22m
Lagos-concentrated, so co-located with the buyers. Absent from ElevenLabs TTS, AWS Transcribe and Google Chirp 3; ElevenLabs' own published STT band for it is 25–50% WER.
49 AI readiness
46m
Commercially dense in trade and diaspora remittances. Absent from AWS Transcribe entirely and capped at Chirp 2 on Google.
49 AI readiness
33m
A non-Latin script, so the text pipeline is a different problem from every other language on this list. Big speaker base, difficult FX and regulatory environment.
49 AI readiness
57m
The largest free corpus covering it carries a CC BY 4.0 headline and still prohibits text-to-speech and voice cloning — which supports an ASR entry and not a TTS one.
49 AI readiness
28m
Unusually well served by open research relative to its size, largely through one sustained local effort.
49 AI readiness
11m
Click consonants make it acoustically distinctive, which is a modelling cost and a differentiator at once.
43 AI readiness
19m
The lingua franca of urban informal commerce from Lagos to Accra to Douala, supported by no major vendor at all, with no tonal orthography problem. The single best first language for a West African voice product.
35 AI readiness
80m±
Ghana has the strongest existing developer community for its own languages on the continent, which makes it the cheapest second market to enter.
35 AI readiness
17m
The largest African language with effectively no commercial speech support anywhere. Its Latin orthography was adopted in 1991, so the text pipeline is cheap — which makes the absence of coverage harder to explain than for Amharic next door.
35 AI readiness
37m
Sits in ElevenLabs' own Moderate band at 25–50% word error rate, which is published evidence that its size has not bought it coverage. A large diaspora in South Africa and the UK puts part of the demand outside Zimbabwe.
35 AI readiness
15m
An official language of South Africa, which means a statutory obligation to serve it and a public procurement route that does not wait on commercial demand. Covered by the NCHLT corpora and by nothing any vendor ships.
35 AI readiness
4.7m
Francophone West Africa's entry point, and the place where French–Wolof code-switching is the norm rather than the exception.
31 AI readiness
12m
Kinshasa is one of the largest cities on the continent and francophone Central Africa is served by essentially nobody. The pivot language here is French, which is a different pipeline from every English-pivot row above it.
31 AI readiness
40m±
The Sahel's main lingua franca, and the clearest case where humanitarian and agricultural demand arrives long before commercial demand. Mande rather than Bantu or Volta-Niger, so none of the transfer learning elsewhere in this catalogue reaches it.
31 AI readiness
15m±
The highest-value regional asset after Hausa: cross-border into Niger, Cameroon and Chad, with strong humanitarian and agricultural demand.
18 AI readiness
20m±
State-government and healthcare procurement is the realistic route in. No vendor covers it and no open corpus of scale exists.
5 AI readiness
11m±
North-east Nigeria and the Lake Chad basin, where NGO and humanitarian demand is real and durable even where commercial demand is not.
5 AI readiness
8.2m±
Named in Nigeria's national language-model programme, which makes it a candidate for publicly funded collection ahead of commercial demand.
5 AI readiness
4.6m±
Niger Delta. Oil-sector and state-government relevance out of proportion to the speaker count, and a cluster rather than one language.
5 AI readiness
2.4m±
A long orthographic and lexicographic tradition relative to its size, which makes text collection unusually cheap here.
5 AI readiness
2.0m±
Benin City and the Edo heartland. Cultural-heritage funding is a more realistic first buyer than any enterprise.
5 AI readiness
1.5m±
Middle Belt. Collection here needs a named funder — a state contract or a research grant — before it is worth starting.
5 AI readiness
1.2m±
Closely related to Yoruba, which makes it the clearest test of whether transfer learning inside a subgroup earns its keep.
5 AI readiness
1.8m±
How the tiers work
Resource tiers
The split is keyed to the 200-hour threshold at which speech recognition becomes production-usable. Past a few hundred hours per language, label quality rather than volume is the binding constraint — in one published study 38.6% of high-error cases traced to noisy ground truth rather than insufficient data.
The catalogue grows from contributor demand rather than from a desk estimate. A language whose speakers organise to record it is worth more than a bigger number with no community behind it.