Data platform
60%
of collection spend on in-domain telephony audio
25%
on studio recording — the scarcest free asset
0%
on replicating corpora already given away free
Most of what a first model needs already exists under permissive licences, and competing on raw hours means competing with two philanthropic programmes for something they are giving away. So the money goes where nobody is giving anything away: noisy 8 kHz telephony audio of real conversations, studio-grade recording, and the consent paperwork that makes either one usable.
Platform targets, not results
Every figure below is a stated goal, not a milestone reached. We would rather publish a plan you can hold us to than a number you cannot check — and the gap between the two is the whole reason this section is labelled.
Languages with a production model
5
Dialects mapped
40
Telephony-channel hours
1,000hrs
Verified speakers
3,500
Studio-grade TTS hours
400hrs
Annotated records
250,000
Consent grants on file
100%
Open cultural records
10,000
What each target means
Hoarding hours is the flat part of a logarithmic curve. One published result puts Kinyarwanda speech recognition at 9.82% WER on 200 hours and 7.14% on 1,400 — a sevenfold increase in data for 2.7 points. Past a few hundred hours per language the binding constraint stops being volume and starts being label quality.
Provenance
Licensing
A licence tag on a model hub is not diligence. Licence metadata is frequently wrong, ShareAlike contaminates anything it is pooled with, and chain of title is a separate question from the licence — one widely used corpus carries a permissive headline and still prohibits text-to-speech outright.
Usable commercially
Google WAXAL
11,000+ hours across 27 African languages
Commercial use, partners retain ownership
African Next Voices
9,000–18,000 hours
CC BY 4.0 — but the South African subset prohibits TTS and voice cloning
Mozilla Common Voice
33,150 hours, all languages
CC0
FLEURS
Speech benchmark, 102 languages
CC BY 4.0
WURA
49.1 GB of text across 20 African languages
Apache 2.0
Sunbird SALT
51-language ASR fine-tune
Apache 2.0
Deliberately not ingested
NaijaVoices
CC BY-NC-SA. Commercial waiver exists but the ShareAlike term contaminates anything it is pooled with.
AfriSpeech-200
Non-commercial. Excellent benchmark, unusable as training data for a product.
Meta MMS
The model itself is CC BY-NC 4.0 — the weights, not just the data.
MENYO-20k
Non-commercial because its sources were never cleared. Chain of title, not licence.
Scraped podcasts and video
Speakers never consented. Not worth the exposure at any volume.
A dataset version’s licence is computed from the most restrictive licence among its sources rather than asserted. That is what mechanically prevents publishing a ShareAlike and non-commercial mixture, which is unpublishable under either.
Knowledge graph
Not Wikidata for Africa. Five entity types — names, places, organisations, languages and dialects — because those are the ones that fix the 40–70% named-entity error rate. It grows as a by-product of the pronunciation layer rather than as a project of its own, and tenant data never enters the shared graph without contractual permission.
Click a node to inspect it, double-click to re-centre. Drag to pan, scroll to zoom. Keyboard: Tab to a node, Enter to inspect, Space to re-centre.
Yoruba
LanguageTonal, Latin script with diacritics. ~45.7m speakers.
PronunciationYOH-ru-bah
Relationships
Contribute
Contributors
Paid per validated unit rather than per submitted one, at rates published on this page. Rate opacity is the incumbent’s choice and it is the cheapest thing a challenger can differentiate on — and a company asking people to record their own voices should be willing to say what it pays.
Record prompted or conversational speech in your own language and dialect. Paid per validated clip, not per submitted one.
Rate
$0.10–0.15 per validated clip · about $3.00–4.50 an hour at 30 usable clips
Select language
And the dialect you actually speak, not the standard one.
Verify contributor
Phone number first. Email-first onboarding excludes the contributors most needed.
Record or upload
Works on a feature phone over a call, or in the browser.
Quality check
SNR, clipping, duration, silence ratio and duplicate hashing, at upload.
Human validation
Two independent passes, a third on disagreement above threshold.
Payment
Mobile money or bank transfer — M-Pesa, MoMo, Moniepoint, Wave — weekly, $5 threshold.
Universities, language associations, state cultural agencies and NGOs. Community corpora are released openly under a named benefit-sharing arrangement rather than licensed.