In-language conversational audio
Unscripted, multi-speaker, overlapping speech in languages with effectively no licensed commercial corpus — from cleared archive and from purpose-built collection.

Two sources feed this line, and a buyer can take either. Cleared archive from radio talk and call-in programming across anglophone and francophone stations, which is existing material and moves at the speed of a rights audit. And purpose-built capture on lavalier — market negotiation, workshop instruction, household conversation — which moves at the speed of a brief.
Purpose-built lots are delivered channel-separated, transcribed, and carrying full speaker metadata, in Nigerian Pidgin and Nigerian English with the code-switching intact rather than cleaned out of the transcript. Consent covers commercial AI training explicitly, including generative models, and is sublicensable onward.
Natural multi-speaker conversation with genuine overlap and interruption is the hardest single requirement in speech data, and the one scripted collection cannot reproduce. Open corpora in this region are overwhelmingly prompted, single-speaker material, because that is what a prompt elicits. It is also, in this region, what an entire broadcast tradition already consists of — which is why the archive line and this one are the same argument approached from two ends.




Specification
- Format
- WAV, 16-bit PCM, 48 kHz
- Channels
- Channel-separated on purpose-built capture; separated where the archive source allows, mix documented where not
- Metadata
- Speaker labels, diarisation, overlap ratio, code-switch events, accent and acoustic conditions
- Transcripts
- JSON, timestamped, speaker-attributed, code-switching preserved
- Languages
- Nigerian Pidgin, Nigerian English, Yoruba, Hausa, Igbo, Lingala, French, Wolof, Twi
- Consent
- Speaker consent naming commercial AI training and generative models; voice likeness addressed separately
Other supply lines

Cleared African archive — radio, broadcast, film
Decades of in-language radio, broadcast and film sitting in archives that were never contracted for machine learning, in a market with no aggregator working it.

Egocentric & real-world capture
Purpose-collected, head- and chest-mounted first-person video from working environments that cannot be collected anywhere else.

Instrumented & sensor-layered capture
The premium tier of embodied data — multi-rig, force-instrumented capture of manual work, stood up against a funded brief.

Enterprise & operational documents
Company archives, contracts, correspondence, forms, technical manuals and printed matter — the document layer that extraction and document-understanding models are starved of outside Europe and North America.

Property & transaction records
Listings, valuations, title histories and completed-deal records from agencies, developers and estate surveyors — the structured layer under African urban property, which no dataset currently describes.

De-identified institutional & financial records
Transaction, statement and ledger records supplied by the institution that holds them, de-identified before they leave the building, and licensed for model training under a data-sharing agreement.