TexaData
Supply lines
Primary line

In-language conversational audio

Unscripted, multi-speaker, overlapping speech in languages with effectively no licensed commercial corpus — from cleared archive and from purpose-built collection.

A barber cutting hair at a stand on the roadside

Two sources feed this line, and a buyer can take either. Cleared archive from radio talk and call-in programming across anglophone and francophone stations, which is existing material and moves at the speed of a rights audit. And purpose-built capture on lavalier — market negotiation, workshop instruction, household conversation — which moves at the speed of a brief.

Purpose-built lots are delivered channel-separated, transcribed, and carrying full speaker metadata, in Nigerian Pidgin and Nigerian English with the code-switching intact rather than cleaned out of the transcript. Consent covers commercial AI training explicitly, including generative models, and is sublicensable onward.

Natural multi-speaker conversation with genuine overlap and interruption is the hardest single requirement in speech data, and the one scripted collection cannot reproduce. Open corpora in this region are overwhelmingly prompted, single-speaker material, because that is what a prompt elicits. It is also, in this region, what an entire broadcast tradition already consists of — which is why the archive line and this one are the same argument approached from two ends.

Multi-speakerCode-switchingChannel-separatedDiarisedCall-in radioField capture
A barber cutting hair at a stand on the roadside
A market seen from above, dense with trading umbrellas
Trading in the open beside a road
A customer in the chair reflected in a barbershop mirror

Specification

Format
WAV, 16-bit PCM, 48 kHz
Channels
Channel-separated on purpose-built capture; separated where the archive source allows, mix documented where not
Metadata
Speaker labels, diarisation, overlap ratio, code-switch events, accent and acoustic conditions
Transcripts
JSON, timestamped, speaker-attributed, code-switching preserved
Languages
Nigerian Pidgin, Nigerian English, Yoruba, Hausa, Igbo, Lingala, French, Wolof, Twi
Consent
Speaker consent naming commercial AI training and generative models; voice likeness addressed separately