TexaData
Solutions
Speech and voice teams

ASR and speech for languages the benchmarks skip

Word error rates in this corridor are not a modelling problem. They are a data problem, and the data was never collected.

A barber cutting hair at a stand on the roadside

The problem

A speech model trained on the open corpora available for West and Central Africa learns a register that does not exist outside a recording prompt. The published material is overwhelmingly read speech: one speaker, one microphone, one sentence at a time, elicited from a script. It is clean, it is single-channel, and it is nothing like a market negotiation, a call-in radio segment or two mechanics talking over an engine.

So models ship with respectable benchmark numbers and fail in deployment on the three things that actually characterise speech here — overlap, interruption, and code-switching mid-clause between Pidgin, English and a local language. The benchmark did not contain them because the collection method could not produce them.

What we put against it

Two sources, and a buyer can take either or both. Cleared archive from radio talk and call-in programming, which is existing material and moves at the speed of a rights audit. And purpose-built capture on lavalier, which moves at the speed of a brief.

Both arrive channel-separated where the source allows, timestamped, speaker-attributed, with diarisation and overlap ratio in the metadata and code-switch events marked rather than normalised away. The transcript preserves the switch because the switch is the hard part.

A barber cutting hair at a stand on the roadside
Trading in the open beside a road
A customer in the chair reflected in a barbershop mirror
A market seen from above, dense with trading umbrellas
A barber at work in the open air, customer under a cape

What this has to satisfy

RequirementHow it is met
Overlapping multi-speaker audio with real interruptionCall-in radio and market capture, channel-separated, overlap ratio reported per asset
Code-switching preserved rather than cleanedSwitch events marked inline in the transcript with the language on either side
Acoustic conditions that match deploymentOpen-air, roadside and workshop capture with the noise floor documented, not removed
Rights that survive a legal reviewSpeaker consent naming commercial AI training and generative models, sublicensable onward

What you can check first

Sample before contract
Published clips with transcript and annotation, downloadable without an NDA
Manifest
38-column CSV per collection, exportable before you commit
Consent chain
Every asset resolves to a consent record and a paid contributor
Languages
Nigerian Pidgin, Nigerian English, Yoruba, Hausa, Igbo, Lingala, French, Wolof, Twi

Supply lines behind this

A solution is marked live only where a line behind it holds cleared material. Nothing sourced to mandate is presented as inventory.

Start a pilot

Fifty delivered hours at production standard, two weeks from spec sign-off.