TexaData
Solutions
Frontier and post-training teams

Post-training and evaluation outside the crawl

The web is not a sample of the world. For this corridor it is a sample of what got typed in English by people with a connection — and a model trained on it inherits that gap as confidence.

A street food seller at her stand, black and white

The problem

A frontier model's coverage of Nigerian Pidgin, Lingala or Twi is downstream of how much of each language happened to be crawled, which is downstream of who was online and writing. The result is a model that produces fluent-sounding text in a language it has mostly seen translated, and evaluates well on benchmarks translated from English into that language — which measures translation, not competence.

Post-training data that would fix it barely exists in licensed form. What does exist is either scraped, and therefore unusable by a laboratory that has to warrant provenance, or it is prompted single-turn material that does not carry the reasoning, the disagreement or the domain knowledge a post-training set needs.

What we put against it

Spontaneous, multi-party, in-language material with the argument intact: call-in radio where a caller and a host disagree, market negotiation where a price is contested and settled, workshop instruction where a skill is transmitted step by step. These are natural reasoning traces in languages that have almost none on record.

Delivered with speaker attribution, timestamps and a written context layer, and — because a laboratory's counsel will ask — with chain of title reconstructed per asset and AI training named in the grant rather than inferred from a general licence.

A street food seller at her stand, black and white
A small barbershop interior lit by a single bulb
A dense city quarter from the air, high-rises beyond
Storm cloud building over a city at dusk
Traders seated behind baskets of produce and dried fish

What this has to satisfy

RequirementHow it is met
In-language material that was not translated from EnglishOriginated and archive audio in the language of use, transcribed in that language
Evaluation sets that measure competence, not translationNative-register material a benchmark can be cut from, with the register documented
Provenance a counsel will acceptPer-asset chain of title, warranted; AI training named explicitly in the grant
Onward sublicensing through tiersGrants written to sublicense, for aggregators reselling to laboratories

What you can check first

Named in the grant
"AI training" appears in the instrument, not inferred from a general licence
Tranche-gated
Nothing is accepted into a tranche until title is reconstructable and warrantable
Open-corpus comparison
The corpora that already exist are named on the open-data page, with the gap stated
Excluded
Third-party music, licensed insert footage, unreleased performers

Supply lines behind this

A solution is marked live only where a line behind it holds cleared material. Nothing sourced to mandate is presented as inventory.

Start a pilot

Fifty delivered hours at production standard, two weeks from spec sign-off.