Post-training and evaluation outside the crawl
The web is not a sample of the world. For this corridor it is a sample of what got typed in English by people with a connection — and a model trained on it inherits that gap as confidence.

The problem
A frontier model's coverage of Nigerian Pidgin, Lingala or Twi is downstream of how much of each language happened to be crawled, which is downstream of who was online and writing. The result is a model that produces fluent-sounding text in a language it has mostly seen translated, and evaluates well on benchmarks translated from English into that language — which measures translation, not competence.
Post-training data that would fix it barely exists in licensed form. What does exist is either scraped, and therefore unusable by a laboratory that has to warrant provenance, or it is prompted single-turn material that does not carry the reasoning, the disagreement or the domain knowledge a post-training set needs.
What we put against it
Spontaneous, multi-party, in-language material with the argument intact: call-in radio where a caller and a host disagree, market negotiation where a price is contested and settled, workshop instruction where a skill is transmitted step by step. These are natural reasoning traces in languages that have almost none on record.
Delivered with speaker attribution, timestamps and a written context layer, and — because a laboratory's counsel will ask — with chain of title reconstructed per asset and AI training named in the grant rather than inferred from a general licence.





What this has to satisfy
| Requirement | How it is met |
|---|---|
| In-language material that was not translated from English | Originated and archive audio in the language of use, transcribed in that language |
| Evaluation sets that measure competence, not translation | Native-register material a benchmark can be cut from, with the register documented |
| Provenance a counsel will accept | Per-asset chain of title, warranted; AI training named explicitly in the grant |
| Onward sublicensing through tiers | Grants written to sublicense, for aggregators reselling to laboratories |
What you can check first
- Named in the grant
- "AI training" appears in the instrument, not inferred from a general licence
- Tranche-gated
- Nothing is accepted into a tranche until title is reconstructable and warrantable
- Open-corpus comparison
- The corpora that already exist are named on the open-data page, with the gap stated
- Excluded
- Third-party music, licensed insert footage, unreleased performers
Supply lines behind this



A solution is marked live only where a line behind it holds cleared material. Nothing sourced to mandate is presented as inventory.
Fifty delivered hours at production standard, two weeks from spec sign-off.
Other solutions

ASR and speech for languages the benchmarks skip
Word error rates in this corridor are not a modelling problem. They are a data problem, and the data was never collected.

First-person work, including the part where it goes wrong
Manipulation policies generalise from the variety of hands, tools and failures they have seen. The open egocentric corpora are a fixed process in an industrial interior, filmed until it works.

Documents as they are actually filed
An extraction model that has only seen clean templates has not seen this corridor's paperwork — the stamp over the field, the third-generation photocopy, the form filled in two hands.