Enterprise & operational documents
Company archives, contracts, correspondence, forms, technical manuals and printed matter — the document layer that extraction and document-understanding models are starved of outside Europe and North America.

Document models trained on Western paperwork fail on African paperwork for ordinary, unglamorous reasons: different form layouts, different letterhead conventions, different table structures, bilingual English–French administrative text, and the specific degradations of paper that has been photocopied, stamped, hole-punched and stored through thirty Lagos rainy seasons. That degradation is not noise to be cleaned out before training. For a document model it is the training signal, and it is exactly what a clean Western corpus cannot teach.
We source from companies, professional firms, publishers, printers and institutional archives, and we digitise at a fixed standard rather than at whatever the holder's own scanner does. Commercially sensitive content is redacted at source, under the holder's direction, before a page is scanned — not blurred afterwards by us, which is a promise a buyer cannot audit.
This is the easiest line in the catalogue to sample honestly, and it is where you will see real material on this site first: a holder's own documents, cleared, digitised to spec and annotated, published with their written agreement rather than their tolerance.




Specification
- Sources
- Company archives, professional firms, publishers and printers, institutional and public-record archives
- Capture
- 400 dpi colour, deskewed, one page per image, TIFF master with a PDF/A compile per document
- Structure
- Page regions, reading order, table structure, stamps and handwriting flagged rather than inferred
- Languages
- English, French, and the bilingual administrative register of the corridor
- Redaction
- Performed at source under the holder's direction, before scanning
- Status
- Sourced to mandate. First cleared tranches sampled publicly on this site.
Other supply lines

Cleared African archive — radio, broadcast, film
Decades of in-language radio, broadcast and film sitting in archives that were never contracted for machine learning, in a market with no aggregator working it.

In-language conversational audio
Unscripted, multi-speaker, overlapping speech in languages with effectively no licensed commercial corpus — from cleared archive and from purpose-built collection.

Egocentric & real-world capture
Purpose-collected, head- and chest-mounted first-person video from working environments that cannot be collected anywhere else.

Instrumented & sensor-layered capture
The premium tier of embodied data — multi-rig, force-instrumented capture of manual work, stood up against a funded brief.

Property & transaction records
Listings, valuations, title histories and completed-deal records from agencies, developers and estate surveyors — the structured layer under African urban property, which no dataset currently describes.

De-identified institutional & financial records
Transaction, statement and ledger records supplied by the institution that holds them, de-identified before they leave the building, and licensed for model training under a data-sharing agreement.