What is already open — and what it does not cover.
Our whole argument is that this material does not exist. That claim is worth nothing unless it survives the corpora that do. Below are the ones a buyer will bring to the meeting, what each is genuinely good for, and the specific thing it does not carry. Where one of them answers your brief, take it — it is free, and we would rather lose the work than sell against a corpus that already solves your problem.






- 01
Egocentric-100K
Build AI · Hugging Face, 2026
- Scale
- 100,405 hours · 2,010,759 clips
- Licence
- Apache 2.0 — commercial training permitted
What it isHead-mounted capture of manual labour on factory floors, at a volume no commissioned programme will match. If a factory floor is what your model needs, take it: it is free and it is enormous.
What it does not carryIndustrial interiors running a fixed process, which is the environment class most heavily covered already. The dataset card carries no consent instrument, no contributor payment record and no per-clip rights document — an Apache licence on a corpus is a statement about the corpus, not a release from the people inside it. There is no language layer.
- 02
EgoStandard
Lightwheel · Hugging Face, 2026
- Scale
- 100,000 hours · 15,000+ tasks and scenes
- Licence
- Open
What it isThe broadest annotated open egocentric corpus, spread deliberately wide across task types.
What it does not carryBreadth of task rather than depth of environment, annotated to Lightwheel's schema rather than yours. The scenes come from the markets Lightwheel already operates in; West and Central African informal work is not among them.
- 03
Ego4D
Meta AI, 2022
- Scale
- 3,670 hours · 923 participants · 74 locations
- Licence
- Restricts commercial use
What it isThe reference benchmark for egocentric video, and the corpus every buyer in this market has already read.
What it does not carryIts licence keeps it out of a commercial training run. It is a yardstick, not supply — and it is four years old.
- 04
WAXAL
Google with Makerere, University of Ghana and Digital Umuganda — February 2026
- Scale
- 11,000+ hours · ~2m recordings · 21 languages
- Licence
- Open
What it isThe largest open African speech collection there has ever been, covering Yoruba, Hausa, Igbo, Lingala and Fulani among others. It is a good piece of work and it changed what the floor looks like in this region.
What it does not carryAbout 1,250 hours are transcribed for ASR and roughly 20 hours are studio recording for TTS. The body of it is prompted, single-speaker material, because that is what a prompt elicits. It holds almost no unscripted multi-speaker exchange with real overlap, interruption and code-switching — the requirement scripted collection cannot reach, and the one our audio line is built for.
Figures are the publishers' own, current at September 2026.

What survives the comparison.
Open egocentric video is abundant and industrial: factory floors under fixed process, annotated to someone else's schema, with no consent instrument behind the people in frame. Open African speech is abundant and prompted: read sentences, one speaker at a time. Neither describes an open-air market at trading hour, a roadside repair that fails twice before the method changes, or four people negotiating over each other in two languages. That is the whole of what we collect, and it is why the catalogue is narrow on purpose.