TexaData

What is already open — and what it does not cover.

Our whole argument is that this material does not exist. That claim is worth nothing unless it survives the corpora that do. Below are the ones a buyer will bring to the meeting, what each is genuinely good for, and the specific thing it does not carry. Where one of them answers your brief, take it — it is free, and we would rather lose the work than sell against a corpus that already solves your problem.

A man working scrap metal at a roadside fabrication yard
Hands working inside an open engine bay
A barber at work in the open air, customer under a cape
Traders seated behind baskets of produce and dried fish
Hands shaping the wall of a large clay pot
A busy street in afternoon haze, buses and stalls
  1. 01

    Egocentric-100K

    Build AI · Hugging Face, 2026

    Scale
    100,405 hours · 2,010,759 clips
    Licence
    Apache 2.0 — commercial training permitted
    What it is

    Head-mounted capture of manual labour on factory floors, at a volume no commissioned programme will match. If a factory floor is what your model needs, take it: it is free and it is enormous.

    What it does not carry

    Industrial interiors running a fixed process, which is the environment class most heavily covered already. The dataset card carries no consent instrument, no contributor payment record and no per-clip rights document — an Apache licence on a corpus is a statement about the corpus, not a release from the people inside it. There is no language layer.

  2. 02

    EgoStandard

    Lightwheel · Hugging Face, 2026

    Scale
    100,000 hours · 15,000+ tasks and scenes
    Licence
    Open
    What it is

    The broadest annotated open egocentric corpus, spread deliberately wide across task types.

    What it does not carry

    Breadth of task rather than depth of environment, annotated to Lightwheel's schema rather than yours. The scenes come from the markets Lightwheel already operates in; West and Central African informal work is not among them.

  3. 03

    Ego4D

    Meta AI, 2022

    Scale
    3,670 hours · 923 participants · 74 locations
    Licence
    Restricts commercial use
    What it is

    The reference benchmark for egocentric video, and the corpus every buyer in this market has already read.

    What it does not carry

    Its licence keeps it out of a commercial training run. It is a yardstick, not supply — and it is four years old.

  4. 04

    WAXAL

    Google with Makerere, University of Ghana and Digital Umuganda — February 2026

    Scale
    11,000+ hours · ~2m recordings · 21 languages
    Licence
    Open
    What it is

    The largest open African speech collection there has ever been, covering Yoruba, Hausa, Igbo, Lingala and Fulani among others. It is a good piece of work and it changed what the floor looks like in this region.

    What it does not carry

    About 1,250 hours are transcribed for ASR and roughly 20 hours are studio recording for TTS. The body of it is prompted, single-speaker material, because that is what a prompt elicits. It holds almost no unscripted multi-speaker exchange with real overlap, interruption and code-switching — the requirement scripted collection cannot reach, and the one our audio line is built for.

Figures are the publishers' own, current at September 2026.

Trading in the open beside a road
Nigeria — the environment class open corpora omit · editorial

What survives the comparison.

Open egocentric video is abundant and industrial: factory floors under fixed process, annotated to someone else's schema, with no consent instrument behind the people in frame. Open African speech is abundant and prompted: read sentences, one speaker at a time. Neither describes an open-air market at trading hour, a roadside repair that fails twice before the method changes, or four people negotiating over each other in two languages. That is the whole of what we collect, and it is why the catalogue is narrow on purpose.