First-person work, including the part where it goes wrong
Manipulation policies generalise from the variety of hands, tools and failures they have seen. The open egocentric corpora are a fixed process in an industrial interior, filmed until it works.

The problem
Egocentric datasets exist, and they are narrow in a way that matters: a known factory, a known tool, a known sequence, repeated. A policy trained on them learns that manipulation is what happens when the tool fits, the light is even and the task completes.
The material that would teach otherwise — improvised tools, worn parts, an unfamiliar grip, a task attempted one way and finished another — has no archive, because nobody filmed it. This is not a licensing gap that can be closed by clearing rights on existing footage. It has to be collected.
What we put against it
Purpose-collected head- and chest-mounted capture from open-air market trading, street food preparation, roadside vehicle repair, artisan workshops and dense two-wheeler traffic — shot to a brief, from the operator's own point of view.
Every session runs a synchronised second angle, clap-synced in frame, because coverage of one subject from two angles is worth materially more than either alone. And every clip carries a written context layer describing task, intent, objects handled and — where it happens — the sequence in which something failed and the operator changed method. The failure events are the point, and they are indexed rather than left for you to find.





What this has to satisfy
| Requirement | How it is met |
|---|---|
| Failure and recovery, not just clean completion | Failure events indexed per clip with the change of method described |
| Tool and object variety outside a fixed process | Improvised tooling and worn parts in ordinary working use |
| Multi-view coverage of the same action | Synchronised second angle every session, clap-synced in frame |
| Task intent, not just pixels | Written context layer per clip — task, intent, objects handled |
What you can check first
- Sample annotation
- A full annotation JSON published for a real asset, with the failure events in it
- Capture spec
- 3840×2160, 30 fps (60 for motion), H.264, no HDR
- Appearance releases
- Every identifiable person on a release; contributors paid and the payment recorded
- Collected to brief
- Commissioned against a written brief rather than sold from stock
Supply lines behind this


A solution is marked live only where a line behind it holds cleared material. Nothing sourced to mandate is presented as inventory.
Fifty delivered hours at production standard, two weeks from spec sign-off.
Other solutions

ASR and speech for languages the benchmarks skip
Word error rates in this corridor are not a modelling problem. They are a data problem, and the data was never collected.

Post-training and evaluation outside the crawl
The web is not a sample of the world. For this corridor it is a sample of what got typed in English by people with a connection — and a model trained on it inherits that gap as confidence.

Documents as they are actually filed
An extraction model that has only seen clean templates has not seen this corridor's paperwork — the stamp over the field, the third-generation photocopy, the form filled in two hands.