Research · QC Layer 2
We hand-checked 240 auto-generated captions in a production robotics dataset. One in three was wrong.
Everyone is benchmarking whether AI can write video labels. We asked whether the labels already shipping in embodied-AI datasets are actually true. Roughly a third failed a human check, the errors concentrate in the verbs, and a calibrated VLM judge catches about half of them with 90% precision for $3 per hour of footage.
Read the study →8 min