What is this problem
Most robot manipulation policies today are trained by imitation learning, which means they need action-labeled demonstrations of a specific embodiment performing a specific task, not the internet-scale text, image, and video corpora that trained large language and vision models.
That data comes from teleoperation rigs, motion-capture suits, handheld grippers, or instrumented human demonstrations, each paired frame-by-frame with the joint angles, forces, or end-effector poses the robot actually needs to reproduce.
Collecting, cleaning, and labeling this data at the volume and diversity required for a policy to generalize across tasks, objects, and environments is its own infrastructure problem, distinct from the model architectures or the hardware itself.
The bottleneck and pain points
Teleoperated demonstration data is expensive per usable hour: it requires trained operators, calibrated rigs or motion-capture setups, and dedicated studio or lab time, and a large share of collected episodes are discarded for failures, ambiguity, or poor state coverage. Most existing datasets are narrow relative to what a general-purpose policy needs (concentrated in a handful of labs, embodiments, and task types), so cross-embodiment and cross-task generalization remains largely unproven outside curated benchmarks.
Human intervention and correction rates during deployment are the honest scorecard here: the more often a human has to step in to unstick or correct a policy, the further the system is from real autonomy, and most fielded systems still intervene far more often than headline demo reels suggest. Quality control and labeling consistency are hard to maintain at scale: action labels, segmentation, and success/failure annotation all require judgment calls that are easy to get wrong or inconsistent across annotators and sites.
The core point for investors is that undifferentiated video volume is not the constraint; simulation and internet video are already abundant. What is scarce, and what actually differentiates data providers, is task-rich, action-aligned, quality-controlled real-world data collected at a cost structure that can scale.