What is this problem
A vision-language-action (VLA) model is a foundation model that takes in camera images (and often other sensor streams) plus a natural-language instruction and directly outputs robot actions (joint commands, end-effector poses, or motor torques) rather than relying on hand-coded perception and planning pipelines built separately for each task.
The bet, borrowed from large language models, is that a single large model pretrained on huge amounts of robot and web data can generalize across many manipulation and locomotion tasks, many scenes, and even different robot embodiments (arms, humanoids, mobile bases), instead of needing a bespoke policy per task and per robot.
If it works, VLAs become the “brain” layer that every hardware platform licenses or builds on, much as operating systems or LLMs became horizontal infrastructure in their own domains. That is why this bottleneck sits at the center of the Brain system and attracts more capital than any other problem in the landscape.
The bottleneck and pain points
Despite rapid progress, generalization is still more promise than proven capability. On independent real-robot evaluation leaderboards such as RoboChallenge, the best current VLA models succeed on roughly half of held-out tasks, while more conservative real-world benchmarks that stress novel objects, scenes, or instructions put the figure closer to 10-15%, far below the near-100% success rate of a human operator on the same tasks.
Performance also degrades sharply once a task, object, or environment falls outside the training distribution, undercutting the core generalization claim.
Training and fine-tuning these models is expensive: they require large volumes of paired vision-language-action data, much of which must come from real or simulated robot teleoperation rather than freely available web text, and compute costs scale with model size much as they do for LLMs.
The field is also extremely crowded: dozens of labs, startups, and hardware companies, from NVIDIA and Tesla to Physical Intelligence, Skild AI, and a wave of Chinese humanoid makers, are pursuing broadly similar transformer-based architectures, so architecture choice alone is unlikely to be defensible.
Durable advantage is more likely to come from proprietary fleets generating real-world interaction data, exclusive embodiment access, or paying customers who provide deployment feedback loops, rather than from the model architecture itself.