Robotics Bottleneck Research机器人瓶颈研究

Main主页 / thesis / brain / generalisable-vla-policies

problem

Generalisable VLA policies可泛化的 VLA 策略

Problem问题 Generalisable VLA policies可泛化的 VLA 策略

Bottleneck瓶颈 Generalizable VLA policies across tasks, scenes, and embodiments跨任务、跨场景、跨本体可泛化的 VLA 策略

Layer层 Foundation Models基础模型

CE score (CE-N)CE 分数(CE-N) 53.6

CE rankCE 排名 #6 / 18

Confidence置信度 High高

Companies mapped关联公司 Link链接

1–5 scale. Budget, solvability and value capture (outlined) enter CE-N; the other three remain context. 1–5 分制。预算、可解性与价值捕获(描边)进入 CE-N,其余三项仅作背景。

P11 Generalisable VLA policies可泛化的 VLA 策略

Maturity成熟度 2.0
Pain痛感 5.0
Budget预算 4.6
Solvability可解性 2.7
Value capture价值捕获 4.8
Timing时机 5.0

What these ratings mean

这些评分代表什么

The input ratings above are stored analyst judgments on a 1–5 scale. This record does not yet contain a source-linked explanation for each rating. Read the scores as provisional judgments while that evidence review remains incomplete.上方输入评级为已存储的分析员判断,采用 1–5 分制。本条目尚未为每个评分提供逐项关联来源的解释。在证据审查完成前,请将评分视为暂定判断。
Maturity成熟度
How established the technology is技术的成熟程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Pain痛点
How severely the problem limits the customer’s task问题对客户任务的限制程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Budget预算
Evidence of willingness and ability to pay支付意愿与支付能力的证据Included in CE-N.计入 CE-N。
Solvability可解性
Feasibility within the assessed scope and time horizon在评估范围与时间内解决问题的可行性Included in CE-N.计入 CE-N。
Value capture价值捕获
Ability of the supplier to retain economic value供应商保留经济价值的能力Included in CE-N.计入 CE-N。
Timing时机
Readiness of the conditions needed for adoption采用所需条件的就绪程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。

Current calculation当前计算方法 · Evidence and rating rules证据与评分规则

What is this problem

A vision-language-action (VLA) model is a foundation model that takes in camera images (and often other sensor streams) plus a natural-language instruction and directly outputs robot actions (joint commands, end-effector poses, or motor torques) rather than relying on hand-coded perception and planning pipelines built separately for each task.

The bet, borrowed from large language models, is that a single large model pretrained on huge amounts of robot and web data can generalize across many manipulation and locomotion tasks, many scenes, and even different robot embodiments (arms, humanoids, mobile bases), instead of needing a bespoke policy per task and per robot.

If it works, VLAs become the “brain” layer that every hardware platform licenses or builds on, much as operating systems or LLMs became horizontal infrastructure in their own domains. That is why this bottleneck sits at the center of the Brain system and attracts more capital than any other problem in the landscape.

The bottleneck and pain points

Despite rapid progress, generalization is still more promise than proven capability. On independent real-robot evaluation leaderboards such as RoboChallenge, the best current VLA models succeed on roughly half of held-out tasks, while more conservative real-world benchmarks that stress novel objects, scenes, or instructions put the figure closer to 10-15%, far below the near-100% success rate of a human operator on the same tasks.

Performance also degrades sharply once a task, object, or environment falls outside the training distribution, undercutting the core generalization claim.

Training and fine-tuning these models is expensive: they require large volumes of paired vision-language-action data, much of which must come from real or simulated robot teleoperation rather than freely available web text, and compute costs scale with model size much as they do for LLMs.

The field is also extremely crowded: dozens of labs, startups, and hardware companies, from NVIDIA and Tesla to Physical Intelligence, Skild AI, and a wave of Chinese humanoid makers, are pursuing broadly similar transformer-based architectures, so architecture choice alone is unlikely to be defensible.

Durable advantage is more likely to come from proprietary fleets generating real-world interaction data, exclusive embodiment access, or paying customers who provide deployment feedback loops, rather than from the model architecture itself.

这是什么问题

视觉-语言-动作(VLA)模型是一类基础模型:它接收摄像头画面(往往还有其他传感器信号)和自然语言指令,直接输出机器人动作——关节指令、末端执行器位姿或电机力矩——而不再依赖为每个任务单独搭建的感知与规划流水线。

其核心赌注借鉴自大语言模型:用海量机器人数据和网络数据预训练出一个大模型,使其能够跨越大量操作与移动任务、跨越不同场景,甚至跨越不同的机器人本体(机械臂、人形机器人、移动底盘)泛化,而不必为每个任务、每台机器人单独训练策略。

如果这一路径走得通,VLA 就会成为几乎所有硬件平台都会调用或基于其构建的”大脑”层,就像操作系统或大语言模型成为各自领域的横向基础设施一样。这正是该瓶颈处于”大脑”系统核心、也是整个图谱中吸引资本最多的问题的原因。

瓶颈与痛点

尽管进展很快,“泛化”目前更多是愿景而非已被证实的能力。在 RoboChallenge 等独立的真实机器人评测榜单上,目前最好的 VLA 模型在留出任务上的成功率大致在五成左右;而更为严格、强调陌生物体、场景或指令的真实世界基准测试给出的数字则更接近 10%-15%,远低于人类操作者接近 100% 的成功率。

一旦任务、物体或环境超出训练分布,模型表现也会明显下滑,这削弱了”泛化”这一核心卖点。

训练和微调这类模型的成本高昂:需要大量配对的视觉-语言-动作数据,其中相当一部分必须来自真实或仿真环境下的机器人遥操作,而非可以低成本获取的网络文本;算力开销也会随模型规模上升,与大语言模型的情况类似。

这个赛道也异常拥挤——从英伟达、特斯拉到 Physical Intelligence、Skild AI,再到一批中国人形机器人公司,数十家实验室、创业公司和硬件厂商都在采用大致相似的 Transformer 架构,因此架构本身很难构成护城河。

真正可持续的优势更可能来自专有机器人机队积累的真实交互数据、独家的本体准入,或是能提供部署反馈闭环的付费客户,而非模型架构本身。