Robotics Bottleneck Research机器人瓶颈研究

Main主页 / thesis / brain / predictive-world-models

problem

Predictive world models预测式世界模型

Problem问题 Predictive world models预测式世界模型

Bottleneck瓶颈 Predictive world models for physical dynamics and consequences面向物理动力学与后果的预测式世界模型

Layer层 World Models世界模型

CE score (CE-N)CE 分数(CE-N) 48.6

CE rankCE 排名 #11 / 18

Confidence置信度 Moderate中

Companies mapped关联公司 Link链接

1–5 scale. Budget, solvability and value capture (outlined) enter CE-N; the other three remain context. 1–5 分制。预算、可解性与价值捕获(描边)进入 CE-N,其余三项仅作背景。

P10 Predictive world models预测式世界模型

Maturity成熟度 2.0
Pain痛感 4.9
Budget预算 4.4
Solvability可解性 2.8
Value capture价值捕获 4.7
Timing时机 4.9

What these ratings mean

这些评分代表什么

The input ratings above are stored analyst judgments on a 1–5 scale. This record does not yet contain a source-linked explanation for each rating. Read the scores as provisional judgments while that evidence review remains incomplete.上方输入评级为已存储的分析员判断,采用 1–5 分制。本条目尚未为每个评分提供逐项关联来源的解释。在证据审查完成前,请将评分视为暂定判断。
Maturity成熟度
How established the technology is技术的成熟程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Pain痛点
How severely the problem limits the customer’s task问题对客户任务的限制程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Budget预算
Evidence of willingness and ability to pay支付意愿与支付能力的证据Included in CE-N.计入 CE-N。
Solvability可解性
Feasibility within the assessed scope and time horizon在评估范围与时间内解决问题的可行性Included in CE-N.计入 CE-N。
Value capture价值捕获
Ability of the supplier to retain economic value供应商保留经济价值的能力Included in CE-N.计入 CE-N。
Timing时机
Readiness of the conditions needed for adoption采用所需条件的就绪程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。

Current calculation当前计算方法 · Evidence and rating rules证据与评分规则

What is this problem

A predictive world model is a learned model of physical dynamics: given the current state of a scene and a candidate action, it predicts what happens next, how an object moves when pushed, how a grasp will succeed or fail, how the ground will respond to a footstep.

Rather than only reacting to what it senses, a robot (or its planner) can use a world model to mentally roll out several candidate actions, compare their predicted outcomes, and choose before acting on real hardware. That’s the robotic analogue of a person imagining a move before making it.

World models can predict in raw sensory space (future video frames) or in a compressed latent representation, and they get used in a few ways: driving model-predictive control, training policies in “imagined” rollouts instead of real trials, or scoring and filtering the outputs of a separate action policy. The appeal is generalization and safety: costly, rare, or dangerous outcomes can be evaluated virtually rather than discovered by breaking hardware.

The bottleneck and pain points

The core problem is that prediction accuracy collapses over longer horizons: small errors in one predicted step compound into the next, so plans built on more than a few steps of rollout quickly become unreliable, and this gets worse the more contact-rich and deformable the interaction is (fabric, fluids, granular media, multi-object contact).

Learning accurate physical dynamics requires enormous amounts of diverse interaction data, but the real-world robot interaction data that exists today is orders of magnitude smaller and narrower than what video- or language-scale models were trained on. There is no internet-scale corpus of “robot pushed object, here’s what happened” the way there is for text or images.

Models trained in one embodiment or one environment also transfer poorly to a different robot morphology, sensor suite, or scene distribution, which undercuts the pitch that a single world model generalizes broadly.

Because of this data gap, the technology that gets funded matters less than who is generating and controlling the interaction data behind it: a world model built on public datasets or simulation alone, without a pipeline of proprietary real-world interaction data or an actual deployed robot fleet generating that data, does not have a defensible moat, however strong the underlying research is.

这是什么问题

预测式世界模型是一种学习出来的物理动力学模型:给定场景当前状态和一个候选动作,它预测接下来会发生什么——物体被推动后如何移动、抓取动作会成功还是失败、脚踏上地面后地面会如何反馈。

机器人(或其规划模块)不必只是被动地对感知做出反应,而是可以借助世界模型在”脑内”推演多个候选动作,比较各自的预测结果,再决定在真实硬件上执行哪一个——这类似于人在动手之前先在脑海中演练一遍。

世界模型既可以在原始感知空间中预测(比如预测未来的视频帧),也可以在压缩后的潜在表征空间中预测,其用途也有几种:驱动模型预测控制、在”想象出来的”推演中训练策略而不必依赖真实试错、或者对另一个动作策略的输出进行打分和筛选。这类方法的吸引力在于泛化能力和安全性——代价高、罕见或危险的后果可以先在虚拟中评估,而不必靠损坏真实硬件来发现。

瓶颈与痛点

核心问题在于预测精度会随着推演步数增加而迅速崩溃:某一步预测的小误差会累积传递到下一步,因此建立在较长推演序列之上的规划很快就变得不可靠,而在接触关系复杂、涉及可变形物体的交互中(如布料、流体、颗粒状介质、多物体接触),这一问题会更加严重。

要学到准确的物理动力学,需要海量且多样的交互数据,但当下真实世界中可获取的机器人交互数据,无论在数量还是覆盖范围上,都比训练视频或语言大模型所用的数据小得多、窄得多——目前并不存在类似”机器人推了某物体、结果如何”这样互联网规模的语料库,不像文本和图像那样。

在某一具身形态或某一环境中训练出的模型,迁移到不同的机器人形态、传感器配置或场景分布时效果也往往很差,这削弱了”单一世界模型可以广泛泛化”这一说法的说服力。

正因为存在这一数据鸿沟,真正值得关注的与其说是模型本身的技术水平,不如说是谁在生产并掌控背后的交互数据:一个仅依赖公开数据集或纯仿真训练出来的世界模型,如果没有专有的真实交互数据管线、也没有实际部署的机器人机队持续产生数据,无论其底层研究多么出色,都很难形成可防御的护城河。