Robotics Bottleneck Research机器人瓶颈研究

Main主页 / thesis / brain / long-horizon-planning-recovery

problem

Long-horizon planning & recovery长时程规划与恢复

Problem问题 Long-horizon planning & recovery长时程规划与恢复

Bottleneck瓶颈 Long-horizon planning, recovery, and uncertainty-aware decision making长时程规划、故障恢复与不确定性感知决策

Layer层 Planning规划

CE score (CE-N)CE 分数(CE-N) 43.9

CE rankCE 排名 #14 / 18

Confidence置信度 Low低

Companies mapped关联公司 Link链接

1–5 scale. Budget, solvability and value capture (outlined) enter CE-N; the other three remain context. 1–5 分制。预算、可解性与价值捕获(描边)进入 CE-N,其余三项仅作背景。

P12 Long-horizon planning & recovery长时程规划与恢复

Maturity成熟度 2.0
Pain痛感 4.8
Budget预算 4.4
Solvability可解性 2.6
Value capture价值捕获 4.5
Timing时机 4.8

What these ratings mean

这些评分代表什么

The input ratings above are stored analyst judgments on a 1–5 scale. This record does not yet contain a source-linked explanation for each rating. Read the scores as provisional judgments while that evidence review remains incomplete.上方输入评级为已存储的分析员判断,采用 1–5 分制。本条目尚未为每个评分提供逐项关联来源的解释。在证据审查完成前,请将评分视为暂定判断。
Maturity成熟度
How established the technology is技术的成熟程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Pain痛点
How severely the problem limits the customer’s task问题对客户任务的限制程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Budget预算
Evidence of willingness and ability to pay支付意愿与支付能力的证据Included in CE-N.计入 CE-N。
Solvability可解性
Feasibility within the assessed scope and time horizon在评估范围与时间内解决问题的可行性Included in CE-N.计入 CE-N。
Value capture价值捕获
Ability of the supplier to retain economic value供应商保留经济价值的能力Included in CE-N.计入 CE-N。
Timing时机
Readiness of the conditions needed for adoption采用所需条件的就绪程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。

Current calculation当前计算方法 · Evidence and rating rules证据与评分规则

What is this problem

Long-horizon planning and recovery is a robot’s ability to string together many steps toward a goal, rather than simply reacting to the current frame: tracking progress across a task, noticing when something has gone wrong partway through, and replanning or recovering instead of stalling or failing silently.

A single pick or step can look robust in isolation; chaining dozens or hundreds of such actions into a multi-stage task (unload a truck, assemble a subassembly, clear a warehouse aisle) is a different problem, because small errors compound at each step and the robot must reason about what to do next given an uncertain, partially observed state.

This is what separates a choreographed demo, run once under a human’s watchful eye, from a system trusted to run unattended for a full shift.

The bottleneck and pain points

Even when individual skills (grasping, navigation, manipulation) have decent per-step success rates, chaining them multiplies failure probability across a long task, so a robot that succeeds 95% of the time on any single step can still fail most attempts at a 50-step job.

Recovery behavior is typically thin: when a grasp slips, an object shifts, or a door doesn’t open as expected, the common response is to retry blindly or halt and wait for a human, rather than diagnose the failure and adapt.

Most published benchmarks and demos run in clean, pre-staged environments that don’t stress-test the messy, adversarial conditions (clutter, occlusion, unexpected obstacles) where recovery actually matters, so reported success rates travel poorly to the field. Genuine uncertainty-aware decision-making (a robot recognizing when its world model is wrong or its confidence is low, rather than confidently executing the wrong action) remains largely unsolved outside narrow, engineered settings.

This is one of the lowest-solvability, lowest-confidence bottlenecks on the board: promising directions exist (world models, hierarchical planners, VLA-based replanning), but little of it has been shown to generalize beyond curated demos or scripted fallback logic, and progress here remains thin and mostly unproven in the field.

这是什么问题

长时程规划与恢复,指机器人能否把许多步骤串联起来去完成一个目标,而不只是对当下这一帧做出反应——在任务推进过程中持续追踪进度,察觉某个环节出了问题,并及时重新规划或恢复,而不是卡住不动或悄无声息地失败。

单看一次抓取或一步动作,机器人可能表现得相当可靠;但把几十上百个这样的动作串成一个多阶段任务(卸一车货、组装一个部件、清空一条仓库通道),性质就完全不同了,因为每一步的小误差会不断累积,机器人必须在状态不确定、观测又不完整的情况下判断下一步该怎么做。

这正是”在人盯着的情况下演示成功一次”和”能无人值守跑完一整班”之间的分水岭。

瓶颈与痛点

即便抓取、导航、操作这些单项技能各自的成功率都不错,把它们串联起来后,失败概率会成倍叠加——单步成功率95%的机器人,在一个50步的任务里也可能大概率整体失败。

恢复能力普遍薄弱:抓取打滑、物体移位、门没能按预期打开时,常见的应对是盲目重试或直接停机等人工介入,而不是先诊断失败原因再做适应性调整。

目前公开的评测和演示大多在整洁、预先布置好的环境中进行,很少真正测试杂乱、带对抗性的现场条件——遮挡、意外障碍物、混乱堆放——而这些恰恰是最考验恢复能力的地方,所以演示中的成功率很难迁移到真实场景。真正具备不确定性感知的决策能力——机器人能意识到自己的世界模型出错了或置信度不够,而不是自信地执行错误动作——在狭窄的工程化场景之外基本仍未解决。

这是整个版图上可解性最低、置信度最低的瓶颈之一:世界模型、分层规划器、基于VLA的重新规划等方向都有一定苗头,但很少有证据表明这些方法能超越精心策划的演示或写死的兜底逻辑而真正泛化,这一领域的进展总体依然稀薄,且大多未在实际现场得到验证。