What is this problem
Long-horizon planning and recovery is a robot’s ability to string together many steps toward a goal, rather than simply reacting to the current frame: tracking progress across a task, noticing when something has gone wrong partway through, and replanning or recovering instead of stalling or failing silently.
A single pick or step can look robust in isolation; chaining dozens or hundreds of such actions into a multi-stage task (unload a truck, assemble a subassembly, clear a warehouse aisle) is a different problem, because small errors compound at each step and the robot must reason about what to do next given an uncertain, partially observed state.
This is what separates a choreographed demo, run once under a human’s watchful eye, from a system trusted to run unattended for a full shift.
The bottleneck and pain points
Even when individual skills (grasping, navigation, manipulation) have decent per-step success rates, chaining them multiplies failure probability across a long task, so a robot that succeeds 95% of the time on any single step can still fail most attempts at a 50-step job.
Recovery behavior is typically thin: when a grasp slips, an object shifts, or a door doesn’t open as expected, the common response is to retry blindly or halt and wait for a human, rather than diagnose the failure and adapt.
Most published benchmarks and demos run in clean, pre-staged environments that don’t stress-test the messy, adversarial conditions (clutter, occlusion, unexpected obstacles) where recovery actually matters, so reported success rates travel poorly to the field. Genuine uncertainty-aware decision-making (a robot recognizing when its world model is wrong or its confidence is low, rather than confidently executing the wrong action) remains largely unsolved outside narrow, engineered settings.
This is one of the lowest-solvability, lowest-confidence bottlenecks on the board: promising directions exist (world models, hierarchical planners, VLA-based replanning), but little of it has been shown to generalize beyond curated demos or scripted fallback logic, and progress here remains thin and mostly unproven in the field.