What is this problem
Multimodal sensing and calibration is about fusing cameras, LiDAR, IMUs, tactile/force sensors, and sometimes radar into a single, temporally and spatially consistent estimate of a robot’s surroundings and its own pose within them. Each sensor has blind spots and failure modes on its own; fusion is what lets a robot trust its perception enough to act on it.
Calibration is the often-invisible layer underneath: knowing precisely how each sensor is positioned and oriented relative to the others (extrinsics) and how its own optics and electronics behave (intrinsics), so that a LiDAR point cloud, a camera pixel, and an IMU reading all describe the same physical event in the same coordinate frame. Without accurate, stable calibration, fusion doesn’t fail loudly. It produces a confidently wrong picture, which is arguably worse for a robot that has to plan motion around it.
This is why sensing is treated as a deployment precondition rather than a differentiator: a robot cannot be given autonomy budget until its senses are trustworthy.
The bottleneck and pain points
The core tension is cost versus robustness: LiDAR ASPs have collapsed from well over RMB10,000 to roughly RMB1,000-2,000 as Chinese suppliers scaled production, but cheaper units generally trade off range, point density, or performance in rain, dust, fog, and direct sunlight: conditions lab bench-testing rarely replicates.
Calibration compounds the problem: extrinsics (the geometric relationship between sensors) drift with thermal cycling, mechanical vibration, and ordinary wear, and can be knocked out entirely by a bumped mount, a dropped payload, or a swapped camera during field service. Yet most fleets lack standardized, automated recalibration tooling, so drift is typically caught only after downstream perception errors surface. Sensors still represent a meaningful share of total robot BOM (commonly cited in the 10-30% range for mobile and humanoid platforms, depending on the sensor suite), which pushes integrators toward cheaper components at the exact moment robustness matters most.
Fusion software has to paper over these imperfections in real time, but garbage-in from an uncalibrated or degraded sensor produces a confident, silently wrong state estimate: a harder failure mode to catch than an outright sensor dropout. The net effect is that many deployments work reliably in controlled pilot environments but degrade in the field, where dust, vibration, temperature swings, and inconsistent lighting are the norm rather than the exception.