Robotics Bottleneck Research机器人瓶颈研究

Main主页 / thesis / senses / multimodal-sensing-calibration

problem

Multimodal sensing & calibration多模态传感与标定

Problem问题 Multimodal sensing & calibration多模态传感与标定

Bottleneck瓶颈 Low-cost, robust multimodal sensing and calibration低成本、鲁棒的多模态传感与标定

Layer层 Sense感知

CE score (CE-N)CE 分数(CE-N) 38.6

CE rankCE 排名 #15 / 18

Confidence置信度 High高

Companies mapped关联公司 Link链接

Manufacture mapped关联制造 Link链接

1–5 scale. Budget, solvability and value capture (outlined) enter CE-N; the other three remain context. 1–5 分制。预算、可解性与价值捕获(描边)进入 CE-N,其余三项仅作背景。

P01 Multimodal sensing & calibration多模态传感与标定

Maturity成熟度 4.0
Pain痛感 4.5
Budget预算 4.0
Solvability可解性 4.2
Value capture价值捕获 4.0
Timing时机 4.5

What these ratings mean

这些评分代表什么

The input ratings above are stored analyst judgments on a 1–5 scale. This record does not yet contain a source-linked explanation for each rating. Read the scores as provisional judgments while that evidence review remains incomplete.上方输入评级为已存储的分析员判断,采用 1–5 分制。本条目尚未为每个评分提供逐项关联来源的解释。在证据审查完成前,请将评分视为暂定判断。
Maturity成熟度
How established the technology is技术的成熟程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Pain痛点
How severely the problem limits the customer’s task问题对客户任务的限制程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Budget预算
Evidence of willingness and ability to pay支付意愿与支付能力的证据Included in CE-N.计入 CE-N。
Solvability可解性
Feasibility within the assessed scope and time horizon在评估范围与时间内解决问题的可行性Included in CE-N.计入 CE-N。
Value capture价值捕获
Ability of the supplier to retain economic value供应商保留经济价值的能力Included in CE-N.计入 CE-N。
Timing时机
Readiness of the conditions needed for adoption采用所需条件的就绪程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。

Current calculation当前计算方法 · Evidence and rating rules证据与评分规则

What is this problem

Multimodal sensing and calibration is about fusing cameras, LiDAR, IMUs, tactile/force sensors, and sometimes radar into a single, temporally and spatially consistent estimate of a robot’s surroundings and its own pose within them. Each sensor has blind spots and failure modes on its own; fusion is what lets a robot trust its perception enough to act on it.

Calibration is the often-invisible layer underneath: knowing precisely how each sensor is positioned and oriented relative to the others (extrinsics) and how its own optics and electronics behave (intrinsics), so that a LiDAR point cloud, a camera pixel, and an IMU reading all describe the same physical event in the same coordinate frame. Without accurate, stable calibration, fusion doesn’t fail loudly. It produces a confidently wrong picture, which is arguably worse for a robot that has to plan motion around it.

This is why sensing is treated as a deployment precondition rather than a differentiator: a robot cannot be given autonomy budget until its senses are trustworthy.

The bottleneck and pain points

The core tension is cost versus robustness: LiDAR ASPs have collapsed from well over RMB10,000 to roughly RMB1,000-2,000 as Chinese suppliers scaled production, but cheaper units generally trade off range, point density, or performance in rain, dust, fog, and direct sunlight: conditions lab bench-testing rarely replicates.

Calibration compounds the problem: extrinsics (the geometric relationship between sensors) drift with thermal cycling, mechanical vibration, and ordinary wear, and can be knocked out entirely by a bumped mount, a dropped payload, or a swapped camera during field service. Yet most fleets lack standardized, automated recalibration tooling, so drift is typically caught only after downstream perception errors surface. Sensors still represent a meaningful share of total robot BOM (commonly cited in the 10-30% range for mobile and humanoid platforms, depending on the sensor suite), which pushes integrators toward cheaper components at the exact moment robustness matters most.

Fusion software has to paper over these imperfections in real time, but garbage-in from an uncalibrated or degraded sensor produces a confident, silently wrong state estimate: a harder failure mode to catch than an outright sensor dropout. The net effect is that many deployments work reliably in controlled pilot environments but degrade in the field, where dust, vibration, temperature swings, and inconsistent lighting are the norm rather than the exception.

这是什么问题

多模态传感与标定,是指将摄像头、激光雷达(LiDAR)、惯性测量单元(IMU)、触觉/力传感器,有时还包括毫米波雷达,融合成对机器人周围环境及自身位姿在时间和空间上一致的估计。每种传感器单独看都有各自的盲区和失效模式;融合正是让机器人能够信任其感知结果、从而据此行动的关键。

标定则是这一切背后常常被忽视的底层能力:需要精确掌握每个传感器相对其他传感器的位置与朝向(外参),以及其自身光学与电子特性的表现(内参),才能让激光雷达点云、摄像头像素和IMU读数在同一坐标系下描述同一个物理事件。标定不准确或不稳定时,融合不会明显地”失灵”,而是给出一幅看似可信却是错误的画面——对于需要据此规划运动的机器人而言,这可能是更危险的情况。

这正是为什么传感被视为部署的前置条件而非差异化因素:机器人只有在其感知可信之后,才能被赋予自主决策的空间。

瓶颈与痛点

核心矛盾在于成本与鲁棒性的取舍:随着中国供应商规模化生产,激光雷达的平均售价已从远高于一万元人民币降至一千至两千元区间,但更便宜的型号通常要在探测距离、点云密度,或雨、雾、沙尘、强光等环境下的表现上做出妥协——而这些工况很少能在实验室台架测试中被真实还原。

标定问题让情况更复杂:外参(传感器之间的几何关系)会随温度循环、机械振动和日常磨损而漂移,一次磕碰、一次货物跌落,或现场维修时更换一个摄像头,都可能让标定完全失效;然而大多数机队缺乏标准化、自动化的重新标定工具,漂移往往要等到下游感知出现错误后才被发现。传感器成本在机器人整体物料成本(BOM)中仍占相当比例——移动机器人和人形机器人常见的引用区间在10%到30%之间,具体取决于传感器套件配置——这使得集成商在最需要鲁棒性的时候,反而更倾向于选择廉价元件。

融合软件需要在实时运行中弥补这些不完美,但一个未校准或性能退化的传感器输入的是”垃圾数据”,融合系统给出的往往是自信却错误的状态估计——这比传感器直接失效更难被察觉。最终结果是:许多部署项目在受控的试点环境中表现良好,但一旦进入现场——沙尘、振动、温度变化和不稳定光照是常态而非例外——性能便会下降。