Robotics Bottleneck Research机器人瓶颈研究

Main主页 / thesis / brain / edge-inference

problem

Edge inference边缘推理

Problem问题 Edge inference边缘推理

Bottleneck瓶颈 Low-latency, power-efficient edge inference低延迟、低功耗的边缘推理

Layer层 Compute计算

CE score (CE-N)CE 分数(CE-N) 63

CE rankCE 排名 #2 / 18

Confidence置信度 Low低

Companies mapped关联公司 Link链接

1–5 scale. Budget, solvability and value capture (outlined) enter CE-N; the other three remain context. 1–5 分制。预算、可解性与价值捕获(描边)进入 CE-N,其余三项仅作背景。

P04 Edge inference边缘推理

Maturity成熟度 4.0
Pain痛感 4.6
Budget预算 4.5
Solvability可解性 4.1
Value capture价值捕获 4.6
Timing时机 4.8

What these ratings mean

这些评分代表什么

The input ratings above are stored analyst judgments on a 1–5 scale. This record does not yet contain a source-linked explanation for each rating. Read the scores as provisional judgments while that evidence review remains incomplete.上方输入评级为已存储的分析员判断,采用 1–5 分制。本条目尚未为每个评分提供逐项关联来源的解释。在证据审查完成前,请将评分视为暂定判断。
Maturity成熟度
How established the technology is技术的成熟程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Pain痛点
How severely the problem limits the customer’s task问题对客户任务的限制程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。
Budget预算
Evidence of willingness and ability to pay支付意愿与支付能力的证据Included in CE-N.计入 CE-N。
Solvability可解性
Feasibility within the assessed scope and time horizon在评估范围与时间内解决问题的可行性Included in CE-N.计入 CE-N。
Value capture价值捕获
Ability of the supplier to retain economic value供应商保留经济价值的能力Included in CE-N.计入 CE-N。
Timing时机
Readiness of the conditions needed for adoption采用所需条件的就绪程度Context only; excluded from CE-N.仅作背景;未计入 CE-N。

Current calculation当前计算方法 · Evidence and rating rules证据与评分规则

What is this problem

Edge inference means running a robot’s perception, control, and policy models directly on its onboard compute, in real time, inside a fixed power, thermal, and weight budget, rather than shipping sensor data to the cloud and waiting for a response.

A mobile or humanoid robot carries its compute and its battery with it, so every watt spent on a chip is a watt not spent on motors, and every gram of heatsink is a gram not spent on payload.

Safety-critical control loops (balance, collision avoidance, force control) and high-frequency actuation loops run at tens to hundreds of hertz and cannot tolerate the latency, jitter, or occasional dropout of a network round-trip to a remote GPU. That makes low-latency, power-efficient onboard inference a hard requirement for autonomy outside of controlled, well-networked environments, not an optimization to add later.

The bottleneck and pain points

The central tension is that the models robotics increasingly wants to run (vision-language-action models and other foundation-model-scale policies) keep growing, while the power envelope of a battery-powered mobile or humanoid robot stays roughly fixed at tens to a couple hundred watts for compute.

Vendor-quoted TFLOPS/TOPS figures are measured under best-case conditions and routinely overstate real sustained throughput once a model is actually running inference at low wattage inside an enclosed, fanless or lightly cooled chassis, where thermal throttling silently cuts performance during sustained operation.

Compute silicon and modules are also frequently announced, sampled, and design-won well ahead of volume shipment, so a chip’s presence in a roadmap or a partner’s dev kit is not evidence of revenue today.

On the cost side, the inference chip and its memory can be one of the largest single line items in a robot’s bill of materials, so price-performance directly pressures robot gross margins.

Finally, porting large models down to efficient edge runtimes requires quantization, pruning, and compiler/toolchain work that lags well behind the pace of model research, creating a persistent gap between what labs publish and what can actually run on a robot.

这是什么问题

边缘推理是指机器人在自身的车载算力上实时运行感知、控制与策略模型,并且必须限定在固定的功耗、散热与重量预算之内,而不是把传感器数据传到云端、等待返回结果。

移动或人形机器人要自带算力和电池,芯片每多耗一瓦,电机可用的功率就少一瓦;散热片每多一克,可用于负载的重量就少一克。

平衡控制、避障、力控等安全关键回路,以及以数十到数百赫兹运行的高频执行回路,都无法容忍网络往返到远端GPU所带来的延迟、抖动或偶发断连。因此,低延迟、低功耗的车载推理是机器人在受控、网络良好环境之外实现自主运行的硬性要求,而不是可以留到后期再优化的选项。

瓶颈与痛点

核心矛盾在于:机器人行业越来越想运行的模型——视觉-语言-动作(VLA)模型及其他基础模型级别的策略——规模持续增长,而电池驱动的移动或人形机器人的算力功耗预算大体固定在几十到一两百瓦。

芯片厂商公布的TFLOPS/TOPS数字通常在最理想条件下测得,一旦模型真正在低功耗、封闭且散热能力有限(无风扇或轻度散热)的机身内运行推理,实际持续吞吐量往往被严重高估,持续运行时的热降频还会在不知不觉中拉低性能。

新一代高性能计算模块也常常在正式量产出货之前很久就被官宣、送样并拿下设计中标,因此某款芯片出现在厂商路线图或合作伙伴的开发套件里,并不能证明它已经带来实际收入。

在成本端,推理芯片及其配套内存往往是机器人物料成本中占比最大的单项之一,价格与性能的取舍直接压缩了机器人的整机毛利。

最后,把大模型压缩到能在边缘高效运行,需要量化、剪枝以及配套的编译器和工具链工作,而这方面的进展长期落后于模型研究本身的节奏,导致实验室发布的成果与机器人实际可运行的能力之间始终存在差距。