What is this problem
Edge inference means running a robot’s perception, control, and policy models directly on its onboard compute, in real time, inside a fixed power, thermal, and weight budget, rather than shipping sensor data to the cloud and waiting for a response.
A mobile or humanoid robot carries its compute and its battery with it, so every watt spent on a chip is a watt not spent on motors, and every gram of heatsink is a gram not spent on payload.
Safety-critical control loops (balance, collision avoidance, force control) and high-frequency actuation loops run at tens to hundreds of hertz and cannot tolerate the latency, jitter, or occasional dropout of a network round-trip to a remote GPU. That makes low-latency, power-efficient onboard inference a hard requirement for autonomy outside of controlled, well-networked environments, not an optimization to add later.
The bottleneck and pain points
The central tension is that the models robotics increasingly wants to run (vision-language-action models and other foundation-model-scale policies) keep growing, while the power envelope of a battery-powered mobile or humanoid robot stays roughly fixed at tens to a couple hundred watts for compute.
Vendor-quoted TFLOPS/TOPS figures are measured under best-case conditions and routinely overstate real sustained throughput once a model is actually running inference at low wattage inside an enclosed, fanless or lightly cooled chassis, where thermal throttling silently cuts performance during sustained operation.
Compute silicon and modules are also frequently announced, sampled, and design-won well ahead of volume shipment, so a chip’s presence in a roadmap or a partner’s dev kit is not evidence of revenue today.
On the cost side, the inference chip and its memory can be one of the largest single line items in a robot’s bill of materials, so price-performance directly pressures robot gross margins.
Finally, porting large models down to efficient edge runtimes requires quantization, pruning, and compiler/toolchain work that lags well behind the pace of model research, creating a persistent gap between what labs publish and what can actually run on a robot.