Robotics Bottleneck Research机器人瓶颈研究

Main主页 / thesis / immune

system

Immune免疫

CE-NCE-N 46.5

Safety testing and certification are normally just a cost, but whoever becomes the trusted certifier gains a real edge competitors can't easily copy.

安全验证与认证:监管由成本转为护城河之处。

Bottlenecks瓶颈 P16 Safety validation & certification安全验证与认证

What it isImmune is the safety-validation layer: monitoring, testing, and certifying robotic systems before and after they operate around people. The systems that matter most here are 'learned' ones (AI models trained on data rather than hand-coded step-by-step) because their behavior is harder to fully test and predict in advance. That's the real problem this system covers (P16, Safety validation & certification).

是什么“免疫”(Immune)是安全验证层:在学习型(而非人工编码)机器人系统部署到人类周围前后,对其进行监控、测试与认证(P16,安全验证与认证)。

ThesisThis is one of the taxonomy's more contrarian bets, and the evidence so far is genuinely mixed, reflected in a weak CE-N rank: 13th of 18, below the midpoint of the board. The idea: safety certification is normally just a cost, but if a single trusted certifier ever becomes the standard everyone must pass, passing it first turns into a real, hard-to-copy advantage. Nobody has become that trusted certifier yet, which is exactly why the rank is weak today. AeroVironment (a maker of military drones and counter-drone systems) shows what the advantage looks like once it does form: its certifications plus real battlefield performance data create a barrier competitors can't simply buy their way past. EHang (a company building autonomous flying passenger vehicles) shows the opposite case: it already holds official flight-safety clearance (airworthiness), but that clearance hasn't translated into revenue. Only four aircraft were delivered in Q1 2026. Until a neutral certifier or evaluation consortium actually exists, the safer read is that most of this system's value remains uncaptured, not mispriced.

论点这里是本分类法中最具争议的押注所在:监管有可能从成本转化为护城河,但前提是要出现一个可信的验证方——而如今尚未出现,“免疫”在 CE-N 排名中位列 18 个瓶颈中的第 13。AeroVironment 展示了护城河成形后的样子:认证叠加实战数据,构成了竞争对手难以靠资金绕过的真正壁垒。EHang 则展示了另一面——认证只是一个无需行使的期权,而非收入驱动力:虽已取得适航许可,2026 年第一季度仅交付四架飞行器。在中立认证方或评估联盟出现之前,更稳妥的判断是:这里的大部分价值仍未被捕获,而非被错误定价。

1–5 scale. Budget, solvability and value capture (outlined) enter CE-N; the other three remain context. 1–5 分制。预算、可解性与价值捕获(描边)进入 CE-N,其余三项仅作背景。

P16 Safety validation & certification安全验证与认证

Maturity成熟度 2.0
Pain痛感 4.8
Budget预算 4.5
Solvability可解性 3.2
Value capture价值捕获 4.3
Timing时机 4.7

Products mapped to this system映射到本系统的产品

Physical-AI capability evaluation and benchmarkingPhysical-AI capability evaluation and benchmarking

Standardized, closed-loop evaluation of real-robot and Physical-AI task performance -- proving how good a policy or robot actually is rather than just that it runs.

1 companies 1 家公司 +1 named in conference materials, unverified +1 项来自会议材料,未经核实

Hardware-in-loop validation and safety hardwareHardware-in-loop validation and safety hardware

Testing robot behavior under controlled faults and degraded states, plus the safety scanners and safety devices that make human-robot collaboration compliant.

Sensors & Vision传感与视觉

1 companies 1 家公司 2 manufacturers 2 家制造商 +1 named in conference materials, unverified +1 项来自会议材料,未经核实

Companies mapped to this system映射到本系统的公司

Manufacture mapped to this system映射到本系统的制造

What the primary evidence shows

None yet, and that absence is worth stating rather than hiding. Unlike most of this taxonomy’s other systems, no live-web research pass has been run on P16, Safety validation & certification directly — the research behind this system comes entirely from a sell-side equity-research triage batch, not from primary company or filing research.

That is itself informative: it means the paragraph below is this system’s whole evidentiary record today, not a summary of a deeper primary layer sitting behind it.

What sell-side research adds

The honest starting point is that almost nobody in sell-side equity research covers the actual subject: no certification body, safety-validation-tooling vendor, or standards organisation for humanoid safety surfaced anywhere in the batch. What does exist is real, and newer than the taxonomy’s original framing.

Three public, reproducible real-robot evaluation venues have gone live since October 2025 — RoboChallenge (Dexmal with Hugging Face, 80,000-plus real-robot task executions across 20-plus models), RoboDojo (HKU Multimedia Lab with roughly 20 universities including Berkeley and Tsinghua), and Lightwheel’s RoboFinals — alongside RobOmni on the tactile side. Evaluation, in other words, now exists and is reproducible; what does not exist is a neutral evaluator. RoboChallenge is run by a VLA-tooling company, RoboFinals by a data vendor, RobOmni by a sensor vendor; only RoboDojo is academic.

The investable gap is therefore not “build a benchmark,” it is “become the venue a buyer trusts” — the role MLPerf, SPEC, and TPC played in compute, all of which were consortia, not vendors.

Downstream, the absence of that trusted venue is already generating measurable friction: J.P. Morgan’s insurance desk frames China’s nascent humanoid P&C insurance line (four insurers now write dedicated humanoid policies) as explicitly hard to price “given limited claims history” — an underwriting problem that is the mirror image of a certifier’s absence. Separately, Morgan Stanley’s coverage of the regulatory angle developing fastest is not a safety standard at all — it is trade policy, via the bipartisan GUARD Act, which would apply FCC Covered-List-style national-security review to Chinese-made humanoid and quadruped robots.

The investor read

P16, Safety validation & certification is a real gap with no obvious owner yet, which is a different risk profile from most of this taxonomy’s other systems. The GUARD Act is a real, checkable near-term catalyst, but it should not be mistaken for the engineering-validation layer the taxonomy is actually pricing — trade policy and safety certification are two different gates that happen to be developing on the same beat.

Until a neutral evaluation consortium exists, or a certification standard appears, treat this problem as an unbuilt opportunity, not a mispriced one.

And given that no primary research pass exists yet, treat this system’s evidence base itself as a to-do, not a settled record — the highest-value next step here is direct research (safety standards, certifying bodies, tooling vendors), not another sell-side triage batch.

一手证据显示什么

目前尚无一手证据,这一点值得如实说明,而非略过不提。与本分类法其他大多数系统不同,P16,安全验证与认证 尚未经历一手网络调研;支撑这一系统的研究完全来自一批卖方股票研究的筛选整理,而非一手公司或财报研究。

这一点本身就很说明问题:意味着下文这一段就是该系统今天全部的证据记录,而不是对其背后某个更深一层一手研究的摘要。

卖方研究补充了什么

最诚实的起点是:卖方股票研究几乎没有人真正覆盖这个主题——在这批研究中,没有发现任何面向人形机器人安全的认证机构、安全验证工具供应商或标准组织。确实存在的东西是真实的,也比本分类法最初的框架更新。

自 2025 年 10 月以来,已有三个公开、可复现的真机评测场上线——RoboChallenge(原力灵机 Dexmal 与 Hugging Face 合办,覆盖 20 多个模型、已完成 8 万次以上真机任务执行)、RoboDojo(香港大学多媒体实验室联合约 20 所高校,含伯克利与清华),以及光轮智能的 RoboFinals——此外还有触觉方向的 RobOmni。换句话说,评测如今已经存在,且可复现;不存在的是中立的评测方。RoboChallenge 由一家 VLA 工具公司运营,RoboFinals 由数据供应商运营,RobOmni 由传感器供应商运营;只有 RoboDojo 属于学术机构。

因此可投资的缺口并非“搭建一个基准测试”,而是“成为买方信任的评测场”——这正是 MLPerf、SPEC 与 TPC 在算力领域曾扮演的角色,而它们都是联盟,而非厂商。

在下游,这一可信评测场的缺失已经在产生可度量的摩擦:摩根大通的保险研究团队指出,中国新兴的人形机器人财产与责任保险业务(现已有四家保险公司提供专属人形机器人保单)明确难以定价,“因为理赔历史有限”——这正是认证方缺位问题的镜像,只不过换成了承保端的困境。另外,摩根士丹利对目前发展最快的监管角度的覆盖,其实完全不是安全标准,而是贸易政策——两党支持的 GUARD Act 法案,若通过将对中国制造的人形与四足机器人施加类似 FCC“受管制清单”式的国家安全审查。

给投资者的启示

P16,安全验证与认证 是一个真实存在、但目前尚无明显归属者的缺口,这与本分类法其他大多数系统的风险画像不同。GUARD Act 是一个真实、可核查的近期催化因素,但不应被误认为本分类法真正试图定价的那个工程验证层——贸易政策与安全认证是两道恰好在同一赛道上同时演进的不同关卡。

在一个中立的评测联盟出现、或一套认证标准问世之前,应把这一问题视为一个尚未被建立的机会,而非一个被错误定价的机会。

既然目前尚无一手研究,也应把这一系统自身的证据基础当作一项待办事项,而非已经定型的记录——眼下价值最高的下一步,是直接研究(安全标准、认证机构、工具供应商),而不是再做一轮卖方筛选。

Contents目录

immune/
└── safety-validation-certification Safety validation & certification安全验证与认证