What the primary evidence shows
None yet, and that absence is worth stating rather than hiding. Unlike most of this taxonomy’s other systems, no live-web research pass has been run on P16, Safety validation & certification directly — the research behind this system comes entirely from a sell-side equity-research triage batch, not from primary company or filing research.
That is itself informative: it means the paragraph below is this system’s whole evidentiary record today, not a summary of a deeper primary layer sitting behind it.
What sell-side research adds
The honest starting point is that almost nobody in sell-side equity research covers the actual subject: no certification body, safety-validation-tooling vendor, or standards organisation for humanoid safety surfaced anywhere in the batch. What does exist is real, and newer than the taxonomy’s original framing.
Three public, reproducible real-robot evaluation venues have gone live since October 2025 — RoboChallenge (Dexmal with Hugging Face, 80,000-plus real-robot task executions across 20-plus models), RoboDojo (HKU Multimedia Lab with roughly 20 universities including Berkeley and Tsinghua), and Lightwheel’s RoboFinals — alongside RobOmni on the tactile side. Evaluation, in other words, now exists and is reproducible; what does not exist is a neutral evaluator. RoboChallenge is run by a VLA-tooling company, RoboFinals by a data vendor, RobOmni by a sensor vendor; only RoboDojo is academic.
The investable gap is therefore not “build a benchmark,” it is “become the venue a buyer trusts” — the role MLPerf, SPEC, and TPC played in compute, all of which were consortia, not vendors.
Downstream, the absence of that trusted venue is already generating measurable friction: J.P. Morgan’s insurance desk frames China’s nascent humanoid P&C insurance line (four insurers now write dedicated humanoid policies) as explicitly hard to price “given limited claims history” — an underwriting problem that is the mirror image of a certifier’s absence. Separately, Morgan Stanley’s coverage of the regulatory angle developing fastest is not a safety standard at all — it is trade policy, via the bipartisan GUARD Act, which would apply FCC Covered-List-style national-security review to Chinese-made humanoid and quadruped robots.
The investor read
P16, Safety validation & certification is a real gap with no obvious owner yet, which is a different risk profile from most of this taxonomy’s other systems. The GUARD Act is a real, checkable near-term catalyst, but it should not be mistaken for the engineering-validation layer the taxonomy is actually pricing — trade policy and safety certification are two different gates that happen to be developing on the same beat.
Until a neutral evaluation consortium exists, or a certification standard appears, treat this problem as an unbuilt opportunity, not a mispriced one.
And given that no primary research pass exists yet, treat this system’s evidence base itself as a to-do, not a settled record — the highest-value next step here is direct research (safety standards, certifying bodies, tooling vendors), not another sell-side triage batch.
一手证据显示什么
目前尚无一手证据,这一点值得如实说明,而非略过不提。与本分类法其他大多数系统不同,P16,安全验证与认证 尚未经历一手网络调研;支撑这一系统的研究完全来自一批卖方股票研究的筛选整理,而非一手公司或财报研究。
这一点本身就很说明问题:意味着下文这一段就是该系统今天全部的证据记录,而不是对其背后某个更深一层一手研究的摘要。
卖方研究补充了什么
最诚实的起点是:卖方股票研究几乎没有人真正覆盖这个主题——在这批研究中,没有发现任何面向人形机器人安全的认证机构、安全验证工具供应商或标准组织。确实存在的东西是真实的,也比本分类法最初的框架更新。
自 2025 年 10 月以来,已有三个公开、可复现的真机评测场上线——RoboChallenge(原力灵机 Dexmal 与 Hugging Face 合办,覆盖 20 多个模型、已完成 8 万次以上真机任务执行)、RoboDojo(香港大学多媒体实验室联合约 20 所高校,含伯克利与清华),以及光轮智能的 RoboFinals——此外还有触觉方向的 RobOmni。换句话说,评测如今已经存在,且可复现;不存在的是中立的评测方。RoboChallenge 由一家 VLA 工具公司运营,RoboFinals 由数据供应商运营,RobOmni 由传感器供应商运营;只有 RoboDojo 属于学术机构。
因此可投资的缺口并非“搭建一个基准测试”,而是“成为买方信任的评测场”——这正是 MLPerf、SPEC 与 TPC 在算力领域曾扮演的角色,而它们都是联盟,而非厂商。
在下游,这一可信评测场的缺失已经在产生可度量的摩擦:摩根大通的保险研究团队指出,中国新兴的人形机器人财产与责任保险业务(现已有四家保险公司提供专属人形机器人保单)明确难以定价,“因为理赔历史有限”——这正是认证方缺位问题的镜像,只不过换成了承保端的困境。另外,摩根士丹利对目前发展最快的监管角度的覆盖,其实完全不是安全标准,而是贸易政策——两党支持的 GUARD Act 法案,若通过将对中国制造的人形与四足机器人施加类似 FCC“受管制清单”式的国家安全审查。
给投资者的启示
P16,安全验证与认证 是一个真实存在、但目前尚无明显归属者的缺口,这与本分类法其他大多数系统的风险画像不同。GUARD Act 是一个真实、可核查的近期催化因素,但不应被误认为本分类法真正试图定价的那个工程验证层——贸易政策与安全认证是两道恰好在同一赛道上同时演进的不同关卡。
在一个中立的评测联盟出现、或一套认证标准问世之前,应把这一问题视为一个尚未被建立的机会,而非一个被错误定价的机会。
既然目前尚无一手研究,也应把这一系统自身的证据基础当作一项待办事项,而非已经定型的记录——眼下价值最高的下一步,是直接研究(安全标准、认证机构、工具供应商),而不是再做一轮卖方筛选。