What the primary evidence shows
That thesis took real damage in a single nine-month window. AgiBot open-sourced more than one million real-world manipulation trajectories — roughly 2,400 hours across 200 task types, with tactile signals, LiDAR, and full-body joint states, including labelled failure cases — for free. Spirit AI open-sourced its top-ranked RoboChallenge policy within 24 hours of taking the leaderboard. NVIDIA gives away Isaac, Cosmos, and the GR00T weights. And China alone had 64 state-subsidised data-collection facilities planned or under construction across at least 27 cities as of early 2026 — a public-sector cost floor near zero for the same labour-arbitrage teleoperation work a private data vendor is trying to sell. Lightwheel, the system’s most prominent data company, discloses a data resale ratio “above 10x” — the same non-exclusive asset sold to ten-plus competing buyers, which is a strong gross-margin story and a weak moat story, closer to a stock-photo library than a flywheel.
The value that did move in this period moved to deployment, not to data. Skild AI’s clearest strategic move was not a dataset — it was buying Zebra Technologies’ Fetch Robotics division to acquire a deployed fleet, then signing channel agreements with ABB Robotics and Universal Robots and landing onto Foxconn’s production lines. Physical Intelligence’s gains came from co-training on partner deployment data, not from a bigger open corpus. Google put a public Gemini Robotics API directly into Apptronik’s Apollo 2, Franka’s Duo, and a Boston Dynamics Atlas fleet. Every one of these is the same argument: the scarce asset is access to a robot that is already doing paid work, which the deploying OEM owns — not a data vendor selling the same trajectories to every buyer. Yet free data alone has not solved the underlying problem either: two independent 2026 benchmarks (RoboDojo, RoboChallenge) put frontier generalist policies at roughly 12% and roughly 50% real-robot success against a human baseline near 100%, so raw demonstration volume was evidently never the binding constraint.
What sell-side research adds
A batch of 16 sell-side reports (Deutsche Bank, Goldman Sachs, Barclays, UBS, CB Insights and others) independently converges on the same starting premise as the primary evidence: data, not model architecture or compute, is the binding constraint. Barclays lists a “training-data gap” as one of five gating factors on humanoid scaling; CB Insights is blunter still — “the real bottleneck is training data, not model capability.” UBS ties data quality and monetisation directly to demand for robotic data-collection centres, an independent second corroboration of the 64-subsidised-centre finding above.
But the sell-side’s own evidence for how leaders are closing the gap cuts against a pure data-vendor thesis, and corroborates the deployment argument from the opposite direction: Apptronik’s “Robot Park,” a 90,000-square-foot data factory, trains Apollo 3 on real Apollo 2 field data; Morgan Stanley frames Tesla’s Optimus programme identically, as a flywheel harvesting task and vision data from Tesla’s own manufacturing footprint and vehicle fleet; Goldman’s China-advantage thesis resolves the same way — 10,000-15,000 actual 2025 deployments generating data at the point of use, not a data marketplace.
Not one of sixteen reports identifies a scaled company monetising data as an arm’s-length product line with disclosed revenue. The two purpose-built synthetic-data players that do surface, Archetype AI and Bifrost, are both early-stage and privately scouted, not covered as investable names.
The investor read
The house view’s conclusion — data infrastructure is picks-and-shovels exposure worth owning across every OEM — survives, and the sell-side batch corroborates it a third independent time (after the free-data glut and the 64 subsidised centres) rather than contradicting it. But its mechanism needs rewriting.
What is scarce is not data collection in general; it is task-specific, failure-inclusive data tied to a deployment a customer already pays for, which sits one layer down from where a pure data or simulation vendor operates. That points back toward the deployment-owning names this taxonomy already prefers — Dexterity, Ambi Robotics, Deep Robotics, Symbotic — rather than toward standalone data or synthetic-data companies.
A pure-play data vendor, on the evidence of this period, is a weaker business than the original picks-and-shovels framing assumed.
一手证据显示什么
这一论点在短短九个月的窗口期内受到了实质性冲击。智元机器人(AgiBot)免费开源了超过一百万条真实世界操作轨迹——约 2400 小时、覆盖 200 种任务类型,包含触觉信号、激光雷达与全身关节状态,且含有标注的失败案例。千寻智能(Spirit AI)在夺得 RoboChallenge 排行榜榜首后 24 小时内即开源了夺冠模型。英伟达免费开放 Isaac、Cosmos 与 GR00T 权重。而截至 2026 年初,仅中国就有 64 个国家补贴的数据采集设施已规划或在建,分布在至少 27 个城市——为同样的遥操作劳动力套利业务提供了近乎为零的公共部门成本底线,而这正是私营数据供应商试图售卖的东西。光轮智能(Lightwheel),该系统内最知名的数据公司,自己披露的数据转售倍数“超过 10 倍”——即同一份非独占资产被卖给十个以上相互竞争的买家,这是一个很好的毛利率故事,却是一个很差的护城河故事,更接近图库网站,而非数据飞轮。
这一时期真正发生价值转移的地方,是部署端,而非数据端。Skild AI 最清晰的战略举动不是一个数据集,而是收购了 Zebra Technologies 旗下的 Fetch Robotics 部门,以获取一支已部署的机器人车队,随后又与 ABB 机器人和 Universal Robots 签订渠道协议,并进入富士康的产线。Physical Intelligence 的进步来自与合作伙伴部署数据的联合训练,而非更大规模的开放语料库。谷歌把公开的 Gemini Robotics API 直接接入了 Apptronik 的 Apollo 2、Franka 的 Duo,以及 Boston Dynamics 的一支 Atlas 车队。这些案例指向同一个论点:稀缺资产是接触一台已经在从事有偿工作的机器人,而这由部署方 OEM 所掌握——而不是一个把同样轨迹数据卖给每一个买家的数据供应商。然而,单靠免费数据本身也没有解决底层问题:两项独立的 2026 年基准测试(RoboDojo、RoboChallenge)显示,前沿通用策略的真实机器人成功率分别约为 12% 与约 50%,而人类基线接近 100%,可见原始示教数据的体量从来都不是那个真正的约束条件。
卖方研究补充了什么
一批 16 份卖方研究报告(德意志银行、高盛、巴克莱、瑞银、CB Insights 等)独立地得出了与一手证据相同的起点判断:真正的约束是数据,而非模型架构或算力。巴克莱把“训练数据缺口”列为人形机器人放量的五大门槛因素之一;CB Insights 说得更直接——“真正的瓶颈是训练数据,而非模型能力”。瑞银把数据质量与变现直接与机器人数据采集中心的需求挂钩,这是对上文“64 个国家补贴中心”这一发现的第二次独立印证。
但卖方研究自身关于领先者“如何”弥合这一缺口的证据,恰恰不利于纯数据供应商的叙事,反而从另一个角度印证了部署端论点:Apptronik 的“机器人园区”(Robot Park)——一座 9 万平方英尺的数据工厂——用 Apollo 2 的真实现场数据训练 Apollo 3;摩根士丹利对特斯拉 Optimus 项目的描述如出一辙:一个从特斯拉自有制造版图与车辆车队中收获任务与视觉数据的飞轮;高盛的中国优势论点也是同样的解法——2025 年 1 万至 1.5 万台的实际部署在使用端产生数据,而非一个数据交易市场。
十六份报告中,没有一份指出存在一家把数据作为独立产品线、按公平交易方式变现、且披露了收入的规模化公司。 出现的两家专注合成数据的公司——Archetype AI 与 Bifrost——都还处于早期阶段,属于私下摸底覆盖,并未被当作可投资标的报道。
给投资者的启示
本表原有的结论——数据基础设施是值得在每家 OEM 敞口中持有的卖铲人生意——依然成立,卖方研究批次是第三次独立印证(继免费数据泛滥与 64 个补贴中心之后),而非反证。但其作用机制需要重写。
真正稀缺的并非泛泛的数据采集,而是与客户已经付费的某个部署绑定、且包含失败案例的任务特定数据,这处于比纯数据或仿真供应商更深一层的位置。这把焦点重新指向本分类法本就偏好的、掌握部署的公司——Dexterity、Ambi Robotics、Deep Robotics、Symbotic——而非独立的数据或合成数据公司。
就这一时期的证据而言,一家纯数据供应商,其生意的质地比最初的卖铲人叙事所假设的要弱。