Robotics Bottleneck Research机器人瓶颈研究

Main主页 / thesis / blood

system

Blood血液

CE-NCE-N 54.6

Robots learn from data and simulation before they ever meet the real world. Whether that training survives contact with reality decides if a robot-learning company actually has a product.

数据、仿真与 sim-to-real:机器人公司体内真正循环之物。

Bottlenecks瓶颈 P06 Robot data collection机器人数据采集 · P08 Simulation & synthetic data仿真与合成数据 · P09 Sim-to-real transferSim-to-real 迁移

What it isBlood is the data layer: everything a robot-learning company needs to train its policy (the model that decides what the robot actually does) before it ever touches a paying customer. That's real-world demonstration collection (P06, Robot data collection), physics simulation and synthetic data (P08, Simulation & synthetic data), and the sim-to-real transfer that decides whether any of it survives contact with a real deployment (P09, Sim-to-real transfer).

是什么“血液”(Blood)是数据层:一家机器人学习公司在触达付费客户之前所需要的一切——真实世界示教数据采集(P06,机器人数据采集)、物理仿真与合成数据(P08,仿真与合成数据),以及决定这一切能否在真实部署中存活下来的 sim-to-real 迁移(P09,Sim-to-real 迁移)。

ThesisOn CE-N, Blood's three bottlenecks (real-world data collection, simulation and synthetic data, and sim-to-real transfer) rank 4th, 7th, and 8th of 18, comfortably in the stronger half of the board. That ranking rests on the taxonomy's original bet that data is a picks-and-shovels business (a supplier that profits no matter which specific robot company wins) with whoever owns the data owning the toll booth every model has to pass through. That toll booth has cracked: free data from AgiBot, Spirit AI, NVIDIA, and state-subsidized Chinese collection facilities has commoditized raw trajectories faster than any vendor can monetize them. The value that actually moved went to deployment, not data: Skild buying Fetch's deployed fleet is the clearest statement that data access is bought, not trained, and the same pattern holds at Physical Intelligence and in Google's Gemini Robotics rollout onto partner fleets. The exposure that survives sits one layer down from pure data plays: task-specific, failure-inclusive data tied to a deployment someone already pays for, which favors deployment-owning names over standalone data or simulation vendors.

论点在 CE-N 上,真实世界数据采集、仿真与合成数据,以及 sim-to-real 迁移分别位列 18 个瓶颈中的第 4、第 7、第 8 位——这正是本分类法最初的判断:数据是一种卖铲人式敞口,谁掌握数据,谁就掌握了每个模型都必须经过的收费站。但这座收费站已经出现裂缝:AgiBot、Spirit AI、NVIDIA 提供的免费数据,以及中国各地国家补贴的数据采集设施,让原始示教数据的商品化速度快过了任何厂商将其变现的能力。真正发生价值转移的方向是部署,而不是数据——Skild 收购 Fetch 已部署车队的举动,最清楚地说明了数据获取是买来的,而不是训练出来的;Physical Intelligence 以及 Google 将 Gemini Robotics 接入合作伙伴车队的路径,呈现出同样的模式。真正留存下来的敞口比纯数据业务更下沉一层:是与客户已经在付费的部署场景绑定的、包含失败案例的任务专属数据——这更利好掌握部署环节的公司,而非独立的数据或仿真厂商。

1–5 scale. Budget, solvability and value capture (outlined) enter CE-N; the other three remain context. 1–5 分制。预算、可解性与价值捕获(描边)进入 CE-N,其余三项仅作背景。

P06 Robot data collection机器人数据采集

Maturity成熟度 3.0
Pain痛感 4.9
Budget预算 4.6
Solvability可解性 3.7
Value capture价值捕获 4.5
Timing时机 4.9

P08 Simulation & synthetic data仿真与合成数据

Maturity成熟度 3.0
Pain痛感 4.6
Budget预算 4.3
Solvability可解性 4.0
Value capture价值捕获 4.4
Timing时机 4.7

P09 Sim-to-real transferSim-to-real 迁移

Maturity成熟度 2.0
Pain痛感 4.8
Budget预算 4.5
Solvability可解性 3.1
Value capture价值捕获 4.6
Timing时机 4.9

Products mapped to this system映射到本系统的产品

Real-world robot / egocentric dataReal-world robot / egocentric data

Companies that capture, curate, or sell real-world robot and egocentric demonstration data -- teleoperation and fleet logs, tactile/visuotactile capture, human-video and skill-capture collection, and multimodal annotation -- used to train robot-learning models.

19 companies 19 家公司 +12 named in conference materials, unverified +12 项来自会议材料,未经核实

Simulation / sim-to-realSimulation / sim-to-real

Companies building physics-simulation platforms, synthetic-data generation, and sim-to-real transfer infrastructure used to train and evaluate robot policies alongside or instead of real-world data.

6 companies 6 家公司 +1 named in conference materials, unverified +1 项来自会议材料,未经核实

Companies mapped to this system映射到本系统的公司

What the primary evidence shows

That thesis took real damage in a single nine-month window. AgiBot open-sourced more than one million real-world manipulation trajectories — roughly 2,400 hours across 200 task types, with tactile signals, LiDAR, and full-body joint states, including labelled failure cases — for free. Spirit AI open-sourced its top-ranked RoboChallenge policy within 24 hours of taking the leaderboard. NVIDIA gives away Isaac, Cosmos, and the GR00T weights. And China alone had 64 state-subsidised data-collection facilities planned or under construction across at least 27 cities as of early 2026 — a public-sector cost floor near zero for the same labour-arbitrage teleoperation work a private data vendor is trying to sell. Lightwheel, the system’s most prominent data company, discloses a data resale ratio “above 10x” — the same non-exclusive asset sold to ten-plus competing buyers, which is a strong gross-margin story and a weak moat story, closer to a stock-photo library than a flywheel.

The value that did move in this period moved to deployment, not to data. Skild AI’s clearest strategic move was not a dataset — it was buying Zebra Technologies’ Fetch Robotics division to acquire a deployed fleet, then signing channel agreements with ABB Robotics and Universal Robots and landing onto Foxconn’s production lines. Physical Intelligence’s gains came from co-training on partner deployment data, not from a bigger open corpus. Google put a public Gemini Robotics API directly into Apptronik’s Apollo 2, Franka’s Duo, and a Boston Dynamics Atlas fleet. Every one of these is the same argument: the scarce asset is access to a robot that is already doing paid work, which the deploying OEM owns — not a data vendor selling the same trajectories to every buyer. Yet free data alone has not solved the underlying problem either: two independent 2026 benchmarks (RoboDojo, RoboChallenge) put frontier generalist policies at roughly 12% and roughly 50% real-robot success against a human baseline near 100%, so raw demonstration volume was evidently never the binding constraint.

What sell-side research adds

A batch of 16 sell-side reports (Deutsche Bank, Goldman Sachs, Barclays, UBS, CB Insights and others) independently converges on the same starting premise as the primary evidence: data, not model architecture or compute, is the binding constraint. Barclays lists a “training-data gap” as one of five gating factors on humanoid scaling; CB Insights is blunter still — “the real bottleneck is training data, not model capability.” UBS ties data quality and monetisation directly to demand for robotic data-collection centres, an independent second corroboration of the 64-subsidised-centre finding above.

But the sell-side’s own evidence for how leaders are closing the gap cuts against a pure data-vendor thesis, and corroborates the deployment argument from the opposite direction: Apptronik’s “Robot Park,” a 90,000-square-foot data factory, trains Apollo 3 on real Apollo 2 field data; Morgan Stanley frames Tesla’s Optimus programme identically, as a flywheel harvesting task and vision data from Tesla’s own manufacturing footprint and vehicle fleet; Goldman’s China-advantage thesis resolves the same way — 10,000-15,000 actual 2025 deployments generating data at the point of use, not a data marketplace.

Not one of sixteen reports identifies a scaled company monetising data as an arm’s-length product line with disclosed revenue. The two purpose-built synthetic-data players that do surface, Archetype AI and Bifrost, are both early-stage and privately scouted, not covered as investable names.

The investor read

The house view’s conclusion — data infrastructure is picks-and-shovels exposure worth owning across every OEM — survives, and the sell-side batch corroborates it a third independent time (after the free-data glut and the 64 subsidised centres) rather than contradicting it. But its mechanism needs rewriting.

What is scarce is not data collection in general; it is task-specific, failure-inclusive data tied to a deployment a customer already pays for, which sits one layer down from where a pure data or simulation vendor operates. That points back toward the deployment-owning names this taxonomy already prefers — Dexterity, Ambi Robotics, Deep Robotics, Symbotic — rather than toward standalone data or synthetic-data companies.

A pure-play data vendor, on the evidence of this period, is a weaker business than the original picks-and-shovels framing assumed.

一手证据显示什么

这一论点在短短九个月的窗口期内受到了实质性冲击。智元机器人(AgiBot)免费开源了超过一百万条真实世界操作轨迹——约 2400 小时、覆盖 200 种任务类型,包含触觉信号、激光雷达与全身关节状态,且含有标注的失败案例。千寻智能(Spirit AI)在夺得 RoboChallenge 排行榜榜首后 24 小时内即开源了夺冠模型。英伟达免费开放 Isaac、Cosmos 与 GR00T 权重。而截至 2026 年初,仅中国就有 64 个国家补贴的数据采集设施已规划或在建,分布在至少 27 个城市——为同样的遥操作劳动力套利业务提供了近乎为零的公共部门成本底线,而这正是私营数据供应商试图售卖的东西。光轮智能(Lightwheel),该系统内最知名的数据公司,自己披露的数据转售倍数“超过 10 倍”——即同一份非独占资产被卖给十个以上相互竞争的买家,这是一个很好的毛利率故事,却是一个很差的护城河故事,更接近图库网站,而非数据飞轮。

这一时期真正发生价值转移的地方,是部署端,而非数据端。Skild AI 最清晰的战略举动不是一个数据集,而是收购了 Zebra Technologies 旗下的 Fetch Robotics 部门,以获取一支已部署的机器人车队,随后又与 ABB 机器人和 Universal Robots 签订渠道协议,并进入富士康的产线。Physical Intelligence 的进步来自与合作伙伴部署数据的联合训练,而非更大规模的开放语料库。谷歌把公开的 Gemini Robotics API 直接接入了 Apptronik 的 Apollo 2、Franka 的 Duo,以及 Boston Dynamics 的一支 Atlas 车队。这些案例指向同一个论点:稀缺资产是接触一台已经在从事有偿工作的机器人,而这由部署方 OEM 所掌握——而不是一个把同样轨迹数据卖给每一个买家的数据供应商。然而,单靠免费数据本身也没有解决底层问题:两项独立的 2026 年基准测试(RoboDojo、RoboChallenge)显示,前沿通用策略的真实机器人成功率分别约为 12% 与约 50%,而人类基线接近 100%,可见原始示教数据的体量从来都不是那个真正的约束条件。

卖方研究补充了什么

一批 16 份卖方研究报告(德意志银行、高盛、巴克莱、瑞银、CB Insights 等)独立地得出了与一手证据相同的起点判断:真正的约束是数据,而非模型架构或算力。巴克莱把“训练数据缺口”列为人形机器人放量的五大门槛因素之一;CB Insights 说得更直接——“真正的瓶颈是训练数据,而非模型能力”。瑞银把数据质量与变现直接与机器人数据采集中心的需求挂钩,这是对上文“64 个国家补贴中心”这一发现的第二次独立印证。

但卖方研究自身关于领先者“如何”弥合这一缺口的证据,恰恰不利于纯数据供应商的叙事,反而从另一个角度印证了部署端论点:Apptronik 的“机器人园区”(Robot Park)——一座 9 万平方英尺的数据工厂——用 Apollo 2 的真实现场数据训练 Apollo 3;摩根士丹利对特斯拉 Optimus 项目的描述如出一辙:一个从特斯拉自有制造版图与车辆车队中收获任务与视觉数据的飞轮;高盛的中国优势论点也是同样的解法——2025 年 1 万至 1.5 万台的实际部署在使用端产生数据,而非一个数据交易市场。

十六份报告中,没有一份指出存在一家把数据作为独立产品线、按公平交易方式变现、且披露了收入的规模化公司。 出现的两家专注合成数据的公司——Archetype AI 与 Bifrost——都还处于早期阶段,属于私下摸底覆盖,并未被当作可投资标的报道。

给投资者的启示

本表原有的结论——数据基础设施是值得在每家 OEM 敞口中持有的卖铲人生意——依然成立,卖方研究批次是第三次独立印证(继免费数据泛滥与 64 个补贴中心之后),而非反证。但其作用机制需要重写。

真正稀缺的并非泛泛的数据采集,而是与客户已经付费的某个部署绑定、且包含失败案例的任务特定数据,这处于比纯数据或仿真供应商更深一层的位置。这把焦点重新指向本分类法本就偏好的、掌握部署的公司——Dexterity、Ambi Robotics、Deep Robotics、Symbotic——而非独立的数据或合成数据公司。

就这一时期的证据而言,一家纯数据供应商,其生意的质地比最初的卖铲人叙事所假设的要弱。

Contents目录

blood/
├── robot-data-collection Robot data collection机器人数据采集
├── simulation-synthetic-data Simulation & synthetic data仿真与合成数据
└── sim-to-real-transfer Sim-to-real transferSim-to-real 迁移