Robotics Bottleneck Research机器人瓶颈研究

Main主页 / roboindex / methodology

RBI methodologyRBI methodology

What the index measures, how it differs from stock valuation, how it's built, and what it doesn't do yet.

  • RESEARCH CUT 2026-08-13研究截点 2026-08-13

Robotics Bottleneck Index (RBI): Methodology

Version 0.2 · Base date 2026-08-13 · Base level 1000 · Rebalance quarterly Research cut 2026-08-12 (heatmap, memo) / 2026-08-13 (landscape) Research use only. Not investment advice. Not a recommendation to buy or sell anything.

This page describes the current calculation. Earlier methodology versions and score-change reviews are retained in the GitHub repository. The website shows the current approved CE-N score only.


1. Why this index exists

Eight robotics indices and ETFs already exist. Every one of them organises by label (is this a robotics company?) and weights by market capitalisation or by committee. That design has a documented failure mode: iShares’ IRBO became ARTY in 2024 and is now an AI-semiconductor fund. A theme index with no falsifiable claim underneath it drifts to whatever is working.

RBI organises by bottleneck instead. The eighteen problems in the underlying heatmap (P01–P18) are grouped into nine body systems, and each system’s target weight is set by how much capital efficiency that bottleneck offers, not by how large its incumbents happen to be.

The index makes one testable claim: capital compounds fastest at the bottlenecks where a budget already exists, the problem is tractable today, and the solver keeps the value. If that is wrong, this index will underperform a cap-weighted robotics basket, and the rebalance log will show exactly when and where the claim broke.

2. What the index actually measures

This is the question worth answering precisely, because the name invites a false comparison. Standard finance “capital efficiency” (ROIC, the ROIC–WACC spread, free cash flow over invested capital) is a company-level accounting metric, computed from a balance sheet and an income statement. CE-N is not that. It is a bespoke problem-level prioritisation score: of the eighteen named engineering/economic bottlenecks a robot has to clear, which one offers the most capital-efficient exposure, not which stock, at what price, sized how.

CE-N is built from techniques borrowed from named, real methodologies: MSCI’s winsorise-then-standardise pattern, the Weighted Sum vs. Weighted Product distinction from multi-criteria decision analysis, and ELECTRE/PROMETHEE-style veto floors (§6 below), but the specific combination, the three chosen axes, and the parameters are this desk’s own construction, stated plainly so they can be argued with. It is closer to a venture-style thesis screen than to a benchmarked financial factor.

Three inputs drive the score, each scored 1–5 by an analyst against a stated definition:

  • Budget: whether a line item exists that someone is already authorised to spend against this problem.
  • Solvability: whether the problem is tractable now, with today’s methods, not a research-agenda bet on tomorrow’s.
  • Value capture: whether whoever solves the problem keeps the value, or hands it to a customer, a platform, or a standard that commoditises it away.

Pain, Timing, and Maturity remain visible as context and do not enter the current calculation. Pain ranges from 4.0 to 5.0 and Timing from 4.1 to 5.0 across the eighteen problems. These narrow ranges do not establish that the axes contain no useful information. Maturity describes progress already achieved; whether it adds information beyond Solvability remains a review question. The CE-N Framework explains the stored calculation, and the scoring review sets out evidence requirements and rating anchors.

3. How this differs from a stock valuation

CE-N and a valuation are two different layers of this site’s data, and conflating them is the single easiest way to misread this index:

  • CE-N asks “which problem.” It never touches a share price, a discount rate, or a company’s financial statements. A problem can score high on CE-N while every public company exposed to it is expensive, illiquid, or doesn’t exist yet. That gap is exactly what §8’s Thesis Gap measures.
  • A discounted-cash-flow model asks “which company, at what price.” That is a separate, company-level layer: every public company carries a valuation_tier (1 = its disclosed statements support a DCF; 2 = real revenue but earnings quality is too volatile for a DCF to mean anything, recast recurring earnings and track a multiple instead; 3 = no disclosed cash-flow base at all, track the disclosure-vs-narrative gap, never a DCF). The eight tier-1 names each carry a sourced dcf_assumptions block and a computed dcf_output (intrinsic value per share, margin of safety). See each company’s page under companies, refreshed quarterly. A company can be tier-1 as a business while a specific line of its revenue still must not be separately valued: CATL is tier-1 (its battery business has real, disclosed cash flow), and this site’s own research still says never attribute a value to CATL’s robotics exposure specifically, because no arm’s-length segment revenue exists for it. Both facts are true at once.

Put simply: CE-N tells you where to look. A DCF, where one is buildable, tells you what a specific company is worth. This index’s weights are driven entirely by the first; the second exists as a separate, per-company research layer this site also publishes.

4. How the index works

Universe. Listed equities in the research file. A company enters the universe only with a completed record: problem tags, primary source, and a last-verified date.

Within-system weighting: equal, not cap-weighted. Cap-weighting inside a thematic sleeve reproduces exactly the failure this index exists to avoid: it would hand the Brain system to NVIDIA and turn the index into a semiconductor proxy. Equal weight forces the bottleneck framing to survive contact with the actual portfolio.

Caps. No single name above 10% of the index. No single system above 25%, regardless of CE-N score. Breaches are resolved by capping and reporting the residual as Thesis Gap, never by substituting in a different name.

The Thesis Gap. Target weights describe where value should accrue on the bottleneck argument. They say nothing about whether a public investor can actually buy that exposure, and on this board, largely, they cannot. Wherever a system’s target weight can’t be absorbed by eligible constituents (no eligible name exists, or every eligible name has hit its cap), the shortfall is reported as Thesis Gap, published as the index’s headline number every quarter. At the current cut it is 25.67%, roughly a quarter of what the bottleneck argument says should be held cannot be bought on public markets today.

Two variants are published side by side, every quarter:

  • RBI-system-ce-T (Thesis): target weights, gap held as cash. Measures the thesis itself.
  • RBI-system-ce-I (Investable): the gap redistributed pro rata across systems with eligible constituents. Measures what an investor could actually own today.

The spread between them is itself a published series, and it is usually the more honest number than either variant alone.

5. Current system weights

Weights are proportional to each system’s CE-N score (§6), the mean of its member problems’ composite score, computed fresh every rebalance, never hand-typed.

SystemProblemsCE-NTarget weight
WorkP1871.215.38%
IntegrationP1760.713.12%
MuscleP13, P1454.111.69%
BloodP06, P08, P0954.611.80%
BrainP04, P10, P11, P1252.311.30%
ImmuneP1646.510.05%
HeartP1547.610.29%
SensesP01, P02, P0341.58.97%
NervesP05, P0734.37.41%

Note the compression: first to last spans only 15.38% to 7.41%. The bottleneck framing is doing most of its work in the exclusions and the Thesis Gap, not in the weights themselves. Read §7 and §4 before over-reading this table on its own.

6. The scoring method

The stored CE-N scores use Budget, Solvability, and Value capture. The calculation has five steps:

  1. Cap the tails. Cap each axis at its 10th and 90th percentiles across the eighteen problems, using linear interpolation between adjacent observations.
  2. Standardise. Subtract the capped axis mean and divide by its population standard deviation. A constant axis contributes zero. This expresses relative position within the current board.
  3. Combine. Average the three standardised inputs with equal weights: composite_z = (z_budget + z_solvability + z_value_capture) / 3.
  4. Rescale. Map to 0–100 using (composite_z + 3) / 6 × 100, clipped to that range.
  5. Apply the floor rule and round. If any of the three raw inputs is below 2.5, multiply the rescaled score by 0.85. Round half up to one decimal place.

The read-only October 2026 audit reproduced all eighteen stored ce_score_v2 values. Fixed endpoints do not make scores directly comparable across research cuts: the means, spreads, and percentile thresholds depend on the selected problems and their ratings.

These transforms do not remove subjective sensitivity. A half-point change to one input moved its problem by up to eight ranks in the diagnostic tests. See the repository review records for diagnostic details. Before a future rebalance, the approved calculation and the candidate stored outputs must be checked again.

6a. Evidence confidence

Evidence quality is assessed separately from the displayed CE-N score. Sources should be checked for independence, relevance, and date; repeated claims do not become independent evidence. Internal confidence-adjustment calculations remain documented in the repository. They are not shown as a second CE-N score and this display change does not alter portfolio calculations.

6b. Two companion signals, deliberately narrow in scope

Two further signals sit alongside CE-N without feeding into it, aimed at the two most subjective axes:

  • value_capture_density_tier (Low/Moderate/High, all eighteen problems) counts how many companies in the book are mapped to each problem and splits the eighteen into three equal groups by that count. A High reading is a caution, not a bonus: it means many funded competitors are chasing the same bottleneck, the exact commoditisation pattern this site’s tactile-sensing and free-data-glut research documents.
  • solvability_benchmark_pct is populated only where a research brief discloses a real, reproducible real-robot benchmark result: currently P11, Generalisable VLA policies only (RoboChallenge real-robot success ≈50%, RoboDojo’s more conservative real-world figure ≈12%, against a human baseline near 100%). Left null everywhere else, deliberately: this is not a general pipeline, and nothing is estimated to fill the field.

Both are advisory context next to value_capture and solvability, never an input to them.

7. Eligibility gates: the exclusions

A name in the universe is excluded if it fails any gate. Exclusions are published with reasoning. Refusing to hold something is a stronger signal than holding it, and the exclusion list is the most defensible part of a published methodology.

GateRuleCurrent exclusions
TradeabilityNo verifiable secondary close, or valuation confidence < 50AeroVironment, Inovance, Unitree
Earnings qualityRecurring earnings must be separable from non-operating items before sizingNot excluded: capped at half weight (Hesai, Horizon Robotics, ESTUN)
ReconciliationShare count, net cash, or debt not reconciledOuster
ValuationPrice already discounts a penetration that has not occurredLeaderdrive: ~470× P/E, ~100× EV/Sales
Governance / controlMaterial internal-control finding, single-customer dependence, or an unresolved related-party structureSymbotic: ICFR material weakness, >84% single-customer revenue, Up-C/GreenBox VIE structure
PurityRobotics undisclosed and immaterial against the reporting segmentHeld at a 2% purity cap, never quoted as a robotics comp: NVIDIA, Rockwell Automation, Tesla, Hikvision, CATL

A gate failure is not a judgement about the underlying business. Intuitive Surgical is the strongest proof on the board that the endgame exists (consumables $1.73bn against systems $685m across 11,710 installed units), and it is still a valuation_tier: 1 name whose current DCF output reads as expensive at today’s price. Passing every gate and being fairly priced are two different questions.

8. Rebalance

Quarterly, on the research cut following each quarter end. Each rebalance publishes:

  1. Every score change, with the evidence that moved it.
  2. Additions and deletions, with the gate that triggered each.
  3. The Thesis Gap, by system, versus the prior quarter.
  4. A scored self-assessment of the prior quarter’s calls, including the misses.

Point 4 is not decoration. An index graded against itself in public, on a fixed schedule, is the cheapest available way to build a timestamped, falsifiable track record before there is a fund, which is the entire strategic purpose of publishing RBI for free. The eight tier-1 companies’ DCF assumptions and outputs (§3) refresh on the same quarterly cadence.

9. What this index does not do

  • It does not include private companies as index constituents. Twenty of twenty-four private names have no live terms; marking them would be fiction. They appear on the site as a Sleeve B tracking roster by system (diligence score, valuation confidence, funding-mark staleness, related-party flags), a read-only reference, explicitly not a tradable index, since there is no committed-dollar amount or daily price discovery to weight against.
  • It does not include the 38 factory and supply-chain nodes. Those are an information asset (component order books lead OEM revenue by several quarters), not a portfolio.
  • It does not adjust for currency, and it does not model transaction costs, borrow, or tax.
  • It is not a fund, does not accept capital, and has no live track record. Base date is 2026-08-13; everything before that date is history, not performance.

10. Known gaps, stated plainly

A methodology that hides its own weak points is marketing.

  1. Small N cuts both ways. Four of nine systems map to a single problem, so their weight is only as stable as one score. The 10th/90th winsorisation band is itself a concession to N=18 being small.
  2. Solvability and value capture are still judgements. The scoring math (§6) bounds how much leverage that subjectivity has over the final number; it does not remove the subjectivity. §6b’s companion signals are a first, deliberately narrow step toward a computed check on both , not a replacement for the underlying scores.
  3. The floor threshold (2.5) and the equal-thirds weighting are stated choices, not derived ones. Nothing in the data forces either; both are published exactly so they can be argued with.
  4. This desk has not re-scored any problem’s raw inputs based on the most recent sell-side research pass. New evidence is catalogued by system in the research briefs and flagged for review, never applied silently.
  5. The eight-name DCF layer is new (2026-08-26) and has known open items per company: see each company’s dcf_assumptions.notes for genuine data-source discrepancies (e.g. CATL’s beta ranging 0.37–0.86 across providers, AeroVironment’s unreconciled debt figure). One name (AeroVironment) currently produces a non-meaningful output because its free cash flow is negative post-acquisition, flagged, not hidden.
  6. A conservative-terminal-growth DCF will read almost every premium-multiple growth name as “overvalued.” That is a documented, deliberate property of the model (see each company’s dcf_output), not a bug: it isolates how much of a price is justified by disclosed near-term cash flow versus priced-in narrative or optionality the model doesn’t capture. Read a large negative margin of safety as a measurement, not a sell signal.

Falsification condition for CE-N specifically, stated in advance: if a future rebalance shows CE-N systematically reordering problems in a direction that tracks which analyst scored them rather than tracking new evidence, the standardisation step is amplifying noise rather than removing it, and the weighting scheme should be revisited first.


Sources: Robotics Industry Heatmap & Investment Intelligence System v0.1; Robotics IC Memo Book; Robotics Landscape 95, research cut 2026-08-13. Methodology references: MSCI Quality Indices Methodology (winsorisation + z-score standardisation); Weighted Sum Model vs. Weighted Product Model, multi-criteria decision analysis literature; outranking-method veto thresholds (ELECTRE/PROMETHEE family) as precedent for the conjunctive floor. Facts, derived calculations, company claims and analyst inferences are kept in separate fields throughout the underlying data. Unknowns are left unknown.

给地图打分

当前 CE-N 采用预算、可解性和价值捕获三个维度。痛点、时机与成熟度保留作背景。 痛点在十八个问题上的范围为 4.0–5.0,时机为 4.1–5.0。范围较窄并不能证明这些维度没有信息; 成熟度是否提供了可解性以外的信息,也需要审查。

计算依次将三个输入限制在各自的第 10 和第 90 百分位之间(线性插值),减去截尾后的均值并除以 总体标准差,然后等权平均。常数维度贡献零。通过 (composite_z + 3) / 6 × 100 映射并限制到 0–100;若任一原始输入低于 2.5,结果乘以 0.85,最后四舍五入至一位小数。

2026 年 10 月的只读审计复现了全部十八个已存储的 CE-N 分数。历史版本保存在 GitHub,网站仅展示当前已批准的 CE-N 分数。固定端点不代表不同研究截点的分数可直接比较: 均值、离散程度和百分位阈值均取决于当期问题集合与评分。

敏感性测试中,单一输入变化 0.5 分,相关问题最多移动八个名次。完整计算见 CE-N 框架,评分证据要求见评分审查。 下一次再平衡前,应重新核对经批准的计算方式与候选输出。

这套方法没有覆盖什么

一个只写方法、不写盲区的研究站是营销;写清楚盲区,才能被人反驳。

  • 分数是基于热力图自身输入的分析性估计。它们不是事实,也不来自市场。
  • 覆盖范围是某个研究截点上的快照,不是实时普查。数字未知就保持未知,不用行业平均数填。
  • 事实、公司口径与分析推断始终分层。公司关于自己的说法,不会因为被反复引用就升格为事实。

构建 RoboIndex

以下是 RoboIndex 的构建规则,和给地图打分的规则分开放。打分回答「这个问题有多好」; 构建回答「什么可以持有、以多大权重持有、什么时候调整」。

合格标准

成分股必须是有可追溯代码的上市证券,机器人业务的敞口也要大到配得上它所代表的那个瓶颈。 未经核验的实体一律排除;待定条目进不了指数,也不会悄悄混进去。

那个缺口

如果某个身体系统的配置额度找不到合格的上市标的来承接,这部分权重就以现金留着, 永远不摊派。这个缺口是一项发现,不是四舍五入的误差; 一个由「大半在非上市公司手里」的论点推导出来的指数,藏起这个缺口就失去了发布的意义。