Robotics Bottleneck Index (RBI): Methodology
Version 0.2 · Base date 2026-08-13 · Base level 1000 · Rebalance quarterly Research cut 2026-08-12 (heatmap, memo) / 2026-08-13 (landscape) Research use only. Not investment advice. Not a recommendation to buy or sell anything.
This page describes the current calculation. Earlier methodology versions and score-change reviews are retained in the GitHub repository. The website shows the current approved CE-N score only.
1. Why this index exists
Eight robotics indices and ETFs already exist. Every one of them organises by label (is this a robotics company?) and weights by market capitalisation or by committee. That design has a documented failure mode: iShares’ IRBO became ARTY in 2024 and is now an AI-semiconductor fund. A theme index with no falsifiable claim underneath it drifts to whatever is working.
RBI organises by bottleneck instead. The eighteen problems in the underlying heatmap (P01–P18) are grouped into nine body systems, and each system’s target weight is set by how much capital efficiency that bottleneck offers, not by how large its incumbents happen to be.
The index makes one testable claim: capital compounds fastest at the bottlenecks where a budget already exists, the problem is tractable today, and the solver keeps the value. If that is wrong, this index will underperform a cap-weighted robotics basket, and the rebalance log will show exactly when and where the claim broke.
2. What the index actually measures
This is the question worth answering precisely, because the name invites a false comparison. Standard finance “capital efficiency” (ROIC, the ROIC–WACC spread, free cash flow over invested capital) is a company-level accounting metric, computed from a balance sheet and an income statement. CE-N is not that. It is a bespoke problem-level prioritisation score: of the eighteen named engineering/economic bottlenecks a robot has to clear, which one offers the most capital-efficient exposure, not which stock, at what price, sized how.
CE-N is built from techniques borrowed from named, real methodologies: MSCI’s winsorise-then-standardise pattern, the Weighted Sum vs. Weighted Product distinction from multi-criteria decision analysis, and ELECTRE/PROMETHEE-style veto floors (§6 below), but the specific combination, the three chosen axes, and the parameters are this desk’s own construction, stated plainly so they can be argued with. It is closer to a venture-style thesis screen than to a benchmarked financial factor.
Three inputs drive the score, each scored 1–5 by an analyst against a stated definition:
- Budget: whether a line item exists that someone is already authorised to spend against this problem.
- Solvability: whether the problem is tractable now, with today’s methods, not a research-agenda bet on tomorrow’s.
- Value capture: whether whoever solves the problem keeps the value, or hands it to a customer, a platform, or a standard that commoditises it away.
Pain, Timing, and Maturity remain visible as context and do not enter the current calculation. Pain ranges from 4.0 to 5.0 and Timing from 4.1 to 5.0 across the eighteen problems. These narrow ranges do not establish that the axes contain no useful information. Maturity describes progress already achieved; whether it adds information beyond Solvability remains a review question. The CE-N Framework explains the stored calculation, and the scoring review sets out evidence requirements and rating anchors.
3. How this differs from a stock valuation
CE-N and a valuation are two different layers of this site’s data, and conflating them is the single easiest way to misread this index:
- CE-N asks “which problem.” It never touches a share price, a discount rate, or a company’s financial statements. A problem can score high on CE-N while every public company exposed to it is expensive, illiquid, or doesn’t exist yet. That gap is exactly what §8’s Thesis Gap measures.
- A discounted-cash-flow model asks “which company, at what price.” That is a separate,
company-level layer: every public company carries a
valuation_tier(1 = its disclosed statements support a DCF; 2 = real revenue but earnings quality is too volatile for a DCF to mean anything, recast recurring earnings and track a multiple instead; 3 = no disclosed cash-flow base at all, track the disclosure-vs-narrative gap, never a DCF). The eight tier-1 names each carry a sourceddcf_assumptionsblock and a computeddcf_output(intrinsic value per share, margin of safety). See each company’s page under companies, refreshed quarterly. A company can be tier-1 as a business while a specific line of its revenue still must not be separately valued: CATL is tier-1 (its battery business has real, disclosed cash flow), and this site’s own research still says never attribute a value to CATL’s robotics exposure specifically, because no arm’s-length segment revenue exists for it. Both facts are true at once.
Put simply: CE-N tells you where to look. A DCF, where one is buildable, tells you what a specific company is worth. This index’s weights are driven entirely by the first; the second exists as a separate, per-company research layer this site also publishes.
4. How the index works
Universe. Listed equities in the research file. A company enters the universe only with a completed record: problem tags, primary source, and a last-verified date.
Within-system weighting: equal, not cap-weighted. Cap-weighting inside a thematic sleeve reproduces exactly the failure this index exists to avoid: it would hand the Brain system to NVIDIA and turn the index into a semiconductor proxy. Equal weight forces the bottleneck framing to survive contact with the actual portfolio.
Caps. No single name above 10% of the index. No single system above 25%, regardless of CE-N score. Breaches are resolved by capping and reporting the residual as Thesis Gap, never by substituting in a different name.
The Thesis Gap. Target weights describe where value should accrue on the bottleneck argument. They say nothing about whether a public investor can actually buy that exposure, and on this board, largely, they cannot. Wherever a system’s target weight can’t be absorbed by eligible constituents (no eligible name exists, or every eligible name has hit its cap), the shortfall is reported as Thesis Gap, published as the index’s headline number every quarter. At the current cut it is 25.67%, roughly a quarter of what the bottleneck argument says should be held cannot be bought on public markets today.
Two variants are published side by side, every quarter:
- RBI-system-ce-T (Thesis): target weights, gap held as cash. Measures the thesis itself.
- RBI-system-ce-I (Investable): the gap redistributed pro rata across systems with eligible constituents. Measures what an investor could actually own today.
The spread between them is itself a published series, and it is usually the more honest number than either variant alone.
5. Current system weights
Weights are proportional to each system’s CE-N score (§6), the mean of its member problems’ composite score, computed fresh every rebalance, never hand-typed.
| System | Problems | CE-N | Target weight |
|---|---|---|---|
| Work | P18 | 71.2 | 15.38% |
| Integration | P17 | 60.7 | 13.12% |
| Muscle | P13, P14 | 54.1 | 11.69% |
| Blood | P06, P08, P09 | 54.6 | 11.80% |
| Brain | P04, P10, P11, P12 | 52.3 | 11.30% |
| Immune | P16 | 46.5 | 10.05% |
| Heart | P15 | 47.6 | 10.29% |
| Senses | P01, P02, P03 | 41.5 | 8.97% |
| Nerves | P05, P07 | 34.3 | 7.41% |
Note the compression: first to last spans only 15.38% to 7.41%. The bottleneck framing is doing most of its work in the exclusions and the Thesis Gap, not in the weights themselves. Read §7 and §4 before over-reading this table on its own.
6. The scoring method
The stored CE-N scores use Budget, Solvability, and Value capture. The calculation has five steps:
- Cap the tails. Cap each axis at its 10th and 90th percentiles across the eighteen problems, using linear interpolation between adjacent observations.
- Standardise. Subtract the capped axis mean and divide by its population standard deviation. A constant axis contributes zero. This expresses relative position within the current board.
- Combine. Average the three standardised inputs with equal weights:
composite_z = (z_budget + z_solvability + z_value_capture) / 3. - Rescale. Map to 0–100 using
(composite_z + 3) / 6 × 100, clipped to that range. - Apply the floor rule and round. If any of the three raw inputs is below 2.5, multiply the rescaled score by 0.85. Round half up to one decimal place.
The read-only October 2026 audit reproduced all eighteen stored ce_score_v2 values.
Fixed endpoints do not make scores directly comparable across research cuts: the means,
spreads, and percentile thresholds depend on the selected problems and their ratings.
These transforms do not remove subjective sensitivity. A half-point change to one input moved its problem by up to eight ranks in the diagnostic tests. See the repository review records for diagnostic details. Before a future rebalance, the approved calculation and the candidate stored outputs must be checked again.
6a. Evidence confidence
Evidence quality is assessed separately from the displayed CE-N score. Sources should be checked for independence, relevance, and date; repeated claims do not become independent evidence. Internal confidence-adjustment calculations remain documented in the repository. They are not shown as a second CE-N score and this display change does not alter portfolio calculations.
6b. Two companion signals, deliberately narrow in scope
Two further signals sit alongside CE-N without feeding into it, aimed at the two most subjective axes:
value_capture_density_tier(Low/Moderate/High, all eighteen problems) counts how many companies in the book are mapped to each problem and splits the eighteen into three equal groups by that count. A High reading is a caution, not a bonus: it means many funded competitors are chasing the same bottleneck, the exact commoditisation pattern this site’s tactile-sensing and free-data-glut research documents.solvability_benchmark_pctis populated only where a research brief discloses a real, reproducible real-robot benchmark result: currently P11, Generalisable VLA policies only (RoboChallenge real-robot success ≈50%, RoboDojo’s more conservative real-world figure ≈12%, against a human baseline near 100%). Leftnulleverywhere else, deliberately: this is not a general pipeline, and nothing is estimated to fill the field.
Both are advisory context next to value_capture and solvability, never an input to them.
7. Eligibility gates: the exclusions
A name in the universe is excluded if it fails any gate. Exclusions are published with reasoning. Refusing to hold something is a stronger signal than holding it, and the exclusion list is the most defensible part of a published methodology.
| Gate | Rule | Current exclusions |
|---|---|---|
| Tradeability | No verifiable secondary close, or valuation confidence < 50 | AeroVironment, Inovance, Unitree |
| Earnings quality | Recurring earnings must be separable from non-operating items before sizing | Not excluded: capped at half weight (Hesai, Horizon Robotics, ESTUN) |
| Reconciliation | Share count, net cash, or debt not reconciled | Ouster |
| Valuation | Price already discounts a penetration that has not occurred | Leaderdrive: ~470× P/E, ~100× EV/Sales |
| Governance / control | Material internal-control finding, single-customer dependence, or an unresolved related-party structure | Symbotic: ICFR material weakness, >84% single-customer revenue, Up-C/GreenBox VIE structure |
| Purity | Robotics undisclosed and immaterial against the reporting segment | Held at a 2% purity cap, never quoted as a robotics comp: NVIDIA, Rockwell Automation, Tesla, Hikvision, CATL |
A gate failure is not a judgement about the underlying business. Intuitive Surgical is the
strongest proof on the board that the endgame exists (consumables $1.73bn against systems $685m
across 11,710 installed units), and it is still a valuation_tier: 1 name whose current DCF
output reads as expensive at today’s price. Passing every gate and being fairly priced are two
different questions.
8. Rebalance
Quarterly, on the research cut following each quarter end. Each rebalance publishes:
- Every score change, with the evidence that moved it.
- Additions and deletions, with the gate that triggered each.
- The Thesis Gap, by system, versus the prior quarter.
- A scored self-assessment of the prior quarter’s calls, including the misses.
Point 4 is not decoration. An index graded against itself in public, on a fixed schedule, is the cheapest available way to build a timestamped, falsifiable track record before there is a fund, which is the entire strategic purpose of publishing RBI for free. The eight tier-1 companies’ DCF assumptions and outputs (§3) refresh on the same quarterly cadence.
9. What this index does not do
- It does not include private companies as index constituents. Twenty of twenty-four private names have no live terms; marking them would be fiction. They appear on the site as a Sleeve B tracking roster by system (diligence score, valuation confidence, funding-mark staleness, related-party flags), a read-only reference, explicitly not a tradable index, since there is no committed-dollar amount or daily price discovery to weight against.
- It does not include the 38 factory and supply-chain nodes. Those are an information asset (component order books lead OEM revenue by several quarters), not a portfolio.
- It does not adjust for currency, and it does not model transaction costs, borrow, or tax.
- It is not a fund, does not accept capital, and has no live track record. Base date is 2026-08-13; everything before that date is history, not performance.
10. Known gaps, stated plainly
A methodology that hides its own weak points is marketing.
- Small N cuts both ways. Four of nine systems map to a single problem, so their weight is only as stable as one score. The 10th/90th winsorisation band is itself a concession to N=18 being small.
- Solvability and value capture are still judgements. The scoring math (§6) bounds how much leverage that subjectivity has over the final number; it does not remove the subjectivity. §6b’s companion signals are a first, deliberately narrow step toward a computed check on both , not a replacement for the underlying scores.
- The floor threshold (2.5) and the equal-thirds weighting are stated choices, not derived ones. Nothing in the data forces either; both are published exactly so they can be argued with.
- This desk has not re-scored any problem’s raw inputs based on the most recent sell-side research pass. New evidence is catalogued by system in the research briefs and flagged for review, never applied silently.
- The eight-name DCF layer is new (2026-08-26) and has known open items per company: see
each company’s
dcf_assumptions.notesfor genuine data-source discrepancies (e.g. CATL’s beta ranging 0.37–0.86 across providers, AeroVironment’s unreconciled debt figure). One name (AeroVironment) currently produces a non-meaningful output because its free cash flow is negative post-acquisition, flagged, not hidden. - A conservative-terminal-growth DCF will read almost every premium-multiple growth name as
“overvalued.” That is a documented, deliberate property of the model (see each company’s
dcf_output), not a bug: it isolates how much of a price is justified by disclosed near-term cash flow versus priced-in narrative or optionality the model doesn’t capture. Read a large negative margin of safety as a measurement, not a sell signal.
Falsification condition for CE-N specifically, stated in advance: if a future rebalance shows CE-N systematically reordering problems in a direction that tracks which analyst scored them rather than tracking new evidence, the standardisation step is amplifying noise rather than removing it, and the weighting scheme should be revisited first.
Sources: Robotics Industry Heatmap & Investment Intelligence System v0.1; Robotics IC Memo Book; Robotics Landscape 95, research cut 2026-08-13. Methodology references: MSCI Quality Indices Methodology (winsorisation + z-score standardisation); Weighted Sum Model vs. Weighted Product Model, multi-criteria decision analysis literature; outranking-method veto thresholds (ELECTRE/PROMETHEE family) as precedent for the conjunctive floor. Facts, derived calculations, company claims and analyst inferences are kept in separate fields throughout the underlying data. Unknowns are left unknown.