From physical compute to the end user: a 4-layer value chain stratified by "Unit of Trade"
The through-line of this research is the production and consumption of tokens. The ultimate goal: from underlying compute to end-user price, see clearly how value is distributed at every link in the chain and who takes the lion's share.
The key change in v0.2 was establishing an organizing principle: each layer uses its "output unit of trade" as the stratification marker. This turns "layer" from a "role label" into a "rigid criterion" — same transaction unit means same layer.
What defines each layer is not "what it does" but "what unit of trade it outputs." The unit of trade converts between layers; the conversion point is the layer boundary.
Top to bottom, arrows label the unit of trade passed between each layer. L3 and L4 each have an internal sub-structure (1P/3P, wholesale/retail) — expanded below.
The bottom layer of hardware reality — infrastructure (data center / power / networking) + the AI chips themselves. Outputs "physical carriers of compute."
Physical space, stable power, high-speed interconnect, and electrical & cooling equipment — the AI factory's "land + nervous system."
| Sub-item | Share | Key Suppliers |
|---|---|---|
| Electrical (UPS / PDU / switchgear / distribution) | 40–50% | Vertiv · Schneider · Eaton · ABB |
| Cooling (mostly liquid) | 15–20% | Vertiv · CoolIT · nVent · Boyd |
| Networking (switches + optical modules + NIC) | 12–18% | Arista · Broadcom · NVIDIA · Coherent · Innolight |
| Building shell | 15–20% | Colocation providers + general contractors |
| Design / management / permits | 5–10% | — |
Turn silicon into AI compute units, sold or self-used in the form of "cards" or "systems."
| Item | Capital Investment | Depreciation Period | Annualized Cost |
|---|---|---|---|
| Infrastructure (DC building) | ~$15M | 25 yr | ~$0.6M |
| GPUs (~750 H100 @ $35k) | ~$26M | 5 yr | ~$5.2M |
| Power (1 MW × $0.08/kWh × 8760h) | — | — | ~$0.7M |
| Total | ~$41M | — | ~$6.5M / year |
Directly owns and operates L1 assets, converting hardware into "machine-hours" sold externally.
Converts GPU Hours into a callable Token API. Internally formed by the convergence of two streams: training and inference.
Both kinds ultimately "output a Token API," but the costs they bear are entirely different — a critical cut that any subsequent "value-chain verification" must make cleanly.
Bears training sunk cost + in-house inference marginal cost. Strong pricing power, must amortize training investment.
Bears only inference cost (open weights are free; licensed weights are paid). Essentially Inference-as-a-Service.
Engines powering 3P: Inferact (vLLM, $800M)RadixArk (SGLang, $400M)NVIDIA TensorRT-LLM
The Token API is consumed at this layer. After entering from L3, it splits into three parallel paths flowing to different endpoints — some continues to circulate as Token API among developers; some is packaged by application-layer companies into "Subscription + Credit" sold to end users.
Via an Aggregator or Token Brokerage, N 1P / 3P APIs are aggregated behind a single interface and then resold at a markup.
An AI Lab (1P) or a 3P provider sells the Token API directly to developers, with no middleman markup.
The Token API enters application-layer companies, where it is wrapped into vertical products by industry. These companies ultimately sell to end users as "Subscription + Credit" — the highest-markup segment of the entire value chain happens here.
| General Chatbot | ChatGPT Claude.ai Gemini Poe |
|---|---|
| Coding | Cursor Windsurf GitHub Copilot Cognition · Devin Replit |
| Legal | Harvey Robin AI Spellbook |
| Presentation / Slides | Gamma Tome Beautiful.ai Prezi Minds |
| Customer Service | Crescendo Decagon Sierra |
| Search | Tavily Perplexity You.com |
| Sales / Marketing | Attentive Postscript Jasper Copy.ai |
| Finance | Surf |
| AI + Payment | Coinbase AgentKit Stripe · Tempo MPP Skyfire Kite AI x402 Protocol |
| Multi-channel IM | Tinyfish Locus |
| Personal Agent | Magic Character.ai Pi |
| Notes / Writing | Notion AI Granola |
The unit of trade defines crisp layer boundaries, but the leading companies frequently swallow up adjacent layers, folding upstream and downstream into a single house. When doing "value-chain verification," you must first separate these companies' financials from any single-layer markup calculation — otherwise gross margins get blended into a single pot.
| Direction | Representatives | Economic Implication |
|---|---|---|
| L3 → L1 AI Lab builds own compute |
OpenAI Stargate · Google TPU+DC · Meta 350k H100 · xAI Colossus | Big labs skip L2 and procure bare metal / build their own data centers directly. L2 New Cloud's bargaining ceiling is pushed down. |
| L2 → L3b Cloud providers move up to API |
AWS Bedrock · Azure OpenAI · GCP Vertex | L2 uses its distribution to package closed-source models as 3P APIs and resell them, blurring the 1P / 3P boundary and absorbing part of L3 profit. |
| L2 + L3b integrated Independent 3P with in-house compute |
Together AI · Fireworks · Groq | A single company straddles L2 / L3b. Token cost cannot be cleanly separated in the books, and external pricing does not expose per-layer markups. |
| L1 → L3 / L4 NVIDIA software tax |
NVIDIA NIM · DGX Cloud | A hardware vendor jumps over L2 and goes straight to inference services and API platforms. CUDA lock-in extends upward. |
| L3a → L4 Path 3 1P self-operated end-product |
OpenAI ChatGPT · Anthropic Claude.ai · Google Gemini app | An AI Lab runs its own end-product, bundling 1P API + application under one roof. Gross margin is not observable from outside. |
The first three sections settled "how the layers divide and who swallowed whom." This section answers a plainer question: how many times larger is the money the whole industry pours in than the money it actually earns back? Data is based on 2026 Q1 earnings (disclosed 2026-04-29) and public commitments; all amounts in USD.
Sources: Q1 2026 earnings calls (CNBC / Fortune / Microsoft Source / Alphabet earnings, Apr 29 2026) · Synergy Research (Apr 2026) · Sacra · CreditSights · Morgan Stanley · The Information.
Hyperscaler CapEx traced a near-straight line upward across 2023–2026. From the ~$155B range in 2023 to a 2026 guidance median of $735B, with ~60–70% annualized growth sustained. All five had Q1 CapEx YoY growth above 65%; Oracle tripled in a single quarter.
| FY | Big 5 combined CapEx | YoY | Context |
|---|---|---|---|
| 2023 | $155B | — | First year after ChatGPT; still normal levels |
| 2024 | $256B | +65% | First acceleration wave |
| 2025 | $443B | +73% | Training compute rolled out at scale |
| 2026E | $725–760B | +64–72% | Current guidance; several raised within the year |
| 2027E | ~$1,050B | +40% | Goldman / Moody's forecast; may cross $1T within the year |
CapEx / Sales ratio at record highs: Oracle 86%, Meta 54%, Microsoft 47%, Alphabet 46%, Amazon 25%. Big 5 free-cash-flow coverage dropped below 1.0× for the first time — CapEx + dividends + buybacks now exceed operating cash flow.
Source: Q1 2026 earnings + Introl / Morgan Stanley estimates.Reorganizing revenue by this token map's 4-layer framework shows "which layer the money enters from." Cloud AI revenue (L2+L3b) is currently the largest ($74.5B), but much of it comes from AI Labs' compute return flow; AI Labs (L3) external revenue ~$57B, Neocloud (L2) ~$10B.
| Layer | Representatives + scope | 2026E ARR | Key watch |
|---|---|---|---|
| L3 · AI Labs | Anthropic + OpenAI + xAI + Mistral + Cohere + Perplexity + others | ~$57B | Anthropic $30B (disclosed Apr, only $9B at year start) overtakes OpenAI (~$25B) for the first time; Top 2 = 96% |
| L2/L3b · Cloud AI | MSFT AI Business + AWS Trainium/Bedrock + Google Cloud GenAI | ~$74.5B | MSFT $37B (incl. OpenAI usage, +123% YoY) · AWS $20B · GCP ~$17.5B (+800% YoY) |
| L2 · Neocloud | CoreWeave + Lambda + Together + Crusoe + Nebius + others | ~$10–25B | CoreWeave $66.8B backlog · Nebius 6× growth; essentially swing capacity backstopping the two layers above |
| Total · gross revenue | — | ~$150–160B | net external ~$110–120B after stripping circular flows |
AI Labs' largest cost is compute, and the compute sellers are their largest shareholders. Investment buys compute commitments, compute commitments prop up valuations, valuations buy more investment — this loop means cloud AI revenue's "gold content" needs a discount. Three main loops:
| MSFT cumulative investment in OpenAI | $13B+ (incl. 27% profit-share equity, ~$135B mark) |
| OpenAI Azure long-term commitment | $250B (~7 years, ~40% of MSFT's $627B RPO) |
| OpenAI 2026E ARR | ~$25B (year-end expectation $30B+) |
| Est. annualized Azure return | ~$10–20B/yr (50–80% of OpenAI revenue) |
MSFT's "AI Business" $37B ARR disclosure scope itself states it "includes model builders' Azure usage" — i.e., OpenAI's training + inference consumption is booked directly into Microsoft's AI revenue.
Source: Microsoft Source (Apr 29 2026) · om.co 10-Q analysis (May 1 2026).| AMZN cumulative investment in Anthropic | $8B (invested) + up to $25B (incl. convertibles), latest valuation $60.6B |
| Anthropic AWS long-term commitment | $100B+ (10 years, 5GW Trainium compute) |
| Anthropic 2026E ARR | $30B (disclosed Apr; $9B at year start → $30B Apr = 3.3× / 4 months) |
| Est. annualized AWS return | ~$10–25B/yr (Morgan Stanley est. 75% of cost is cloud) |
AWS self-disclosed Trainium business $20B ARR, highly synced with Anthropic's commitment curve. A significant share of AWS's $364B backlog comes from Anthropic long-term contracts.
Source: Anthropic announcement (Feb 2026) · VentureBeat (Apr 2026) · Bloomberg.| Google investment in Anthropic | $10B invested + up to $40B milestones (incl. cloud credits) |
| Anthropic Google Cloud commitment | $200B / 5 years (5GW TPU compute, starting 2027) |
| Share of GCP $462B backlog | ~40% (per $200B/$462B) |
| Est. annualized GCP return (mature) | ~$40B/yr (starts 2027; 2026 ~$2–5B) |
GCP backlog doubled within Q1 2026 to $462B, driven precisely by Anthropic's $200B commitment (disclosed 2026-05-05). Google Cloud GenAI revenue +800% YoY, operating margin doubling from 17.8% → 32.9%, behind the same lock-in.
Source: The Information (May 5 2026) · Google Q1 2026 earnings call · BigGo Finance.Circular-flow total: Anthropic + OpenAI alone have locked ~$550B in long-term compute commitments to the Big 3 clouds (AWS $100B + GCP $200B + Azure $250B), nearly half of the Big 3's (incl. Oracle) combined ~$2 trillion RPO. Annualized "double-counted" revenue est. ~$28B (range $21–35B), ~18% of gross AI revenue.
| Cycle | Peak-year CapEx | CapEx / Revenue | Asset utilization | Outcome |
|---|---|---|---|---|
| Fiber bubble 1998–2000 | $120B/yr | ~2× | ~2% (measured 2002) | WorldCom / Global Crossing etc. ~80% of companies bankrupt |
| Cloud buildout 2010–2015 | $80B/yr | ~2.7× | rising | Demand caught up; AWS/Azure/GCP became leaders |
| AI buildout 2024–2026 | $1,050B/yr (2026E) | ~6.6× gross / ~9.4× net | ~97.5% (current GPU) | Undetermined |
Three data differences decide whether "this time is different" holds: (1) Speed — AI buildout is 3× faster than cloud buildout; (2) Concentration — 5 companies control ~60% of spend (vs the fiber era spread across dozens of telcos); (3) Utilization — GPU 97.5% vs fiber's 2% post-crash, the strongest "demand is real" evidence today, but also means no buffer.
The Big 5 have already issued $108B in debt (3.4× the historical average); Morgan Stanley/JPM estimate $1.5T cumulative issuance across 2025–2027 to fund this wave. CapEx/Sales sits at historical extremes; future increments come from debt, not cash flow.
Anthropic's $200B GCP contract doesn't book until 2027; OpenAI's 7-year $250B is still in its first half. Labs' return-flow share of cloud AI revenue is projected to rise from 18% to 25–30%, meaning "real revenue" growth lags the disclosed figures.
Anthropic's 80× annual growth with training cost only 1/4 of OpenAI's means the efficiency frontier is shifting left. If frontier-model training cost structure enters a "train once, beats four prior" phase, 2027 CapEx ROI gets repriced.
OpenAI 900M WAU (was 800M in 2025-10); M365 Copilot paid seats 15M → 20M in 6 months; Anthropic Fortune 100 penetration 70%, $1M+ annual-spend customer count doubled in 2 months to 1000+; Menlo Ventures measured enterprise LLM spend at 2.4× in half a year. Demand is shifting from "pilot" to "line-item IT budget."
$10–25B ARR vs $40–65B CapEx is an inverse spread. CoreWeave's $66.8B backlog has Microsoft as its main customer (once ~62% of revenue) — essentially an off-balance-sheet extension of hyperscalers. Neocloud utilization / GPU spot price / hyperscaler in-house capacity are the early signals for whether this round is overheating.
E5 treated Neocloud as a single ARR blob, but the same $10–25B mixes two very different kinds of money: renting bare GPU compute (compute rental, per GPU-hour) and selling tokens (inference API, per million tokens). This line decides whether a company looks more like a "commoditized GPU landlord" or a "sticky, high-margin software" business.
| Company | Rent GPUs · compute | Sell tokens · API | Scope / source |
|---|---|---|---|
| CoreWeave | ~95%+ | minimal | Multi-year reserved + enterprise contracts; no standalone API; Microsoft once ~62% of revenue |
| Crusoe / Lambda | ~90%+ | small | Energy-tied / developer-driven bare compute rental; H100 spot ~$3.9/GPU-hr |
| Nebius | compute-dominated | Token Factory (not broken out) | Q1 2026 total revenue $399M (+684% YoY); 6-K reports a single "AI cloud" segment, inference revenue not disaggregated |
| Together AI | ~60–70% | ~30–40% | ~$1B ARR; Sacra / Contrary estimate; the most token-leaning neocloud |
| Q1 2026 total revenue | $399.0M (vs $50.9M, +684% YoY) |
| Segment scope | Nebius AI cloud / Avride / TripleTen; the latter two "contributed only limitedly to group revenue" |
| Within AI cloud | Single revenue line (GPU compute + storage + software services), inference vs compute not split |
| Token Factory (launched 2025/11) | Zero mentions in the entire 6-K |
Implication: for neoclouds, "inference / Token Factory" remains narrative > financials — companies will tell the inference story on earnings calls, but in the revenue structure it is still too small to break out. Renting GPUs is still the base.
Source: Nebius Group 6-K Q1 2026 (SEC EDGAR, filed Apr 2026) · Sacra · Contrary Research.Direction call: the closer to "renting GPUs," the more like a landlord (low margin, lives and dies with GPU spot price, bleeds the moment utilization dips); the closer to "selling tokens," the more like software (high margin, API stickiness). The industry trend is migrating toward the token end — Microsoft internal data: ~50% of GPU customers now access compute via AI APIs, only 20–25% still lock in reserved bare metal; inference is projected at ~2/3 of all 2026 AI compute (only 1/3 in 2023). Together is the only clearly token-leaning neocloud, the root of why its valuation narrative is more "software-like" than pure CoreWeave. Upstream engine-author companies (vLLM→Inferact $800M · SGLang→RadixArk $400M) entering managed-API hosting will pull value further from "renting GPUs" toward the "selling tokens" layer.
Scope & disclaimer: AI-specific CapEx is not separately disclosed by any company, estimated at the industry-standard 75%; circular flows are external estimates (companies do not disclose "OpenAI's spend on Azure"); some long-term commitments (Anthropic-GCP $200B) do not start booking until 2027. Core figures are as of 2026-05-11; §4·F (Neocloud revenue structure + Nebius 6-K) added 2026-06-28. Next review: 2026 Q2 earnings (~2026-07-29).
§4·F argued that the more a neocloud "rents GPUs," the more it lives and dies with the GPU spot price — so here we pull the H100 $/GPU-hr out and track it weekly, split into spot / neocloud / hyperscaler (same chip, up to a 20×+ spread). The past-year storyline: a three-year slide bottomed in autumn 2025, then 2026 turned into a structural reversal — but it's the contract / tight tiers that reversed while marketplace spot stayed low. The divergence itself is the signal this panel watches.
Method & disclaimer: Once a week we take public on-demand single-H100 quotes from getDeploying, bucket providers into the three tiers, and take the median (self-built index, base 100 = first live week 2026-06-29). H100 composite avg = equal-weight mean of the three tier medians (dark-red line); its stat card shows price, index and week-over-week change (live points only, never across the seed/live seam). Dashed = reconstructed history (2025-07 to 2026-06, rebuilt from Silicon Data tier medians + spot-floor reports); solid = weekly live capture. The two baskets differ in methodology, so a level jump at the seam is expected — read each segment's trend, don't compare absolute levels across the seam. Dotted = published anchors of two paid industry indices (Silicon Data H100 blended index; SemiAnalysis 1yr-reserved contract index) — these are not scraped, just a few manually entered points disclosed in their blog/newsletter, overlaid to calibrate our self-built live line. Their methodology is not fully comparable to our three tiers.