The Inference Economy
Training built the AI industry's first chapter. Serving models continuously to agents and users is writing the second.
The first phase of the AI build-out was about training: enormous clusters, months-long runs, and a small number of labs competing to build the largest models. The second phase is about what happens after the model exists. Every query, every agent step, every voice interaction is inference, and inference runs all day, every day, for every user.
Share of AI compute: training versus inference. Training: 2022 70%, 2023 62%, 2024 55%, 2025 45%, 2026 38%. Inference: 2022 30%, 2023 38%, 2024 45%, 2025 55%, 2026 62%.
Why now
- Agents multiply the number of model calls per task by an order of magnitude or more.
- Reasoning models spend more compute per answer by design.
- Consumer scale means billions of daily interactions.
- Cost per token has fallen sharply, which increases usage rather than reducing spend.
What changes
Hardware: inference favours chips optimised for memory bandwidth and efficiency, opening room for custom silicon and alternative architectures alongside the accelerators that dominate training. Location: latency-sensitive workloads pull capacity toward users, while batch workloads can follow cheap power. Pricing: every layer, from chips to cloud to applications, is moving toward usage-based models.
- Cost per million tokens
- Falling
- Provider price lists
- Tokens per agent task
- Rising
- Vendor disclosures
- Inference-specific silicon announcements
- Rising
- Company announcements
- Edge and regional inference sites
- Rising
- Project announcements
Who benefits, who is at risk
Beneficiaries: efficient inference hardware makers, serving infrastructure providers, regions with cheap power, and application companies whose costs fall as prices drop. At risk: companies whose margins depend on inference costs staying high, and buyers who provisioned training-class capacity for inference-class workloads.
What happens next?
- Custom inference silicon gains share as workloads standardise.
- Regional and edge inference capacity expands.
- Application pricing shifts fully to usage and outcomes.
Related topics
Sources & references
- 01Model provider pricing pages and usage disclosures — Frontier labsprimary
- 02Semiconductor company earnings commentary — Public filingscompany