Skip to content
Tech/ Data Story

The Inference Economy

Training built the AI industry's first chapter. Serving models continuously to agents and users is writing the second.

By
· Updated 5 min read
inLinkedIn𝕏Post
Demo contentThis piece is launch placeholder editorial. Its analysis is illustrative and its charts use labelled demo data. It has not passed the full Parallax Nexus verification process. See How We Use AI.

The first phase of the AI build-out was about training: enormous clusters, months-long runs, and a small number of labs competing to build the largest models. The second phase is about what happens after the model exists. Every query, every agent step, every voice interaction is inference, and inference runs all day, every day, for every user.

Share of AI compute: training versus inferenceIllustrative
Editorial estimate, illustrative
0%25%50%75%100%20222023202420252026
Source: Parallax Nexus editorial estimate. Illustrative demo data.

Share of AI compute: training versus inference. Training: 2022 70%, 2023 62%, 2024 55%, 2025 45%, 2026 38%. Inference: 2022 30%, 2023 38%, 2024 45%, 2025 55%, 2026 62%.

Why now

  • Agents multiply the number of model calls per task by an order of magnitude or more.
  • Reasoning models spend more compute per answer by design.
  • Consumer scale means billions of daily interactions.
  • Cost per token has fallen sharply, which increases usage rather than reducing spend.

What changes

Hardware: inference favours chips optimised for memory bandwidth and efficiency, opening room for custom silicon and alternative architectures alongside the accelerators that dominate training. Location: latency-sensitive workloads pull capacity toward users, while batch workloads can follow cheap power. Pricing: every layer, from chips to cloud to applications, is moving toward usage-based models.

What the desk tracks
Cost per million tokens
Falling
Provider price lists
Tokens per agent task
Rising
Vendor disclosures
Inference-specific silicon announcements
Rising
Company announcements
Edge and regional inference sites
Rising
Project announcements

Who benefits, who is at risk

Beneficiaries: efficient inference hardware makers, serving infrastructure providers, regions with cheap power, and application companies whose costs fall as prices drop. At risk: companies whose margins depend on inference costs staying high, and buyers who provisioned training-class capacity for inference-class workloads.

What happens next?

  • Custom inference silicon gains share as workloads standardise.
  • Regional and edge inference capacity expands.
  • Application pricing shifts fully to usage and outcomes.

Sources & references

  1. 01Model provider pricing pages and usage disclosuresFrontier labsprimary
  2. 02Semiconductor company earnings commentaryPublic filingscompany
Published 5 September 2026 · Updated 13 September 2026 · Report a correction · How we use AI
inLinkedIn𝕏Post