Inference — running a trained AI model to answer a query — has collapsed in price at a fixed performance level: GPT-3.5-level inference cost fell more than 280-fold between late 2022 and late 2024, per Stanford HAI's 2025 AI Index Report. The same report finds hardware costs falling roughly 30 percent yearly. That decline is applied AI's biggest economic fact.
What is inference, and why does it cost money?
Training a model is a one-time capital expense; inference is the recurring one. Every prompt runs real hardware: accelerator chips doing the arithmetic, memory shuttling the model's weights and the conversation's tokens, and electricity for both. A model with hundreds of billions of parameters cannot be reloaded for each request, so providers keep resident capacity warm — GPUs that are idle still cost money, which is why pricing is designed to smooth utilization across customers.
That is also why the bill is denominated in tokens, the chunks of text the model reads and writes. Input tokens are cheaper than output tokens because generating text requires the model to compute every next token in sequence, while reading a prompt can be parallelized and, increasingly, cached. None of this is exotic — it is capacity economics, the same shape as any compute utility.
Two costs sit inside every served request and they scale differently. Compute scales with the number of tokens processed and with the model's active parameters; memory scales with the model's total size, because all weights must be resident whether or not they fire. Much of the past three years of inference engineering — sparser models, quantized weights, paged key-value caches — is an attack on the memory term, because memory is the resource that has historically set how many simultaneous users a given cluster can serve.
How fast are costs actually falling?
The reference point is Stanford HAI's 2025 AI Index, the eighth edition of the annual report, whose stated scope includes novel estimates of inference costs alongside new analyses of AI hardware trends. Its headline finding: the inference cost for a system performing at the level of GPT-3.5 dropped over 280-fold in roughly two years — from late 2022 to late 2024 — holding capability constant and letting the models shrink.
Two supporting curves run underneath. The report puts hardware cost reductions at about 30 percent per year, and energy efficiency gains at roughly 40 percent per year. Multiply better silicon, leaner models and cheaper serving software, and the price of a fixed unit of intelligence falls faster than any single component — a compounding effect, not a one-off price war.
A caution on labels: that 280-fold figure measures cost at a fixed capability level, not the price of frontier models. The most capable systems remain expensive precisely because they use more parameters, more tokens and more compute per query. Cheap intelligence and frontier intelligence are two different markets moving at two different speeds.
What is actually driving the decline?
The report's analysts attribute the collapse to several stacked forces, and the mechanics are worth understanding separately:
- Hardware economics — accelerator cost per unit of compute falling roughly 30 percent a year, per the AI Index, with successive chip generations.
- Energy efficiency — roughly 40 percent annual improvement in work done per watt, which compounds directly into serving cost.
- Distillation — small models trained to mimic large ones deliver GPT-3.5-class quality at a fraction of the parameter count.
- Quantization — running weights at lower numeric precision, which trades negligible quality for large memory and speed savings.
- Competition — multiple providers selling overlapping capability tiers, with open-weight alternatives putting public price discipline on the market.
The forces compound rather than add. A distilled model that needs a quarter of the parameters, running on silicon that is 30 percent cheaper per year per unit of work, served with a cache that halves effective input cost, multiplies out to far more than any single improvement — which is how a 280-fold result becomes plausible in roughly 24 months rather than a decade of any one trend.
None of the mechanisms is proprietary in principle, which is why the curve has held across providers rather than depending on any single vendor's pricing strategy. Distillation and quantization are published techniques; hardware economics follow the accelerator market; caching is a systems discipline. The decline is an industry learning curve, not a promotion.
How should a team pick a tier without overpaying?
The practical discipline follows from the curve. Because cost at a fixed capability level keeps falling, the cheapest acceptable model today is a moving target — capability tiers that were premium-priced two years ago now sit near the bottom of public price lists, per the AI Index's framing of capability-adjusted cost. Teams that benchmark their own workload against several tiers quarterly capture the decline automatically; teams that pin to one model inherit its price trajectory.
The measurement that matters is cost per completed task, not per token. A cheaper model that requires longer prompts, more retries or more output tokens to finish the same job can be more expensive end to end. Rigorous evaluations on real workloads — with pass rates, token counts and retry rates recorded — turn the falling price curve into actual savings rather than an interesting statistic.
How do providers charge for inference?
The unit is almost always dollars per million tokens, split into input and output rates — output typically several times input. Around that base, providers have built discounting machinery that rewards predictable usage: cached input, where repeated prefixes of a prompt are stored and re-sold cheaply; batch processing, where non-urgent jobs run at reduced rates on off-peak capacity; and tiered latency, where slower or faster responses carry different prices.
For builders, the lever hierarchy is fixed. Model choice dominates — dropping one capability tier usually beats any optimization. Prompt caching comes next for applications with repeated context, then batching for anything offline. Token discipline, such as trimming system prompts and truncating history, remains the unglamorous third lever that shows up in every serious engineering write-up.
What still keeps AI bills high?
Unit prices fall; usage does not. Three structural costs resist the curve. Output-heavy workloads pay the expensive token every time, and content generation or long-form synthesis cannot dodge that ratio. Agentic systems multiply inference by looping — a task that calls a model twenty times inherits twenty bills, plus retries. And long context re-reads large prompts, which caching softens but does not eliminate.
The second-order effect is the oldest one in energy economics: when a resource gets radically cheaper, consumption rises to meet it. Teams that priced an AI feature as a premium add-on in 2023 are now embedding model calls into core loops — voice interfaces, code review, document pipelines — because per-call cost no longer forbids it. The 280-fold decline, in other words, is not just a savings story; it is the reason AI features stopped being a line item and started being an assumption.

