The Price of AI Performance Is Falling 47 Percent a Quarter
Epoch AI's new study finds the cost of hitting a fixed benchmark score has dropped about 13x a year since 2023, fastest in math, slowest in games, and only for buyers who keep switching models.
In January 2025, OpenAI's o3 could score 75% on GPQA Diamond, a multiple-choice exam of PhD-level physics, chemistry and biology, for an estimated 30 cents per question. Just under 18 months later, OpenAI's GPT-5.6 Luna matched that score for $0.0004 per question.
That is a 725-fold drop. Epoch AI, the research group that published the comparison on September 22, likens it to a new car's sticker price falling from $50,000 to $69. The example opens "The plunging price of thought," a report by Luke Emberson and David Roodman that tries to put one number on how fast AI is getting cheaper.
Their number: the cost of reaching a given level of performance has fallen about 47% per quarter since 2023, or roughly 13x per year.
What Epoch measured
The report does not track the price per token. Token prices are easy to compare but misleading, because reasoning models can burn very different amounts of tokens on the same question. Epoch instead asks what it costs to reach a fixed score on a benchmark, and tracks the cheapest model that can do it at each point in time.
Five benchmarks carry the headline: OTIS Mock AIME, Chess Puzzles, FrontierMath Tiers 1-3, GPQA Diamond and Mystery Game Puzzles. To get cost curves without rerunning every model at every budget, Epoch followed a method from the federal Center for AI Standards and Innovation (CAISI), which uses one high-budget transcript to estimate how many questions a model would answer before hitting a lower spending limit.
Epoch also filled an obvious gap. Benchmark data is biased toward expensive, maximum-effort runs, so the team ran low-effort settings and small efficiency-focused models, including open-weight models hosted on rented hardware. It checked that approach against five open-weight models also sold by API and found its costs within 30% of API pricing.
Math falls fastest, games slowest
The 47% figure is an average, and the spread is informative. Math got cheaper fastest, at 50-52% per quarter, or 16-19x per year. Game-based puzzles fell slowest among the five, at 39-43% per quarter, or 7-10x per year. FrontierMath Tiers 1-3, Epoch's own hard math benchmark, led the table at 53% per quarter. Chess Puzzles were lowest at 43%.
Coding was the notable laggard. SWE-bench Verified, a secondary benchmark in the study, showed a decline of 27.5% per quarter. That is still fast by any normal standard, but it suggests the price of real software work is not collapsing as quickly as the price of exam answers.
The report also finds that timing matters. When a performance level is brand new, its cost falls about 66% per quarter, or 75x per year. Two years later, the decline slows to about 32% per quarter, or 4.7x per year. Epoch offers one possible explanation: the lab that first reaches a new level can charge a premium for a short time, then competitors, open and closed, catch up and the price falls sharply.
The comparison that will get quoted
Epoch compares this with other technologies. In log terms, it says the decline is four times faster than DNA sequencing (1.84x per year, 2001-2025), six times faster than compute (1.51x per year, 1940-2001), 18 times faster than lithium-ion batteries (1.16x per year, 1991-2024) and 54 times faster than US residential electricity (1.05x per year, 1892-1973). The authors write that no other general-purpose technology "appears to have gotten so cheap so fast."
There is outside support for the direction. Mert Demirer, Andrey Fradkin and Nadav Tadelis, writing in the Summer 2026 Journal of Economic Perspectives, used OpenRouter data and concluded that the price of intelligence has fallen roughly a thousandfold, with open-source models about 90% cheaper than comparable closed ones. Epoch added a citation to that paper on September 23. Economist Alex Tabarrok, writing on Marginal Revolution, drew his own conclusion: a given level of intelligence now needs far less inference spending, not just better models.
The caveats are real
Epoch lists its own limits, and they matter for anyone planning a budget around the headline.
First, benchmaxxing. Labs may train for known benchmarks, so scores can rise faster than real usefulness. Mystery Game Puzzles, which keeps its game secret, is the check on this, and its decline rate sits below average at 44.0% per quarter without a model, or 38.7% under Epoch's fitted model.
Second, the headline assumes a buyer who always moves to the cheapest model that can do the job. Epoch says directly that real users do not switch models that often and do not get the full savings.
Third, the data is thin. It covers barely three years, not every model is tested on every benchmark, and reasonable averaging choices produce a range of about 42.9% to 58.0% per quarter. The authors say they cannot credibly compute standard errors. On September 23 they also added a caveat that the "price of thought" is a moving bundle of services, which makes comparisons with electricity or sequencing less clean.
Epoch is also not a neutral observer of one benchmark here. It runs FrontierMath. That does not make the numbers wrong, and the code and data are public on GitHub, but it is worth knowing.
What it means if you buy AI
The practical lesson is in the gap between the frontier and the typical buyer. The savings go mostly to teams that re-evaluate models often and move work to cheaper tiers once they match the old flagship. If your stack is fixed to one model you picked a year ago, you are likely paying last year's price for work a newer, cheaper model can now do.
It also changes how to read a new flagship's price. The premium on a new top score is real but short-lived, and Epoch's data says the fastest declines come in the first quarters after a level is reached. For work that does not need the very top of the leaderboard, waiting a quarter or two is a pricing strategy.
The exception is coding, where the curve is flatter. Agent-heavy engineering budgets should not assume math-style price declines.
