ANTHROPIC

AI token prices plummet, but why actual costs increase: the story of Jevons paradox

Bùi Đăng MinhTuesday, August 11, 20264 min read
AI token prices plummet, but why actual costs increase: the story of Jevons paradox

In recent weeks, if we pay attention to the prices of AI models, we see a rather strange trend: almost all major companies are racing to reduce prices, but operating these models themselves is not correspondingly cheaper. OpenAI just cut the price of the smallest and fastest model in its GPT-5.6 family, called Luna, by 80%, bringing the price down to $0.20 per million input tokens and $1.20 per million output tokens, much lower than before. The mid-range model Terra is also down 20%, to $2 per million token input and $12 output. Token here, simply put, is the smallest unit that a language model uses to read and create text, which can be roughly understood as a word or part of a word. AI companies charge by the number of tokens processed, similar to how carriers charge by data capacity.

[​IMG]
[​IMG]

OpenAI recently made a move to reduce model prices What's really curious is why there was such a sharp decrease. Just over a week before OpenAI's move, Google also launched two cheaper models, Gemini 3.6 Flash and 3.5 Flash-Lite. Anthropic went in a different direction, not cutting prices directly, but replacing its cheapest model, the Opus 4.8, with the more powerful Claude 5.0 but keeping the same price. Overall, this is clearly a price race driven by competitive pressure from Chinese models, specifically Moonshot AI's Kimi K3 and DeepSeek V4 Flash, open source models that are significantly cheaper than their US competitors, forcing OpenAI and others to react so as not to lose market share in the low-cost segment.

Paradox called Jevons

This is the part I find most interesting when reading about this topic: the AI ​​industry is operating strictly according to an economic law called Jevons paradox, named after British economist William Stanley Jevons, who first observed this phenomenon in the 19th century while studying the coal industry. This paradox states that as a resource becomes more efficient or cheaper to use, the total consumption of that resource, instead of decreasing, often increases, as users tend to use more to compensate. Applying to AI: the cheaper the price per token, the more motivated users and businesses are to stuff AI into more processes, call the model more times, build more automated tasks, and the actual total cost can still end up increasing, even though the unit price per token decreases.

afIp9cBOoF08xbSS-SP590-Jevons’Paradox-revised.png.jpeg
afIp9cBOoF08xbSS-SP590-Jevons’Paradox-revised.png.jpeg

Jevon Paradox: more economical energy consumption does not reduce energy consumption, on the contrary Actual data is showing that this is true. A recently cited Goldman Sachs report estimates that AI systems that are capable of acting autonomously, meaning AI that not only answers a question but plans itself, calls tools, and executes a series of multiple steps to complete a task, can increase token demand by up to 24 times compared to conventional question-and-answer AI usage. The reason is simply because an agentic task does not just generate one answer, but may have to "think" through many intermediate steps, calling the same model many times to test, fix errors, iterate, each step consuming tokens. Even Sam Altman, CEO of OpenAI, has publicly admitted that token costs are becoming "a big problem," as overspending on AI has become a subject of online ridicule.

Why do companies still accept this game?

Perhaps what's worth mentioning here is not that AI companies are "unreasonably losing capital", but that they are betting on the Jevons paradox itself for growth. The logic here is: if the price is low enough to attract more users and use cases, total revenue can still increase even as profit margins per token decrease, as long as the growth rate in usage volume exceeds the rate of price decline. This is also the reason why some businesses are starting to turn to Chinese open source models or cheaper options to keep budgets under control, instead of continuing to rely entirely on closed American models. But the price to pay for this strategy is not small. There are specific cases that show the downside of "using AI as much as possible" without control: there was a company that accidentally spent up to 500 million dollars on Claude in just one month, simply because it did not set usage limits for employees. On a larger scale, many organizations, including Microsoft, Meta and Amazon, are said to be having to curb the trend of employees abusing AI for every trivial task, jokingly called "tokenmaxxing", after realizing that operating costs are silently ballooning without a corresponding increase in work efficiency. The takeaway from reading this story is: the cheaper price of AI does not mean that using AI will be cheaper. It's just like a more fuel-efficient car doesn't necessarily cost you less on gas during the year, because you'll probably drive more. With AI, where each token is virtually free, the real question is no longer “how much does each model call cost,” but “are we calling the model in a controlled manner and with a clear purpose.” That is probably the real management problem behind the ongoing price reduction race.

Nguồn / Original source: Tinh tế