Why AI Token Prices Keep Falling While the Bill Still Grows
Published October 7, 2026
Gemini 4 Argon launched at $2 per million input tokens, half of Claude Opus 5.5. Why the price per token keeps falling while total spend keeps rising.

On 30 September 2026, Google launched Gemini 4 Argon at $2 per million input tokens and $10 per million output tokens. It is roughly half what Claude Opus 5.5 costs, and access starts narrow: vetted cyber defenders in Google's Fairwind Program, then enterprise customers.
For most developers that is the whole story, and a good deal. But it is not the number that decides what companies actually pay. The token price is falling because of a mechanism that will not stop. What will not stop is the bill.
What actually determines the price of one token
API price is not set from training cost. It comes from four things that interact.
First, compute per token. A model using a mixture of experts architecture does not run all of its parameters on every token, so cost per token drops even though the total parameter count is unchanged.
Second, hardware quality. A newer generation of chip processes more tokens per dollar than the one it replaces, so a distributor holding new chips can offer a lower price without changing the model at all.
Third, quantization. A model whose weights are cut from 16 bits to 8 or 4 bits needs less memory and less data movement per token. Capability drops slightly; cost drops a lot.
Fourth, and least often mentioned, competition. When several labs release models of comparable ability, price becomes a transaction tool. Margin is traded away for volume.
Why the bill still rises
This consequence is nearly unavoidable and rarely discussed openly. If the price per token falls by half, the same work costs a quarter of what it did. So anyone who needs twice the tokens, or ten times as many, sees their bill rise.
Token demand grows for three reasons.
First, agents. A request that used to take 1,000 tokens can now run a long exchange with several different models, each carrying its own context. That context is reprocessed every turn.
Second, larger context windows. Gemini 4 Argon raises its output limit from 64,000 to 1 million tokens. A bigger limit creates its own demand: if a model can take 1 million tokens, users will fill it.
Third, longer reasoning. A model that thinks longer before answering spends more tokens on the same answer. That is a choice rather than a defect, but the cost sits somewhere invisible.
A short history of token prices
Modern API pricing began around 2020, when language models moved from research into products. For the first two years price was calculated to cover GPU cost, and a customer had no option but to accept it.
Between 2022 and 2024 the direction reversed. Smaller models became good enough for most tasks, and 4-bit quantization went from experiment to default at many providers. Prices fell several fold in two years.
From 2025 the pattern shifted again. Competition moved into the frontier tier itself rather than only cheap models. Google, OpenAI and Anthropic began selling models of comparable capability at lower prices. The fall happening now is different in nature: not a technical saving, but a business decision to take share.
Numbers nobody can forecast
Nobody can honestly predict token prices three years out. The direction can be stated: down. The rate cannot, because it depends on things nobody can see, including whether an architectural breakthrough arrives, whether quantization moves further, and whether hardware improves faster than it has.
Engineering will keep making tokens cheaper, and there is a physical limit. At some point the time it takes to move one token through memory becomes a floor that current hardware cannot pass.
Numbers you can work out yourself
Compare price per million tokens across frontier models with published prices:
| Model | Input per 1M tokens | Output per 1M tokens |
|---|---|---|
| Gemini 4 Argon | $2 | $10 |
| GPT-6.1 Sol | $2 | $10 |
| Claude Opus 5.5 | $4 | $20 |
The Claude figure is derived from Google's statement that Argon costs about half of it, so it is not directly confirmed. The Gemini and GPT figures are published.
Now do the arithmetic. A team processing 50 billion input tokens a month at $2 per million pays $100,000 a month. If tokens rise to 250 billion -- which happens as soon as agents enter the daily workflow -- the bill becomes $500,000.
Price per token is not rising as fast as the bill. The bill rises because volume rises, and volume rises because capability rises.
What this means for readers in Indonesia
API cost is one of the most concrete parts of building a product. In Indonesia the calculation is more open because most users write in Indonesian, while most frontier models are trained on data dominated by English.
Two practical consequences. First, models of equal capability can differ sharply in cost if trained better on the target language. Second, quantization tuned for English can degrade disproportionately on non-English text, because tokenization cuts words differently.
For a team choosing a model for an Indonesian language application, the number to compare is not price per million tokens. It is cost per correct answer. A cheaper model that fails to understand the question in its own language ends up more expensive, because it gets asked again.
Three things to do now to reduce the bill:
- Cache frequent answers. Many questions are identical. Caching turns a repeated request into zero tokens.
- Trim the context you do not need. Most applications send documents the model never reads. Cutting them before sending reduces the bill directly.
- Route by task. Not every task needs a frontier model. Sending easy work to a small model is where the large saving is.
Analysis: whether this price decline keeps going
Three things decide whether the mechanism still has fuel in it.
First, hardware efficiency. The increase in tokens per dollar from one chip generation to the next has historically been slowing. If that pattern holds, price per token stops falling at some point and what remains is margin competition.
Second, the floor on quantization. Cutting weights from 8 bits to 4 bits saves a lot. The next cut saves far less, and the quality lost becomes noticeable. Practically there is little headroom left.
Third, whether a different business model appears. Almost all of this traffic is billed per token today. A different model, such as a per device licence, shifts the incentive completely.
The conclusion: token prices will fall again, but not as fast as in 2024 and 2025. And while volume grows faster than price falls, the bill per user rises as well.
| Factor | Remaining headroom | Price direction |
|---|---|---|
| 8-bit to 4-bit quantization | large | falls |
| 4-bit to 2-bit quantization | small | falls |
| Next generation chip efficiency | narrowing | falls |
| Competition between models | unlimited | falls |
The distribution is designed on purpose
Customers do not buy tokens. They buy the ability to finish a task. Here is the same work priced three ways, assuming 100 million tokens a month per user:
| Usage | Price | Bill per user |
|---|---|---|
| 50M tokens, cheaper model | $2 per million | $100 per month |
| 10M tokens, frontier model | $2 per million | $20 per month |
| 200M tokens, frontier model | $10 per million | $2,000 per month |
Look at the direction. The most active user, the one who needs the most expensive capability most often, pays the most. That gradient is what keeps vendor incentives honest: the better the model, the more people need it.
Why this eventually consumerises
Two things can stop the rising bill. The first is a model good enough for the median user. If a mid tier model reaches 90 percent of frontier capability, most people see no difference and the majority migrate.
The second is efficiency on the user side. Shorter prompts, caching, and batched processing reduce the tokens a task needs. Large companies already do this. Small teams do not.
If both happen, price per token falls further while volume grows more slowly. That is the scenario that improves totals. It takes years, and nothing speeds it up.
Related tools
Free browser tools that apply to this topic.
Share this article
Share to
Related articles

October 7, 2026
AI Agents That Browse on Their Own, and the Controls That Failed
899 AI requests to a Canadian government site and a cancelled GPT-6.1 release. The root cause was not the model but the controls around it.

October 7, 2026
501 Billion Parameters Just Beat 2 Trillion. It Is Not the Tech
Reflection AI released Beam with 501 billion parameters and beat a 2 trillion model on some tasks. The parameter count is not what made the difference.

October 7, 2026
Why Countries Now Pay for Their Own Model, and Why It Is Hard
South Korea is preparing $3.5 billion for a national AI model. Four things have to work out before a programme like that succeeds.



