2026-08-12
The Price Fell And The Bill Rose
Published token prices have fallen by two thirds on a like for like tier, and one scheduled increase was withdrawn before it took effect. Bills are still rising, because the unit is defined by your supplier and the volume is defined by you.
Anthropic's published pricing documentation carries a sentence that ought to unsettle anyone who has written an AI budget. Its models from one recent generation onward, it records, use a newer tokenizer that produces approximately 30 per cent more tokens for the same text. The company attributes the change to improved performance across a wide range of tasks, and that is a reasonable thing for it to have done. The price per token did not move. The token did.
That is tokenomics in one line. Every forecast, every business case and every vendor comparison in this market is denominated in a unit whose size is set by the counterparty and can change with a model release.
What the record says
The optimistic case is real, and the vendors' own price lists support it. On the same Anthropic pricing page, a now retired flagship is listed at fifteen dollars per million input tokens and seventy five dollars per million output tokens. Every model that succeeded it on that line is listed at five and twenty five. That is a two thirds reduction on a like for like tier, on published rates, visible to anyone who opens the page.
The mid tier tells the sharper story. Its current model launched at two dollars and ten dollars as introductory pricing due to expire on 31 August 2026. The page now records that this is the standard price, and that the scheduled increase to three and fifteen on 1 September will not occur. A price rise announced, then withdrawn. Whatever else is true of this market, nobody in it currently feels able to charge more.
The strongest case that this is a solved problem
That case deserves stating properly, because it is not a weak one.
Compute gets cheaper, model efficiency improves, four well capitalised laboratories are competing for the same enterprise budget, and the published direction of travel is unambiguously down. On this reading the rational executive response is to do nothing clever. Wait, buy later, and let the price curve do the work. Anyone optimising a token bill today is solving a problem that deflation will solve for them by the next budget cycle.
The argument is correct about the rate card and wrong about the invoice.
Why the unit will not hold still
Four things move underneath the price, and only one of them is the price.
The first is the size of the unit itself, which is the tokenizer point above. Thirty per cent more tokens for identical text is, in every respect that matters to a finance director, a thirty per cent price increase that appears nowhere on the price list.
The second is that a single token has many prices. Anthropic's documentation sets a cache hit at a tenth of the base input rate, a five minute cache write at 1.25 times, a one hour write at double, batch processing at half on both input and output, and United States only inference at 1.1 times across every category. A priority speed setting on the flagship runs at ten and fifty against a standard five and twenty five, for the same model. The page states that these multipliers stack. The same text, sent by two competent teams, can differ in cost by close to an order of magnitude before anyone has argued about which model to use.
The third is that the meter is no longer only tokens. Its managed agent product bills session runtime at eight cents per session hour on top of tokens. Code execution is free alongside web search and web fetch, and otherwise runs at five cents per container hour after 1,550 free hours a month. Web search is ten dollars per thousand searches. A pricing model described in public as per token has quietly become multi dimensional, and the dimensions that are not tokens are the ones that scale with autonomy rather than with text.
The fourth is volume, and it is the one that dwarfs the others.
The part that should make every executive sit up
At Google I/O in May, Sundar Pichai put Google's monthly token processing at more than 3.2 quadrillion. Two years earlier the same figure was 9.7 trillion a month. A year earlier it was roughly 480 trillion. That is not a market where unit prices falling by two thirds produces a smaller bill. It is a market where consumption is outrunning deflation by orders of magnitude, because reasoning models, tool calls and agents that run for an hour unattended consume tokens at rates that a chat window never did.
Set that beside the worked example Anthropic publishes for its own agent product. A one hour coding session on its flagship model, consuming fifty thousand input tokens and fifteen thousand output tokens, comes to seventy and a half cents, of which eight cents is session runtime rather than text. Enable prompt caching so that forty thousand of those input tokens are cache reads, and the identical session costs fifty two and a half cents. A quarter off the bill, with no change to the model, the prompt, or the work performed.
That is the number to hold on to. The controllable variable is not the rate card, and it was never going to be. It is the engineering configuration, and it sits with people who have historically not been asked to think about gross margin.
What follows
Five consequences, none of them exotic.
Forecast in work completed rather than tokens consumed. Tokens per resolved support ticket, per reviewed contract, per closed sprint item. A budget line denominated in tokens is a budget line denominated in a unit your supplier defines.
Read any contract, chargeback model or internal recharge that quotes a token rate as carrying an unwritten clause about unit size. If a model upgrade can change the tokens required for the same text by thirty per cent, that clause is worth negotiating explicitly rather than discovering.
Treat caching, batching, model tiering and inference geography as financial controls, because that is what the multipliers make them. The gap between a well configured deployment and a careless one is larger than the gap between two vendors' headline rates.
Instrument consumption per task before you optimise price per token. Falling rates against rising volume can hide a doubling in cost per outcome, and the rate card will look like good news throughout.
And ask every supplier the two questions that the price list does not answer. What is my unit, and who decides how big it is.
The industry will go on announcing lower prices, and those announcements will go on being true. Whether your bill follows them down depends on a unit you do not control and a volume you very largely do. Only one of those is being tracked in most organisations, and it is the wrong one.
