Daily.Brief

AI

How AI model pricing collapsed, and who actually pays for inference

#002May 11, 20268 min readBy Joseph

In early 2023 the most capable model you could rent cost about sixty dollars per million output tokens. By 2026 a model that is better on every benchmark costs a small fraction of that, and the cheapest usable models cost cents. That is the fastest price decline of any major input in modern economic history, and it changes who wins.

Why the price fell

Three things happened at once. The chips got faster and, per unit of work, cheaper: each generation of Nvidia hardware roughly doubled the useful output per dollar. The models got more efficient, with techniques like distillation, where a big model teaches a small one to do most of what it does at a fraction of the size. And the labs started competing on price, because once two models are close in quality the only lever left is the invoice.

The result is that inference, which means running a trained model to answer a question, went from a scarce, expensive thing to something closer to electricity. You still pay for it. You just stop thinking about it.

Where the money went

Follow the dollar. A company paying for tokens sends money to a lab. The lab spends most of it on cloud compute, which means it sends most of it to Microsoft, Amazon, Google or a specialised GPU host. Those companies send a large share of it to Nvidia for chips, and Nvidia sends a large share of that to TSMC for manufacturing.

At every step in that chain, the margins are wildly different. Nvidia's gross margin has been above 70 percent. The cloud providers make healthy but normal margins. The labs, by most reporting, lose money on the frontier models and make it back, if at all, on scale. The company that sells the shovels made the money. The companies digging are still hoping.

Price per million output tokens, best available model at each date
Price per million output tokens, best available model at each date15 USD30 USD45 USD60 USD2023 H12023 H22024 H12024 H22025 H12025 H22026 H13 USD
Approximate list price in US dollars for a frontier-class model over time, log-scale thinking applies: each step is a large multiple. Based on published API pricing from OpenAI, Anthropic and Google.

Who pays now

Increasingly, not the user. The price of a chat with a model has fallen below what most people would notice, so the labs bundle it into subscriptions or give it away and charge businesses instead. The real paying customers in 2026 are companies running models inside their own products: customer service, coding tools, document processing, search. They buy tokens by the billion and negotiate prices that never appear on a public price list.

The other payer is the investor. Every large lab has raised money at valuations that only make sense if inference eventually becomes a very large, very profitable business. That money is subsidising today's prices. If the investors are right, the subsidy ends when scale arrives. If they are wrong, prices have to go up, and the products built on cheap tokens get more expensive.

What this means for a business built on AI

If your product is a thin layer over someone else's model, your costs will keep falling and so will your competitors'. That is good for users and bad for margins. The businesses that hold up are the ones where the model is one ingredient and the value is in the data, the workflow, or the relationship with the customer.

My read: token prices keep falling for another two or three years, then flatten as the labs need to show profits. The window where you can build something on nearly free intelligence is open now. It will not stay open forever, and the companies that treat cheap inference as permanent will be the ones surprised when the bill arrives.

Subscribe

The next one, in your inbox.

Free, about 7:00 AM Eastern. English or Chinese, daily or weekly.

Free. No spam. One click to leave, and you can change language or frequency from any email.