GPU rents doubled to $8.08 an hour while token prices fell
Tomasz Tunguz puts the reconciliation down to efficiency. For anyone pricing an AI product on token costs, the useful number is gross profit per GPU-hour.
Renting a GPU costs roughly twice what it did six months ago. Tunguz, a partner at Theory Ventures, puts the move from $4.40 to $8.08 per GPU-hour.
Over the same period, the price of inference fell. Both things are true, which is what makes the pricing decision awkward.
His explanation is efficiency. Models are getting cheaper to run faster than silicon is getting dearer to rent.
The figures he cites: Claude Opus 5.5 costs 40% less to run than its predecessor, and OpenAI cut the price of one model by 80% in July and another 50% in September.
On the same account, a benchmark threshold that cost $0.55 to clear eighteen months ago now clears for $0.0015. That is a 377-fold reduction.
The cost pressure on the other side is physical. Tunguz notes that every input to a data centre buildout is rising, from concrete to copper to credit, and that electricity is the binding constraint.
He points to Oracle invoking force majeure on its New Mexico campus last week after the natural-gas pipeline feeding the site slipped by six months.
Demand is not helping. Inference and model companies posting record growth attract more capital, which they spend bidding for the same limited GPU-hours.
The financing backdrop has also changed shape. Tunguz reports that the 10-year Treasury's correlation with tech valuations, historically −0.50, has flipped to +0.39 over the past two years, with rates climbing from 3.63% to 5.18% while the NASDAQ rose roughly 78%.
In other words, the market has stopped repricing tech on the cost of money and started betting the growth arithmetic works.
The metric he says matters is gross profit dollars per GPU-hour. It asks one question: are your efficiency gains outpacing your compute costs?
Microsoft says it generates 90% more tokens per GPU than a year ago, concentrated in its smaller models. If that pace holds, the two curves are running neck and neck.
For a founder, the practical point is that input costs and output prices are moving in opposite directions at once. Pricing a product off today's token price assumes the efficiency curve keeps paying for the rent curve.
That assumption is worth writing down as an assumption. Then track gross profit per GPU-hour monthly, and check whether the cheaper small model does the job before paying frontier rates for it.