Token prices keep falling, yet H100 GPU rental prices hold steady or climb, according to data from Ornn, Silicon Data and Bloomberg as of August 2026. The pattern was described by a16z as a textbook Jevons paradox in the AI market. Cheaper tokens unlock agents, automation and new applications, so total volume grows faster than unit cost falls. For hardware suppliers, this is the outcome that keeps demand scarce and expensive.
Why cheaper tokens support hardware demand
Ornn, Silicon Data and Bloomberg track two opposite price lines: inference tokens moving down and H100 rentals refusing to follow. The August 2026 snapshot shows no parallel decline in compute rents despite cheaper output. That divergence matters because inference revenue depends on volume multiplied by price. As long as consumption rises fast enough, lower token prices do not reduce total spending on infrastructure.
The mechanism runs through usage intensity rather than through chip performance alone. Lower token costs make it economical to run AI agents, broader automation and applications that were previously too expensive. Each task then consumes many more tokens across planning, tool calls, checks and retries. Agentic AI burns through tokens at a staggering rate, so even limited adoption can produce large compute loads.
The source frames this as a classic efficiency effect: falling cost per unit expands total use. A central uncertainty is how much demand comes from humans versus the systems themselves. Part of measured compute demand could be artificially inflated by machine-generated traffic between models and tools. In that case, modest growth in human usage could still trigger outsized hardware needs.
What this means for companies buying AI capacity
For buyers of AI services, cheaper tokens reduce the cost of each request but increase the incentive to automate larger workflows. Companies running support triage, document processing, sales research or internal assistants can afford longer agent runs and more frequent checks. Small firms gain access to use cases that once required careful rationing of calls. Large firms see the bigger shift in aggregate volume, where many teams and processes generate continuous token consumption.
The risk sits in the assumption behind the whole chain: usage must grow fast enough to offset falling token prices. If demand flattens, pressure moves from chip makers and memory suppliers to energy providers and cloud companies. Buyers should therefore separate per-token pricing from total cost of ownership, including retries, supervision and infrastructure. It is also important not to read every efficiency gain as permanent savings, since expanded scope can absorb the difference.
A practical marker to watch is whether inference volume continues to outpace token price declines in industry datasets. A second signal is revenue durability at major model providers after reports that OpenAI annualized revenue might be lower than previously reported moved US stocks. If volume growth slows while rents stay high, procurement terms and capacity planning will change first.
