Ways to think about token pricing in the AI supply crunch

I argue that the current AI supply crunch is unstable and likely temporary. With massive data center investments and improving inference efficiency, foundation models risk becoming low-margin commodity infrastructure rather than retaining sustainable pricing power. The market will eventually settle into a new equilibrium, but predicting the exact outcome remains impossible due to unknown variables in future use cases and training costs.
At one extreme, there are two or three giant minds that run half of everything and have massive pricing power, and at the other extreme LLMs look like databases - there'll be millions of them, some very big and some very small, and the value is in what you build on top.
- nekusar
1. Slot machine
You put in minutes of time combined with token costs. You pull the lever and HOPE something good comes back.
And it gets tantalizingly close, but for 100% gotta pull the arm again!
- typ
Tokens will surely become commodities, but models don't need to.
The point is that electricity becomes a commodity, rather than, say, nuclear reactors or gas turbines also need to be commoditized.
Contrary to most people in the software-minded circle, I am not very excited about running local LLMs. Most of what I have seen involves selling wrappers that call APIs on top of an LLM AI, similar to how traditional SaaS has generated revenue from new technologies. So, this would certainly make those people who are working on that sensitive about the pricing. LLM pricing is treated as a continuous cost in the COGS. They don't really use AI to create much value (consumer surplus) other than replacing the existing boring products, like just building the good old CRUD apps, but faster.
What I find REAL interesting about the potential of LLM AI instead is to create new technology out of it, or revolutionize something old to be order-of-magnitude more efficient. In this regard, the expenses on those tokens are more akin to an upfront CapEx.
Cheaper tokens would surely be nice. But if what we are talking about is, like, solving self-driving, curing cancer, or making air-conditioners 100x efficient, the narrow focus on running a cheaper model in my home so I can write my SaaS apps for my get-rich-quick business looks really unpromising and a waste of time for human civilization in a near-singularity horizon.
- mips_avatar
I know there's a lot of reasons to think that everyone will just use AI inference in the cloud, but I think if everyone had access to a dgx gb400 class machine with 512gb of hbm4 vram and 1tb of lpddr8x a lot of people are going to be running finetunes of models locally. Like the dgx gb300 is $94k now, but i bet this class of machine will come down to $20kish in the next 2-3 years.