Open-weight models make older Nvidia GPUs cheaper to run than new ones
The Economics of Open-Weight Inference
A new paper challenges the assumption that each new Nvidia generation makes the previous one obsolete. Comparing eleven open-weight and eight closed models, it finds the cheapest open-weight option completes a task at about one-fifth the cost of a comparable closed model. Self-hosting on rented A100s costs $0.12–$0.35 per million output tokens and beats the H100 on sparse models. Rental data show the A100 holds 80% of its one-month price at five years, versus 44–60% for Hopper and Blackwell.
The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old.
- hungryhobbit
Dark grey text on a black background: it's like they are trying not to let anyone read their article!
- augment_me
I think an interesting point is that hardware as of today still has no utility value after its reported lifetime has elapsed, which prevents neolabs and smaller labs from getting older HW clusters as the banks are not willing to give out loans against them. There is no agreed upon pricing for "expired" A100 clusters or similar.
This is clearly not true, and we are starting to see compute markets, but only for rental prices/H, not for the hardware itself. I feel like there is some artificial moat being built here to stimulate sales of new hardware, because an H100 at 1/16th the price will have comparable dollar/FLOP as Vera Rubin.
- jcmontx
The problem, IMO, with open-weight models is that you accustom to the capabilities of frontier models too quickly; and downgrading to an open-weight "frontier minus 2" or "frontier minus 3" model is often painful, since they feel way less useful than their newer closed-weights counterpart. To be honest, I don't know any companies using OW models at a large scale for their operations (agents or chat assistants).