How many GPUs does 1 trillion tokens per month need?

How many GPUs is 1M/B/T tokens?

How many GPUs does 1 trillion tokens per month need?

Cedana's Tokens to GPUs Calculator estimates that serving 1 trillion tokens per month on Llama 3.3 70B requires about 367 H100 GPUs, with a plausible range of 211 to 853. The calculation converts monthly tokens to average throughput, adjusts for token mix and utilization, then divides by per-GPU speed. Sensitivity analysis shows H100 speed and utilization assumptions dominate the result. The tool also covers A100, V100, and B200, and warns that estimates can be off by 2x in either direction.

The result is a planning estimate. Expect an error of 2x in each direction for measured presets, and more for estimated presets. Measure your model on your hardware before you buy.
  1. DiabloD3

    This calculator is kinda useless. Different models work differently, you cannot generically define them by parameter count, and the models it does list seem to just be it prefilling the param counts for you.

    It also misses most cards being used for inference, only limiting it to relatively recent Nvidia cards. Even if this was meant purely for datacenters, it would be useless for AMD customers.

    And as for different models being different, it also doesn't understand kv cache size, how many concurrent sessions you need to manage, the compute cost overhead of kv cache quantization, nor the compute cost overhead of model quantization. It also doesn't know what MTP nor DFlash is, and cannot increase your effective tps to match.

    As an example: compare Unsloth's quant of Qwen 3.8 Q4_K_M (4.5 BPW), a decent smaller model, without MTP, you'd get a baseline of, say, around 3-6 TPS per 100W. With the built in MTP model, you're closer to 6-12 TPS per 100W; then, switch quants to Byteshape's IQ4-XS (3.84 BPW) and use their suggested Dflash drafter, you're now in the realm of 15-30 TPS

    ... yet if I scribble into the calculator "27B, 27B active, 4-bit, 1B per one day", it seems to be the low end of my non-MTP figure: 1B per day = 11574 per second, it recommends 3x B200s, each B200 is 1200W, so 11574 / (1200*3) = 3.215, yet real world results would be 2x to 10x higher.

    Also, quoting from the website, "A100 and V100 results for the large models are theoretical.". The small scale inference peo […]

  2. dannyw

    I’m sorry but this looks like vibe coded marketing slop that’s highly inaccurate.

    For one, there is zero consideration of prompt/KV caching, which we all now is basically essential especially for workloads at scale.

    Secondly, it seems to base all benchmarks off a batch size of 1.

    Nobody running a cluster of B200s or H100s is doing inference with batch sizes of 1.

    And even worse, the calculator assumes you run models with a context window of 0 tokens? That affects how many GPUs massively.

    I’m not nitpicking over small details or intentional simplifications here, but the estimates this is giving is horrendously inaccurate by a few multiples.

More from this day

2026-10-07