AMD MI355X beats B300 on performance per dollar for Kimi K3

Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

AMD MI355X beats B300 on performance per dollar for Kimi K3

Wafer runs Kimi K3, a 2.8T-parameter model, on a single node of AMD MI355X GPUs, achieving 952 tok/s aggregate and 118 tok/s single-stream decode. This outperforms a two-node B200 setup and offers better performance per dollar than NVIDIA's B300. The post details fixes for ROCm-specific issues, including a missing top-k renorm function and an MLA prefill kernel shape mismatch, which improved prefill speed by 2-3x.

But that’s exactly the point: Kimi K3 at its size is one of the first models we’ve seen where the MI355X’s focus on HBM capacity gives it a practical, measurable edge over the B200.
  1. kelnos

    I wish they wouldn't call them "open source models". They aren't open source. They didn't publish the training data. They didn't publish the tools they used to train the model.

    They published the weights. It's an "open weight model", a term that it seems nearly everyone has agreed is appropriate. Why is this company not using it?

  2. logicallee

    This part sounds like AI assisted setting this up and benchmarking it:

    >The fix was trivially simple: zero-pad the head count 12→16, run the fast kernel, and extract the real 12 heads from the output.

    I've recently used a frontier AI (ChatGPT 5.6 Sol on ultra) to set up a much smaller local model, and the performance optimizations it introduced left the model totally incoherent. (The model just repeats a single character, etc.)

    When I see a line like the one I just quoted, it leaves me wondering if the setup is still coherent like a stock install of Kimi K3 on supported hardware.

    Did they run any benchmarks on it to see if it is still correct?

  3. inferencecoder

    Wafer is making themselves synonymous with slop in the inference space. Exaggerated unfair comparisons in all their results, twitter hype posts with alarm emojis etc.

    > $2.50/GPU-hr for the MI355X, $6.00 for the B300, and $4.25 for the B200.

    This is not an accurate price comparison for real terms.

  4. GuestFAUniverse

    How is the capex on 8 * MI354X even remotely justified at less than $10/h?

    Even without the base system, power and every other expenses:

    365d * 24h * $2.95 = $25842/a invoicable.

    That doesn't add up within one year, that doesn't add up in three years and it is questionable that it brings in the money during the lifetime of the device?

  5. jpgvm

    If you do good work you should at least take the time to review the slop that details that work for slopiness. Otherwise it's hard to take it seriously. Especially the prefill section.

More from this day

2026-08-02