Self-Hosting Kimi K3: 20% Higher Cost for 20% Better Task Resolution

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

Self-Hosting Kimi K3: 20% Higher Cost for 20% Better Task Resolution

I explored self-hosting Kimi K3 on an 8×B300 node, finding it costs 20% more hardware than our GLM-5.2 setup but resolves 24 percentage points more tasks. While token throughput drops and task times increase, the significant quality gain suggests that for heavy coding workloads, owning infrastructure might outweigh the volatility of API token bills. The key lies in balancing peak usage against developer experience.

The better unit of value delivery is a task successfully completed.
  1. walrus01

    Quoting from the article:

    > In our runs, K3 served 16 concurrent sessions (GLM-5.2 managed 24). Aggregate token throughput is about 30% lower (122 vs 170 tok/s at 16 users), and median task time is about 50% longer (38 vs 26 minutes). That makes K3 roughly 8 times slower than our Claude Code baseline. However, K3 makes up for it in quality, resolving 86.4% of tasks, 24 percentage points above both GLM-5.2 and Opus 4.8 (62.5% for both).

    I think there's some less tangible advantages to self-hosting something on the scale of Kimi K3 that can't be quantified in a specific number like token/s or percentage of problems solved. Such as:

    a) data privacy/sovereignty from a wide range of possible perspectives, from medical to personal to "we can't have our data go to the USA" for some Canadians and Europeans.

    b) being able to give it information security/network security tasks and red team scenarios without triggering claude or openai refusals.

    c) being able to give it information security/network security tasks with zero risk of getting your account banned or investigated by anthropic or openai.

  2. ktosobcy

    Recently I started playing with LM Studio and local models and found out that `gemma-4-26b-a4b` is suprisingly capable. I don't need elaborate akin to "create complete app to do X" or "refactor the whole codebase of bazzilion of LOC" but rather "how to go about doing thing X" or for language study (explaining nuances of phrasal verbs or subtelties of vocabulary in other languagues) and darn -- the results are rather good (to the point that I use it mostly nowadays)

  3. Lord-Jobo

    >which box to buy

    > spark, costs less than a conference trip.

    I know putting actual prices regionally localizes your article and temporally, with how prices are so unstable. But analysis of “what to buy” without actual prices is borderline meaningless.

    Overall, good article, very interesting to see a real deployment that’s actually attainable and not just a subscription to a big 3 token plan.

  4. joshstrange

    I could not focus on the article with all the noise in the background. It was annoying in the header but then it continued down the page. If you want to do this on your marketing pages have at it but for a blog/news style page? Reader mode was the only way to restore sanity.

  5. michalpleban

    I would love to see such comparisons but with quantized versions, because quantization allows running models on smaller hardware with some quality loss. I am running Qwen3.6-35B-A3B quantized to int4 on an A6000 card that was otherwise just sitting around idle. It works up to a degree, but I would love to see benchmarks comparing different quantizations of several models, especially in quality. This is an important dimension in deciding whether to buy a GPU is worth it, and it is missing from this (otherwise comprehensive) article.

More from this day

2026-07-29