Kimi K3 Ranks Second on AA-Briefcase Despite High Costs and Slow Speeds

Kimi K3: second only to Fable 5 on AA-Briefcase

Kimi K3 Ranks Second on AA-Briefcase Despite High Costs and Slow Speeds

Kimi K3 achieves the second-highest score on our AA-Briefcase benchmark, trailing only Fable 5 with strong analytical capabilities. However, this performance comes at a steep price, costing over ten dollars per task and averaging nearly an hour to complete. While it outperforms models like GPT-5.6 Sol in correctness, its presentation quality and speed lag behind competitors, raising questions about its practical efficiency.

Kimi K3 is second only to Fable 5 on AA-Briefcase, but costs more than Opus 4.8 to run while averaging nearly an hour per task.
  1. eugene3306

    What harness do they use for testing?

    Back in 2025 it was common to test models in a different harnesses.

    I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc.

    Why did it come out of fashion ?

    EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%

  2. am17an

    It's expensive now, I expect once it is with inference providers it will be really dirt cheap. Then it would be truly be a "bicycle for the mind", which Fable promised to be except it proved to be too capricious for that.

  3. thecopy

    When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?

More from this day

2026-07-22