Kimi K3 Ranks Second on AA-Briefcase Despite High Costs and Slow Speeds
Kimi K3: second only to Fable 5 on AA-Briefcase

Kimi K3 achieves the second-highest score on our AA-Briefcase benchmark, trailing only Fable 5 with strong analytical capabilities. However, this performance comes at a steep price, costing over ten dollars per task and averaging nearly an hour to complete. While it outperforms models like GPT-5.6 Sol in correctness, its presentation quality and speed lag behind competitors, raising questions about its practical efficiency.
Kimi K3 is second only to Fable 5 on AA-Briefcase, but costs more than Opus 4.8 to run while averaging nearly an hour per task.
- eugene3306
What harness do they use for testing?
Back in 2025 it was common to test models in a different harnesses.
I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc.
Why did it come out of fashion ?
EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%
- am17an
It's expensive now, I expect once it is with inference providers it will be really dirt cheap. Then it would be truly be a "bicycle for the mind", which Fable promised to be except it proved to be too capricious for that.
- thecopy
When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?