MiMo-v2.6-Pro: Intelligence, Performance & Price Analysis

MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

MiMo-v2.6-Pro: Intelligence, Performance & Price Analysis

Artificial Analysis evaluates MiMo-v2.6-Pro across its Intelligence Index v4.3.2, which aggregates ten benchmarks including AA-Briefcase, GDPval-AA, AutomationBench, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam. The analysis covers capability indexes, openness, cost per intelligence task, token usage, context window, output speed, latency, and end-to-end response time, with comparisons against other models.

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
  1. segmondy

    Mimo2.5 is really good, but tended to loop too much for my taste. Locally, Pro2.5 wasn't much better. I would reach for it for one shots, hopefully they sorted it out with v2.6, it's a model that's slept on by many. I found that most people that used it did so because it was free. It's a top model worth exploring if you have never given it a go.

  2. Gareth321

    OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.

  3. egeres

    It feels suspicious that MiMo-V2.6

    Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)

  4. dom96

    It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.

    KillSwitch-Bench 1.0

    Claude Opus 5 66.9

    GPT-6 Astra 57.9

    Claude Fable 5.1 46.7

    MiMo-V2.6-Pro 38.8

    Muse Spark 1.3 36.5

    1 - https://bench.killswitch-lang.org/

  5. tensegrist

    where's the flash model? it's out already isn't it

More from this day

2026-09-22