Mercury 2.5 LLM hits 770 tokens per second
Artificial Analysis benchmarks Mercury 2.5, a new LLM that reaches 770 output tokens per second. The model is evaluated on the Intelligence Index v4.3.2, a suite of 10 tests including AA-Briefcase, GDPval-AA, and Humanity's Last Exam. The analysis also covers cost per task, token usage, context window, latency, and end-to-end response time, offering a comprehensive view of the model's performance and price.
Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).
- bearjaws
If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.
I've used it on a few for fun projects and its decent but the speed is crazy to watch.
- the_arun
Chat Jimmy clocks at 17K tokens per sec burning LLM into the Chip - https://chatjimmy.ai/ - Source: https://theashishmaurya.medium.com/taalas-the-startup-that-p...
- jjcm
I'm still sad that we haven't seen a new Taalas style chip a la https://chatjimmy.ai/. Smaller models are good enough now to make that insane burst of tokens so useful.
- nextaccountic
At some point the bottleneck becomes tool calling.. and as such, it's preferably if the model is co-hosted (in the same datacenter, at least) with your code repository and all other reference/context it needs (full documentation for most ecosystems, maybe even a copy of common crawl to minimize web fetch usage, etc)
- freakynit
I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max. Even GPT-OSS-20B performs way better than this in my own attempts to use it.
I really really wanted to use this because it offers incredible speeds and pricing combinations. But nop.. I still am not using it.. not even for basic tasks.