M5 Ultra Mac Studio: The Dream Machine for Local AI Agents

M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents

After four days of testing the M5 Ultra Mac Studio with 256 GB RAM, I found it a game-changer for local AI agents. Compared to the M3 Ultra, prompt processing is 2.5× faster and token generation ~70% faster, enabling smooth multi-turn agentic loops. Running Qwen3.8-Flash-Next locally now powers my daily assistants, eliminating cloud costs and privacy concerns.

It was all written by me, the old-fashioned human way. But the entire research stack, deep-linking between notes, and keeping track of new features and betas were all performed by my agents, running locally on the Mac Studio, for a total cost of $0.
  1. simonw

    The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

    Qwen3.8 27B tokens/sec generation speed

    Prompt size 8K 64K 128K 256K

    RTX 5090 PC 59 51 44 n/a

    M5 Ultra 48 39 32 24

    M3 Ultra 31 23.5 20 15

    A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...

  2. srcreigh

    This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

    I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

    It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

    The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

    An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

  3. sajithdilshan

    On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.

    That’s like 12 years worth of OpenAI Pro subscriptions

  4. tempoponet

    While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.

    This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.

  5. ApolloFortyNine

    The model being tested is 18k as configured.

    I didn't expect this to make the 5090 to look like a good deal.

  6. akozak

    "a total cost of $0" Uhh ... how much is that hardware?

  7. hamiltont

    Once you hit the memory you need, generation speed is mainly set by bandwidth, and every Ultra from M1 thru M3 has ~800 GB/s. IMO best ROI for most people is 'cheapest used Ultra with enough RAM'

    I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)

  8. liuliu

    When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.

More from this day

2026-09-21