Hands-On with the AMD Ryzen AI Halo: Local AI at 45 Tokens Per Second

I tested the AMD Ryzen AI Halo to see how it handles local AI tasks. The processor delivered impressive speeds, generating text at 45 tokens per second without needing cloud connections. This hands-on review explores its real-world performance for running large language models directly on your machine, offering a glimpse into the future of private, on-device artificial intelligence.
Local AI at 45 tokens per second means you can run powerful models privately, right on your desktop, without waiting for the cloud.
- chmod775
This is more of an ad, not a review, and reads like the author has hardly any experience with the things he's trying out. That Z Image Turbo diffusion model would've also run on many consumer GPUs and with way higher performance for a fraction of the price. Misleading.
- jmward01
It seems like there is a very health space for an MOE targeted GPU where it has essentially an 5070ti with 16gb ish GDDR7 but then also has 128 GB LPDDR5x (or even just DDR5 as expansion dimms on it?). Putting this into the same card would likely reduce the transfer hit when a cache miss happened and the gpu had to load from the slower LPDDR5x. No need to have PCIe 5x16 limiting memory transfer if it is on the card. MOE models could then get near native performance and even models where the active parameters + context didn't quite fit the thrashing would be less of a problem. Not UMA but gets the UMA 'lots of system memory to play with' benefit.
- syntaxing
Highly recommend lemonade server if you have a strix halo desktop. Been using Qwen3.6-35B @ Q_8 as my main driver and it’s been great with 60 TPS for generation. I occasionally use the 27B @ q6 but only get 20-25 TPS for generation with MTP.
- netinstructions
Unfortunately, the table of models and tokens per second (TPS) and time to first token (TTFT) is not helpful without specifying the quantization of the model.
- cmrdporcupine
Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people.
My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case.
Once people can run something like GLM 5.2 at lower quants (512GB could do a passable job), then I think the story changes.
Whether we ever see DRAM as cheap as it was ever again, I don't know.