I run a local LLM server on an M4 Pro Mac mini — here's my setup

My local model setup on an M4 Pro Mac Mini

I run a local LLM server on an M4 Pro Mac mini — here's my setup

A developer details how they run a local LLM server on an M4 Pro Mac mini with 48GB RAM, using Qwen3.6-35B-A3B and Gemma-4-E4B models via oMLX, with Tailscale for network access. They explain the benefits: cost predictability, latency, offline capability, no rate limits, and data privacy. The post also demystifies MoE model memory requirements and shows how to swap models easily.

Cloud APIs are rented land.

More from this day

2026-09-01