llama.cpp: Run LLMs Locally with Pi Coding Agent

The official llama.cpp website highlights pairing it with Pi, a local coding agent, via the pi-llama plugin for automatic model discovery and zero-config setup. llama.cpp is optimized for any hardware, from laptops to clusters, with hand-tuned kernels for various GPUs and CPUs, including Apple Silicon, NVIDIA RTX, AMD, Intel, and more.
Files stay on your machine, requests never leave it.