TurboFieldfare - Gemma 4 26B inference on 2 GB RAM for Apple Silicon

Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac

TurboFieldfare - Gemma 4 26B inference on 2 GB RAM for Apple Silicon

TurboFieldfare is an open-source Swift and Metal runtime that enables running the Gemma 4 26B-A4B model on any Apple Silicon Mac with as little as 2 GB of RAM. By streaming only necessary expert weights from the SSD while keeping the core model in memory, it bypasses traditional hardware limitations. This innovative approach allows users to run powerful local AI models on 8 GB Macs without loading the full 14.3 GB dataset. The project includes a native Mac app, CLI, and an OpenAI-compatible server, offering a complete solution for efficient local inference on resource-constrained devices.

הזיכרון הפך ליקר. אז נתתי למודל עם 26 מיליארד פרמטרים תקציב של כ-2 ג'יגה-בייט בלבד.

עוד מהיום הזה

2026-07-30