Needle 2 - 14MB Agentic LLM for Tiny Devices

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

Needle 2 - 14MB Agentic LLM for Tiny Devices

Needle 2 is a 45M-parameter language model compressed to a 14MB binary, designed for on-device tool calling and structured extraction on phones, wearables, smart home devices, and robots. It runs in just 28MB of RAM, achieving 500+ tokens/sec on a Raspberry Pi 5 and fitting on microcontrollers like ESP32-S3. Built on Simple Attention Network and Cactus Quants, it delivers performance comparable to models 5-70x larger while using 7-85x less compute per token. Apache 2.0 licensed, with weights on Hugging Face and a GitHub repo for easy deployment.

The model's footprint is tiny and the performance never lets us down.

More from this day

2026-08-10