Cactus Needle 2 - 14MB Agentic LLM for Tiny Devices
Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Cactus Needle 2 is a 45M-parameter open model for tool calling, device use, and structured extraction. It runs as a single 14MB binary in just 28MB of RAM, hitting 500 tokens/sec on a Raspberry Pi 5 and working on sub-$200 phones, wearables, and even microcontrollers like ESP32-S3. Trained with Cactus Quants for lossless 2-bit quantization, it matches larger models like FunctionGemma 270M and Apple FM on function-calling benchmarks while being 5× to 70× smaller. Apache 2.0 licensed, with weights on Hugging Face and a repo to get started.
The Pebble Index Ring has no screen. So when you speak to it, the action just has to happen, every time, with or without internet connection. We run Cactus Needle locally in the app, instead of relying on the cloud. The model's footprint is tiny and the performance never lets us down.