Needle 2: A 45M-Parameter Tool-Calling Model That Runs in 28MB of RAM as a 14MB Binary

Cactus Compute released Needle 2, an open 45-million-parameter model purpose-built for tool calling that ships as a single 14MB binary and can run a complete session in just 28MB of RAM. This makes it one of the smallest viable tool-calling models available, designed explicitly for edge deployment, embedded systems, and environments where larger runtimes like llama.cpp are still too heavy. For developers building agents on constrained hardware — IoT devices, mobile, edge servers — Needle 2 offers a practical path to structured function-calling without cloud dependency. The model's open release means it can be fine-tuned, quantized further, or integrated into custom inference pipelines. This is a meaningful data point in the ongoing miniaturization of capable AI: tool-calling, once a capability exclusive to 7B+ models, is now feasible at sub-100M parameter scales.
Read original source ↗Part of the 2026-08-15 briefing→