Needle 2
14MB on-device tool-calling model
AI Agents / Orchestration · OSS/MIT · 10.1k
What's Good
Cactus Compute built a tool-calling model that fits in 14MB. Runs entirely on-device—no API calls, no latency, no costs per call. Purpose-built for structured function calling, not general chat. The small footprint makes it embeddable anywhere.
The Catch
Not a chat LLM. Needle does one thing: parse user intent into tool calls. You still need the tools and the orchestration around it. Do not expect conversation or reasoning.
Verdict
Tiny on-device tool-caller. One job, done locally.
Embed
[](https://stackgems.com/gems/needle)