Edge AI is undergoing a massive paradigm shift. While the industry focuses on 3B+ parameter models running on high-end NPUs, Cactus Compute introduces Needle 2: a 14MB agentic LLM specifically engineered for the 21 billion connected IoT devices lacking dedicated AI accelerators. This includes sub-$200 smartphones, microcontrollers, wearables, and VR headsets. By compressing 45 million parameters into a 2-bit architecture, Needle 2 delivers high-speed structured extraction and tool calling in 28MB of RAM.
Technical Architecture & Compute Efficiency
Unlike conventional transformers that demand massive memory bandwidth, Needle 2 leverages Simple Attention Networks. This architectural decision drastically reduces the computational load per token.
- Parameter Compression: 45M parameters compressed at 2-bit precision, resulting in a single 14MB binary.
- Memory Footprint: Operates entirely within a 28MB RAM session limit.
- FLOP Reduction: A conventional transformer of this width/depth requires 164 MFLOPs per token. Even aggressively shrunk transformers require 87 MFLOPs. Needle 2 processes a token in just 70 MFLOPs.
This 7x to 85x reduction in MFLOPs translates directly to preserved milliwatt-hours—a critical requirement for always-on smart home devices and battery-constrained wearables.
Real-World Performance Benchmarks
Needle 2's decode speeds outpace many local inference engines on hardware strictly bound by thermal and power limits. If you are building hardware applications, you can review our portfolio of high-performance computing projects to see how we integrate such technologies.
- Microcontrollers & SBCs: 500 tokens/sec decode speed on a Raspberry Pi 5.
- VR/AR Devices: 400 to 1,500 tokens/sec on Meta Quest 3S and Apple Vision Pro.
- Budget Smartphones: 300 to 700 tokens/sec on sub-$200 devices (e.g., Samsung A-Series).
Functional Capabilities: Tool Calling & Structured Extraction
Needle 2 is not designed for open-ended creative prose; it is a highly specialized agentic LLM built for routing, device control, and data parsing. It competes directly with models like LFM2.5 230M and Apple Foundation Model, despite being 5x to 70x smaller.
1. In-Place Schema Extraction
Developers can bypass standard conversational outputs by passing a JSON schema directly to the model. Needle 2 maps messy conversational text into strictly typed parameters. This allows the model to function as a hyper-efficient text classifier, summarization engine, or API payload generator without relying on free-range token decoding.
2. Local Fine-Tuning
Using the open-source Python package, teams can fine-tune Needle 2 on standard Macs or PCs. The automated data-generation pipeline requires only a few prompt-completion samples, adapting the model to proprietary tool vocabularies in minutes.
3. Hybrid Cloud Escalation
Every local response generated by Needle 2 includes a learned confidence score (based on the Cactus Hybrid technique). If the confidence falls below a specified threshold, the system automatically escalates the prompt to a larger cloud model (like DeepSeek-v4-Flash), ensuring enterprise-grade reliability without constant cloud compute costs.
Final Verdict: A New Standard for TinyML
Needle 2 proves that device-side intelligence does not require a 3-billion parameter model. By rigidly constraining the use-case to tool-calling and structured parameter mapping, Cactus has built one of the most efficient on-device LLMs on the market. For teams building native hardware integrations, you can explore our custom engineering services to properly architect your local AI pipelines. To test the model's capabilities firsthand, check out the official Cactus Needle 2 release and playground.