Context--
CategoryLarge Language Model
About Cactus Needle 2
Cactus Needle 2 is a 45M-parameter, 14 MB agentic model for tiny devices. It focuses on tool calling, device control, schema-constrained JSON, and structured extraction; Cactus reports 28 MB session RAM, 800+ tok/s prefill, and 500+ tok/s decode on Raspberry Pi 5. Its 256-token sliding attention window is not a total context or output limit; those ceilings are not established by the model documentation.
Capabilities
tool usestructured outputjson schemadevice controledge inferenceoffline
Input Modalities
text
Output Modalities
text
Technical Details
- API Identifier
- Cactus-Compute/needle2
- Category
- Large Language Model
Tags
edge-llmtool-callingstructured-outputtiny-modelon-device-aiapache-2
Benchmarks
Performance scores for Cactus Needle 2 across standard benchmarks.
Mobile Actions Ordered Strict Exact Matchcactus-compute
63.7%
BFCL v4 Single-Turn Overallcactus-compute
42.6%
Pricing and access
Apache 2.0 model artifact; runtime cost depends on local hardware, not token API pricing.
Cactus Computestartup
On-device AI with cloud fallback
6 MCP Servers2 ModelsFounded 2023San Francisco, United States
View full profileAI Models by Cactus Compute
AI models from the same organization.
| Model | Context Window | Price (In / Out per M) |
|---|---|---|
| NeedleCurrent | -- | -- / -- |
MCP Servers by Cactus Compute
Connect this tool to AI assistants via the Model Context Protocol.