Cactus Compute logo

Cactus Needle 2

Needlev2Current
byCactus ComputeCactus Compute(startup)
Released August 11, 2026
Context--
CategoryLarge Language Model

About Cactus Needle 2

Cactus Needle 2 is a 45M-parameter, 14 MB agentic model for tiny devices. It focuses on tool calling, device control, schema-constrained JSON, and structured extraction; Cactus reports 28 MB session RAM, 800+ tok/s prefill, and 500+ tok/s decode on Raspberry Pi 5. Its 256-token sliding attention window is not a total context or output limit; those ceilings are not established by the model documentation.

Capabilities

tool usestructured outputjson schemadevice controledge inferenceoffline

Input Modalities

text

Output Modalities

text

Technical Details

API Identifier
Cactus-Compute/needle2
Category
Large Language Model

Tags

edge-llmtool-callingstructured-outputtiny-modelon-device-aiapache-2

Benchmarks

Performance scores for Cactus Needle 2 across standard benchmarks.

Mobile Actions Ordered Strict Exact Matchcactus-compute
63.7%
BFCL v4 Single-Turn Overallcactus-compute
42.6%

Pricing and access

Apache 2.0 model artifact; runtime cost depends on local hardware, not token API pricing.

Cactus Compute logo

On-device AI with cloud fallback

6 MCP Servers2 ModelsFounded 2023San Francisco, United States
View full profile

AI Models by Cactus Compute

AI models from the same organization.

ModelContext WindowPrice (In / Out per M)
NeedleCurrent---- / --

MCP Servers by Cactus Compute

Connect this tool to AI assistants via the Model Context Protocol.