Cactus Compute logo

Cactus Needle 2

Needlev2Current
byCactus ComputeCactus Compute(startup)
Released August 11, 2026
Context256 tokens
Price (In / Out)Free / Free
CategoryLarge Language Model
Max Output256 tokens

About Cactus Needle 2

Cactus Needle 2 is a 45M-parameter, 14 MB agentic model for tiny devices. It focuses on tool calling, device control, schema-constrained JSON, and structured extraction; Cactus reports 28 MB session RAM, 800+ tok/s prefill, and 500+ tok/s decode on Raspberry Pi 5.

Capabilities

tool usestructured outputjson schemadevice controledge inferenceoffline

Input Modalities

text

Output Modalities

textjson

Technical Details

API Identifier
Cactus-Compute/needle2
Category
Large Language Model
Context Window
256 tokens
Max Output Tokens
256 tokens

Tags

edge-llmtool-callingstructured-outputtiny-modelon-device-aiapache-2

Benchmarks

Performance scores for Cactus Needle 2 across standard benchmarks.

Mobile Actions Ordered Strict Exact Matchcactus-compute
63.7%
BFCL v4 Single-Turn Overallcactus-compute
42.6%

Pricing

Token pricing for Cactus Needle 2 API usage.

Input Tokens

Free

per million tokens

Output Tokens

Free

per million tokens

Pricing Calculator

Input cost$0.00
Output cost$0.00
Estimated monthly cost$0.00

Apache 2.0 model artifact; runtime cost depends on local hardware, not token API pricing.

Competing Models

Same pricing tier — direct alternatives to Cactus Needle 2

Efficiency
Cactus Compute logo

On-device AI with cloud fallback

6 MCP Servers2 ModelsFounded 2023San Francisco, United States
View full profile

AI Models by Cactus Compute

Large language models from the same organization.

ModelContext WindowPrice (In / Out per M)
NeedleCurrent--Free / Free

MCP Servers by Cactus Compute

Connect this tool to AI assistants via the Model Context Protocol.