Context128K tokens
Training CutoffNot publicly specified in queued source
CategoryLarge Language Model
Max Output32.8K tokens
About MiniCPM
MiniCPM is OpenBMB’s ultra-efficient open language-model family for edge and end-device deployment. The MiniCPM4 and MiniCPM4.1 lines focus on fast local reasoning, while MiniCPM-SALA extends the family toward sparse/linear attention and million-token context research.
Capabilities
textcodereasoninglocal inference
Input Modalities
text
Output Modalities
text
Technical Details
- API Identifier
- OpenBMB/MiniCPM
- Category
- Large Language Model
- Context Window
- 128,000 tokens
- Max Output Tokens
- 32,768 tokens
Tags
open-weightedge-ailocal-llmreasoningopenbmb
Benchmarks
Performance scores for MiniCPM across standard benchmarks.
MiniCPM-SALA standard benchmark averageofficial-github-readme · May 2026
76.5
MiniCPM-SALA long-context averageofficial-github-readme · May 2026
39
MiniCPM-SALA 2048K extrapolation scoreofficial-github-readme · May 2026
81.6
MiniCPM4.1 reasoning decoding speedupofficial-github-readme · May 2026
3
MiniCPM4 Jetson AGX Orin decoding speedup vs Qwen3-8Bofficial-github-readme · May 2026
7
Pricing and access
Open-weight GitHub and Hugging Face model family. There is no fixed vendor API price; runtime cost depends on the host, hardware, or inference provider.
OpenBMBopen source
Open-source efficient AI models and agent infrastructure
1 ModelsFounded 2022China
View full profileAI Models by OpenBMB
AI models from the same organization.
| Model | Context Window | Price (In / Out per M) |
|---|---|---|
| VoxCPMCurrent | 4K | -- / -- |