OpenBMB's latest 2.6 billion parameter dense reasoning model is available as open source under the Apache 2.0 license. The software outperformed the 3 billion parameter Granite 4.2 by 4 points on the Artificial Analysis Intelligence Index v4.2, where it matched the performance of Qwen3.5 9B. MiniCPM5-2B supports a 131,000 token context window, accepts text-only input, and offers a non-hallucination rate of 78% on the AA-Omniscience test by abstaining from 71% of questions.

The release includes the full data pipeline—covering web, code, and RL components—and training recipes. Serving support launched on day one via vLLM, which includes tool calling support, and SGLang, which showed decode speeds over 250 tokens per second per user on Nvidia RTX 5090 cards for coding tasks. The model scored 20 on the Agentic Index, 10 times higher than the score of 2 recorded for Mistral 3 3B or LFM2.5-2.6B.

Sign in to suggest edits

Key sources

  1. SOURCE@openbmb“average score of 53.9, covering coding, math, long-context understanding, tool use, and agentic tasks”x.com
  2. SUPPORT@artificialanlys“Non-Hallucination Rate of 78%”x.com
  3. SUPPORT@vllm_project“Tool Calling support via vLLM’s minicpm5 parser”x.com
  4. SUPPORT@sgl_project“over 250 tok/s/user decode at bs=1 on coding tasks with DSpark enabled on a 5090”x.com
  5. SOURCE@openbmb“On the Agentic Index, it scores 20, compared with 2 for LFM2.5-2.6B, Granite 4.2 3B, and Mistral 3 3B”x.com
  6. SUPPORT@artificialanlys“MiniCPM5-2B scored 23 on the recently updated Artificial Analysis Intelligence Index v4.1.1”x.com
Markdown