Instructions to use Caldalis/MiniCPM5-2B-Math with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Caldalis/MiniCPM5-2B-Math with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Caldalis/MiniCPM5-2B-Math") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Caldalis/MiniCPM5-2B-Math") model = AutoModelForCausalLM.from_pretrained("Caldalis/MiniCPM5-2B-Math", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Caldalis/MiniCPM5-2B-Math with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Caldalis/MiniCPM5-2B-Math" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Caldalis/MiniCPM5-2B-Math", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Caldalis/MiniCPM5-2B-Math
- SGLang
How to use Caldalis/MiniCPM5-2B-Math with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Caldalis/MiniCPM5-2B-Math" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Caldalis/MiniCPM5-2B-Math", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Caldalis/MiniCPM5-2B-Math" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Caldalis/MiniCPM5-2B-Math", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Caldalis/MiniCPM5-2B-Math with Docker Model Runner:
docker model run hf.co/Caldalis/MiniCPM5-2B-Math
MiniCPM-Math Harness | GGUF | MiniCPM5-2B | MiniCPM GitHub
English | 䏿–‡
Highlights
MiniCPM5-2B-Math is a math and proof model built on MiniCPM5: a dense 2.5B model based on MiniCPM5-2B and post-trained for long-form mathematical reasoning and proof writing. It runs on a single consumer GPU or a laptop and keeps the native 128K context.
- Strong for its size. As a plain single call with up to 128K tokens of thinking, 92.3 on AIME 2025, 89.8 on AIME 2026 and 81.8 on HMMT February 2026 (avg@16), and 100 / 93.3 / 90.9 with a 16-sample vote. Overall, it outperforms the original general-purpose MiniCPM5-2B on AIME and HMMT.
- Writes checkable proofs. On IMO-ProofBench Basic, an average single call with up to 128K tokens writes a complete proof for 21 of the 30 problems, and the proof swarm for 23, as graded blind by an AI grader.
- Comes with a proof swarm. The open-source MiniCPM-Math Harness runs parallel provers, adversarial verifiers, brute-force checks and repair loops around this model on any OpenAI-compatible server. It is designed around the failure modes we measured.
- Small enough to own. BF16 weights are 5.0 GB; GGUF Q4_K_M is 1.6 GB. Standard
LlamaForCausalLMarchitecture, so no custom code is needed.
Evaluation
AIME and HMMT
| Benchmark | Metric | MiniCPM5-2B | MiniCPM5-2B-Math |
|---|---|---|---|
| AIME 2025 | avg@16 | 87.1 | 92.3 |
| vote (maj@16) | 93.3 | 100.0 | |
| AIME 2026 | avg@16 | 89.8 | 89.8 |
| vote (maj@16) | 93.3 | 93.3 | |
| HMMT Feb 2026 | avg@16 | 68.4 | 81.8 |
| vote (maj@16) | 75.8 | 90.9 |
avg@16 is the mean accuracy of single samples, i.e. of one plain API call; vote is the majority answer of the 16
samples (reproduce with the harness:
minicpm-math solve -n 16 --max-tokens 131072 --no-early-stop --no-force-finish).
Proofs: IMO-ProofBench Basic
All 30 problems. Single call: 16 independent calls per problem with up to 128K tokens each. Proof swarm: the
harness' standard profile with 128K-token provers and refiners. Every proof and grade is
published.
| Complete proofs | Mean score (0–7) | |
|---|---|---|
| single call (avg@16) | 21.0 / 30 | 5.48 |
| proof swarm | 23 / 30 | 5.83 |
Use it with the proof swarm
pip install "minicpm-math[math] @ git+https://github.com/Caldalis/MiniCPM-Math-harness"
vllm serve Caldalis/MiniCPM5-2B-Math --max-model-len 131072 --port 8000
minicpm-math prove "Prove that for every integer n, n^5 - n is divisible by 30." --base-url http://127.0.0.1:8000/v1
The proof-swarm numbers above used --max-tokens 131072; the default profile gives each call 32K tokens, which is
faster.
The swarm returns the best proof together with how it was graded (accepted only if every verifier gives at
least 6/7), and keeps a full trace of every call. See the harness README
for profiles, answer-mode voting, and benchmarking.
Quickstart
Recommended sampling:
temperature=0.9, top_p=0.95, min_p=0.0, with thinking enabled. Proofs for hard problems can need 64K–120K output tokens. Give the server enough context (--max-model-len 131072) and the request a largemax_tokens.
vLLM
pip install "vllm>=0.21"
vllm serve Caldalis/MiniCPM5-2B-Math --max-model-len 131072 --port 8000
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Caldalis/MiniCPM5-2B-Math",
"messages": [{"role": "user", "content": "Prove that the square root of 2 is irrational."}],
"max_tokens": 16384, "temperature": 0.9, "top_p": 0.95,
"chat_template_kwargs": {"enable_thinking": true}
}'
Measured with vLLM 0.24.0 on RTX 4090s: about 155 tok/s single-stream decode on one GPU with short inputs. On long proof runs across five 4090s with up to 10 concurrent requests, the median request decoded at about 118 tok/s, about 900 tok/s in total.
SGLang
pip install "sglang[srt]>=0.5.16"
python -m sglang.launch_server --model-path Caldalis/MiniCPM5-2B-Math --context-length 131072 --port 30000
llama.cpp
llama-server -hf Caldalis/MiniCPM5-2B-Math-GGUF:Q8_0 --jinja -ngl 99 -c 98304 -np 4 -kvu --port 8000
Port 8000 is where the harness looks by default (llama-server alone listens on 8080); with this server, run the
harness with --profile quick. On an Apple M5 laptop (16 GB) with Q8_0: about 46 tok/s for one stream, about
125 tok/s total with 4 slots. llama.cpp's default min_p=0.05 can encourage repetition, so send min_p=0.0 (the
harness does this for you).
Transformers
pip install -U "transformers>=5.6" accelerate torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Caldalis/MiniCPM5-2B-Math"
tokenizer = AutoTokenizer.from_pretrained(model_id)
# On Apple Silicon (MPS), also pass attn_implementation="eager": PyTorch's MPS SDPA kernel (checked with
# torch 2.11) mishandles this model's grouped-query attention and degrades the output. CUDA and CPU are fine.
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Find all positive integers n such that n^2 - 19n + 99 is a perfect square."}]
inputs = tokenizer.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True, enable_thinking=True,
return_dict=True, return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=8192, do_sample=True, temperature=0.9, top_p=0.95)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Prompting tips
- For problems with a final answer, end the prompt with "Please reason step by step, and put your final answer within \boxed{}."
- For proofs, ask explicitly for a complete and rigorous proof, and to prove both directions when the problem asks you to "find all" or "determine".
- With thinking enabled, the reasoning ends with
</think>, followed by the final answer. Servers with a reasoning parser return the two parts separately; otherwise split on</think>.
License
The model weights are released under the Apache-2.0 license (see
LICENSE), the same license as the base model,
MiniCPM5-2B.
- Downloads last month
- 47