Note
SGLang upstream does not yet ship MiniCPM 4 support, so we maintain it on the tc-mb/sglang fork under the minicpm branch. Same branch as MiniCPM 4.1 — pull once and you can swap between 4 and 4.1 freely. (SALA lives on a separate minicpm_sala branch.)
git clone https://github.com/tc-mb/sglang.git
cd sglang
git checkout minicpm
pip install --upgrade pip
pip install -e "python[all]"# 8B
python -m sglang.launch_server \
--model-path openbmb/MiniCPM4-8B \
--port 30000 --trust-remote-code --dtype bfloat16 --context-length 32768
# 0.5B (edge-friendly)
python -m sglang.launch_server \
--model-path openbmb/MiniCPM4-0.5B \
--port 30000 --trust-remote-code --dtype bfloat16 --context-length 32768OpenAI-compatible API on http://localhost:30000/v1.
from openai import OpenAI
client = OpenAI(api_key="EMPTY", base_url="http://localhost:30000/v1")
resp = client.chat.completions.create(
model="openbmb/MiniCPM4-8B",
messages=[{"role": "user", "content": "Write an article about Artificial Intelligence."}],
temperature=0.7, top_p=0.8, max_tokens=512,
)
print(resp.choices[0].message.content)- Branch in use:
tc-mb/sglang @ minicpm. - MiniCPM 4 does not support the
enable_thinkingtoggle. For hybrid reasoning use MiniCPM 4.1. - For very long inputs see MiniCPM-SALA.