Repository navigation
Conversation
openbmb/MiniCPM5-2B ships a different chat template from MiniCPM5-1B: it renders the reasoning of past assistant turns instead of dropping them, and falls back to an empty <think> block when a past turn carries no reasoning. Both models keep the same generation prompt, tool-call syntax and reasoning tags, so 2B already resolves to the MiniCPM5 chat format and needs no parser change. Add the template and cover the rendering difference from both sides, so neither model can drift into the other's behaviour unnoticed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Hi @cyxu0401, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Closing this. MiniCPM5-2B works with llama.cpp as-is: its tokenizer is byte-identical to 1B, and its chat template ships inside the GGUF, so no upstream change is required. This was only regression coverage, not a fix, and it isn't worth maintainer time right now. |
Follow-up to #23384 (MiniCPM5-1B), now that
openbmb/MiniCPM5-2B is public.
Adds
models/templates/openbmb-MiniCPM5-2B.jinja(fetched verbatim from theHF repo) and covers it in
tests/test-chat.cpp.2B's chat template is not the same as 1B's. The tool-call and reasoning
syntax is identical, so 2B is already picked up correctly by the existing
detection in
common/chat.cppand resolves to the same chat format. Whatdiffers is how past assistant turns are rendered:
<think>\n\n</think>block instead.Nothing covered either behaviour. Since the format is selected by substring
matching on the template text, it is easy to break by accident, so the tests
assert both halves: that 2B still resolves to the shared format, and that the
generation prompt diverges from 1B in the way described above. A negative
assertion was added on the 1B side so the two cannot silently converge.
This is test-only coverage plus a fixture — no runtime code is touched, and
models/templates/is not read at runtime.Independent of the tokenizer-note PR; either can land first.
Testing
test-chatpasses. I checked the new assertions actually execute bytemporarily inverting one and confirming it fails rather than passing
vacuously.