fix(chat): only prefill past reasoning the model cannot decline

The prefill landed on every model whose template ignores enable_thinking. On-device that made
LFM2.5 strictly worse: handed a pre-closed empty reasoning span it reasons straight past it and
emits the reasoning UNTAGGED into the answer, where before it at least parsed into a collapsible
block. 500 tokens without reaching an answer.

The two families were never the same case. gpt-oss declares no reasoning tags because reasoning is
a channel the format separates structurally, so starting the turn past it is not something the
model can decline. LFM2.5 declares <think>/</think>: the span is the model's own to open and close,
so handing it an empty closed one is a suggestion, and this model was not trained on that
convention. Liquid ships a separate non-reasoning checkpoint rather than an off switch.

So the probe now reads that distinction off the tags the model itself declares — no model names,
no template string matching. Tags declared and the flag inert means the request cannot be honoured:
report none, disable the control, say why. No tags means the prefill binds.

Measured on-device, three models:
  LFM2.5-8B-A1B   none      reasoning tagged into reasoning_content, answer clean ("4")
  gpt-oss-120b    prefill   no reasoning, direct answer (the retired hardcoded path's behaviour)
  Qwen3.6-35B-A3B template   unchanged, no reasoning, direct answer
This commit is contained in:
Helldez 2026-07-19 21:08:29 +02:00
parent aa6a7fdafd
commit 4f5d4c2246
6 changed files with 69 additions and 30 deletions

View file

@ -11,12 +11,14 @@ Semantic Versioning.
variable `enable_thinking` and stopped there — but that variable is only a *request* to the variable `enable_thinking` and stopped there — but that variable is only a *request* to the
model's chat template, and many templates never read it. LFM2.5's is one: the rendered prompt came model's chat template, and many templates never read it. LFM2.5's is one: the rendered prompt came
out byte-identical either way, the model reasoned on, and nothing reported that the setting had out byte-identical either way, the model reasoned on, and nothing reported that the setting had
been dropped. The engine now renders the template at load to find out which case a model is in, been dropped. The engine now renders the template at load and reports what it found as `think_ctl`
and where the flag is inert it closes the reasoning span in the prompt instead, using llama.cpp's on `BMOE_READY` (see `docs/telemetry.md`). Where the flag is inert but reasoning is a *structural*
own continuation hook so the markers for each family come from upstream's per-template handler section of the format, the turn now starts past that section, built by llama.cpp's own
rather than from this engine. Models that support neither are reported as such continuation hook so every family's markers come from upstream rather than from this engine.
(`think_ctl` on `BMOE_READY`, see `docs/telemetry.md`) and the app shows the Thinking switch Where the model owns its reasoning span and simply cannot be asked to skip it — LFM2.5 — that is
disabled with the reason, instead of leaving a control that does nothing. reported instead of papered over, and the app shows the Thinking switch disabled with the reason.
Measured, not assumed: handing LFM2.5 a pre-closed empty reasoning span makes it reason *untagged
into the answer*, worse than leaving the setting alone, so the engine does not do it.
### Changed ### Changed
- **The harmony/gpt-oss marker strings are gone from the decode path.** Priming gpt-oss to answer - **The harmony/gpt-oss marker strings are gone from the decode path.** Priming gpt-oss to answer

View file

@ -62,14 +62,29 @@ void add_no_think_prefill(common_chat_templates_inputs & inputs) {
ThinkControl probe_think_control(const common_chat_templates * tmpls) { ThinkControl probe_think_control(const common_chat_templates * tmpls) {
if (tmpls == nullptr) return ThinkControl::Template; if (tmpls == nullptr) return ThinkControl::Template;
try { try {
const std::string off = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/false).prompt; const common_chat_params off_p = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/false);
const std::string & off = off_p.prompt;
const std::string on = apply_probe(tmpls, /*enable_thinking=*/true, /*prefill=*/false).prompt; const std::string on = apply_probe(tmpls, /*enable_thinking=*/true, /*prefill=*/false).prompt;
if (on != off) return ThinkControl::Template; if (on != off) return ThinkControl::Template;
const std::string prefilled = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/true).prompt; const std::string prefilled = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/true).prompt;
if (prefilled != off) return ThinkControl::Prefill; if (prefilled == off) return ThinkControl::None; // nothing reaches this model at all
return ThinkControl::None; // The prefill changes the prompt — but that alone does not mean the model will honour it,
// and the difference is visible in whether the model declares reasoning tags.
//
// Tags declared: reasoning is a span the MODEL opens and closes at will. Handing it one that
// is already closed and empty is a suggestion, and a model not trained on that convention
// reasons straight past it — measured on LFM2.5-8B-A1B, which ignores the closed span and
// reasons into the answer instead (issue #82). Worse than leaving it alone, because the
// reasoning arrives untagged. Report it as uncontrollable rather than making it worse.
//
// No tags declared: reasoning is structural — a channel or section the format itself
// separates — so the prefill does not ask the model to skip anything, it places the model
// past the reasoning section entirely. That the model cannot ignore (harmony/gpt-oss).
if (!off_p.thinking_start_tag.empty() || !off_p.thinking_end_tag.empty()) return ThinkControl::None;
return ThinkControl::Prefill;
} catch (const std::exception & e) { } catch (const std::exception & e) {
std::fprintf(stderr, "bmoe: thinking-control probe failed (%s); assuming the template honours it\n", e.what()); std::fprintf(stderr, "bmoe: thinking-control probe failed (%s); assuming the template honours it\n", e.what());
return ThinkControl::Template; return ThinkControl::Template;

View file

@ -38,10 +38,16 @@ void add_no_think_prefill(common_chat_templates_inputs & inputs);
// analyzer performs; `common_chat_params::supports_thinking` cannot answer the question, // analyzer performs; `common_chat_params::supports_thinking` cannot answer the question,
// because per handler it is a hardcoded literal reporting "this model can reason", not // because per handler it is a hardcoded literal reporting "this model can reason", not
// "this template reads the variable". // "this template reads the variable".
// 2. with the prefill applied — if that changes the prompt, the handler implements the // 2. with the prefill applied — if that leaves the prompt untouched, nothing reaches this model
// continuation hook and the span can be closed ahead of generation. // and the request cannot be honoured.
// Neither: the request cannot be honoured on this model, and saying so is better than offering a // 3. otherwise the prefill lands, and whether the model must OBEY it is read off the reasoning
// control that does nothing. // tags it declares: a declared span is the model's own to open and close, so a pre-closed
// empty one is only a suggestion (LFM2.5 ignores it — see probe_think_control); no declared
// span means the format separates reasoning structurally and the prefill places the model
// past it, which it cannot ignore.
//
// Saying "this model cannot be silenced" is better than offering a control that does nothing — and
// far better than one that makes the output worse.
// //
// Fails open to Template (the pre-existing behaviour: pass the flag and let the template decide), // Fails open to Template (the pre-existing behaviour: pass the flag and let the template decide),
// because a probe that itself failed is no evidence that the flag is inert. // because a probe that itself failed is no evidence that the flag is inert.

View file

@ -93,12 +93,17 @@ that seam unit-testable without a model.
A second translation unit, `thinking_control.cpp`, crosses the same boundary for "thinking off". A second translation unit, `thinking_control.cpp`, crosses the same boundary for "thinking off".
`enable_thinking` is only a *request* to the template, and many templates never read it, so the `enable_thinking` is only a *request* to the template, and many templates never read it, so the
engine renders the template to find out (three renders at open, no model names involved) and, where engine renders the template to find out (three renders at open, no model names involved) and, where
the flag is inert, asks for a **continuation** instead: the `continue_final_message` field of the flag is inert *and* reasoning is a structural section of the format, asks for a **continuation**
`common_chat_templates_inputs`, plus a synthetic trailing assistant message, makes llama.cpp's own instead: the `continue_final_message` field of `common_chat_templates_inputs`, plus a synthetic
per-template handler emit that family's "reasoning is over" span into the prompt. This is why no `<think>` or trailing assistant message, makes llama.cpp's own per-template handler emit that family's "reasoning
harmony channel marker appears anywhere in `core/` — the markers stay upstream, where a submodule is over" span into the prompt. This is why no `<think>` or harmony channel marker appears anywhere in
bump keeps them current. `tests/think_control_test.cpp` pins the behaviour against the vendored `core/` — the markers stay upstream, where a submodule bump keeps them current.
templates, again with no model.
Whether the continuation is *binding* is read off `common_chat_params::thinking_start_tag`/
`thinking_end_tag`: a model that declares a reasoning span owns it, so a pre-closed empty one is a
suggestion it can decline (LFM2.5 does), while a model that declares none separates reasoning
structurally and cannot. Both facts come from the loaded model, never from its name.
`tests/think_control_test.cpp` pins all of it against the vendored templates, again with no model.
Unlike the public-C-API streaming seam, `common` is **not a stable API** — it can change Unlike the public-C-API streaming seam, `common` is **not a stable API** — it can change
between upstream versions. So a submodule bump may require updating this chat glue in between upstream versions. So a submodule bump may require updating this chat glue in

View file

@ -271,8 +271,15 @@ model's own chat template, not from a list of model names:
| value | meaning | what the UI should do | | value | meaning | what the UI should do |
|---|---|---| |---|---|---|
| `template` | the chat template reads `enable_thinking` (Qwen3 and most reasoning models) | offer the toggle | | `template` | the chat template reads `enable_thinking` (Qwen3 and most reasoning models) | offer the toggle |
| `prefill` | it does not, so the engine closes the reasoning span in the prompt instead (LFM2.5, gpt-oss) | offer the toggle | | `prefill` | it does not, but reasoning is a structural section the prompt can start past (harmony/gpt-oss) | offer the toggle |
| `none` | neither is available — the model reasons on every turn | show the control disabled, and say why | | `none` | the model reasons on every turn and cannot be asked not to (LFM2.5) | show the control disabled, and say why |
The `prefill` / `none` split is decided by the reasoning tags the model declares, not by its name.
A model that declares a `<think>`-style span owns that span: handing it one already closed and empty
is a suggestion, and a model not trained on the convention reasons past it — measured on LFM2.5,
which then emits its reasoning *untagged into the answer*, worse than not asking at all. A model
that declares no tags separates reasoning structurally (a channel), and starting the turn past that
section is not something it can decline.
`BMOE_DONE` carries the end-of-generation summary (the one-shot mode's `generation:` / `BMOE_DONE` carries the end-of-generation summary (the one-shot mode's `generation:` /
`moe-stream:` text lines are not emitted in session mode). `n_prompt` is the tokens actually `moe-stream:` text lines are not emitted in session mode). `n_prompt` is the tokens actually

View file

@ -110,21 +110,25 @@ int main() {
} }
// LFM2.5 — the model from issue #82. Its template never mentions enable_thinking, so both // LFM2.5 — the model from issue #82. Its template never mentions enable_thinking, so both
// renders are identical; the handler does implement the continuation hook, so the span can // renders are identical. The continuation hook does land (the span below really is closed),
// be closed in the prompt. // but LFM2.5 declares <think>/</think>, so the closed span is only a suggestion to a model
// that owns its own reasoning tags — and measured on-device it reasons straight past it,
// untagged, into the answer. Reported as uncontrollable rather than made worse.
{ {
auto t = load(BMOE_TMPL_LFM25); auto t = load(BMOE_TMPL_LFM25);
expect("lfm2.5: flag is inert in the template", render(t.get(), /*think*/ true, /*prefill*/ false) == expect("lfm2.5: flag is inert in the template", render(t.get(), /*think*/ true, /*prefill*/ false) ==
render(t.get(), /*think*/ false, /*prefill*/ false)); render(t.get(), /*think*/ false, /*prefill*/ false));
expect("lfm2.5: probed as prefill", expect("lfm2.5: probed as none (declares its own reasoning tags)",
bmoe::detail::probe_think_control(t.get()) == bmoe::ThinkControl::Prefill); bmoe::detail::probe_think_control(t.get()) == bmoe::ThinkControl::None);
expect_span_closed("lfm2.5: prefilled span is closed", t.get(), "</think>"); // The prefill itself is well-formed — this is why "none" is a statement about the
// MODEL, not about a mechanism that failed to render.
expect_span_closed("lfm2.5: the prefill would close the span correctly", t.get(), "</think>");
} }
// gpt-oss / harmony. Its template ignores the flag too, and it exposes no thinking tags at // gpt-oss / harmony. Its template ignores the flag too, and it declares NO thinking tags:
// all — a tag-based detector would give up here. Probing the continuation instead sees the // reasoning is a channel the format separates structurally, so priming past it is not
// mechanism that does exist: the handler primes the answer channel. This is the case the // something the model can decline. This is the case the engine used to carry as two
// engine used to carry as two hardcoded harmony marker strings in the decode path. // hardcoded harmony marker strings in the decode path.
{ {
auto t = load(BMOE_TMPL_GPTOSS); auto t = load(BMOE_TMPL_GPTOSS);
expect("gpt-oss: probed as prefill", expect("gpt-oss: probed as prefill",