fix(chat): only prefill past reasoning the model cannot decline

The prefill landed on every model whose template ignores enable_thinking. On-device that made
LFM2.5 strictly worse: handed a pre-closed empty reasoning span it reasons straight past it and
emits the reasoning UNTAGGED into the answer, where before it at least parsed into a collapsible
block. 500 tokens without reaching an answer.

The two families were never the same case. gpt-oss declares no reasoning tags because reasoning is
a channel the format separates structurally, so starting the turn past it is not something the
model can decline. LFM2.5 declares <think>/</think>: the span is the model's own to open and close,
so handing it an empty closed one is a suggestion, and this model was not trained on that
convention. Liquid ships a separate non-reasoning checkpoint rather than an off switch.

So the probe now reads that distinction off the tags the model itself declares — no model names,
no template string matching. Tags declared and the flag inert means the request cannot be honoured:
report none, disable the control, say why. No tags means the prefill binds.

Measured on-device, three models:
  LFM2.5-8B-A1B   none      reasoning tagged into reasoning_content, answer clean ("4")
  gpt-oss-120b    prefill   no reasoning, direct answer (the retired hardcoded path's behaviour)
  Qwen3.6-35B-A3B template   unchanged, no reasoning, direct answer
This commit is contained in:
Helldez 2026-07-19 21:08:29 +02:00
parent aa6a7fdafd
commit 4f5d4c2246
6 changed files with 69 additions and 30 deletions

View file

@ -93,12 +93,17 @@ that seam unit-testable without a model.
A second translation unit, `thinking_control.cpp`, crosses the same boundary for "thinking off".
`enable_thinking` is only a *request* to the template, and many templates never read it, so the
engine renders the template to find out (three renders at open, no model names involved) and, where
the flag is inert, asks for a **continuation** instead: the `continue_final_message` field of
`common_chat_templates_inputs`, plus a synthetic trailing assistant message, makes llama.cpp's own
per-template handler emit that family's "reasoning is over" span into the prompt. This is why no `<think>` or
harmony channel marker appears anywhere in `core/` — the markers stay upstream, where a submodule
bump keeps them current. `tests/think_control_test.cpp` pins the behaviour against the vendored
templates, again with no model.
the flag is inert *and* reasoning is a structural section of the format, asks for a **continuation**
instead: the `continue_final_message` field of `common_chat_templates_inputs`, plus a synthetic
trailing assistant message, makes llama.cpp's own per-template handler emit that family's "reasoning
is over" span into the prompt. This is why no `<think>` or harmony channel marker appears anywhere in
`core/` — the markers stay upstream, where a submodule bump keeps them current.
Whether the continuation is *binding* is read off `common_chat_params::thinking_start_tag`/
`thinking_end_tag`: a model that declares a reasoning span owns it, so a pre-closed empty one is a
suggestion it can decline (LFM2.5 does), while a model that declares none separates reasoning
structurally and cannot. Both facts come from the loaded model, never from its name.
`tests/think_control_test.cpp` pins all of it against the vendored templates, again with no model.
Unlike the public-C-API streaming seam, `common` is **not a stable API** — it can change
between upstream versions. So a submodule bump may require updating this chat glue in

View file

@ -271,8 +271,15 @@ model's own chat template, not from a list of model names:
| value | meaning | what the UI should do |
|---|---|---|
| `template` | the chat template reads `enable_thinking` (Qwen3 and most reasoning models) | offer the toggle |
| `prefill` | it does not, so the engine closes the reasoning span in the prompt instead (LFM2.5, gpt-oss) | offer the toggle |
| `none` | neither is available — the model reasons on every turn | show the control disabled, and say why |
| `prefill` | it does not, but reasoning is a structural section the prompt can start past (harmony/gpt-oss) | offer the toggle |
| `none` | the model reasons on every turn and cannot be asked not to (LFM2.5) | show the control disabled, and say why |
The `prefill` / `none` split is decided by the reasoning tags the model declares, not by its name.
A model that declares a `<think>`-style span owns that span: handing it one already closed and empty
is a suggestion, and a model not trained on the convention reasons past it — measured on LFM2.5,
which then emits its reasoning *untagged into the answer*, worse than not asking at all. A model
that declares no tags separates reasoning structurally (a channel), and starting the turn past that
section is not something it can decline.
`BMOE_DONE` carries the end-of-generation summary (the one-shot mode's `generation:` /
`moe-stream:` text lines are not emitted in session mode). `n_prompt` is the tokens actually