mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
fix(chat): only prefill past reasoning the model cannot decline
The prefill landed on every model whose template ignores enable_thinking. On-device that made
LFM2.5 strictly worse: handed a pre-closed empty reasoning span it reasons straight past it and
emits the reasoning UNTAGGED into the answer, where before it at least parsed into a collapsible
block. 500 tokens without reaching an answer.
The two families were never the same case. gpt-oss declares no reasoning tags because reasoning is
a channel the format separates structurally, so starting the turn past it is not something the
model can decline. LFM2.5 declares <think>/</think>: the span is the model's own to open and close,
so handing it an empty closed one is a suggestion, and this model was not trained on that
convention. Liquid ships a separate non-reasoning checkpoint rather than an off switch.
So the probe now reads that distinction off the tags the model itself declares — no model names,
no template string matching. Tags declared and the flag inert means the request cannot be honoured:
report none, disable the control, say why. No tags means the prefill binds.
Measured on-device, three models:
LFM2.5-8B-A1B none reasoning tagged into reasoning_content, answer clean ("4")
gpt-oss-120b prefill no reasoning, direct answer (the retired hardcoded path's behaviour)
Qwen3.6-35B-A3B template unchanged, no reasoning, direct answer
This commit is contained in:
parent
aa6a7fdafd
commit
4f5d4c2246
6 changed files with 69 additions and 30 deletions
17
docs/seam.md
17
docs/seam.md
|
|
@ -93,12 +93,17 @@ that seam unit-testable without a model.
|
|||
A second translation unit, `thinking_control.cpp`, crosses the same boundary for "thinking off".
|
||||
`enable_thinking` is only a *request* to the template, and many templates never read it, so the
|
||||
engine renders the template to find out (three renders at open, no model names involved) and, where
|
||||
the flag is inert, asks for a **continuation** instead: the `continue_final_message` field of
|
||||
`common_chat_templates_inputs`, plus a synthetic trailing assistant message, makes llama.cpp's own
|
||||
per-template handler emit that family's "reasoning is over" span into the prompt. This is why no `<think>` or
|
||||
harmony channel marker appears anywhere in `core/` — the markers stay upstream, where a submodule
|
||||
bump keeps them current. `tests/think_control_test.cpp` pins the behaviour against the vendored
|
||||
templates, again with no model.
|
||||
the flag is inert *and* reasoning is a structural section of the format, asks for a **continuation**
|
||||
instead: the `continue_final_message` field of `common_chat_templates_inputs`, plus a synthetic
|
||||
trailing assistant message, makes llama.cpp's own per-template handler emit that family's "reasoning
|
||||
is over" span into the prompt. This is why no `<think>` or harmony channel marker appears anywhere in
|
||||
`core/` — the markers stay upstream, where a submodule bump keeps them current.
|
||||
|
||||
Whether the continuation is *binding* is read off `common_chat_params::thinking_start_tag`/
|
||||
`thinking_end_tag`: a model that declares a reasoning span owns it, so a pre-closed empty one is a
|
||||
suggestion it can decline (LFM2.5 does), while a model that declares none separates reasoning
|
||||
structurally and cannot. Both facts come from the loaded model, never from its name.
|
||||
`tests/think_control_test.cpp` pins all of it against the vendored templates, again with no model.
|
||||
|
||||
Unlike the public-C-API streaming seam, `common` is **not a stable API** — it can change
|
||||
between upstream versions. So a submodule bump may require updating this chat glue in
|
||||
|
|
|
|||
|
|
@ -271,8 +271,15 @@ model's own chat template, not from a list of model names:
|
|||
| value | meaning | what the UI should do |
|
||||
|---|---|---|
|
||||
| `template` | the chat template reads `enable_thinking` (Qwen3 and most reasoning models) | offer the toggle |
|
||||
| `prefill` | it does not, so the engine closes the reasoning span in the prompt instead (LFM2.5, gpt-oss) | offer the toggle |
|
||||
| `none` | neither is available — the model reasons on every turn | show the control disabled, and say why |
|
||||
| `prefill` | it does not, but reasoning is a structural section the prompt can start past (harmony/gpt-oss) | offer the toggle |
|
||||
| `none` | the model reasons on every turn and cannot be asked not to (LFM2.5) | show the control disabled, and say why |
|
||||
|
||||
The `prefill` / `none` split is decided by the reasoning tags the model declares, not by its name.
|
||||
A model that declares a `<think>`-style span owns that span: handing it one already closed and empty
|
||||
is a suggestion, and a model not trained on the convention reasons past it — measured on LFM2.5,
|
||||
which then emits its reasoning *untagged into the answer*, worse than not asking at all. A model
|
||||
that declares no tags separates reasoning structurally (a channel), and starting the turn past that
|
||||
section is not something it can decline.
|
||||
|
||||
`BMOE_DONE` carries the end-of-generation summary (the one-shot mode's `generation:` /
|
||||
`moe-stream:` text lines are not emitted in session mode). `n_prompt` is the tokens actually
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue