mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
fix(chat): only prefill past reasoning the model cannot decline
The prefill landed on every model whose template ignores enable_thinking. On-device that made
LFM2.5 strictly worse: handed a pre-closed empty reasoning span it reasons straight past it and
emits the reasoning UNTAGGED into the answer, where before it at least parsed into a collapsible
block. 500 tokens without reaching an answer.
The two families were never the same case. gpt-oss declares no reasoning tags because reasoning is
a channel the format separates structurally, so starting the turn past it is not something the
model can decline. LFM2.5 declares <think>/</think>: the span is the model's own to open and close,
so handing it an empty closed one is a suggestion, and this model was not trained on that
convention. Liquid ships a separate non-reasoning checkpoint rather than an off switch.
So the probe now reads that distinction off the tags the model itself declares — no model names,
no template string matching. Tags declared and the flag inert means the request cannot be honoured:
report none, disable the control, say why. No tags means the prefill binds.
Measured on-device, three models:
LFM2.5-8B-A1B none reasoning tagged into reasoning_content, answer clean ("4")
gpt-oss-120b prefill no reasoning, direct answer (the retired hardcoded path's behaviour)
Qwen3.6-35B-A3B template unchanged, no reasoning, direct answer
This commit is contained in:
parent
aa6a7fdafd
commit
4f5d4c2246
6 changed files with 69 additions and 30 deletions
14
CHANGELOG.md
14
CHANGELOG.md
|
|
@ -11,12 +11,14 @@ Semantic Versioning.
|
||||||
variable `enable_thinking` and stopped there — but that variable is only a *request* to the
|
variable `enable_thinking` and stopped there — but that variable is only a *request* to the
|
||||||
model's chat template, and many templates never read it. LFM2.5's is one: the rendered prompt came
|
model's chat template, and many templates never read it. LFM2.5's is one: the rendered prompt came
|
||||||
out byte-identical either way, the model reasoned on, and nothing reported that the setting had
|
out byte-identical either way, the model reasoned on, and nothing reported that the setting had
|
||||||
been dropped. The engine now renders the template at load to find out which case a model is in,
|
been dropped. The engine now renders the template at load and reports what it found as `think_ctl`
|
||||||
and where the flag is inert it closes the reasoning span in the prompt instead, using llama.cpp's
|
on `BMOE_READY` (see `docs/telemetry.md`). Where the flag is inert but reasoning is a *structural*
|
||||||
own continuation hook so the markers for each family come from upstream's per-template handler
|
section of the format, the turn now starts past that section, built by llama.cpp's own
|
||||||
rather than from this engine. Models that support neither are reported as such
|
continuation hook so every family's markers come from upstream rather than from this engine.
|
||||||
(`think_ctl` on `BMOE_READY`, see `docs/telemetry.md`) and the app shows the Thinking switch
|
Where the model owns its reasoning span and simply cannot be asked to skip it — LFM2.5 — that is
|
||||||
disabled with the reason, instead of leaving a control that does nothing.
|
reported instead of papered over, and the app shows the Thinking switch disabled with the reason.
|
||||||
|
Measured, not assumed: handing LFM2.5 a pre-closed empty reasoning span makes it reason *untagged
|
||||||
|
into the answer*, worse than leaving the setting alone, so the engine does not do it.
|
||||||
|
|
||||||
### Changed
|
### Changed
|
||||||
- **The harmony/gpt-oss marker strings are gone from the decode path.** Priming gpt-oss to answer
|
- **The harmony/gpt-oss marker strings are gone from the decode path.** Priming gpt-oss to answer
|
||||||
|
|
|
||||||
|
|
@ -62,14 +62,29 @@ void add_no_think_prefill(common_chat_templates_inputs & inputs) {
|
||||||
ThinkControl probe_think_control(const common_chat_templates * tmpls) {
|
ThinkControl probe_think_control(const common_chat_templates * tmpls) {
|
||||||
if (tmpls == nullptr) return ThinkControl::Template;
|
if (tmpls == nullptr) return ThinkControl::Template;
|
||||||
try {
|
try {
|
||||||
const std::string off = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/false).prompt;
|
const common_chat_params off_p = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/false);
|
||||||
|
const std::string & off = off_p.prompt;
|
||||||
const std::string on = apply_probe(tmpls, /*enable_thinking=*/true, /*prefill=*/false).prompt;
|
const std::string on = apply_probe(tmpls, /*enable_thinking=*/true, /*prefill=*/false).prompt;
|
||||||
if (on != off) return ThinkControl::Template;
|
if (on != off) return ThinkControl::Template;
|
||||||
|
|
||||||
const std::string prefilled = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/true).prompt;
|
const std::string prefilled = apply_probe(tmpls, /*enable_thinking=*/false, /*prefill=*/true).prompt;
|
||||||
if (prefilled != off) return ThinkControl::Prefill;
|
if (prefilled == off) return ThinkControl::None; // nothing reaches this model at all
|
||||||
|
|
||||||
return ThinkControl::None;
|
// The prefill changes the prompt — but that alone does not mean the model will honour it,
|
||||||
|
// and the difference is visible in whether the model declares reasoning tags.
|
||||||
|
//
|
||||||
|
// Tags declared: reasoning is a span the MODEL opens and closes at will. Handing it one that
|
||||||
|
// is already closed and empty is a suggestion, and a model not trained on that convention
|
||||||
|
// reasons straight past it — measured on LFM2.5-8B-A1B, which ignores the closed span and
|
||||||
|
// reasons into the answer instead (issue #82). Worse than leaving it alone, because the
|
||||||
|
// reasoning arrives untagged. Report it as uncontrollable rather than making it worse.
|
||||||
|
//
|
||||||
|
// No tags declared: reasoning is structural — a channel or section the format itself
|
||||||
|
// separates — so the prefill does not ask the model to skip anything, it places the model
|
||||||
|
// past the reasoning section entirely. That the model cannot ignore (harmony/gpt-oss).
|
||||||
|
if (!off_p.thinking_start_tag.empty() || !off_p.thinking_end_tag.empty()) return ThinkControl::None;
|
||||||
|
|
||||||
|
return ThinkControl::Prefill;
|
||||||
} catch (const std::exception & e) {
|
} catch (const std::exception & e) {
|
||||||
std::fprintf(stderr, "bmoe: thinking-control probe failed (%s); assuming the template honours it\n", e.what());
|
std::fprintf(stderr, "bmoe: thinking-control probe failed (%s); assuming the template honours it\n", e.what());
|
||||||
return ThinkControl::Template;
|
return ThinkControl::Template;
|
||||||
|
|
|
||||||
|
|
@ -38,10 +38,16 @@ void add_no_think_prefill(common_chat_templates_inputs & inputs);
|
||||||
// analyzer performs; `common_chat_params::supports_thinking` cannot answer the question,
|
// analyzer performs; `common_chat_params::supports_thinking` cannot answer the question,
|
||||||
// because per handler it is a hardcoded literal reporting "this model can reason", not
|
// because per handler it is a hardcoded literal reporting "this model can reason", not
|
||||||
// "this template reads the variable".
|
// "this template reads the variable".
|
||||||
// 2. with the prefill applied — if that changes the prompt, the handler implements the
|
// 2. with the prefill applied — if that leaves the prompt untouched, nothing reaches this model
|
||||||
// continuation hook and the span can be closed ahead of generation.
|
// and the request cannot be honoured.
|
||||||
// Neither: the request cannot be honoured on this model, and saying so is better than offering a
|
// 3. otherwise the prefill lands, and whether the model must OBEY it is read off the reasoning
|
||||||
// control that does nothing.
|
// tags it declares: a declared span is the model's own to open and close, so a pre-closed
|
||||||
|
// empty one is only a suggestion (LFM2.5 ignores it — see probe_think_control); no declared
|
||||||
|
// span means the format separates reasoning structurally and the prefill places the model
|
||||||
|
// past it, which it cannot ignore.
|
||||||
|
//
|
||||||
|
// Saying "this model cannot be silenced" is better than offering a control that does nothing — and
|
||||||
|
// far better than one that makes the output worse.
|
||||||
//
|
//
|
||||||
// Fails open to Template (the pre-existing behaviour: pass the flag and let the template decide),
|
// Fails open to Template (the pre-existing behaviour: pass the flag and let the template decide),
|
||||||
// because a probe that itself failed is no evidence that the flag is inert.
|
// because a probe that itself failed is no evidence that the flag is inert.
|
||||||
|
|
|
||||||
17
docs/seam.md
17
docs/seam.md
|
|
@ -93,12 +93,17 @@ that seam unit-testable without a model.
|
||||||
A second translation unit, `thinking_control.cpp`, crosses the same boundary for "thinking off".
|
A second translation unit, `thinking_control.cpp`, crosses the same boundary for "thinking off".
|
||||||
`enable_thinking` is only a *request* to the template, and many templates never read it, so the
|
`enable_thinking` is only a *request* to the template, and many templates never read it, so the
|
||||||
engine renders the template to find out (three renders at open, no model names involved) and, where
|
engine renders the template to find out (three renders at open, no model names involved) and, where
|
||||||
the flag is inert, asks for a **continuation** instead: the `continue_final_message` field of
|
the flag is inert *and* reasoning is a structural section of the format, asks for a **continuation**
|
||||||
`common_chat_templates_inputs`, plus a synthetic trailing assistant message, makes llama.cpp's own
|
instead: the `continue_final_message` field of `common_chat_templates_inputs`, plus a synthetic
|
||||||
per-template handler emit that family's "reasoning is over" span into the prompt. This is why no `<think>` or
|
trailing assistant message, makes llama.cpp's own per-template handler emit that family's "reasoning
|
||||||
harmony channel marker appears anywhere in `core/` — the markers stay upstream, where a submodule
|
is over" span into the prompt. This is why no `<think>` or harmony channel marker appears anywhere in
|
||||||
bump keeps them current. `tests/think_control_test.cpp` pins the behaviour against the vendored
|
`core/` — the markers stay upstream, where a submodule bump keeps them current.
|
||||||
templates, again with no model.
|
|
||||||
|
Whether the continuation is *binding* is read off `common_chat_params::thinking_start_tag`/
|
||||||
|
`thinking_end_tag`: a model that declares a reasoning span owns it, so a pre-closed empty one is a
|
||||||
|
suggestion it can decline (LFM2.5 does), while a model that declares none separates reasoning
|
||||||
|
structurally and cannot. Both facts come from the loaded model, never from its name.
|
||||||
|
`tests/think_control_test.cpp` pins all of it against the vendored templates, again with no model.
|
||||||
|
|
||||||
Unlike the public-C-API streaming seam, `common` is **not a stable API** — it can change
|
Unlike the public-C-API streaming seam, `common` is **not a stable API** — it can change
|
||||||
between upstream versions. So a submodule bump may require updating this chat glue in
|
between upstream versions. So a submodule bump may require updating this chat glue in
|
||||||
|
|
|
||||||
|
|
@ -271,8 +271,15 @@ model's own chat template, not from a list of model names:
|
||||||
| value | meaning | what the UI should do |
|
| value | meaning | what the UI should do |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `template` | the chat template reads `enable_thinking` (Qwen3 and most reasoning models) | offer the toggle |
|
| `template` | the chat template reads `enable_thinking` (Qwen3 and most reasoning models) | offer the toggle |
|
||||||
| `prefill` | it does not, so the engine closes the reasoning span in the prompt instead (LFM2.5, gpt-oss) | offer the toggle |
|
| `prefill` | it does not, but reasoning is a structural section the prompt can start past (harmony/gpt-oss) | offer the toggle |
|
||||||
| `none` | neither is available — the model reasons on every turn | show the control disabled, and say why |
|
| `none` | the model reasons on every turn and cannot be asked not to (LFM2.5) | show the control disabled, and say why |
|
||||||
|
|
||||||
|
The `prefill` / `none` split is decided by the reasoning tags the model declares, not by its name.
|
||||||
|
A model that declares a `<think>`-style span owns that span: handing it one already closed and empty
|
||||||
|
is a suggestion, and a model not trained on the convention reasons past it — measured on LFM2.5,
|
||||||
|
which then emits its reasoning *untagged into the answer*, worse than not asking at all. A model
|
||||||
|
that declares no tags separates reasoning structurally (a channel), and starting the turn past that
|
||||||
|
section is not something it can decline.
|
||||||
|
|
||||||
`BMOE_DONE` carries the end-of-generation summary (the one-shot mode's `generation:` /
|
`BMOE_DONE` carries the end-of-generation summary (the one-shot mode's `generation:` /
|
||||||
`moe-stream:` text lines are not emitted in session mode). `n_prompt` is the tokens actually
|
`moe-stream:` text lines are not emitted in session mode). `n_prompt` is the tokens actually
|
||||||
|
|
|
||||||
|
|
@ -110,21 +110,25 @@ int main() {
|
||||||
}
|
}
|
||||||
|
|
||||||
// LFM2.5 — the model from issue #82. Its template never mentions enable_thinking, so both
|
// LFM2.5 — the model from issue #82. Its template never mentions enable_thinking, so both
|
||||||
// renders are identical; the handler does implement the continuation hook, so the span can
|
// renders are identical. The continuation hook does land (the span below really is closed),
|
||||||
// be closed in the prompt.
|
// but LFM2.5 declares <think>/</think>, so the closed span is only a suggestion to a model
|
||||||
|
// that owns its own reasoning tags — and measured on-device it reasons straight past it,
|
||||||
|
// untagged, into the answer. Reported as uncontrollable rather than made worse.
|
||||||
{
|
{
|
||||||
auto t = load(BMOE_TMPL_LFM25);
|
auto t = load(BMOE_TMPL_LFM25);
|
||||||
expect("lfm2.5: flag is inert in the template", render(t.get(), /*think*/ true, /*prefill*/ false) ==
|
expect("lfm2.5: flag is inert in the template", render(t.get(), /*think*/ true, /*prefill*/ false) ==
|
||||||
render(t.get(), /*think*/ false, /*prefill*/ false));
|
render(t.get(), /*think*/ false, /*prefill*/ false));
|
||||||
expect("lfm2.5: probed as prefill",
|
expect("lfm2.5: probed as none (declares its own reasoning tags)",
|
||||||
bmoe::detail::probe_think_control(t.get()) == bmoe::ThinkControl::Prefill);
|
bmoe::detail::probe_think_control(t.get()) == bmoe::ThinkControl::None);
|
||||||
expect_span_closed("lfm2.5: prefilled span is closed", t.get(), "</think>");
|
// The prefill itself is well-formed — this is why "none" is a statement about the
|
||||||
|
// MODEL, not about a mechanism that failed to render.
|
||||||
|
expect_span_closed("lfm2.5: the prefill would close the span correctly", t.get(), "</think>");
|
||||||
}
|
}
|
||||||
|
|
||||||
// gpt-oss / harmony. Its template ignores the flag too, and it exposes no thinking tags at
|
// gpt-oss / harmony. Its template ignores the flag too, and it declares NO thinking tags:
|
||||||
// all — a tag-based detector would give up here. Probing the continuation instead sees the
|
// reasoning is a channel the format separates structurally, so priming past it is not
|
||||||
// mechanism that does exist: the handler primes the answer channel. This is the case the
|
// something the model can decline. This is the case the engine used to carry as two
|
||||||
// engine used to carry as two hardcoded harmony marker strings in the decode path.
|
// hardcoded harmony marker strings in the decode path.
|
||||||
{
|
{
|
||||||
auto t = load(BMOE_TMPL_GPTOSS);
|
auto t = load(BMOE_TMPL_GPTOSS);
|
||||||
expect("gpt-oss: probed as prefill",
|
expect("gpt-oss: probed as prefill",
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue