fix(chat): honour "thinking off" on templates that ignore enable_thinking

Turning Thinking off set the template variable enable_thinking and stopped there. That
variable is only a request to the model's own chat template, and a template is free to
ignore it. LFM2.5's never reads it, so the rendered prompt was byte-identical with thinking
on and off, the model reasoned anyway, and nothing reported that the setting had been
dropped (#82).

Detection is measured, not assumed: at open() the template is rendered with the flag on and
off and the two prompts compared, then rendered once more with a continuation. That answers
the only question that matters — does this template react — for any model in any language.
common_chat_templates_support_enable_thinking cannot answer it: per handler it is a
hardcoded literal reporting "this model can reason", not "this template reads the variable".

Enforcement uses llama.cpp's continuation hook: a synthetic trailing assistant message with
continue_final_message makes upstream's per-template handler render that family's own
"reasoning is over" span into the prompt. The model resumes at the first token of its answer
with the reasoning already behind it. This is what a template that implements the toggle
natively does (Qwen3 renders <think></think> for enable_thinking=false), so it needs no
cooperation from the template and no sampler — it works on the greedy path the byte-identity
gates run on.

Because the span comes from upstream's handler, the engine names no markers of its own. That
retires the two hardcoded harmony strings in the decode path: priming gpt-oss to answer
without reasoning was a literal "<|start|>assistant" suffix test and a literal
"<|channel|>final<|message|>" appended to the prompt. gpt-oss now takes the same generic path
as every other family and resumes at the same point, so a submodule bump that changes those
markers needs no engine change.

Models where neither mechanism exists are reported rather than fought: BMOE_READY gains
think_ctl (template | prefill | none) and the app shows the Thinking switch disabled, with
the reason, instead of offering a control that does nothing.

Deliberately not done: forcing the reasoning block closed on logits. Measured on-device in
the closed PR #83, it made LFM2.5 strictly worse — the model reopened the block, then
abandoned the tags and reasoned in plain prose into the answer. Suppression belongs in the
prompt, before the model commits to reasoning, not mid-generation.

tests/think_control_test.cpp pins all three regimes against the vendored templates with no
model: Qwen3 as template, LFM2.5 and gpt-oss as prefill with the span asserted CLOSED, a
plain template as none, and fail-open when the probe cannot run.
This commit is contained in:
Helldez 2026-07-19 20:26:22 +02:00
parent 266b92e6ce
commit aa6a7fdafd
15 changed files with 443 additions and 36 deletions

View file

@ -30,6 +30,11 @@ data class UiState(
val ioMode: String? = null, // effective read mode reported by the engine (direct / buffered)
val cpuTempC: Double? = null, // SoC/CPU temperature (°C), sampled while generating (battery fallback)
val sessionSig: String? = null, // signature of the loaded session (AppSettings.sessionSignature)
// How the loaded model can honour "Thinking off", reported once at BMOE_READY: "template" (its
// chat template reads the flag), "prefill" (it does not, so the engine closes the reasoning span
// in the prompt), or "none" (neither — the model always reasons, and the switch is hidden rather
// than left there doing nothing). Null until a session reports it. See docs/telemetry.md.
val thinkControl: String? = null,
val transcript: List<ChatTurn> = emptyList(), // committed turns; the in-flight answer is `answer`
val streaming: Boolean = true, // is the loaded session using the MoE streamer (vs mmap baseline)?
) {

View file

@ -112,8 +112,10 @@ class RunService : Service() {
RunBus.update {
// ioMode is re-sniffed from the new session's stderr, so clear it: it describes the
// session being replaced, and the sniffer's first-writer-wins guard would keep it.
// thinkControl is a property of the model being loaded, so it goes the same way — the
// incoming session reports its own at BMOE_READY.
it.copy(state = EngineState.LOADING, error = null, sessionSig = sig, answer = "", summary = "",
transcript = emptyList(), streaming = streaming, ioMode = null)
transcript = emptyList(), streaming = streaming, ioMode = null, thinkControl = null)
}
thread(name = "bmoe-session") { runSession(argv, model, myEpoch, dying) }
@ -199,7 +201,8 @@ class RunService : Service() {
val t = line.trim()
when {
t.startsWith("BMOE_READY ") -> {
RunBus.setState(EngineState.READY)
val ctl = Regex(""""think_ctl":"([a-z_]+)"""").find(t)?.groupValues?.get(1)
RunBus.update { it.copy(state = EngineState.READY, thinkControl = ctl) }
main.post { notify("Model ready") }
pending?.let { p -> pending = null; sendGenerate(p) } ?: scheduleIdleUnload()
}

View file

@ -11,6 +11,7 @@ import androidx.compose.ui.Modifier
import androidx.compose.ui.text.font.FontWeight
import androidx.compose.ui.unit.dp
import androidx.compose.ui.unit.sp
import androidx.lifecycle.compose.collectAsStateWithLifecycle
/**
* Groups every tunable the engine exposes. Changes are applied to [current] live and
@ -19,6 +20,12 @@ import androidx.compose.ui.unit.sp
@OptIn(ExperimentalMaterial3Api::class)
@Composable
fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack: () -> Unit) {
// Reported by the loaded session at BMOE_READY. "none" means this model reasons no matter what
// it is asked, so the Thinking switch is shown disabled with the reason rather than left there
// pretending to work (#82). Null = nothing loaded yet, so nothing is claimed either way.
val ui by RunBus.state.collectAsStateWithLifecycle()
val thinkingLocked = ui.thinkControl == "none"
Scaffold(
topBar = {
TopAppBar(
@ -151,10 +158,18 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
Section("Prompt") {
SwitchRow(
"Thinking",
"Let a reasoning model think before answering; its reasoning shows in a " +
"collapsible block above the reply. Off tells the model to skip thinking. " +
"No effect on models that don't reason.",
current.thinking,
if (thinkingLocked)
"This model always reasons — it offers no way to turn thinking off, so the " +
"switch is disabled here instead of being ignored silently. Its reasoning " +
"still shows in a collapsible block above the reply."
else
"Let a reasoning model think before answering; its reasoning shows in a " +
"collapsible block above the reply. Off tells the model to skip thinking. " +
"No effect on models that don't reason.",
// Locked reads ON, not OFF: the model reasons on every turn, and that is what
// the switch should be showing whatever the stored preference says.
checked = current.thinking || thinkingLocked,
enabled = !thinkingLocked,
) { onChange(current.copy(thinking = it)) }
}