openclaw/docs/concepts
Ayaan Zaidi 98e86147b1
fix(agents): stop repeating a progress message when the turn ends right after it (#162367)
Related: #110565, #115104, #119605

## What Problem This Solves

Fixes: a chat gets two nearly identical messages when a turn sends a progress update with `message(final=false)` and then ends without a final answer.

In our setup a group turn sent "started the run, I'll report back" as progress and stopped empty. Settled-turn finalization then made an extra model call, which restated the same progress, and both messages landed 12 s apart.

## User Impact

If the last thing a turn does is post progress to the current conversation and then it stops empty or with `NO_REPLY`, that progress message is the reply. No second model call, no duplicate. Turns that did more work after or alongside the progress, or that sent nothing, still get finalization.

## Why This Change Was Made

Provenance:

- Settled-turn finalization (#110565, closes #108738) exists so users are never left with no reply after tools ran but the model said nothing. Fallback text after an empty finalizer came later in #133520; #138645 keeps that fallback private in message-tool-only groups.
- `final:false` vs `final:true` (#105365, then the shared contract in #119605) keeps progress from ending a run early and keeps a later failure visible after progress.
- #115104 (fixing #111764) made marked progress *permit* finalization. #111764's caveat was progress sent early in a long tool run, followed by an empty final completion, which was swallowed silently. #127070 applied the same rule to queued follow-ups.

Root cause: none of these checks look at order. A progress message sent as the turn's **last** tool batch was written after the model had seen every other tool result. When the model then stops empty, the finalizer gets the same transcript and no new information, so all it can do is restate the progress. That is the duplicate. The #111764 case is different: there, tool work came after the progress, and that case is unchanged.

Fix: the decision happens when the runtime settles the attempt (`completeEmbeddedAttemptResult`), not per assistant message. The embedded subscriber only records, per provider turn (reset on `turn_start`), whether every tool that finished in that turn was a complete, non-error, non-partial `final:false` text send to the current source, and it keeps that fact for the latest turn that ran tools. At settlement, if that holds, the run ended normally (`terminal.kind === "ok"`), and the terminal assistant response stopped with empty text or `NO_REPLY`, the attempt result marks that latest progress send final and sets `sourceReplyDeliveryState: "delivered"`. Every downstream consumer (finalization gate, terminal resolution, follow-up delivery, run-entry terminal reply) then sees the same completed-reply fact an explicit `final:true` produces. No consumer special case, no prompt or tool-text change. Native-async fragments do not split a turn: an async send followed by a later tool turn, or async work and progress in the same response, still finalizes. Progress reactions, partial or errored sends, closing text, and work in the same batch keep today's behavior.

Options considered:
- Give the finalizer the sent text and accept `NO_REPLY`: this still costs a model call, depends on the model's judgment, and needs a new "empty finalizer is fine" path next to the #133520 fallback.
- Tighten the `final` tool description: models already end empty after progress, so this cannot guarantee anything.

Hermes (`agent/turn_empty_response.py`) handles this the same way. When an empty response follows visible content with only housekeeping tools after it, Hermes reuses that content as the final answer and marks it as already shown so it is not re-sent. It nudges once only when substantive tools ran after the content.

Scope: this changes only the embedded runner. Codex app-server turns can plausibly hit the same duplicate: Codex marks `final:false` in `dynamic-tools.ts` and can complete a turn with no agent message. But "nothing ran after the progress" there depends on item order across native Codex items (shell, apply_patch) that never pass through the OpenClaw dynamic-tool bridge. That makes the Codex event projector a separate owner, so it is a follow-up.

No overlap with Pash/Sarah changes: `a2bbcbf406` and #154217 touch unrelated hunks in these files, and intentional `NO_REPLY` after a delivered message stays honored.

## Bounded cost

- The new path only removes a model call: when the latest tool batch was only complete progress sends and the model then stops empty, settled finalization is skipped (0 extra calls instead of up to 2).
- Every other path is unchanged: settled finalization keeps its existing cap of 2 tool-free attempts (`MAX_EMPTY_SETTLED_FINALIZATION_ATTEMPTS`, covered by `settled-turn-finalization.test.ts`). The finalizer has no tools, so it cannot send progress and cannot re-enter this path.
- The settlement check runs once per attempt result and does not schedule retries, runs, or model calls.
- Tested: `attempt-result.test.ts` drives the real `runAgentLoop` with the real subscriber, then the real `completeEmbeddedAttemptResult` and `resolveSettledToolTerminalContinuationInstruction`. Progress-last followed by an empty terminal gives a delivered reply and no finalizer. Two cases still finalize: an async progress send, then an empty tail, then a read, then an empty terminal; and async read plus progress in one response, then an empty terminal. Each failed on the head before its fix with `expected 'delivered' to be 'missing'`. The QA scenario asserts 2 model requests (no finalization request) for progress-then-empty and 3 for the write-then-empty control.

## Evidence

QA lab, mock-openai, isolated gateway child, qa-channel group with `visibleReplies: message_tool`. New scenario `group-progress-then-empty-finalization`:

- Before (origin/main + scenario): fail. The progress room got `["Started the run, I will report back. QA-GROUP-PROGRESS-OK", "Still running, I will report back. QA-GROUP-PROGRESS-OK"]`.
- After: pass. The progress room got 1 post and made 2 model requests (no finalization request). The control room (write, then empty stop, nothing sent) still ran finalization, made 3 model requests, and posted `QA-GROUP-WRITE-FINALIZED-OK` once.

Settlement regressions in `attempt-result.test.ts` (real agent loop): on the previous per-message head, the async-continuation case fails with `expected 'delivered' to be 'missing'`. Subscriber batch-fact table: only-progress is true; same-batch work, a later batch, a reaction, and a partial/errored send are false. Focused suites pass: attempt-result (43), suppression (32), tools handler, settled-tool evidence, settled-turn finalization, attempt-execution-phase, and lifecycle.

Real gateway (qa-channel, round 3 head): the scenario passes with 1 progress post and 2 requests, and the write control finalizes with 3 requests. The scenario is the only CI-runnable check of the consumer chain after the attempt result: run-loop finalization admission, terminal resolution, payload building, and channel delivery. The loop tests stop at the delivery fact, so a consumer that ignored it would still pass them, but here it would post a second message.

### Live model, Telegram Test Server group

Setup: a real user in the QA group, `messages.groupChat.visibleReplies: message_tool`, live OpenAI `gpt-5.5` through a pass-through proxy, isolated gateway. The prompt asked the model to post a `final:false` progress update saying the report run started and then stop. In both runs below the model did that unsteered: one successful `message(final:false)`, then an empty stop. In each run the bot's transient status draft (`Working`) appears and is deleted after delivery; that is existing progress-draft behavior and the same before and after.

| Run | Bot posts that remain | Model requests (main agent) | Finalizer |
|---|---|---|---|
| Before, origin/main `7dd6ab7` | 2: `Monthly report run has started — I’ll report back here when it finishes.` (t=13.1 s), then `Done.` (t=17.7 s) | 3 | ran (`settled post-tool turn lacked a final answer`) |
| After, `e27810b` | 1: `Monthly report run started. I’ll report back here when it finishes.` (t=13.8 s) | 2 | not run |

Control on `e27810b` (finalizer must still run after real work): the model posted `Checking the clock now.` with `final:false`, then ran `exec date -u`. The proxy replaced only the model's next response with an empty turn, which is the incident's "ends empty" shape; the finalizer request itself went to the live model. The finalizer ran and the group received `The UTC time is Thu Oct 1 13:08:11 UTC 2026.`

An earlier variant that told the model to end with `NO_REPLY` produced one post on both refs: on main the live finalizer also answered `NO_REPLY` twice, and the private fallback stayed private per #138645. Main still spent 2 extra model calls there; the fix spends none.

LOC vs merge base: production +47/-3 (`src/agents` subscriber + attempt-result), docs +5/-1, tests and QA fixtures +441.

Test cost (`pnpm test <file> --maxWorkers=1`, wall, head e27810b, shared loaded host): `attempt-result.test.ts` 46.4 s for 44 tests, with the 3 new real-loop cases at 0.2 to 0.6 s each; `embedded-agent-subscribe.subscribe-embedded-agent-session.source-reply-suppression.test.ts` 53.4 s for 32 tests, with the 5 batch-fact cases at 4 to 160 ms each. Most of the wall time is import and transform. The new cases use no timers, sleeps, or gateway boots. No attributable CI timing exists yet for this head.

CI on 959c8b0 (run 36873121248, after one rerun of failed jobs):
- `gateway-timeout-recovery-subagent.e2e.test.ts`: fails 2 of 2 locally on clean origin/main. Known flake #161139, with fixture fix #162260 open. The scenario never uses the message tool.
- `src/gateway/server.chat-steer-direct-command.test.ts` (skills watchers left open): fails the same way on four unrelated PRs today: `steipete/session-store-agent-db-drain`, `chore/chrome-devtools-mcp-1.10.1`, `chore/npm-12-1-20261001`, `codex/repair-child-completion`.
- `node-worker-launch-wire.e2e.test.ts` and `telegram-model-picker-prepared-gateway.e2e.test.ts`: red on the first attempt and green on rerun; both pass locally on head and on main.
- `src/auto-reply/reply/dispatch-from-config.secrets.test.ts` "quiet group link": still red. Only the file's first case fails, at its default 1 s `vi.waitFor`; that case took 1437 ms in CI while the following cases took about 60 ms. The same three-file shard passes locally on the head, on head merged with current main, with `CI=true`, and twice on main under `taskpolicy -b`. The test drives the `secrets` tool, which never touches the message-tool progress path, the new `turn_start` reset, or attempt-result settlement. I could not reproduce it on main or find it on another PR's CI from the last two days (93 failed node jobs scanned), so this one is not proven unrelated.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-02 02:12:10 +08:00
..
active-memory
model-providers fix: keep custom model routes on their declared tool surface (#156017) 2026-09-24 04:36:01 -07:00
qa-e2e-automation fix(qa): explain missing async capture support on older hosts (#159277) 2026-09-26 18:41:58 -07:00
active-memory.md docs: fix Active Memory troubleshooting section links (#156899) 2026-09-25 19:17:31 +08:00
agent-bindings.md refactor(routing): simplify channel route selection (#148848) 2026-09-16 22:38:22 -07:00
agent-loop.md fix(agents): stop repeating a progress message when the turn ends right after it (#162367) 2026-10-02 02:12:10 +08:00
agent-runtimes.md feat: enable automatic Code Mode for preferred models (#155614) 2026-09-22 04:19:15 -07:00
agent-workspace.md fix(agents): reject invalid selectors before routing (#153387) 2026-09-20 01:08:48 -07:00
agent.md
architecture.md docs(gateway): document scheduler ownership and timer census (#160477) 2026-09-28 14:25:31 +00:00
compaction.md fix: resume compacted turns after a Gateway restart (#160128) 2026-09-30 16:45:05 -07:00
context-engine.md fix: start background context maintenance after durable turns (#151936) 2026-09-18 14:15:10 -07:00
context.md
decision-models.md feat: optionally omit tools on conversational turns (#153340) 2026-10-01 09:33:15 -07:00
delegate-architecture.md
dreaming.md perf(memory): bound workspace provenance and dreaming reads (#159190) 2026-09-26 21:42:44 -07:00
experimental-features.md feat: optionally omit tools on conversational turns (#153340) 2026-10-01 09:33:15 -07:00
features.md
main-session.md fix(sessions): notify Home about new sessions by default (#152068) 2026-09-18 16:24:59 -07:00
managed-worktrees.md perf(worktrees): empty-workspace sessions build and clone an empty template (#159544) 2026-09-27 03:07:52 -07:00
mantis-slack-desktop-runbook.md fix(qa): validate Slack run artifacts and approval identities (#117344) 2026-09-21 08:39:57 -07:00
mantis.md fix(qa): Mantis before-after exits when a lane command stalls (#111051) 2026-09-16 17:00:24 -06:00
markdown-formatting.md
memory-architecture.md docs(memory): clarify episodic transcript recall via Active Memory (#161179) 2026-09-29 21:03:10 +08:00
memory-builtin.md perf(memory): floor watcher polling fallback and report it once (#162271) 2026-10-01 05:10:56 +00:00
memory-honcho.md
memory-provenance.md perf(memory): bound workspace provenance and dreaming reads (#159190) 2026-09-26 21:42:44 -07:00
memory-search.md
memory.md
messages.md fix(sessions): bind terminal failure receipts to their lifecycle (#162297) 2026-09-30 23:43:45 -07:00
model-failover.md fix: explain saved provider safety refusals (#162469) 2026-10-01 01:30:57 -07:00
model-providers.md docs(providers): expand llmman guidance and add hybrid inference (#139606) 2026-09-19 10:07:03 -07:00
models.md fix(models): apply downloaded catalogs without a Gateway restart (#158000) 2026-10-01 13:27:55 +08:00
multi-agent.md fix(doctor): preserve channel ownership across updates (#150555) 2026-09-17 04:03:27 -07:00
multi-user.md fix(reactions): worker-owned writes, final-dispatch route recheck, single-slot mirrors (#162434) 2026-10-01 01:42:29 -07:00
oauth.md fix: restart abandoned provider sign-ins with shared OAuth handling (#160574) 2026-09-28 19:40:37 -07:00
parallel-specialist-lanes.md fix: isolate subagent concurrency per session (#156381) 2026-09-23 06:46:57 -07:00
personal-agent-benchmark-pack.md
presence.md fix(ui): include signed-in user in Online roster (#160618) 2026-09-28 20:20:59 +00:00
progress-drafts.md fix(progress): failed commands crowd out plans during active drafts (#156649) 2026-09-25 15:53:32 +05:30
qa-e2e-automation.md
queue-steering.md fix: steered message no longer replaces a question whose model request failed (#162439) 2026-10-01 20:42:48 +08:00
queue.md fix: Web UI turns still ask to try again after a stall (#162490) 2026-10-01 19:14:54 +08:00
retry.md fix: avoid replaying bodyless HTTP client errors (#160584) 2026-09-28 23:27:36 +05:30
session-attachment.md fix(tui): preserve resolved session ownership (#155872) 2026-09-22 10:57:56 -07:00
session-pruning.md fix(agents): prune stale tool results for opted-in custom OpenAI providers (#159100) 2026-10-01 14:18:43 +08:00
session-search.md fix(search): find session titles and unfinished message queries (#162517) 2026-10-01 02:16:38 -07:00
session-state.md fix(sessions): compare fresh session facts when recording state events (#161553) 2026-09-30 05:41:01 +00:00
session-tool.md improve: sessions_send delivers a peer's reply once instead of running automatic agent ping-pong (#162227) 2026-10-01 17:28:00 +00:00
session.md fix(sessions): keep one agent's session cleanup and /stop from killing another agent's work (#161736) 2026-09-30 03:10:12 -07:00
soul.md
standing-intents.md
storage-locations.md feat: back up to external disks and Cloudflare R2 with storage locations (#161913) 2026-10-01 02:05:44 -07:00
streaming.md fix(telegram): restore progress when reply hooks suppress previews (#161546) 2026-10-01 06:41:47 -07:00
subagent-yield-handoff.md fix(agents): waiting subagents post a false "yielded without a continuation source" warning (#162306) 2026-10-01 04:01:04 -07:00
system-prompt.md fix(codex): restore persona on remote app-server connections (#162156) 2026-09-30 16:05:21 -07:00
timezone.md
typebox.md
typing-indicators.md
usage-tracking.md fix(ui): respect known zero usage costs (#143033) 2026-09-20 16:38:06 -07:00
user-model.md fix: invite GitHub visitors without a public email (#162666) 2026-10-01 17:12:53 +01:00