openclaw/packages
Olli 79fbe5890e
fix(compaction): stop retrying empty length-stopped summaries (#160650)
## What Problem This Solves
Compaction can retry the same summary request after a model stops at its output limit without visible text, wasting requests while leaving the user without a useful explanation.

- Stop compaction retrying an unchanged summary request when the model stops at `length` without visible text.
- Return an actionable output-budget error while preserving the transcript and selected provider.

Thanks to @Olli0103 for the fix and to @Jackten for the report and the provider probe that isolated the `finish_reason: "length"` / reasoning-only boundary.

## Evidence
- Focused Vitest run on current `main`: 3 files, 69 tests passed.
- `AgentSession` regression exercises the compaction path with a synthetic reasoning-only response: exactly one summary request, unchanged session entries and selected model, then a successful subsequent turn.
- Live local-model proof on the current PR head (`6cf7dd4334`): a one-off test driver used the real `AgentSession` and delegated both requests to the real `streamSimple` Ollama transport. With an output budget of one token, `gpt-oss:20b` returned a length-stopped summary with no visible text. The session made **one** summary request, retained its entries and `ollama` provider, and answered the next conversation turn. The driver only counted requests; it did not synthesize either model response. Its separate Vitest run passed 1/1.

  ```text
  LIVE_COMPACTION_PROOF {"provider":"ollama","model":"gpt-oss:20b","summaryRequests":1,"historyPreserved":true,"selectedProviderPreserved":true,"conversationRequests":1,"nextTurnAnswered":true}
  Test Files  1 passed (1)
  Tests       1 passed (1)
  ```

- This was a local provider-backed session test, not a deployed Gateway test. No provider or Gateway deployment was changed. A UI screenshot would not show the backend retry count or preserved transcript.

### Maintainer end-to-end proof (CLI + Gateway, mock OpenAI-compatible provider)
Setup: a real isolated `openclaw agent --local` session per run (`--profile p156230base` / `p156230cand`, task-owned `OPENCLAW_HOME`/`STATE_DIR`/`CONFIG_PATH`, nothing under `~/.openclaw`), default `safeguard` compaction, and one local mock `/v1/chat/completions` server behind two providers:
- `big/big-model`: `contextWindow` 1,000,000, normal replies; its summary requests return a valid structured summary.
- `small/small-reasoner`: `contextWindow` 32,768, `maxTokens` 1,024, `reasoning: true`, `thinkingFormat: "deepseek"`. Its summary requests stream only `reasoning_content` and stop with `finish_reason: "length"` and zero visible text, matching the reporter's probe.

Scaling compared with the incident: the context window and `maxTokens` are 4x smaller at the same ratio (131,072/4,096 → 32,768/1,024). History is about 127k prompt tokens (8 seed turns of 48k characters) rather than 381k. Each summary request takes 1 s instead of about 100 s.

Flow: seed the session on the big model, then send `--model small/small-reasoner --thinking high` on the same `--session-id`. Preflight compaction is required, as in the incident. Then send one more turn back on the big model.

| Run | Summary requests on switch turn | Identical bodies | Compaction `durationMs` | Result |
|---|---|---|---|---|
| Base (merged head with the three source files reverted to `origin/main`) | **6** (3 × body A, 3 × body B) | yes, 3x each | 4,660 | `Preflight compaction required but failed: … model returned no summary text` |
| Candidate (`2385310310f`) | **2** (body A once, then the oversized-message fallback body B once) | no | 1,066 | `Preflight compaction required but failed: … summary output budget (819 tokens) was exhausted without visible text; reduce thinking or increase the selected model's maxTokens before retrying` |

The same cancel-and-preserve path ran in both: `compaction-safeguard … cancelling compaction to preserve history` and `outcome=failed reason=guard_blocked`. In both runs the next big-model turn answered (`BIG-REPLY-*`) with the full prompt still present (508,757 characters versus 508,613 before the switch), so history and the session were intact. At the reporter's roughly 100 s per request, the base pattern costs about 6 full requests plus backoff per compaction attempt, while the candidate stops after one request per distinct body.

Control (large-context compaction is unchanged): the same seeded session, then `openclaw sessions compact agent:main:explicit:proof-156230-control` through an isolated Gateway (ports 21633/21634) on `big/big-model`. Base and candidate each made **1** summary request, returned `compacted: true`, and the next turn's prompt dropped from 508,623 to 239,240 characters in both runs.

Merged with current `origin/main`. Focused suites passed after the merge: `compaction.test.ts`, `compaction.summary-format.test.ts`, `agent-session-loop-correctness.test.ts`, `compaction.summarize-fallback.test.ts`, and `compact.hooks.test.ts` (5 shards). There are no storage or schema changes.

Maintainer intent: Pash (`pash-openai`) and Sarah (`sjf-oa`) have no authorship on the changed files.

## Scope
This addresses the deterministic empty-summary retry path. It does not address the original reporter's local-model performance or any unproven provider-selection cause.

Fixes #156230

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-29 08:22:33 +05:30
..
acp-core refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
agent-core fix(compaction): stop retrying empty length-stopped summaries (#160650) 2026-09-29 08:22:33 +05:30
ai feat(anthropic): support Claude Sonnet 5.5 (#160847) 2026-09-29 02:49:52 +00:00
gateway-client fix(gateway): handle connection bursts behind shared IPs (#160814) 2026-09-29 00:48:30 +00:00
gateway-protocol improve: show worker runtime install progress on paired session hosts (#160110) 2026-09-29 02:17:32 +00:00
llm-core feat(anthropic): support Claude Sonnet 5.5 (#160847) 2026-09-29 02:49:52 +00:00
markdown-core fix: streamed replies split Markdown tables that fit one message (#160701) 2026-09-29 03:41:47 +05:30
media-core refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
media-generation-core refactor(packages): deslop shared packages third pass (#159933) 2026-09-27 20:53:42 -07:00
media-understanding-common refactor(packages): deslop shared packages third pass (#159933) 2026-09-27 20:53:42 -07:00
memory-host-sdk refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
mermaid-renderer fix: restore bundled Mermaid parser license notices (#157461) 2026-09-25 14:21:38 -07:00
model-catalog-core feat(anthropic): support Claude Sonnet 5.5 (#160847) 2026-09-29 02:49:52 +00:00
net-policy refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
normalization-core perf(gateway): throttle sidebar narration updates (#160002) 2026-09-28 07:11:52 -07:00
plugin-package-contract refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
plugin-sdk refactor(packages): deslop shared packages third pass (#159933) 2026-09-27 20:53:42 -07:00
retry fix(gateway): handle connection bursts behind shared IPs (#160814) 2026-09-29 00:48:30 +00:00
sdk refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
session-url-contract refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
terminal-core fix(doctor): update refuses when lint warns about an unusable agent database (#160184) 2026-09-28 16:14:57 -07:00
tool-call-repair refactor(packages): deslop shared packages second pass (#159279) 2026-09-28 14:37:02 -07:00
workboard-contract refactor(packages): deslop shared packages third pass (#159933) 2026-09-27 20:53:42 -07:00
tsconfig.json perf(lint): bound the package fallback type graph (#160281) 2026-09-29 00:27:31 +07:00