mirror of
https://github.com/openclaw/openclaw.git
synced 2026-10-03 09:39:25 +00:00
## What Problem This Solves
Compaction can retry the same summary request after a model stops at its output limit without visible text, wasting requests while leaving the user without a useful explanation.
- Stop compaction retrying an unchanged summary request when the model stops at `length` without visible text.
- Return an actionable output-budget error while preserving the transcript and selected provider.
Thanks to @Olli0103 for the fix and to @Jackten for the report and the provider probe that isolated the `finish_reason: "length"` / reasoning-only boundary.
## Evidence
- Focused Vitest run on current `main`: 3 files, 69 tests passed.
- `AgentSession` regression exercises the compaction path with a synthetic reasoning-only response: exactly one summary request, unchanged session entries and selected model, then a successful subsequent turn.
- Live local-model proof on the current PR head (`6cf7dd4334`): a one-off test driver used the real `AgentSession` and delegated both requests to the real `streamSimple` Ollama transport. With an output budget of one token, `gpt-oss:20b` returned a length-stopped summary with no visible text. The session made **one** summary request, retained its entries and `ollama` provider, and answered the next conversation turn. The driver only counted requests; it did not synthesize either model response. Its separate Vitest run passed 1/1.
```text
LIVE_COMPACTION_PROOF {"provider":"ollama","model":"gpt-oss:20b","summaryRequests":1,"historyPreserved":true,"selectedProviderPreserved":true,"conversationRequests":1,"nextTurnAnswered":true}
Test Files 1 passed (1)
Tests 1 passed (1)
```
- This was a local provider-backed session test, not a deployed Gateway test. No provider or Gateway deployment was changed. A UI screenshot would not show the backend retry count or preserved transcript.
### Maintainer end-to-end proof (CLI + Gateway, mock OpenAI-compatible provider)
Setup: a real isolated `openclaw agent --local` session per run (`--profile p156230base` / `p156230cand`, task-owned `OPENCLAW_HOME`/`STATE_DIR`/`CONFIG_PATH`, nothing under `~/.openclaw`), default `safeguard` compaction, and one local mock `/v1/chat/completions` server behind two providers:
- `big/big-model`: `contextWindow` 1,000,000, normal replies; its summary requests return a valid structured summary.
- `small/small-reasoner`: `contextWindow` 32,768, `maxTokens` 1,024, `reasoning: true`, `thinkingFormat: "deepseek"`. Its summary requests stream only `reasoning_content` and stop with `finish_reason: "length"` and zero visible text, matching the reporter's probe.
Scaling compared with the incident: the context window and `maxTokens` are 4x smaller at the same ratio (131,072/4,096 → 32,768/1,024). History is about 127k prompt tokens (8 seed turns of 48k characters) rather than 381k. Each summary request takes 1 s instead of about 100 s.
Flow: seed the session on the big model, then send `--model small/small-reasoner --thinking high` on the same `--session-id`. Preflight compaction is required, as in the incident. Then send one more turn back on the big model.
| Run | Summary requests on switch turn | Identical bodies | Compaction `durationMs` | Result |
|---|---|---|---|---|
| Base (merged head with the three source files reverted to `origin/main`) | **6** (3 × body A, 3 × body B) | yes, 3x each | 4,660 | `Preflight compaction required but failed: … model returned no summary text` |
| Candidate (`2385310310f`) | **2** (body A once, then the oversized-message fallback body B once) | no | 1,066 | `Preflight compaction required but failed: … summary output budget (819 tokens) was exhausted without visible text; reduce thinking or increase the selected model's maxTokens before retrying` |
The same cancel-and-preserve path ran in both: `compaction-safeguard … cancelling compaction to preserve history` and `outcome=failed reason=guard_blocked`. In both runs the next big-model turn answered (`BIG-REPLY-*`) with the full prompt still present (508,757 characters versus 508,613 before the switch), so history and the session were intact. At the reporter's roughly 100 s per request, the base pattern costs about 6 full requests plus backoff per compaction attempt, while the candidate stops after one request per distinct body.
Control (large-context compaction is unchanged): the same seeded session, then `openclaw sessions compact agent:main:explicit:proof-156230-control` through an isolated Gateway (ports 21633/21634) on `big/big-model`. Base and candidate each made **1** summary request, returned `compacted: true`, and the next turn's prompt dropped from 508,623 to 239,240 characters in both runs.
Merged with current `origin/main`. Focused suites passed after the merge: `compaction.test.ts`, `compaction.summary-format.test.ts`, `agent-session-loop-correctness.test.ts`, `compaction.summarize-fallback.test.ts`, and `compact.hooks.test.ts` (5 shards). There are no storage or schema changes.
Maintainer intent: Pash (`pash-openai`) and Sarah (`sjf-oa`) have no authorship on the changed files.
## Scope
This addresses the deterministic empty-summary retry path. It does not address the original reporter's local-model performance or any unproven provider-selection cause.
Fixes #156230
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
|
||
|---|---|---|
| .. | ||
| acp-core | ||
| agent-core | ||
| ai | ||
| gateway-client | ||
| gateway-protocol | ||
| llm-core | ||
| markdown-core | ||
| media-core | ||
| media-generation-core | ||
| media-understanding-common | ||
| memory-host-sdk | ||
| mermaid-renderer | ||
| model-catalog-core | ||
| net-policy | ||
| normalization-core | ||
| plugin-package-contract | ||
| plugin-sdk | ||
| retry | ||
| sdk | ||
| session-url-contract | ||
| terminal-core | ||
| tool-call-repair | ||
| workboard-contract | ||
| tsconfig.json | ||