openclaw/docs/reference/transcript-hygiene.md
Peter Steinberger 84c6eb8fc5
improve: sessions_send delivers a peer's reply once instead of running automatic agent ping-pong (#162227)
* refactor(agents): replace peer announcement loops with one reply delivery

Deliver delayed peer results once to the requester and return waited replies inline without a continuation. Preserve child completion custody and same-session generation-bound channel delivery. Use exact-key saved-route lookup, retire legacy peer control tokens from runtime decisions, and retain historical display suppression.

Simplify the announce owner, preserve required missing child output, and keep wake retries within the remaining operation budget. Update session guidance and the opt-in live peer scenario.

Candidate regression scenarios pass. Broad focused tests and changed-lane validation remain incomplete under host load; original-code regression proof is still pending. No live provider test was run.

* refactor(subagents): consolidate completion custody and lifecycle owners

Share terminal-effect persistence and requester-wake settlement adapters while preserving their distinct authority, frozen-wave, and publication guards. Remove unused preparation and accessor paths and reuse existing parameter contracts.

Propagate unknown SQLite write outcomes before scheduling interrupted completion recovery during restart drain. The focused candidate regression passes; original-code failure proof and the broader validation matrix remain pending.

* test(agents): align completion assertions with reply ownership

Assert that an ACP target already owned by its requester keeps task completion and starts no detached reply. Keep isolated cron fallback routing and participant-store checks while asserting no detached wait.

Remove assertions on the retired duplicate steerMessage field; retain complete source-reply content, ordering, exclusions, visibility and batch settlement checks. Remove the redundant inner same-session return while retaining the enclosing branch exit.

Validation: scoped Codex review at P2, formatting and diff checks passed. Two attempts on the coordinator Testbox lease failed during file sync before installation or tests. Baseline regression proof, the full matrix and the changed-lane gate remain pending.

* test(agents): inject the gateway caller into same-session delivery routing tests

* test(subagents): route direct requester-yield callers through prepared cron authority

* refactor(subagents): keep the requester wake currency assertion module-private

* fix(agents): fence failed and delayed subagent cleanup

Restore the previous resident row when same-ID registration persistence fails.
Carry collector cleanup authority into attachment removal, and bind lifecycle
grace timers to their captured registration and generation.

Regression coverage exercises failed writes, retirement during cleanup, and
same-ID successors without changing storage or timer durations.

* fix(agents): retain paused state in compact registry reads

Read pauseReason from existing stored JSON for ordinary and private-parent
records so cold projections preserve the resident paused status. Derive compact
SQL row types from the query and shrink the removed-assertion allowance.

No schema, serialized representation, or migration changes are required.

* fix(agents): preserve spawn cleanup and request scope

Route fork lookup failures through provisional-child cleanup. Reuse the spawn
admission owner for visible work with the resolved requester agent identity.
Pass trusted-creation transport timeouts through the argument the transport reads.

Consolidate in-process dispatch while preserving creation custody, source
fencing, per-interface deadlines, and signed fallback behavior.

* refactor(agents): simplify subagent registry and spawn owners

Remove unused registry APIs, callback declarations, presentation adapters,
registration inputs, and bootstrap scaffolding. Share registration ACK control
flow while retaining both durable writes and the original settlement error owner.

Complete keyed capability lookups, reuse heartbeat and timeout owners, carry
resolved model references, and remove obsolete ACP transcript work and relay
options. Preserve legacy stored attachment safety and protocol-v4 fallbacks.

* refactor(agents): simplify completion consumers and registry publication

Move full-registry fixture writes into test support and require named production mutations. Consolidate publication subscriptions while preserving projection ordering and persistence observer exception isolation. Remove duplicate completion retry scaffolding, runtime barrels, test-only exports, duplicate types, and single-caller adapters.

* chore(agents): shrink retired test API assertion allowance

* test(gateway): preserve real registry operations in agent fixtures

* test(subagents): preserve named publication fixture preconditions

Name all retired fixture rows when replacing describe/list facts. Warm the compact cache before its initial publication so the existing ordering assertion still detects an unintended SQL reload.

* fix(plugin-sdk): preserve capability store compatibility

Retain the public record-or-lookup store contract, including normalized first-match identity lookup and canonical depth fallback for partial records. Keep migrated internal callers on keyed lookups. Restore the base handoff declarations and requester query/caller linkage so the existing SDK declaration gate stays unchanged.

* test(gateway): publish fixture memory ownership

Publish the staged in-memory owner through commitOwnership before asserting describe and list projections. Named durable writes no longer trigger the whole-index rebuild that accidentally discovered the fixture row. Preserve the existing assertions.

* refactor(agents): tighten completion projections and ownership facts

Project only formatter-consumed child fields without mutating authoritative rows, and carry the session resolver's required requesterOwned boolean directly. Preserve formatting, nested identity references, authority checks, and public signatures while resolving the full core lint findings.

* test(agents): retire obsolete benchmark registry mock

Remove the full-registry writer no-op from the memory benchmark's typed module mock after the production writer cutover. Preserve the named-write mock and benchmark behavior.

* test(agents): await requester settlement in live fixtures

Child cleanup can publish before the requester delivery acknowledgement. Join persistence publications until the requester wake settles, preserving the existing delivery assertions and deadlines.

* refactor(agents): retire hidden sessions_send message aliases

Require the canonical message argument already declared by the public tool schema. Remove alias-only reasoning stripping and keep canonical message indentation unchanged.

* fix(agents): decode compact subagent metadata canonically

Replace the separate compact decoder with the canonical codec and session-list projector. Keep bounded SQL metadata projection, private envelopes, and duplicate-key last-value semantics without transferring retained results. Add real-reader duplicate-key and retained-content regression coverage. Schema, persisted bytes, and update behavior are unchanged.

* refactor(agents): remove redundant delivery comments

* perf(agents): restore compact reader after benchmark regression

Revert 2c0fd630c7 as required by the performance gate. Across five scoped 2000-row reads per revision, the base median was 227.31 ms and the candidate median was 429.25 ms (88.84% slower). Retain the sessions_send alias retirement and comment cleanup; leave the duplicate-key reader fix for separately approved work.

* test(agents): relocate sessions_send alias regression coverage

Keep all four rejection cases in the assembled-tool suite where retired preparation coverage freed space. The oversized sessions test file shrinks again; no assertions, mocks, timeouts, skips, or line-cap baselines change.

* fix(agents): retain sessions_send record input validation

Use the shared isRecord guard from the retired alias normalizer instead of a new type assertion. Preserve canonical missing-message errors for non-record input and satisfy the assertion safety gate without a suppression.

* refactor(agents): keep sessions_send replies on the requester key

Remove the legacy key-only DM-to-main reply remap and its routing exceptions. Use the resolved caller key for reply context, provenance, watches and followup preparation. Preserve exact-incarnation and authorization checks.

* test(agents): bind native followup custody in admission fixture

Keep the key-only DM requester authority test on the native child completion path. Settle the existing completion owner before mock acceptance and reset queued mock implementations between cases.

* fix(agents): reconcile subagent callers after main merge

Use the phase-aware publication API in newer main callers and explicit changed IDs in the prepared-read fixture. Keep retained regressions in their current sibling owners and place requester adoption in the lifecycle controller without weakening guards. Preserve heartbeat narrowing after shared-owner integration.

* test(agents): clean up merged fixture imports and names

* test(agents): refresh single-reply prompt snapshots

Regenerate prompt fixtures for the intentional sessions_send description change. The four-file delta contains only the description and derived sizes and hashes. Reproduced the hosted drift on Testbox, then passed prompt:snapshots:check and the snapshot-only changed checks.

* test(agents): await owned descendant settlement

* fix(agents): retain the leaf followup owner contract

Restore the standalone completion-owner interface and implements check from main. Deriving that contract from the implementation class creates a type-only cycle through the cohort projection. Runtime behavior and the exposed owner methods remain unchanged.

* fix(agents): disambiguate the sessions reply target resolver

Give the async sessions_send reply lookup a distinct exported name from the general synchronous outbound session resolver. Update its sole production caller and direct tests without an alias or behavior change.

* test(agents): remove the retired steer fixture argument

* test(agents): drop the retired sessions_send delivery mode from the follow-up yield live test
2026-10-01 17:28:00 +00:00

14 KiB

summary read_when title
Reference: provider-specific transcript sanitization and repair rules
You are debugging provider request rejections tied to transcript shape
You are changing transcript sanitization or tool-call repair logic
You are investigating tool-call id mismatches across providers
Transcript hygiene

OpenClaw applies provider-specific fixes to transcripts before a run (building model context). These are in-memory adjustments used to satisfy strict provider requirements. Runtime transcript state stays in SQLite; provider-specific assistant-prefill stripping happens only while constructing outbound payloads.

Scope includes:

  • Runtime-only prompt context staying out of user-visible transcript turns
  • Tool call id sanitization
  • Tool call input validation
  • Tool result pairing repair
  • Turn validation / ordering
  • Thought signature cleanup
  • Thinking signature cleanup
  • Image payload sanitization
  • Blank text-block cleanup before provider replay
  • Incomplete reasoning-only length-turn cleanup before provider replay
  • User-input provenance tagging (for inter-session routed prompts)
  • Empty assistant error-turn removal for provider replay

If you need transcript storage details, see Session management deep dive.


Failed attempts and recovery

Text-only assistant errors are buffered until the logical run settles. Recovery discards their partial text because the recovered reply supersedes it. Terminal failure persists the last attempt's partial text and error.

Tool calls, displayable non-text content, and attachment facts are persisted immediately, before dependent tool results or the recovered reply. These fact rows omit the error and use a replayable stop reason so provider replay retains the calls. For mixed text/fact messages, partial text and the error remain buffered separately; terminal settlement does not duplicate facts or usage. This uses existing assistant-row shapes and requires no database migration.

Global rule: runtime context is not user transcript

Runtime/system context can be added to the model prompt for a turn, but it is not end-user-authored content. OpenClaw keeps a separate transcript-facing prompt body for Gateway replies, queued followups, ACP, CLI, and embedded OpenClaw runs. Stored visible user turns use that transcript body instead of the runtime-enriched prompt.

For legacy sessions that already persisted runtime wrappers, Gateway history surfaces apply a display projection before returning messages to WebChat, TUI, REST, or SSE clients.


Where this runs

The embedded runner selects and applies transcript policy:

  • Policy selection: src/agents/transcript-policy.ts (resolveTranscriptPolicy, keyed on provider, modelApi, and modelId)
  • Sanitization/repair application: sanitizeSessionHistory in src/agents/embedded-agent-runner/replay-history.ts

Legacy JSONL validation and import belong to openclaw doctor --fix; the embedded runner does not repair or reopen file-backed runtime transcripts.


Global rule: image sanitization

Image payloads are always sanitized to prevent provider-side rejection due to size limits (downscale/recompress oversized base64 images). This also helps control image-driven token pressure for vision-capable models: lower max dimensions reduce token usage, higher dimensions preserve detail.

Implementation:

  • sanitizeSessionMessagesImages in src/agents/embedded-agent-helpers/images.ts
  • sanitizeContentBlocksImages in src/agents/tool-images.ts
  • Max image side is configurable via agents.defaults.imageMaxDimensionPx (default: 1200)
  • Blank text blocks are removed while this pass walks replay content. Assistant turns that become empty are dropped unless they own opaque provider replay state; user and tool-result turns that become empty receive a non-empty omitted-content placeholder.

Global rule: malformed tool calls

Assistant tool-call blocks missing both input and arguments are dropped before model context is built. This prevents provider rejections from partially persisted tool calls (for example, after a rate limit failure).

Completed call/result pairs remain history when their tool is disabled, removed, or unavailable in the current catalog. Their names still require valid syntax; malformed calls, ambiguous pairing, and synthetic missing-result repairs do not grant this exception. Replaying a completed pair does not advertise or authorize the tool for a new call.

Implementation:

  • sanitizeToolCallInputs in src/agents/session-transcript-repair.ts
  • Applied in sanitizeSessionHistory (src/agents/embedded-agent-runner/replay-history.ts)

Global rule: tool result pairing

Tool results are paired to tool-call occurrences within each assistant turn before provider-specific call IDs are rewritten. Provider-generated IDs may repeat on later turns, so a result adjacent to a repeated call stays with that occurrence. A displaced result is moved only when exactly one unresolved occurrence can own it; ambiguous extras are dropped and missing occurrences receive synthetic error results.

Implementation: sanitizeToolUseResultPairing in src/agents/session-transcript-repair.ts

When switching models, provider replay moves delayed asynchronous tool results next to their originating call before removing the source model's async metadata. Call and result IDs are trimmed before matching, so surrounding whitespace does not turn a real result into a synthetic missing-result error. This projection runs in packages/ai/src/transcript-transform.ts and leaves stored history intact.


Global rule: incomplete or silent reasoning-only turns

Assistant turns are omitted from the in-memory replay copy when they contain only thinking or redacted-thinking content after either of these events:

  • The provider output limit ends the turn with incomplete reasoning state.
  • Silent-reply cleanup removes the turn's only visible NO_REPLY text.

The silent-reply cleanup prevents hidden reasoning from merging into a later assistant tool-use turn when strict providers rebuild the conversation.

Empty length turns remain unchanged, as do length turns with visible text, tool calls, or unknown content blocks. Silent-reply turns with tool calls or unknown content blocks also remain unchanged. Stored transcripts are not rewritten.

Implementation: normalizeAssistantReplayContent in src/agents/embedded-agent-runner/replay-history.ts


Global rule: inter-session input provenance

When an agent sends a prompt into another session via sessions_send (including a delayed reply delivered to the requester), OpenClaw persists the created user turn with message.provenance.kind = "inter_session".

OpenClaw also prepends a same-turn [Inter-session message] ... isUser=false marker before the routed prompt text so the active model call can distinguish foreign session output from external end-user instructions. This marker includes the source session, channel, and tool when available. The transcript still uses role: "user" for provider compatibility, but the visible text and provenance metadata both mark the turn as inter-session data.

During context rebuild, OpenClaw applies the same marker to older persisted inter-session user turns that only have provenance metadata.


Provider matrix (current behavior)

OpenAI / OpenAI Codex

  • Image sanitization only.
  • Drop orphaned reasoning signatures (standalone reasoning items without a following content block) for OpenAI Responses/Codex transcripts, and drop replayable OpenAI reasoning after a model route switch.
  • Preserve replayable OpenAI Responses reasoning item payloads, including encrypted empty-summary items, so manual/WebSocket replay keeps required rs_* state paired with assistant output items.
  • Native ChatGPT Codex Responses follows Codex wire parity by replaying prior Responses reasoning/message/function payloads without prior item IDs while preserving session prompt_cache_key.
  • OpenAI Responses-family replay preserves canonical call_*|fc_* same-model reasoning pairs, but deterministically normalizes malformed or overlong call_id/function-call item ids before pi-ai payload conversion.
  • Tool result pairing repair may move real matched outputs and synthesize Codex-style aborted outputs for missing tool calls.
  • No turn validation or reordering; no thought signature stripping.

OpenAI-compatible Chat Completions

  • Historical assistant thinking/reasoning blocks are stripped before replay so local and proxy-style OpenAI-compatible servers do not receive prior-turn reasoning fields such as reasoning or reasoning_content.
  • Current same-turn tool-call continuations keep the assistant reasoning block attached to the tool call until the tool result has been replayed.
  • Custom/self-hosted model entries with reasoning: true preserve replayed reasoning metadata.
  • Provider-owned exceptions can opt out when their wire protocol requires replayed reasoning metadata.

Google (Generative AI / Gemini CLI / Antigravity)

  • Tool call id sanitization: strict alphanumeric.
  • Tool result pairing repair and synthetic tool results.
  • Turn validation (Gemini-style turn alternation).
  • Google turn ordering fixup (prepend a tiny user bootstrap if history starts with assistant).
  • Antigravity Claude: normalize thinking signatures; drop unsigned thinking blocks.

Anthropic / Minimax (Anthropic-compatible)

  • Prefix-binding Claude models, such as Fable 5.1, persist runtime-context carriers as hidden custom messages immediately after their user turn and replay them in place. Inline inbound metadata on older user turns is also retained. This model-scoped append-only policy includes Bedrock, Vertex, and Foundry routes. Carriers contain only the delimited context body; the shared instruction lives once in the stable system prompt. Carriers remain user-role context and are excluded from chat history and compaction summarization. Other Claude models and Anthropic-compatible models keep transient carriers, avoiding repeated cache-read charges and context use for old carriers when nothing binds the prefix.
  • Tool result pairing repair and synthetic tool results.
  • Turn validation (merge consecutive user turns to satisfy strict alternation). For prefix-binding models on the Messages API, append-only replay keeps consecutive user turns separate instead, so a command turn followed by a prompt replays with the same per-turn timestamp stamps the active turn was signed over; Bedrock Converse still merges them.
  • Trailing assistant prefill turns are stripped from outgoing Anthropic Messages payloads when thinking is enabled, including Cloudflare AI Gateway routes.
  • Pre-compaction assistant thinking signatures are stripped before provider replay when a session has been compacted. On prefix-binding models, compaction changes the signed prefix (summarized content replaces the original), so replaying the original signatures can cause Anthropic to reject the request with "Invalid signature in thinking block". The thinking text is preserved as an unsigned block and then handled by the rule below.
  • Thinking blocks with missing, empty, or blank replay signatures are stripped before provider conversion. If that empties an assistant turn, OpenClaw keeps turn shape with non-empty omitted-reasoning text.
  • Older thinking-only assistant turns that must be stripped are replaced with non-empty omitted-reasoning text so provider adapters do not drop the replay turn.

Amazon Bedrock (Converse API)

  • Empty assistant stream-error turns and legacy fallback placeholders are dropped from the in-memory replay copy. This avoids invalid empty ContentBlocks and synthetic assistant prefill without rewriting the stored transcript.
  • Zero-usage empty stop turns are dropped too; billed silent replies and errors with real assistant content retain their existing replay handling.
  • Pre-compaction assistant thinking signatures are stripped before Converse replay when a session has been compacted, for the same reason as Anthropic above.
  • Claude thinking blocks with missing, empty, or blank replay signatures are stripped before Converse replay. If that empties an assistant turn, OpenClaw keeps turn shape with non-empty omitted-reasoning text.
  • Older thinking-only assistant turns that must be stripped are replaced with non-empty omitted-reasoning text so the Converse replay keeps strict turn shape.
  • Replay filters OpenClaw delivery-mirror and gateway-injected assistant turns.
  • Image sanitization applies through the global rule.

Mistral (including model-id based detection)

  • Tool call id sanitization: strict9 (alphanumeric, length 9).

OpenRouter Gemini

  • Thought signature cleanup: strip non-base64 thought_signature values (keep base64).

OpenRouter Anthropic

  • Trailing assistant prefill turns are stripped from verified OpenRouter OpenAI-compatible Anthropic model payloads when reasoning is enabled, matching direct Anthropic and Cloudflare Anthropic replay behavior.

Everything else

  • Image sanitization only.

Historical behavior (pre-2026.1.22)

Before the 2026.1.22 release, OpenClaw applied multiple layers of transcript hygiene:

  • A transcript-sanitize extension ran on every context build and could:
    • Repair tool use/result pairing.
    • Sanitize tool call ids (including a non-strict mode that preserved _/-).
  • The runner also performed provider-specific sanitization, which duplicated work.
  • Additional mutations occurred outside the provider policy, including stripping <final> tags from assistant text before persistence, dropping empty assistant error turns, and trimming assistant content after tool calls.

This complexity caused cross-provider regressions (notably openai-responses call_id|fc_id pairing). The 2026.1.22 cleanup removed the extension, centralized logic in the runner, and made OpenAI no-touch beyond image sanitization.