openclaw/docs
Ayaan Zaidi d4131afa00
fix(active-memory): recalls on claude-cli never reuse the prompt cache (#160768)
Closes #160672

## What Problem This Solves

Fixes: every Active Memory recall on the `claude-cli` runtime writes the whole OpenClaw system prompt into Claude's prompt cache again, even for two recalls of the same agent seconds apart.

## User Impact

User impact: repeated Active Memory recalls of the same agent on `claude-cli` now send a byte-identical system prompt, so the recall's first API call can read it from Claude's cache instead of paying for a 10k+ token cache write each time. Normal `claude-cli` turns are unchanged.

## Why This Change Was Made

Each recall gets a new session key (`<parent>:active-memory:<per-run hash>`) so its session entry stays unique. The Runtime line rendered that key into the system prompt, and Claude CLI receives the whole system prompt as one `appendSystemPrompt`, so the per-run key forced a full cache rewrite. The recall's session key is the only value that differed between two recalls.

One-shot CLI dispatch runs (today only Active Memory recall) now carry the Runtime facts line in their only user turn, through the same `UserPromptSubmit` context Claude CLI already uses for other per-turn facts. This reuses the existing relocatable Runtime region that Chat Completions routes already move into the first user message. The recall model still sees its agent, session, model and channel. Recall session keys, storage and cleanup are unchanged. Normal resumable CLI turns keep the Runtime line in the system prompt.

Thanks @ndakota79 for the detailed cache-counter report and cause analysis.

No overlap with Pash/Sarah changes.

## Evidence

Isolated Gateway (task-owned home, state, config, port) with agents `alpha` and `beta` on `agentRuntime: claude-cli`, Active Memory `mode: always`, and a fake `claude` executable that records the `initialize.appendSystemPrompt` and the `UserPromptSubmit` context it receives. Two `chat.send` turns on `agent:alpha:main` about 60 s apart, then one on `agent:beta:main`.

- Base (`7f0781faa67`): the two alpha recalls' system prompts differ in exactly one line:
  ```
  -Runtime: agent=alpha | session=agent:alpha:main:active-memory:e10ad4c9dc5e | ...
  +Runtime: agent=alpha | session=agent:alpha:main:active-memory:9945b3fb179b | ...
  ```
- Candidate: both alpha recalls send the same system prompt (14,210 bytes, sha256 `f7a45b89f4e84e5f…`). Each recall's user turn carries its own `Runtime: agent=alpha | session=agent:alpha:main:active-memory:<hash> | ... | channel=webchat | ...`. The candidate recall prompt equals the base recall prompt minus the Runtime line.
- Controls: the beta recall still gets its own prompt (its working directory differs) and its own `agent=beta` Runtime facts in the turn. The alpha and beta main (non-recall) turns send the same system prompt and user turn as on base, with the Runtime line still in the system prompt.
- Regression: `src/agents/cli-runner/prepare.test.ts` "keeps per-run helper session identities out of the reusable system prompt" fails on base (system prompts differ) and passes with the fix. `prepare.test.ts` (194), `cli-backend-dispatch.test.ts` (40) and `helpers.system-prompt.test.ts` (15) pass.
- Test cost: the new test runs in 136 ms. `pnpm test src/agents/cli-runner/prepare.test.ts --maxWorkers=1`: 194 passed, 83.3 s wall (Vitest 65.9 s); the same command filtered to the new test: 23.1 s wall (Vitest 5.8 s).

Not verified: a real Claude Code run. The proof checks the bytes OpenClaw sends to the CLI (the `initialize` system prompt and the `UserPromptSubmit` response); it does not show Claude Code's cache counters. The Runtime facts use the same `UserPromptSubmit` additional-context path that already carries `before_prompt_build` hook context (including Active Memory's own recalled context) on `claude-cli` turns.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-29 04:52:53 +05:30
..
.generated fix: keep Codex chats working during slow model discovery (#160363) 2026-09-28 11:43:38 -07:00
.i18n fix(openai): retire the Sora video provider after the API shutdown (#159543) 2026-09-27 22:52:14 +00:00
announcements
assets docs: remove stale showcase section (#158575) 2026-09-25 21:11:35 -06:00
automation fix(cron): named-session jobs run in the wrong workspace (#159444) 2026-09-28 14:54:53 -07:00
channels fix(slack): progress card links open the visible work session instead of the main session (#160088) 2026-09-28 22:21:00 +00:00
ci fix(ci): pack hybrid hourly tooling within the row cap 2026-09-28 15:04:46 -07:00
cli feat(plugins): show installed accounts and credential status (#160092) 2026-09-28 15:18:42 -07:00
concepts fix: streamed replies split Markdown tables that fit one message (#160701) 2026-09-29 03:41:47 +05:30
diagnostics perf(cli): keep client diagnostics off shared state 2026-09-24 12:56:28 +00:00
gateway perf(nodes): reuse warm workers so node session turns start as fast as local ones (#160265) 2026-09-28 16:07:44 -07:00
help fix(test): include all runners for package directories (#160653) 2026-09-28 14:36:06 -07:00
images
install fix(node-host): updates never install on Bun-only macOS and Linux hosts (#160575) 2026-09-28 14:59:44 -07:00
maturity refactor: remove Tasks and TaskFlow runtime (#159179) 2026-09-27 10:40:29 -07:00
nodes perf(nodes): reuse warm workers so node session turns start as fast as local ones (#160265) 2026-09-28 16:07:44 -07:00
platforms fix(ios): preserve setup after bootstrap preparation refuses (#160218) 2026-09-28 09:25:24 -07:00
plugins fix(http): adapt rejection transport to Bun's Node-compatible closure (#160329) 2026-09-28 15:49:05 -07:00
providers feat: speak with Gemini 3.8 Flash TTS (#157331) 2026-09-28 15:23:34 -07:00
reference fix(active-memory): recalls on claude-cli never reuse the prompt cache (#160768) 2026-09-29 04:52:53 +05:30
releases fix: exclude repository instructions from public docs sync (#158253) 2026-09-25 13:15:24 -06:00
security fix(proxy): keep WebChat and Codex loopback calls direct (#154013) 2026-09-20 21:42:01 -07:00
snippets/plugin-publish chore(deps): refresh dependencies with seven-day cutoff (#158298) 2026-09-26 20:42:53 -07:00
specs
start fix(onboarding): stop claiming inference is ready after Skip for now (#160713) 2026-09-28 15:17:40 -07:00
tools feat: speak with Gemini 3.8 Flash TTS (#157331) 2026-09-28 15:23:34 -07:00
web fix(ui): open voice setup before history admission (#160686) 2026-09-28 15:50:18 -07:00
agent-runtime-architecture.md perf(agents): keep file edits from blocking other chats (#158443) 2026-09-26 03:55:55 -05:00
AGENTS.md fix: exclude repository instructions from public docs sync (#158253) 2026-09-25 13:15:24 -06:00
auth-credential-semantics.md fix(auth): keep session account selection after OAuth re-login (#149591) 2026-09-27 18:07:54 +05:30
ci.md fix(e2e): size first-hop budgets from measured update times 2026-09-28 13:22:12 -07:00
date-time.md
docs.json feat(video): add Kie AI, Z.AI, and Novita video generation (#160080) 2026-09-28 12:48:22 +00:00
docs_map.md
index.md docs: remove stale showcase section (#158575) 2026-09-25 21:11:35 -06:00
logging.md fix: restore steering across queued and active turns (#158699) 2026-09-26 13:29:49 -07:00
network.md
openclaw-agent-runtime.md
prose.md
vps.md
whatsapp-openclaw.jpg