openclaw/docs/cli/doctor.md
Peter Steinberger 5b90a73498
fix(usage): prevent refresh OOMs on large SQLite sessions (#156937)
Closes #156898

## What Problem This Solves

Large SQLite sessions can repeatedly exhaust the usage-refresh worker's 512 MiB heap, leaving usage totals missing or stale after an otherwise successful `sessions.usage` response.

## User Impact

Large sessions can finish refreshing without increasing the worker limit. Doctor reports bounded, per-session refresh failures and successful refreshes clear their warnings. No configuration, rollup format, database schema, or migration changes are required. Thanks @Conan-Scott for the detailed report and allocation control.

## Why This Change Was Made

The reader eagerly decoded the entire requested range before aggregation's 128-record batches. It now reads at most 1,024 rows and 8 MiB of decoded JSON per page, with one lookahead event (an individual oversized identity event is still accepted). A first pass retains compact navigation facts; selected bodies then feed the existing aggregation in ancestry order. Append validation, reset/leaf fallback, checkpoint checks, and conditional publication remain in place. Navigation metadata still scales with event count; payload bodies do not. The existing incognito host-frame reader is unchanged.

Refresh failures use the existing asynchronous core plugin-state storage, capped at 256 session facts, without storing raw error or transcript content. The cache owner clears each fact after successful publication; Doctor displays unresolved failures.

## Evidence

All heavy validation ran on Blacksmith Testbox (`blacksmith-testbox`, profile `openclaw-check`). Primary measurements and typechecks: lease `tbx_01m38g2xckkz1apk1smezhftg1`, Node 24.19.0, [run](https://github.com/openclaw/openclaw/actions/runs/35942606687). Final lint, guards, paging tests, and benchmark replay: lease `tbx_01m38k1hj8p9r5ay910z20s2sv`, [run](https://github.com/openclaw/openclaw/actions/runs/35946363096). Both leases stopped after validation.

The opt-in `node --import tsx scripts/bench-usage-refresh-memory.ts` fixture uses 24,000 events, 6,000 identity rows and 18,000 production zstd rows: 1,186,945,771 decoded bytes (1.105 GiB), with each event below 4 MiB. Both runs use the real refresh worker and its production 512 MiB limit; the baseline substitutes the original reader/scanner from `57f912b5f7` before executing the same harness with `--expect-oom`.

| Measurement | Original reader | Bounded reader |
| --- | --- | --- |
| Outcome | `ERR_WORKER_OUT_OF_MEMORY` | Complete, exact reference rollup |
| Sampled peak worker heap | 489.8 MiB | 228.9 MiB (53.3% lower) |
| Refresh duration | Failed after 2.39 s | Completed in 4.64 s |

The final-source replay completed in 4.40 s at 234.4 MiB sampled peak, again with exact rollup equality and below the asserted 384 MiB bound.

Heap sampling uses `Worker.getHeapStatistics()` every 10 ms; these are observed peaks, not complete allocation profiles. Failure time is not a throughput comparison. The complete rollup matched the bounded JSONL worker reference, independently checked for 24,000 records, 240,000 tokens, and 24,000 synthetic cost units. Rollup SHA-256: `f0f184c9b958a74211b6d66678bb1808497d07a96b6b8a84cfb1caabfdc9a856`.

- `node scripts/run-vitest.mjs src/infra/session-cost-usage-worker-refresh.test.ts src/infra/session-cost-usage-cache.worker.test.ts src/commands/doctor-usage-cost-cache.test.ts src/infra/session-cost-usage-worker-io.test.ts --maxWorkers=1`: 27 tests passed, 65.98 s total including cold worker compilation. New paging file: 5 tests / 7 ms. Changed worker file: 8 tests / 14.93 s; new failure→Doctor→recovery test: 1.76 s.
- `node scripts/check-changed.mjs`: all selected checks completed successfully across the initial run and resumed native-plan commands. Core types and all 25 test-typecheck shards passed; script/root-test types, lint, formatting, storage/import-cycle/security guards also passed. The initial run caught optional `Worker.resourceLimits` typing; lint caught parameter reassignment and benchmark brace/untyped-throw issues. All were corrected and their failed checks replayed successfully.
- Final `node scripts/run-vitest.mjs src/infra/session-cost-usage-worker-refresh.test.ts --maxWorkers=1`: 5 passed / 7 ms test time, 22.28 s wrapper wall including cold worker compilation.
- Final `node --import tsx scripts/bench-usage-refresh-memory.ts`: passed, same rollup hash as the paired measurement.
- Independent Codex review: no actionable P0–P2 findings.

Initial proof attempts exposed two fixture/setup issues, both corrected before collecting the measurements: missing Testbox checkout dependencies and a synthetic session row that bypassed canonical admission. The final fixture creates the session through its owner. Two later native checksum sync attempts timed out before validation; publishing the reviewed branch and warming a fresh lease recovered native targeted sync. No production data or Gateway deployment was used; this proves the synthetic allocation defect and rollup equality, not attribution of all reported Gateway RSS growth.
2026-09-24 02:50:28 +00:00

6.3 KiB

summary read_when title
CLI reference for `openclaw doctor` (health checks + guided repairs)
You have connectivity/auth issues and want guided fixes
You updated and want a sanity check
Doctor CLI

openclaw doctor

Health checks and quick fixes for the gateway, channels, plugins, skills, model routing, local state, and config migrations. Use it whenever something is not behaving as expected and you want one command to explain what is wrong.

When run for a managed Gateway, Doctor compares active official plugins with the OpenClaw package referenced by the installed service. This check still works when the Gateway is stopped or unreachable. When an older Gateway is still running, Doctor reports its version separately from the post-restart version. If the service package cannot be identified, Doctor reports restart readiness as unknown instead of treating the plugin set as compatible.

When Gateway status reports degraded SecretRef owners, doctor prints a Secret runtime degradation warning with every cold or stale owner, affected config path, redacted reason, and the openclaw secrets reload retry command.

When channel ingress events are dead-lettered, doctor names each affected channel account and points to openclaw channels dead-letters list for inspection and recovery.

Doctor warns when a registry-owned project clone is partial or shallow. It names the clone, shallow state, and partial-clone config keys, including URL-keyed remote twins. It prints manual repair commands; --fix does not fetch or repack these clones. Agent workspaces and manually registered checkouts are excluded.

When the Gateway has exporter health facts, doctor reports the latest trusted per-signal state and transport under Telemetry exporters. The summary is redacted and does not include endpoint values, headers, certificates, payloads, or raw errors.

Doctor reports sessions whose usage-cost cache refresh failed, since their totals may be incomplete. Check the Gateway logs and request usage again to retry. The bounded failure history keeps the latest 256 sessions across restarts; a successful refresh clears that session's warning. --fix does not clear a warning before the session has refreshed successfully.

Related:

Doctor pages

This page is an index. openclaw doctor is documented on seven pages, one per reader job. Open the page that matches your task.

Page Read it when
Run doctor Pick a posture, copy a working example, or look up what an option does.
Gateway and service recovery The Gateway service, remote target, Control UI assets, or Gateway token needs repair.
Lint and post-upgrade modes You want read-only findings for a CI gate, or post-upgrade plugin compatibility probes.
Structured health check contract You are writing a doctor check or a plugin-backed health check.
Legacy state migration A file-to-SQLite migration is blocked and needs manual reconciliation.
SQLite maintenance and session migration You are compacting a database, or importing, validating, or recovering session history.
Other checks and repairs You want the inventory of every remaining check and repair, from Nix mode to channels.

Where each section moved

Every section heading from the previous single-page version keeps its anchor here, so an existing link such as /cli/doctor#session-sqlite-migration still resolves. Each entry points at the page that now holds the content.