Commit graph

1 commit

Author SHA1 Message Date
Peter Steinberger
5b90a73498
fix(usage): prevent refresh OOMs on large SQLite sessions (#156937)
Closes #156898

## What Problem This Solves

Large SQLite sessions can repeatedly exhaust the usage-refresh worker's 512 MiB heap, leaving usage totals missing or stale after an otherwise successful `sessions.usage` response.

## User Impact

Large sessions can finish refreshing without increasing the worker limit. Doctor reports bounded, per-session refresh failures and successful refreshes clear their warnings. No configuration, rollup format, database schema, or migration changes are required. Thanks @Conan-Scott for the detailed report and allocation control.

## Why This Change Was Made

The reader eagerly decoded the entire requested range before aggregation's 128-record batches. It now reads at most 1,024 rows and 8 MiB of decoded JSON per page, with one lookahead event (an individual oversized identity event is still accepted). A first pass retains compact navigation facts; selected bodies then feed the existing aggregation in ancestry order. Append validation, reset/leaf fallback, checkpoint checks, and conditional publication remain in place. Navigation metadata still scales with event count; payload bodies do not. The existing incognito host-frame reader is unchanged.

Refresh failures use the existing asynchronous core plugin-state storage, capped at 256 session facts, without storing raw error or transcript content. The cache owner clears each fact after successful publication; Doctor displays unresolved failures.

## Evidence

All heavy validation ran on Blacksmith Testbox (`blacksmith-testbox`, profile `openclaw-check`). Primary measurements and typechecks: lease `tbx_01m38g2xckkz1apk1smezhftg1`, Node 24.19.0, [run](https://github.com/openclaw/openclaw/actions/runs/35942606687). Final lint, guards, paging tests, and benchmark replay: lease `tbx_01m38k1hj8p9r5ay910z20s2sv`, [run](https://github.com/openclaw/openclaw/actions/runs/35946363096). Both leases stopped after validation.

The opt-in `node --import tsx scripts/bench-usage-refresh-memory.ts` fixture uses 24,000 events, 6,000 identity rows and 18,000 production zstd rows: 1,186,945,771 decoded bytes (1.105 GiB), with each event below 4 MiB. Both runs use the real refresh worker and its production 512 MiB limit; the baseline substitutes the original reader/scanner from `57f912b5f7` before executing the same harness with `--expect-oom`.

| Measurement | Original reader | Bounded reader |
| --- | --- | --- |
| Outcome | `ERR_WORKER_OUT_OF_MEMORY` | Complete, exact reference rollup |
| Sampled peak worker heap | 489.8 MiB | 228.9 MiB (53.3% lower) |
| Refresh duration | Failed after 2.39 s | Completed in 4.64 s |

The final-source replay completed in 4.40 s at 234.4 MiB sampled peak, again with exact rollup equality and below the asserted 384 MiB bound.

Heap sampling uses `Worker.getHeapStatistics()` every 10 ms; these are observed peaks, not complete allocation profiles. Failure time is not a throughput comparison. The complete rollup matched the bounded JSONL worker reference, independently checked for 24,000 records, 240,000 tokens, and 24,000 synthetic cost units. Rollup SHA-256: `f0f184c9b958a74211b6d66678bb1808497d07a96b6b8a84cfb1caabfdc9a856`.

- `node scripts/run-vitest.mjs src/infra/session-cost-usage-worker-refresh.test.ts src/infra/session-cost-usage-cache.worker.test.ts src/commands/doctor-usage-cost-cache.test.ts src/infra/session-cost-usage-worker-io.test.ts --maxWorkers=1`: 27 tests passed, 65.98 s total including cold worker compilation. New paging file: 5 tests / 7 ms. Changed worker file: 8 tests / 14.93 s; new failure→Doctor→recovery test: 1.76 s.
- `node scripts/check-changed.mjs`: all selected checks completed successfully across the initial run and resumed native-plan commands. Core types and all 25 test-typecheck shards passed; script/root-test types, lint, formatting, storage/import-cycle/security guards also passed. The initial run caught optional `Worker.resourceLimits` typing; lint caught parameter reassignment and benchmark brace/untyped-throw issues. All were corrected and their failed checks replayed successfully.
- Final `node scripts/run-vitest.mjs src/infra/session-cost-usage-worker-refresh.test.ts --maxWorkers=1`: 5 passed / 7 ms test time, 22.28 s wrapper wall including cold worker compilation.
- Final `node --import tsx scripts/bench-usage-refresh-memory.ts`: passed, same rollup hash as the paired measurement.
- Independent Codex review: no actionable P0–P2 findings.

Initial proof attempts exposed two fixture/setup issues, both corrected before collecting the measurements: missing Testbox checkout dependencies and a synthetic session row that bypassed canonical admission. The final fixture creates the session through its owner. Two later native checksum sync attempts timed out before validation; publishing the reviewed branch and warming a fresh lease recovered native targeted sync. No production data or Gateway deployment was used; this proves the synthetic allocation defect and rollup equality, not attribution of all reported Gateway RSS growth.
2026-09-24 02:50:28 +00:00