mirror of
https://github.com/openclaw/openclaw.git
synced 2026-10-03 01:29:56 +00:00
## What Problem This Solves
Intermittent tasks repeatedly recreate healthy Workers after the existing 30–60 second idle timeout, paying isolate startup costs and leaving resident memory behind after teardown. Existing metrics cannot distinguish which scripts are starting and exiting.
## User Impact
Hot task pools retain one Worker for five minutes of inactivity after an idle-retired Worker is promptly needed again. Node Code Mode retains its existing one-worker cache for five minutes. Operators can graph `openclaw_worker_started_total{script}` and `openclaw_worker_retired_total{script,reason}` through the existing diagnostics heartbeat.
## Why This Change Was Made
The existing retirement owner now owns idle timers and the single warm slot. Other slots keep their ordinary timeout; native cleanup, cancellation, pressure retirement, and rotation retain their existing ownership barriers. Existing idle GC collects released payloads in place. Code Mode still checks the runtime entry and heap limit before reuse.
The existing Worker resource registry owns cumulative lifecycle counts, including bundled direct Worker constructors. Counts advance on successful creation and confirmed native exit, survive exporter restart, and use bounded labels. Eval/unknown Workers remain `other`; nested Workers and V8-native threads remain outside the parent registry. Memory projection moves into a focused Prometheus module to keep the existing service below its growth ratchet.
No new configuration, database changes, dependencies, or managed-service environment changes. Updating installs the new runtime behavior through the normal restart; no migration or operator-state mutation is needed.
## Evidence
Real production-owner calls, ten verified synthetic tasks at 70-second simulated idle intervals (630 seconds observed), with real Worker creation and native teardown:
| Script | Starts before → after | Starts/min before → after | Task-run wall before → after |
| --- | ---: | ---: | ---: |
| `cron-stream-matcher.worker.js` | 10 → 2 | 0.952 → 0.190 | 485 → 111 ms |
| `git-operation.worker.js` | 10 → 2 | 0.952 → 0.190 | 3,076 → 824 ms |
| `code-mode-node.worker.js` | 10 → 1 | 0.952 → 0.095 | 858 → 187 ms |
The original source fails both new warm-window regressions; the pool creates a third Worker and Code Mode creates a second pool. In-place GC proof releases a 138 MiB payload to a 7.4 MiB heap while reusing the Worker.
### Allocator experiment and limits
Requested the largest advertised Blacksmith Linux class, 32 vCPUs, through a temporary proof-only workflow branch. The resulting guest exposes **8 online CPUs**, so this does **not** fulfill ≥32-vCPU validation or prove production memory behavior. Node 24.19.0, glibc 2.39, 32 concurrent Workers, 1,000 identical tasks touching 8 MiB each:
| Fresh process | Worker starts | Wall | RSS minus aggregate heapTotal increase after all Workers exit |
| --- | ---: | ---: | ---: |
| Create/terminate each task | 1,000 | 8.11 s | +639.59 MiB |
| Same churn, `MALLOC_ARENA_MAX=2` | 1,000 | 7.54 s | +423.86 MiB |
| Reuse 32 Workers | 32 | 3.08 s | +512.05 MiB |
The arena control reduces final residue 33.7%; reuse reduces it 19.9%. The churn proxy's post-100 slope is 102.63 KiB/task, versus 19.26 KiB/task with the arena control. A comparable retained-worker proxy slope is **not interpretable**: V8 heap capacity shrinks 992 MiB while RSS remains about 716.5 MiB, producing negative subtraction values and a false slope. Retained raw RSS grows only 0.152 MiB from tasks 400–1,000. These short exploratory results do not establish a many-core native-memory slope or allocator-contention tradeoff; consequently this PR does not set `MALLOC_ARENA_MAX`. Live deployment/production attribution remains a follow-up.
### Validation
Provider: `blacksmith-testbox`; profile: `openclaw-check`; lease: `tbx_01m36makaczyy9pds4yt5t2pem`; [workflow run](https://github.com/openclaw/openclaw/actions/runs/35834466664). One active remote command at a time, normal synchronization throughout. All test/typecheck/build work ran remotely.
- `node scripts/check-changed.mjs --base HEAD -- <all 25 changed paths>`: passed in 43m18s, including all 25 core-test typecheck graphs, core/extension production and test types, lint, documentation, SDK boundaries, dead exports, and import cycles.
- `pnpm test <file> --maxWorkers=1` for every changed test file, then `node scripts/run-vitest.mjs` for the five affected lifecycle/transport siblings: passed; 153 tests across 11 files. Final combined remote command: 59.2s. Original-source targeted runs failed both new regressions for the intended extra Worker/pool creation; candidate runs pass.
- The production-owner before/after benchmark command passed, including synthetic output assertions and native-exit cleanup for all three scripts. The allocator benchmark completed all three fresh-process cases with identical checksums.
- `pnpm build` followed by `OPENCLAW_LOCAL_CHECK=0 node --import tsx scripts/profile-extension-memory.mts --extension telegram --skip-combined --concurrency 1`: passed in 3m54s combined. All 95 plugin distributions built; bootstrap and 53 native control-plane module checks passed. Telegram isolated import exited cleanly, 160.09 MiB peak RSS (113.85 MiB above the profiler's empty-process baseline).
- Independent Codex review: scoped-clean, no actionable P0–P2 findings. `git diff --check`: passed. Net production growth: 146 lines; no test-only production seam.
Measured single-worker test command cost (seconds, including setup):
| Changed file | Wall |
| --- | ---: |
| `src/infra/worker-task-pool.test.ts` | 20.33 |
| `src/infra/worker-cpu.test.ts` | 7.48 |
| `src/logging/diagnostic-memory.test.ts` | 6.58 |
| `src/agents/code-mode-node.lifecycle.test.ts` | 5.21 |
| `extensions/diagnostics-prometheus/src/service.event-loop.test.ts` | 1.81 |
| `extensions/telegram/src/telegram-ingress-worker.test.ts` | 5.18 |
The two added behavior regressions use a fake idle clock; the real-worker pool case took 45ms and the Code Mode lifecycle case 35ms. CI timing will be updated after the head run exists. Crabbox reported an external runner-portal sync timeout after successful changed-check, final-test, and build commands; their underlying commands and reported run status succeeded.
Setup failures were diagnosed rather than counted as product failures: the transport checkout initially lacked `tsx` (fixed with remote frozen install), has no cgroup-v2 `cpu.max`, and lacks `origin/main` (changed checks use the actual original `HEAD` base plus exact changed paths). The initial benchmark fixture awaited an unrelated/unreferenced Worker exit; the corrected fixture uses the pool's resource-release boundary. The first changed gate identified line-cap growth; moving idle timing into its owner and extracting memory projection resolved it without an exception.
|
||
|---|---|---|
| .. | ||
| assets | ||
| skills/discord | ||
| src | ||
| test | ||
| account-inspect-api.ts | ||
| action-runtime-api.ts | ||
| activities-api.ts | ||
| api.ts | ||
| channel-config-api.ts | ||
| channel-plugin-api.ts | ||
| config-api.ts | ||
| config-doctor-api.ts | ||
| contract-api.ts | ||
| directory-contract-api.ts | ||
| doctor-contract-api.ts | ||
| index.ts | ||
| openclaw.plugin.json | ||
| package.json | ||
| README.md | ||
| runtime-api.ts | ||
| runtime-setter-api.ts | ||
| secret-contract-api.ts | ||
| security-audit-contract-api.ts | ||
| security-contract-api.ts | ||
| session-key-api.ts | ||
| setup-entry.test.ts | ||
| setup-entry.ts | ||
| setup-plugin-api.ts | ||
| subagent-hooks-api.ts | ||
| test-api.ts | ||
| thread-binding-api.ts | ||
| transcripts-source-api.ts | ||
| tsconfig.json | ||
OpenClaw Discord
Official OpenClaw channel plugin for Discord servers, channels, DMs, slash commands, and app events.
Install from OpenClaw:
openclaw plugins install @openclaw/discord
Configure a Discord bot token and the channels or servers OpenClaw should handle. The plugin lets OpenClaw agents receive Discord messages and respond through the configured Discord app.