From 0fbded75d298e9bf5cd94d63bd1b58d6bdfd1c88 Mon Sep 17 00:00:00 2001 From: RoboClaw Date: Sun, 20 Sep 2026 21:18:58 -0700 Subject: [PATCH] docs: keep update requests owned through acceptance (#154111) * docs: keep update requests owned through acceptance Co-authored-by: steipete <58493+steipete@users.noreply.github.com> Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com> * docs: notify Team only around actual update downtime Keep non-action outcomes private and summarize the exact accepted main range after verified recovery. Preserve update ownership and distinguish rollback from successful deployment. Co-authored-by: steipete <58493+steipete@users.noreply.github.com> Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com> * docs: own confirmed updater repairs through acceptance Qualify defects at the failing path, use isolated worktrees and subagents where useful, and land tested reviewed repairs without bypassing protection. Co-authored-by: steipete <58493+steipete@users.noreply.github.com> Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com> --------- Co-authored-by: steipete <58493+steipete@users.noreply.github.com> Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com> --- .agents/skills/openclaw-update/SKILL.md | 10 +++++-- .agents/skills/update-team-server/SKILL.md | 34 ++++++++++++++++------ 2 files changed, 33 insertions(+), 11 deletions(-) diff --git a/.agents/skills/openclaw-update/SKILL.md b/.agents/skills/openclaw-update/SKILL.md index 79d3b2a4ff52..cf850fa2a841 100644 --- a/.agents/skills/openclaw-update/SKILL.md +++ b/.agents/skills/openclaw-update/SKILL.md @@ -20,7 +20,7 @@ Read the owner runbook and choose the matching workflow: | Other managed deployments | Deployment-specific skill/runbook and existing owner, including containers and declarative installations. | | Standard package or single-user source | [Updating](../../../docs/install/updating.md) and the [update CLI contract](../../../docs/cli/update.md), using the owning CLI on the target host/profile. | -Follow the selected workflow's commands, locks, backup/migration requirements, recovery, and cleanup. Do not apply standard CLI updates inside a separately managed release or running container, mutate files used by a live Gateway, or overwrite dirty checkouts. +Follow the selected workflow's commands, locks, backup/migration requirements, recovery, and cleanup. For Team deployments, the linked Team skill also owns notification timing and content; report non-action outcomes and blockers privately to the requesting user, not to Team. Do not apply standard CLI updates inside a separately managed release or running container, mutate files used by a live Gateway, or overwrite dirty checkouts. “Latest” preserves the configured channel; explicitly requested “latest main” selects the development source flow (`--channel dev` for the standard CLI). Preview the target and report the actual serving SHA: development preflight can select an older buildable commit. Do not auto-accept unapproved downgrades or plugin capability changes. @@ -32,6 +32,12 @@ Task cancellation, session aborts, and queue clearing can discard work; do not u If the owner lacks controlled interruption, report or repair that gap within scope. Do not invent force flags, bypass persistence/compatibility gates, or repeat the interruption approval question. +## Follow through + +Retain ownership of an explicit update request across busy deferral, failure, rollback, and recovery until the selected owner's full acceptance or a concrete external/access blocker requiring outside action. Respect user pauses and cancellations. Before ending with work pending, establish an active completion/observation path or supported continuation through that owner; return without waiting for another prompt. If no continuation path exists, state that blocker instead of promising one. An existing cadence or restored old process alone does not complete the request. + +Keep deployment-specific recovery and target selection with the linked owner. Qualify suspected issues at the actual failing path, improve updater guidance for confirmed durable gaps, and repair confirmed defects within the authorized scope. Use isolated worktrees and subagents where useful; test and review repairs, then land them through normal PR review, CI, and `scripts/pr` landing. Continue through verified update acceptance after landing; persistence does not authorize bypassing locks, safeguards, or branch protection. + ## Verify -Use the owner's verification to prove the intended version/commit is serving, RPC/health and configured channels work, and UI/assets load through actual ingress when enabled. Inspect startup/recovery errors without exposing secrets. A build, handoff acknowledgement, or healthy old process is not completion. Report the serving version and remaining issues. +Use the owner's verification to prove the intended version/commit is serving, RPC/health and configured channels work, and UI/assets load through actual ingress when enabled. Inspect startup/recovery errors without exposing secrets. Read the native deployment result separately from an observer or wrapper exit status: a successful observer can report a deferred or failed update. A build, handoff acknowledgement, or healthy old process is not completion. Keep detailed receipts private; report the result, next action or exact blocker in one to three short, friendly lines, with a short SHA only when useful and no repeated logs or full hashes. diff --git a/.agents/skills/update-team-server/SKILL.md b/.agents/skills/update-team-server/SKILL.md index 49f72f78a47f..b7b34b3df7c0 100644 --- a/.agents/skills/update-team-server/SKILL.md +++ b/.agents/skills/update-team-server/SKILL.md @@ -1,6 +1,6 @@ --- name: update-team-server -description: "Update the operator-configured Team server unattended through its canonical deployment owner; verify serving code, supported migrations, session continuity, and recovery without duplicating the hourly scheduler." +description: "Update the operator-configured Team server unattended through its canonical deployment owner; verify serving code, supported migrations, session continuity, and recovery without duplicating the configured scheduler." --- # Update Team server @@ -9,16 +9,32 @@ Keep Team current automatically. Routine deployments, controlled interruptions, ## Resolve the owner -Read the operator-provided private deployment runbook before acting. Resolve and verify the host, access, canonical deployment command, service, hourly timer, lock, journal, release pointers, backup destination, and recovery contract from that configuration. Never guess access or copy private connection details, credentials, state, or receipts into public output. +Read the operator-provided private deployment runbook before acting. Resolve and verify the host, access, canonical deployment command, service, sole configured cadence, lock, journal, release pointers, backup destination, and recovery contract from that configuration. Never guess access or copy private connection details, credentials, state, or receipts into public output. -Verify installed owner capabilities against their current source; this skill does not install a migration phase. Missing access or a safe capability is a concrete blocker: repair through the existing owner within authority, or report what remains unavailable. Never invent success or bypass a denial. Keep source/PR repairs in separate worktrees; do not delay an otherwise safe prepared update for them. +Verify installed owner capabilities against their current source; this skill does not install a migration phase. Missing access or a safe capability is a concrete blocker: repair through the existing owner within authority, or report what remains unavailable. Never invent success or bypass a denial. Qualify suspected issues against the actual failing path before treating them as defects. Improve this updater guidance when confirmed issues expose a durable gap, and repair confirmed defects within the authorized update scope. Use isolated worktrees and subagents where useful; test and review repairs, then land them through normal CI, PR review, and `scripts/pr` landing without bypassing branch protection. Keep ownership through verified update acceptance; landed repairs alone do not complete the update. Do not delay an otherwise safe prepared update for unrelated repairs. ## Deploy through one owner -1. Keep the existing hourly timer enabled and active as the sole cadence. Never pause it for proof or create another scheduler/deployer. Inspect the active invocation, lock, and journal; observe an active owner instead of duplicating it. Resolve retained journals through canonical recovery before requesting a new deployment. Do not clear failed status to manufacture idleness. +1. Preserve the operator-configured sole cadence and verify its intended state. A retired host timer may intentionally remain disabled when an authorized Gateway job owns the schedule. Never enable a legacy timer, pause the active cadence for proof, or create another scheduler/deployer. Inspect the active invocation, lock, and journal; observe an active owner instead of duplicating it. Resolve retained journals through canonical recovery before requesting a new deployment. Do not clear failed status to manufacture idleness. 2. When idle, request the configured updater service using its documented command. The canonical owner alone controls deployment, Gateway lifecycle, rollback, and recovery. Do not substitute direct restarts, partial build overlays, or an in-place source build. -3. Keep the incumbent serving while the owner freezes official upstream `main`, builds the complete release off-path, validates it, and seals it. Check runtime-user disk/quota headroom, not just host free space. -4. Use the owner's genuine, unexpired maintenance authority bound to the incumbent generation. Allow controlled interruption after its configured drain budget; active agents and PTYs are not indefinite vetoes. Pending terminal persistence still blocks interruption. Never relabel DRAINING as READY. Bound shutdown, migration, startup, and verification separately; the drain budget is not total downtime. +3. Keep the incumbent serving while the owner freezes official upstream `main`, builds the complete release off-path, validates it, and seals it. After a new instruction to update to latest main following an interruption or outage, let the native controller freeze official `main` once for that request after canonical recovery permits a new deployment. Follow that recorded target through acceptance; do not chase moving main with externally sampled `--sha` assertions or repeat safe target-mismatch refusals. Check runtime-user disk/quota headroom, not just host free space. +4. Use the owner's genuine, unexpired maintenance authority bound to the incumbent generation. Use the configured graceful drain, then the owner's documented bounded-interruption mode when authorized; active agents and PTYs are not indefinite vetoes. Apply existing interruption authorization rather than repeatedly choosing a defer-only policy or asking again. Never bulk-cancel agents, abort sessions, clear queues, manually replay turns, or bypass persistence or locks to force idleness. Pending terminal persistence still blocks interruption. Never relabel DRAINING as READY. Bound shutdown, migration, startup, and verification separately; the drain budget is not total downtime. + +## Retain ownership through acceptance + +Keep an explicit update request open through busy deferral, failure, rollback, and recovery until full native acceptance or a concrete external/access blocker requiring outside action. Respect a user pause or cancellation. A restored incumbent is recovery, not completion of the requested update. + +Inspect the exact invocation and reconcile its journal before continuing through the same owner. Before ending a turn with work pending, establish an active observation/completion path or a supported continuation through the existing owner, and return with acceptance or the exact blocker without another user prompt. A cadence alone is not proof that this request will resume. If no continuation path is available, report that specific blocker; do not promise unattended follow-through. Preserve safeguards and evidence rather than retrying blindly or creating another scheduler. + +## Notify only around actual downtime + +Keep Team notifications to two concise notices for an actual interruption: immediately before the canonical owner begins planned downtime, after preparation and cutover gates permit it; and once service is verified back. Starting an update request, preflight, or drain is not itself downtime. Coordinate with the owner's existing notification path so observers do not duplicate notices. If no safe notification boundary is available, report that limitation privately rather than announce speculative downtime or bypass the owner. + +Do not post to Team for preflight failures, blocked or deferred attempts, no-op attempts, diagnosis, or other non-actions. Report their outcome, next action, and concrete blockers privately to the requesting user. Quiet Team reporting does not end ownership of the update or waive follow-through, acceptance, recovery, or safety requirements. + +For an accepted update, the back-online notice includes concise highlights of changes landed on official `main` between the previous accepted serving commit and the newly accepted serving commit. Resolve both endpoints from the owner's accepted serving records and inspect that exact Git range; do not summarize from request time, a failed candidate, a release label, or moving `main`. Mention only changes included in the accepted range. If that history cannot be verified, say the highlights are unavailable rather than invent them. + +If service returns through rollback or same-release recovery, say it is restored on the previous version and the requested update has not succeeded; do not advertise candidate changes as deployed. Send the back-online notice only after the owner's applicable recovery/readiness verification, keep unresolved update blockers private, and continue the requested update through the canonical owner unless paused, canceled, or externally blocked. Keep detailed receipts, logs, private access details, and full hashes out of Team notices. ## Cross schemas safely @@ -35,10 +51,10 @@ Read [database contracts](https://docs.openclaw.ai/reference/database-schemas) a ## Verify and recover -Require the exact invocation's successful deployment receipt, matching new serving/build SHA, stable process generation, RPC, health/startup/readiness, configured channels, unchanged protected policy/identities, and original-witness verification. Reuse the owner's real model-marker receipt; do not send duplicate marker turns. Require journal resolution and an active hourly timer. Supervisor success, a skip, healthy old code, or same-release recovery is not a new deployment. +Require the exact invocation's successful deployment receipt, matching new serving/build SHA, stable process generation, RPC, health/startup/readiness, configured channels, unchanged protected policy/identities, and original-witness verification. Reuse the owner's real model-marker receipt; do not send duplicate marker turns. Require journal resolution and the intended state of the configured sole cadence. Read the native result, not just the observer process status: observer exit `0` can wrap native exit `75`/deferred and is not accepted deployment. Supervisor success, a skip, healthy old code, rollback, or same-release recovery is not a new deployment. -Bracket live checks with generation and owner-phase checks; never run ordinary RPCs across an active fence or pause the timer for a quiet proof window. Preserve failed outcomes and unresolved journals; do not delete evidence or retry blindly. Use only the owner's compatibility-checked recovery and cleanup, preserving referenced releases, backups, ordinary sessions, unrelated state, and dirty workspaces. +Bracket live checks with generation and owner-phase checks; never run ordinary RPCs across an active fence or pause the configured cadence for a quiet proof window. Preserve failed outcomes and unresolved journals; do not delete evidence or retry blindly. Use only the owner's compatibility-checked recovery and cleanup, preserving referenced releases, backups, ordinary sessions, unrelated state, and dirty workspaces. Let [Gateway restart recovery](https://docs.openclaw.ai/gateway/restart-recovery) resume eligible work; do not duplicate it manually. PTYs end, unsaved work may be lost, and recovery budgets/quarantine remain: neither universal recovery nor exactly-once execution is promised. -Report the invocation outcome, observed serving SHA, migration/readiness and continuity proof, recovery phase, and any exact blocker privately. No credential rotation, release publication, security-policy weakening, or unrelated mutations are authorized. +Keep detailed invocation, serving SHA, migration/readiness, continuity, and recovery receipts private. Report progress and results in one to three short, friendly lines: what happened, what happens next or the exact blocker, and a short SHA only when useful. Light humor is welcome when it does not obscure a failure; omit repeated logs and full hashes. No credential rotation, release publication, security-policy weakening, or unrelated mutations are authorized.