From 0fc181d08277a2d1034235a85ca6a181a4de605f Mon Sep 17 00:00:00 2001 From: Peter Steinberger Date: Fri, 2 Oct 2026 03:16:24 -0500 Subject: [PATCH] docs(agents): trim root AGENTS.md under the 20K bootstrap cap (#163325) * docs(agents): trim root AGENTS.md under the 20K bootstrap limit OpenClaw injects a workspace's AGENTS.md into agent context capped at 20,000 chars; above that it keeps a 9K head, a policy digest, and a 3K tail. Sessions whose workspace is an openclaw checkout or managed worktree therefore saw most of "One owner", "Runtime and code safeguards", and "Product and validation" only as fragments. Root AGENTS.md was 27,913 chars; it is now 18,900. Every rule is kept, either condensed in place or moved verbatim to its owner with a pointer: - red main, CI spend, pre-existing reds -> openclaw-pr-maintainer landing.md - screenshot completion gate -> openclaw-pr-maintainer media.md - test failure policy -> openclaw-testing SKILL.md - cyber-classifier interruption procedure -> security-triage SKILL.md Schema-check placement, test budgets, and main-drift rules already live in docs/reference/database-schemas.md, docs/help/testing/writing-tests.md, and landing.md, so root keeps one-line pointers. The Team update rule is owned by the update-team-server skill. * docs(agents): restore trimmed rules and move remaining detail to owners Autoreview flagged clauses the first trim condensed away. Restore them in root (complete-module reads, maintainer approval for untrusted execution, contributor credit, release lock mirrors, Night Watch, audit approval scope, updater rollback, and others) and move the capability ladder, execution gotchas, CODEOWNERS governance, and the full cleanup rule verbatim to the SDK guide, scripts guide, and maintainer skill. * docs(agents): keep isolation, migration-path, fixture, and rerun conditions --- .../skills/openclaw-pr-maintainer/SKILL.md | 17 ++ .../references/landing.md | 7 + .../references/media.md | 6 +- .agents/skills/openclaw-testing/SKILL.md | 14 ++ .agents/skills/security-triage/SKILL.md | 10 ++ AGENTS.md | 169 ++++++++---------- scripts/AGENTS.md | 10 ++ src/plugin-sdk/AGENTS.md | 13 ++ 8 files changed, 147 insertions(+), 99 deletions(-) diff --git a/.agents/skills/openclaw-pr-maintainer/SKILL.md b/.agents/skills/openclaw-pr-maintainer/SKILL.md index ed369ba790ff..c8760a16e311 100644 --- a/.agents/skills/openclaw-pr-maintainer/SKILL.md +++ b/.agents/skills/openclaw-pr-maintainer/SKILL.md @@ -84,6 +84,14 @@ proof with the limitation stated. Explicit live requests, external API contracts and changes whose risk requires authenticated execution keep their required proof. Never describe mocks, skipped checks, or an older head as live evidence. +## CODEOWNERS review + +`CODEOWNERS` routes review; check live GitHub enforcement. Restricted/security +paths and material product, behavior, security, or ownership changes need +listed-owner involvement. For ownership/review governance, verified active +organization-admin direction also qualifies; repository admin/bypass alone does +not. Neither route waives enforced reviews. + ## Review and publish Before committing or landing nontrivial code, run `$autoreview` and resolve @@ -126,3 +134,12 @@ Required pre-merge proof stays mandatory. Preserve requested deliverables, explicit retention requests, unfinished source, recovery state needed by unfinished operations, active owners, credentials, agent state, and shared dependencies. Optional follow-ups do not keep a completed task open. + +This applies to every task, not only PR work: report routine findings in +chat/stdout. Create files only for deliverables or concrete tool/proof/recovery +needs; state their purpose and reuse them. After the requested outcome and +required verification are complete, remove task-owned proof, scratch, and +finished worktrees through their native lifecycle. Do not require an archive or +evidence handoff. Preserve requested deliverables, explicit retention requests, +unfinished source and recovery state, live owners, credentials, agent state, +shared dependencies, and unknown ownership. diff --git a/.agents/skills/openclaw-pr-maintainer/references/landing.md b/.agents/skills/openclaw-pr-maintainer/references/landing.md index 7517c9298a29..d62ad91ae90f 100644 --- a/.agents/skills/openclaw-pr-maintainer/references/landing.md +++ b/.agents/skills/openclaw-pr-maintainer/references/landing.md @@ -68,6 +68,13 @@ Keep images embedded and video URLs on their own lines for GitHub playback. Feature-detect format/size support rather than assuming a particular CLI release. Do not disclose private desktop content, identifiers, model routes, or secrets. +## Red main and inherited failures + +- **Red `main` is an emergency, not a queue item.** When CI on `main` is red, fix it in a fresh worktree, prove it with the exact failing check plus the focused tests (locally or on a Testbox), and push that exact fix directly to `main` — no PR, no auto-merge wait — so every open PR unblocks at once. This applies to deterministic breaks and to flaky tests seen red on `main` alike; open a PR only if the direct push is refused by a ruleset, and say so. Reruns do not clear a red `main`, and PRs whose merge ref went red need a fresh push after the fix lands. +- Pushes to `main` run only `security-fast`; full `main` CI runs hourly. +- **CI is expensive; do not spend it on proof.** Prove changes with focused local runs (Crabbox/Testbox for suites too heavy for the host) and trust those results. Do not dispatch workflow runs, rerun jobs, re-push, or update branches just to obtain or confirm green. Breaking `main` is an acceptable cost; fix it forward. +- **Pre-existing reds do not block landing.** A PR whose only failing checks also fail on current `main` (same test, in files the PR does not touch) may land through the native landing workflow's admin exception; name the inherited failure in the PR. A tiny PR-caused failure (lint, types, a stale test expectation) may land the same way only if the lander pushes the fix to `main` immediately afterward. Anything larger stays blocked. + ## Review, prepare, merge For main-targeted PRs, prefer the native sequence; adapt as needed. diff --git a/.agents/skills/openclaw-pr-maintainer/references/media.md b/.agents/skills/openclaw-pr-maintainer/references/media.md index 1caef518af22..b042a589ac58 100644 --- a/.agents/skills/openclaw-pr-maintainer/references/media.md +++ b/.agents/skills/openclaw-pr-maintainer/references/media.md @@ -11,4 +11,8 @@ inspected polished MP4; never commit generated media. - Preferred PR/issue media upload: when the command help exposes `--attach`, use the repeatable flag on `gh issue create`, `gh issue edit`, `gh issue comment`, and the matching `gh pr` commands. Example: `gh pr comment --repo openclaw/openclaw --body-file --attach `. - `gh --attach` video rules: accepted extensions are `.mp4`, `.mov`, and `.webm`; the local maximum is 100 MB, while GitHub's account limit may be lower. Do not add `#alt` to a video path. `gh` inserts a bare URL so GitHub renders a player, and the uploaded asset cannot be deleted. - Compatibility fallback: if the installed `gh` lacks `--attach`, upload directly with `curl -s "https://uploads.github.com/user-attachments/assets?name=&content_type=&repository_id=$(gh api repos// --jq .id)" -X POST -H "Authorization: Bearer $(gh auth token)" -H "Accept: application/json" --data-binary @`, then embed the returned `.url`: images as `![alt](url)`, video as a bare line so GitHub renders a player. For that endpoint, 422 = unsupported type and 404 = bad repo id/no push. Use `content_type` `video/mp4`, `video/quicktime`, or `video/webm`; `![]()` does not render the player. Transcode Playwright webm via `ffmpeg -i in.webm -c:v libx264 -pix_fmt yuv420p out.mp4` for broad playback. Both paths use the same drag-drop CDN, inherit repository visibility, and need no browser/computer use. Non-media artifacts: Crabbox artifact publishing plus the manifest URL. -- Upload failure does not waive the [root UI screenshot completion/landing gate](../../../../AGENTS.md#product-and-validation). An approved artifact-store fallback counts for required screenshots only when attached as images in chat and embedded inline in the PR; a manifest URL alone does not count. Successful chat attachment delivery needs no browser rendering check or user confirmation. Verify PR rendering as required by the root gate. Otherwise report an explicit delivery blocker and do not merge or claim completion unless the user explicitly waives that destination. +- Upload failure does not waive the [screenshot completion gate](#screenshot-completion-gate). An approved artifact-store fallback counts for required screenshots only when attached as images in chat and embedded inline in the PR; a manifest URL alone does not count. Successful chat attachment delivery needs no browser rendering check or user confirmation. Verify PR rendering as required by the gate. Otherwise report an explicit delivery blocker and do not merge or claim completion unless the user explicitly waives that destination. + +## Screenshot completion gate + +For changes to visual appearance or rendered UI states, attach inspected, sanitized before/after screenshots in the originating chat AND upload/embed them in the GitHub PR body or a PR comment linked from its body. Successful chat attachment delivery satisfies the chat requirement; do not verify rendering in the user's browser, ask the user to confirm visibility, or block landing because the agent cannot inspect the chat. Verify that the PR images render before merging or claiming completion. Local files, private inspection, chat-only delivery, logs, tests, review approval, and artifact-manifest links are not substitutes. Galleries and videos supplement, never replace, the required screenshots. Known attachment or upload failures remain delivery blockers; repair them or report the blocker. Only an explicit user waiver can exempt a destination. diff --git a/.agents/skills/openclaw-testing/SKILL.md b/.agents/skills/openclaw-testing/SKILL.md index e554250ada8a..d4511749fedb 100644 --- a/.agents/skills/openclaw-testing/SKILL.md +++ b/.agents/skills/openclaw-testing/SKILL.md @@ -119,3 +119,17 @@ For prompt snapshot drift that passes on macOS, reproduce in CI's Linux/Node environment before regenerating; a local pass cannot override failing CI bytes. Fix related failures and rerun the affected proof. Route unrelated failures with evidence rather than broadening this task automatically. + +### Test failure policy + +Treat test failures as defects and make a bounded, best-effort attempt to +reproduce them (same shard order first), identify their cause, and fix the owning +fixture, shared state, ordering, or product. When a safe fix is established, add +a regression and document the cause; cite another owner's fix when applicable. +If reasonable investigation cannot establish or complete a safe fix, record the +original failure, attempted reproductions, evidence, and remaining uncertainty in +the PR, then continue under the normal CI and review gates. The unresolved +failure alone must not block landing or trigger an extra approval request. Never +claim a passing replay proves a fix. Do not rerun, re-push, or refresh merely to +get green, or conceal failures with retries, longer timeouts, weaker assertions, +broader mocks, or altered baselines. diff --git a/.agents/skills/security-triage/SKILL.md b/.agents/skills/security-triage/SKILL.md index ef1c9db53bd8..dc40f50bc351 100644 --- a/.agents/skills/security-triage/SKILL.md +++ b/.agents/skills/security-triage/SKILL.md @@ -138,6 +138,16 @@ gh search issues --repo openclaw/openclaw --match title,body,comments -- "" ``` +## Cyber Classifier Interruptions + +For permitted defensive work interrupted by a cyber classifier, use an approved, +available Daybreak Blue route only for that bounded task, then restore the +previous model. Preserve the approved identity, scope, permissions, sandbox, and +review gates. Never switch models to bypass access denials, disallowed requests, +or refused tool actions; if that route also refuses, stop and report. Follow +[OpenAI's guidance](https://learn.chatgpt.com/docs/cyber-safety) for approved +model access. + ## Decision Notes - “fixed on main, unreleased” is usually not a close. diff --git a/AGENTS.md b/AGENTS.md index 2a5050bf87f6..08a9a27cc13c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,22 +9,21 @@ Update instructions at their owner instead of adding competing rules here. ## Design priorities -- **One owner per responsibility.** An owner makes a decision or changes authoritative state. Callers consume its operations and recorded facts. Adapters translate contracts; caches and projections derive from the owner with an explicit invalidation lifecycle. Different transports can need different adapters, but not competing owners for the same responsibility. -- **Small core, capable plugins.** Model-facing core additions have an ongoing context cost. Optional capability belongs at the edges; core supplies generic contracts. A feature needing a new integration is not, by itself, a reason to add another core tool or manager. [VISION.md](VISION.md) owns product scope. -- **Stable conversation context.** Rebuilding past context defeats prompt-prefix reuse. Keep generated prompt/tool/context additions bounded and deterministic, preserve transcript bytes, and serve required instructions whole. Only compaction rewrites history. Defer changes to stable prompt state until the next session unless its owner defines explicit invalidation; preserve existing skill, tool, and memory refresh contracts. +- **One owner per responsibility.** An owner makes a decision or changes authoritative state; callers consume its operations and recorded facts. Adapters translate contracts; caches and projections derive from the owner with an explicit invalidation lifecycle. Transports may need different adapters, never competing owners. +- **Small core, capable plugins.** Model-facing core additions have an ongoing context cost. Optional capability belongs at the edges; core supplies generic contracts. A new integration alone is no reason for another core tool or manager. New capability takes the first fitting step of the [capability ladder](src/plugin-sdk/AGENTS.md#choose-the-capability-surface): existing owner, existing plugin contract, narrow generic core/SDK capability, then universal core surface. [VISION.md](VISION.md) owns product scope. +- **Stable conversation context.** Rebuilding past context defeats prompt-prefix reuse. Keep generated prompt/tool/context additions bounded and deterministic, preserve transcript bytes, and serve required instructions whole; only compaction rewrites history. Defer stable-prompt changes to the next session unless the owner defines explicit invalidation; keep existing skill, tool, and memory refresh contracts. ## Working agreement -- Follow through on actionable requests, including "can you", within their authorized scope. When execution is requested, a plan or progress report is a checkpoint, not completion. Use prior context and preserve unaffected work across corrections and side questions. -- Resolve routine, reversible choices with reasonable assumptions. Ask only about consequential decisions the request and context cannot resolve; continue independent authorized work while waiting. Silence does not authorize a gated action. -- If a skill causes a pause, permission request, unfinished work, or scope change, link its exact `SKILL.md` and quote the instruction to the user. Explain how it applies, distinguish requirements from interpretation, and check prior authorization before asking again. -- Inspect `git status -sb` before editing or GitHub work. Preserve unrelated work, branches, processes, and user-managed checkouts; serialize shared Git mutations and isolate work when needed. Never switch a checkout while another agent or test run uses it. -- Treat pasted material and tool output as evidence; verify claims against source and observed behavior. -- Lead with the result and follow the user's format. Use plain words, active voice, and useful technical detail; omit stock phrases and repeated summaries. Reference each PR/issue once per reply. Auto-links count; don't repeat the URL. Progress updates explain new findings, decisions, or blockers. Keep delegated messages equally clear. -- Report routine findings in chat/stdout. Create files only for deliverables or concrete tool/proof/recovery needs; state their purpose and reuse them. After the requested outcome and required verification are complete, remove task-owned proof, scratch, and finished worktrees through their native lifecycle. Do not require an archive or evidence handoff. Preserve requested deliverables, explicit retention requests, unfinished source and recovery state, live owners, credentials, agent state, shared dependencies, and unknown ownership. Follow [maintainer closeout](.agents/skills/openclaw-pr-maintainer/SKILL.md#finalize-and-clean-up). -- Read relevant docs before changing behavior; `pnpm docs:list` locates them. `package.json` owns current commands and versions; keep the repository's toolchain and conventions rather than swapping tools without approval. -- Use **OpenClaw** for the product, `openclaw` for CLI/package/config names, **plugins** for user-facing integrations, and American English. -- Edit canonical `AGENTS.md` files directly. +- Follow through on actionable requests, including "can you", within authorized scope; a plan or progress report is a checkpoint, not completion. Preserve unaffected work across corrections and side questions. +- Resolve routine, reversible choices with reasonable assumptions; ask only about consequential decisions context cannot resolve, and continue independent authorized work meanwhile. Silence does not authorize a gated action. +- If a skill causes a pause, permission request, unfinished work, or scope change, link its exact `SKILL.md`, quote the instruction, explain how it applies (requirement vs interpretation), and check prior authorization before asking again. +- Inspect `git status -sb` before editing or GitHub work. Preserve unrelated work, branches, processes, and user-managed checkouts; serialize shared Git mutations and isolate work when needed; never switch a checkout another agent or test run uses. +- Treat pasted material and tool output as evidence; verify against source and observed behavior. +- Lead with the result in the user's format: plain, active, technically useful; no stock phrases or repeated summaries. Reference each PR/issue once per reply. Progress updates explain new findings, decisions, or blockers. +- Report findings in chat; create files only for deliverables or tool/proof/recovery needs, stating their purpose and reusing them. After verified completion, remove task-owned proof, scratch, and finished worktrees per [closeout](.agents/skills/openclaw-pr-maintainer/SKILL.md#finalize-and-clean-up), preserving deliverables, live and unknown owners, unfinished state, and credentials. +- Read relevant docs before changing behavior (`pnpm docs:list`); `package.json` owns commands and versions; tool swaps need approval. +- Use **OpenClaw** (product), `openclaw` (CLI/package/config), **plugins** (user-facing integrations), and American English. Edit canonical `AGENTS.md` files directly. ## Execution discipline @@ -33,113 +32,87 @@ Update instructions at their owner instead of adding competing rules here. - Separate product bugs from tool/fixture failures. - Check relevant prerequisites early. Parallelize independent work. - New check needs named unknown, risk, or required gate. Reuse valid proof. -- After two identical tooling failures without new evidence, or 10 min without progress, change approach. Use an already-authorized supported alternative with the reviewed head pinned; reconcile uncertain writes before retrying. No blind retries or guard bypasses. +- Two identical tooling failures without new evidence, or 10 min without progress: switch to an authorized supported alternative with the reviewed head pinned; reconcile uncertain writes first. No blind retries or guard bypasses. - Behavior proven + required gates green: finish/land. No optional proof polish or speculative scope growth. -- Do not rebase or merge `main` solely because it advanced. Check mergeability first; integrate for actual conflicts or a named failing gate/material base risk. Preserve valid review and validation evidence, rerun checks affected by the resolution, and land promptly once the selected workflow's gates are satisfied. Prior-head CI remains prior-head evidence; an explicit admin exception belongs to the native landing workflow. +- Check mergeability first; don't rebase or merge `main` just because it advanced. Integrate for conflicts, a named failing gate, or material base risk, then rerun affected checks. Prior-head CI stays prior-head evidence; admin exceptions belong to the native landing workflow ([landing](.agents/skills/openclaw-pr-maintainer/references/landing.md)). - Time pressure never waives gates. Report concrete blockers. - Visual change: inspected before/after screenshots. Behavior-only fix: direct boundary proof. No checkbox demos. ## One owner, complete cutover -1. **Intent:** reproduce defects through the actual entry point before editing when feasible. Read complete affected modules, owners, callers, siblings, tests, history, and dependency contracts until the intended user outcome and violated invariant are supported by evidence. Before restoring a missing path, check why it was removed (`git log -p -S `): isolation may be intentional, and a retired alias may be a completed migration. Record concrete reproduction gaps. -2. **Owner:** account for relevant decisions and state writers across creation, updates, reads, recovery, and cleanup. Choose the existing code, plugin, or maintained solution that absorbs the change. A new owner needs a missing responsibility; fix invalid or leaked state at its producer. -3. **Cutover:** migrate all affected internal/bundled callers together. Remove superseded code, duplicate policy/state, wrappers, registrations, exports, tests, and docs. Every retained path needs a cited contract. Workers sharing an owner agree on one interface and cutover plan. -4. **Proof:** exercise the intended user flow and relevant siblings; trace references to confirm retired paths are unreachable. Done means one owner serves the flow, old paths are removed or justified, and observed results or remaining gaps are recorded in existing task/PR evidence. Helper tests or a wrapper around competing implementations alone are insufficient. +1. **Intent:** reproduce defects through the actual entry point before editing when feasible. Read complete affected modules, owners, callers, siblings, tests, history, and dependency contracts until the intended outcome and violated invariant are evidenced. Before restoring a missing path, check why it was removed (`git log -p -S `); isolation may be intentional, a retired alias a completed migration. +2. **Owner:** account for decisions and state writers across creation, updates, reads, recovery, and cleanup. Choose the existing code, plugin, or maintained solution that absorbs the change; a new owner needs a missing responsibility. Fix invalid or leaked state at its producer. +3. **Cutover:** migrate all affected internal/bundled callers together; remove superseded code, duplicate policy/state, wrappers, registrations, exports, tests, and docs. Retained paths need a cited contract; workers sharing an owner agree on one interface and plan. +4. **Proof:** exercise the user flow and relevant siblings; trace references to confirm retired paths are unreachable. Done = one owner serves the flow, old paths removed or justified, results or gaps recorded in task/PR evidence. Helper tests or wrappers around competing implementations are insufficient. -- Prefer smaller, simpler production code; explain necessary growth. Keep coherent nearby repairs together and record unrelated work as follow-ups. No extra report or tracking system is required. -- Delegate independent evidence or implementation lanes when parallel work reduces time or improves verification. Give each lane a clear responsibility and completion condition; keep simple or tightly coupled work with the lead. The lead stays hands-on, verifies consequential conclusions, and coordinates shared-checkout safety. -- Retained compatibility needs an explicit user request or a public API/config/SDK/data, stable-tag upgrade, security/migration, dependency, or observed-production contract, plus a migration/removal path. Main, beta, and nightly code alone are not shipped contracts. - -### Choose the capability surface - -For new capability, use the first path that expresses the actual requirement: - -1. Extend the existing owner or use an existing command, skill, plugin, or supported integration. -2. Use an existing plugin contract. Prefer bundle plugins for skills, MCP servers, and configuration; use code plugins when runtime hooks, providers, channels, or tools are needed. Keep vendor behavior with its vendor plugin and feature behavior with its feature owner. -3. If the contract is missing, define a narrow generic core/SDK capability and move existing bundled implementations and callers onto it together. Repeated independent requests for the same capability trigger this contract review, not another parallel manager or hook. -4. Add universal core surface only when the need is fundamental and existing extension points cannot express it. Explain the gap and ongoing cost; a new hook needs a concrete consumer. - -For example, a new channel action should first use the shared message action -contract. A setup screen needing plugin metadata should use the manifest or -lightweight artifact, not load the plugin's execution runtime. +- Prefer smaller, simpler production code; explain necessary growth. Keep coherent nearby repairs together; record unrelated work as follow-ups. +- Delegate independent lanes when that saves time or improves verification, each with a clear responsibility and completion condition; keep simple or tightly coupled work with the lead, who stays hands-on, verifies consequential conclusions, and owns shared-checkout safety. +- Retained compatibility needs an explicit user request or a public API/config/SDK/data, stable-tag upgrade, security/migration, dependency, or observed-production contract, plus a migration/removal path. Main, beta, and nightly code are not shipped contracts. ## Runtime and code safeguards - Plugins use documented `openclaw/plugin-sdk/*` contracts, manifest metadata, and public/local barrels, never core internals or another plugin's private files. Dependencies follow runtime ownership. -- Runtime consumes canonical config/state. Doctor/migration owners normalize legacy shapes; plugin repairs stay plugin-owned. A change invalidating existing config includes its matching migration. Startup may invoke the same approved Doctor transforms; do not add independent compatibility readers. -- OpenClaw state and caches use SQLite, not new JSON/JSONL/sidecar stores. Files are for named user artifacts, imports/exports, attachments, logs, backups, or external-tool contracts. -- Use Kysely for ordinary SQLite access; raw SQL is limited to schema, migrations, bootstrap, and justified primitives. Write transactions are synchronous: finish asynchronous planning first, then reread authoritative rows before writing. No Promise or `await` in a transaction callback. -- Database access runs in worker threads, never on the Gateway main thread: readers use the read-only worker scope, writers the SQLite worker broker, and the main thread only awaits results and installs published facts. Boot admission, migrations, Doctor/CLI one-shots, and lock primitives are the only synchronous exceptions. Existing synchronous main-thread access is legacy: never add more, and migrate any such path you touch. -- Schema-version, integrity, canonical-index, and table-existence checks belong to database open/admission and the migration owner after migrations; runtime paths carry admitted schema facts with the handle and never re-query them. Cache freshness probes must reflect foreign commits on the next unpinned use while preserving actual SQLite snapshot semantics and retained schema facts. Existing per-call checks are legacy: never add more, and migrate any you touch. -- Privileged actions require current owner-held authority. Revalidate after awaited work and immediately before side effects; tokens, signatures, expiry, and matching IDs alone do not prove live authority. -- Core owns shared message tools, action vocabulary, and dispatch. Channels own their account, security, conversation, and transport contracts. Preserve typed command/approval/URL/action distinctions until encoding; never infer product commands from raw strings. -- Carry prepared facts through hot paths. Reuse process-stable plugin metadata and lifecycle-owned caches; do not repeatedly load registries or freshness-poll files. Preserve lazy module boundaries and verify relevant builds on the authorized host. -- Keep APIs narrow, valid states explicit, and TypeScript ESM/types strict. Prefer real types or `unknown`; no `@ts-nocheck`. Suppressions need an intentional, explained exception. Reuse schema/coercion owners; avoid duplicate guards, speculative helpers, and naming-only wrappers. -- Static-analysis fixes strengthen the real type/runtime contract or remove the unsafe operation; do not conceal it with casts, widening, marker types, or property probes. New lint rules need a meaningful invariant and a clean owner scope. -- Comments explain non-obvious ownership, lifecycle, ordering, cleanup, platform, and dependency constraints, not syntax. Do not edit `node_modules` or generated artifacts by hand, or change formatter settings for a local expression; regenerate owned outputs. +- Runtime consumes canonical config/state; Doctor/migration owners normalize legacy shapes (plugin repairs stay plugin-owned), and config-invalidating changes ship their migration. +- State and caches use SQLite, not new JSON/JSONL/sidecar stores; files are for named user artifacts, imports/exports, attachments, logs, backups, or external-tool contracts. +- Kysely for ordinary SQLite; raw SQL only for schema, migrations, bootstrap, and justified primitives. Write transactions are synchronous: plan async work first, reread authoritative rows, then write; no Promise or `await` in transaction callbacks. +- Database access runs in worker threads (read-only worker scope, SQLite writer broker), never on the Gateway main thread, which only awaits results and installs published facts. Exceptions: boot admission, migrations, Doctor/CLI one-shots, lock primitives. Existing main-thread access is legacy: never add more; migrate what you touch. +- Schema/integrity/index checks run once at admission and in the migration owner; runtime carries admitted facts and never re-queries them, and cache freshness probes observe foreign commits as [database schemas](docs/reference/database-schemas.md) defines. Per-call checks are legacy: never add; migrate when touched. +- Privileged actions need current owner-held authority: revalidate after awaited work and right before side effects. Tokens, signatures, expiry, and matching IDs alone do not prove live authority. +- Core owns shared message tools, action vocabulary, and dispatch; channels own account, security, conversation, and transport contracts. Keep typed command/approval/URL/action distinctions until encoding; never infer commands from raw strings. +- Carry prepared facts through hot paths; reuse process-stable plugin metadata and lifecycle-owned caches; never repeatedly load registries or freshness-poll files. Preserve lazy module boundaries and verify relevant builds on the authorized host. +- Narrow APIs, explicit valid states, strict ESM/types. Real types or `unknown`; no `@ts-nocheck`; suppressions need an explained exception. Reuse schema/coercion owners; no duplicate guards, speculative helpers, or naming-only wrappers. +- Static-analysis fixes strengthen the real contract or remove the unsafe operation, never hide it with casts, widening, marker types, or property probes. New lint rules need a real invariant and clean owner scope. +- Comments explain non-obvious ownership, lifecycle, ordering, cleanup, platform, and dependency constraints, not syntax. Regenerate generated outputs; never hand-edit them or `node_modules`, or tweak formatter settings per expression. ## Product and validation -- Defaults should produce a working, understandable result. Prioritize silent failures. Each action has a visible outcome or recorded intentional non-outcome; errors explain the next useful step. -- **Updates always work.** `openclaw update` finishes best effort on every install. Any change touching update, Doctor, service lifecycle, config/state migration, or plugin loading states its update behavior: the installed updater runs first and cannot be patched, so candidate-side fixes key on markers shipped drivers already set, and existing operator state is the input. Recoverable hiccups become recorded warnings; back up before mutating and let rollback restore it; refuse only for concrete data at risk, naming the reason and leaving the previous Gateway running. Timeouts and budgets are generous, derived from measured state, and sized for old, slow hardware. Proof: a published-driver × candidate cell. -- Prompts, tools, and results describe available capabilities accurately and give enough context for the next useful action; avoid unnecessary model round trips. Inject cross-tool references from the enabled tool set and remove stale model-facing arguments instead of hidden compatibility. New optional features need discovery paths. -- Security is a product tradeoff, not a goal to maximize restrictions. Weigh concrete risk and likely impact against user effort, lockouts, and lost capability. Prefer the least restrictive effective safeguard; bounded, understood risk can be acceptable for a substantial usability benefit. Keep risky paths explicit and operator-controlled within the existing trust model and approval boundaries, and explain the tradeoff instead of inventing extra gates. -- Tests must protect meaningful behavior; skip tests for reversible, low-impact changes that merely mirror the implementation. Regressions fail on the original defect; shared-state failures use the original order. Review tests for value and duplication. -- Treat test failures as defects and make a bounded, best-effort attempt to reproduce them (same shard order first), identify their cause, and fix the owning fixture, shared state, ordering, or product. When a safe fix is established, add a regression and document the cause; cite another owner's fix when applicable. If reasonable investigation cannot establish or complete a safe fix, record the original failure, attempted reproductions, evidence, and remaining uncertainty in the PR, then continue under the normal CI and review gates. The unresolved failure alone must not block landing or trigger an extra approval request. Never claim a passing replay proves a fix. Do not rerun, re-push, or refresh merely to get green, or conceal failures with retries, longer timeouts, weaker assertions, broader mocks, or altered baselines. -- **Red `main` is an emergency, not a queue item.** When CI on `main` is red, fix it in a fresh worktree, prove it with the exact failing check plus the focused tests (locally or on a Testbox), and push that exact fix directly to `main` — no PR, no auto-merge wait — so every open PR unblocks at once. This applies to deterministic breaks and to flaky tests seen red on `main` alike; open a PR only if the direct push is refused by a ruleset, and say so. Reruns do not clear a red `main`, and PRs whose merge ref went red need a fresh push after the fix lands. -- Pushes to `main` run only `security-fast`; full `main` CI runs hourly. -- **CI is expensive; do not spend it on proof.** Prove changes with focused local runs (Crabbox/Testbox for suites too heavy for the host) and trust those results. Do not dispatch workflow runs, rerun jobs, re-push, or update branches just to obtain or confirm green. Breaking `main` is an acceptable cost; fix it forward. -- **Pre-existing reds do not block landing.** A PR whose only failing checks also fail on current `main` (same test, in files the PR does not touch) may land through the native landing workflow's admin exception; name the inherited failure in the PR. A tiny PR-caused failure (lint, types, a stale test expectation) may land the same way only if the lander pushes the fix to `main` immediately afterward. Anything larger stays blocked. -- Every test spends CI time on every PR. New or changed tests state their measured cost in the PR (`pnpm test --maxWorkers=1` wall, and CI seconds once the run exists) and stay within the budgets in [writing tests](docs/help/testing/writing-tests.md): no real timers, sleeps, or polling; no per-test Gateway or process boots when a suite-level fixture exists; no new serial config or worker pin; no broad barrel imports. A test that needs seconds must prove a contract that no cheaper layer can, and long end-to-end compositions belong in the release-only tier, not per-PR CI. -- Select proof for the touched contract and complete the chosen workflow's required gates within user/host limits. Command references do not mandate unrelated suites. Reuse valid proof; rerun for changed inputs or missing coverage. Docs-only work needs docs sanity and `git diff --check`. Report unrun checks and gaps. -- Prove user-visible behavior through the real flow when feasible; external API changes need live contract proof. A covering isolated mock-Gateway harness is valid channel boundary proof; live channel proof is stronger. State concrete capture or execution blockers. -- **Visual screenshot completion/landing gate:** For changes to visual appearance or rendered UI states, attach inspected, sanitized before/after screenshots in the originating chat AND upload/embed them in the GitHub PR body or a PR comment linked from its body. Successful chat attachment delivery satisfies the chat requirement; do not verify rendering in the user's browser, ask the user to confirm visibility, or block landing because the agent cannot inspect the chat. Verify that the PR images render before merging or claiming completion. Local files, private inspection, chat-only delivery, logs, tests, review approval, and artifact-manifest links are not substitutes. Galleries and videos supplement, never replace, the required screenshots. Known attachment or upload failures remain delivery blockers; repair them or report the blocker. Only an explicit user waiver can exempt a destination. -- Before committing or landing nontrivial code, obtain fresh review through the permitted workflow and resolve actionable findings unless the user opts out. Tests protect observable contracts; a helper test can pass while the registered entry point never calls it. +- Defaults produce a working, understandable result. Prioritize silent failures: every action has a visible outcome or recorded intentional non-outcome; errors name the next step. +- **Updates always work** (`openclaw update` finishes best effort everywhere). Changes to update, Doctor, service lifecycle, migrations, or plugin loading state their update behavior. The installed updater runs first and can't be patched, so candidate fixes key on markers shipped drivers set; existing operator state is the input. Recoverable hiccups become recorded warnings; back up before mutating and let rollback restore it; refuse only for named data risk with the previous Gateway kept running. Timeouts and budgets are generous, derived from measured state, and sized for slow hardware. Proof: a published-driver × candidate cell. +- Prompts, tools, and results describe available capabilities accurately, with context for the next useful action and no unnecessary model round trips; inject cross-tool references from the enabled tool set, drop stale model-facing arguments, and give new optional features discovery paths. +- Security is a product tradeoff: weigh concrete risk against user effort, lockouts, and lost capability; prefer the least restrictive effective safeguard, keep risky paths explicit and operator-controlled within the existing trust model and approval boundaries, and explain tradeoffs instead of inventing gates. +- Tests protect meaningful behavior, not reversible low-impact changes that mirror the implementation; review them for value and duplication. Regressions fail on the original defect; shared-state failures use the original order. +- Test failures are defects: reproduce, fix the owner with a regression, or record evidence and continue (unresolved alone doesn't block landing). Never mask failures or claim a passing replay proves a fix ([policy](.agents/skills/openclaw-testing/SKILL.md#test-failure-policy)). +- **Red `main` is an emergency; CI is expensive.** Push proven red-`main` fixes directly to `main`; prove changes locally or on Testbox, never rerun or re-push just for green; inherited-only reds land via the native admin exception, named in the PR ([policy](.agents/skills/openclaw-pr-maintainer/references/landing.md#red-main-and-inherited-failures)). +- New/changed tests follow the [writing tests](docs/help/testing/writing-tests.md) cost budget: PRs state `pnpm test --maxWorkers=1` wall time and CI seconds; no real timers, sleeps, polling, per-test Gateway/process boots when a suite-level fixture exists, new serial config or worker pins, or broad barrel imports. Seconds-long tests must prove a contract no cheaper layer can; long end-to-end compositions go to the release-only tier. +- Select proof for the touched contract, reuse valid proof (rerun for changed inputs or missing coverage), and finish the workflow's required gates within user/host limits; report unrun checks and gaps. Prove user-visible behavior through the real flow when feasible; external APIs need live contract proof; an isolated mock-Gateway harness is valid channel boundary proof. Docs-only: docs sanity and `git diff --check`. +- **Visual changes** need inspected, sanitized before/after screenshots in chat and embedded in the PR before merge or completion ([gate](.agents/skills/openclaw-pr-maintainer/references/media.md#screenshot-completion-gate)). +- Before committing or landing nontrivial code, get fresh review through the permitted workflow and resolve actionable findings unless the user opts out. ### Execution gotchas -These commands apply on the host permitted by the task and its workflow; they do -not authorize local execution or a broader test plan. - -- For fs-safe dependency trouble, follow [on-demand vendoring instructions](scripts/vendor-fs-safe.md); keep vendor contents local and registry dependencies as the default. -- Restore missing dependencies in a trusted normal checkout with `pnpm install`, then retry once before diagnosing a code defect. Never reconcile a shared/worktree install while other jobs use it. -- Run the CLI through `pnpm openclaw ...` or `pnpm dev`, never `node --import tsx src/index.ts`; the supported wrappers own build freshness and process setup. -- Use installed `oxfmt` for formatting and the repository's `tsgo` lanes for typechecking. Inspect scope with `pnpm changed:lanes --json`; use targeted tests/checks. When avoiding worktree reconciliation, use `node scripts/check-changed.mjs` or `node scripts/run-vitest.mjs` with ready dependencies. Host restrictions still apply. +Run the CLI via `pnpm openclaw ...` or `pnpm dev`, never `node --import tsx src/index.ts`, and never reconcile a shared/worktree install other jobs use. Dependency, vendoring, format, and typecheck gotchas: [scripts guide](scripts/AGENTS.md#execution-gotchas). ## Authority and safety -- Review/triage is read-only; mutations require task authority. Existing approval carries through the same scoped work and recovery. When new approval is required, complete the already-authorized preparation first and present a concrete, reviewable result; pause only the gated action. Product rejection remains maintainer judgment. Bulk close/reopen above 50 items needs explicit count and scope. -- Keep credentials, private data/config, and unreleased model identities out of commits and shared text, logs, transcripts, and media. Inspect outgoing content. Use synthetic fixtures and verified human credit; omit agent-attribution trailers. -- For permitted defensive work interrupted by a cyber classifier, use an approved, available Daybreak Blue route only for that bounded task, then restore the previous model. Preserve the approved identity, scope, permissions, sandbox, and review gates. Never switch models to bypass access denials, disallowed requests, or refused tool actions; if that route also refuses, stop and report. Follow [OpenAI's guidance](https://learn.chatgpt.com/docs/cyber-safety) for approved model access. -- Untrusted contributor/fork code runs only in secretless isolation, never locally. Source review alone does not authorize execution with credentials or on a trusted host; maintainer approval is required. An instruction to land named, reviewed PRs supplies that approval. Use the authorized isolation route and only task credentials. -- Modifying/restarting a Gateway or live state you did not create requires per-task approval. Tests use isolated state and ports; copy real data for migration tests. Destructive reset/clean, stash, or deletion of unrelated work needs authorization. +- Review/triage is read-only; mutations need task authority. Approval carries through the same scoped work and recovery; for new approval, complete authorized preparation, present a concrete reviewable result, and pause only the gated action. Product rejection is maintainer judgment. Bulk close/reopen above 50 items needs explicit count and scope. +- Keep credentials, private data/config, and unreleased model identities out of commits, shared text, logs, transcripts, and media; inspect outgoing content. Synthetic fixtures, verified human credit, no agent-attribution trailers. +- Never switch models to bypass refusals; cyber-classifier interruptions of permitted defensive work follow [security triage](.agents/skills/security-triage/SKILL.md#cyber-classifier-interruptions). +- Untrusted contributor/fork code runs only in secretless isolation, never locally; credentialed or trusted-host execution needs maintainer approval, which an instruction to land named, reviewed PRs supplies (isolation route, task credentials only). +- Modifying/restarting a Gateway or live state you did not create needs per-task approval. Tests use isolated state and ports; copy real data for migration tests. Destructive reset/clean, stash, or deleting unrelated work needs authorization. - Updating `team.openclaw.ai` must only happen by negotiating with Night Watch on `stable.openclaw.ai`, never directly. -- Explicit repair-and-land authority includes internal scheduling, database admission, and lifecycle implementation decisions. The agent owns design selection, risk assessment, and verification; do not request renewed approval for implementation decisions within that scope. -- Bug fixes within the authorized task do not need renewed approval, including compatible SDK changes needed to restore intended behavior. Ask again for new configuration options, breaking public contracts, intentional changes to schemas, durability, retention, or permissions beyond the bug fix, paid services, or destructive actions. Preserve FIFO ordering, live-authority and integrity checks, and settlement of write-capable work. -- Protocol/version bumps, dependency patches/overrides/vendor changes, paid services, releases, and publishing need explicit approval; fix/ship authority does not imply release authority. Advisory workflows require an explicit request for that security action. -- Extended-stable is one line: the trailing completed month relative to `main`'s version. Older `.33+` lines retire when `main` advances another month; publishing a retired line needs an explicit maintainer decision, not a routine guard bypass. -- Baseline, snapshot, ignore, and expected-failure exceptions need approval, except narrow scanner qualifications for verified synthetic test fixtures. Agents may add these within the authorized task without asking when matching is bound to exact fixture bytes and source location, scanning and verification stay enabled, and proof rejects changed inputs. Broad suppressions and uncertain or real credentials still require approval. Exact shrink-only ratchet updates are maintenance. -- `CODEOWNERS` routes review; check live GitHub enforcement. Restricted/security paths and material product, behavior, security, or ownership changes need listed-owner involvement. For ownership/review governance, verified active organization-admin direction also qualifies; repository admin/bypass alone does not. Neither route waives enforced reviews. -- Complete the authorized workflow's review/merge gates; resolve substantive findings or explain rejections. Address failures under the best-effort test-failure policy above; distinguish proven unrelated failures from unexplained ones and cite an owning fix when known. Verify remote outcomes before success or cleanup; uncertain writes require reconciliation, not blind retries. -- Stage only intended files and use concise Conventional Commits with verified author/writer identities. Preserve contributor credit; team-session credit requires consented, verified humans and its canonical backlink. A bare URL grants no public mutation authority. Keep PR bodies current with problem, solution, impact, and evidence; use body files/heredocs for shell-sensitive text. +- Repair-and-land authority covers internal scheduling, database admission, and lifecycle design, risk, and verification; in-task bug fixes (including compatible SDK fixes) need no renewed approval. Ask again for new config options, breaking public contracts, intentional schema/durability/retention/permission changes beyond the bug fix, paid services, or destructive actions. Preserve FIFO ordering, live-authority/integrity checks, and write-capable work settlement. +- Protocol/version bumps, dependency patches/overrides/vendoring, paid services, releases, and publishing need explicit approval; fix/ship authority is not release authority. Advisory workflows need an explicit security request. +- Extended-stable is one line, the trailing completed month relative to `main`'s version; older `.33+` lines retire when `main` advances a month, and publishing a retired line needs an explicit maintainer decision, not a guard bypass. +- Baseline, snapshot, ignore, and expected-failure exceptions need approval, except narrow scanner qualifications for verified synthetic fixtures bound to exact bytes and location (scanning on, changed inputs rejected). Exact shrink-only ratchet updates are maintenance. +- `CODEOWNERS` routes review; restricted/security paths and material changes need listed-owner involvement ([review governance](.agents/skills/openclaw-pr-maintainer/SKILL.md#codeowners-review)). +- Complete the workflow's review/merge gates; resolve substantive findings or explain rejections. Verify remote outcomes before success or cleanup; reconcile uncertain writes, never retry blindly. +- Stage only intended files; concise Conventional Commits with verified author/writer identity. Preserve contributor credit; team-session credit needs consented, verified humans and the canonical backlink. A bare URL grants no public mutation authority. Keep PR bodies current with problem, solution, impact, and evidence; use body files for shell-sensitive text. ## Read when relevant -Read matching guides in full and follow their narrower task-specific pointers. -Commands and implementation detail stay with these owners. +Read matching guides in full; commands and details stay with them. -- **Product/design:** [VISION.md](VISION.md). -- **Plugins/discovery/SDK:** [plugins](extensions/AGENTS.md), [loader](src/plugins/AGENTS.md), [SDK](src/plugin-sdk/AGENTS.md). The SDK guide owns public boundary expansion, including callers outside these trees. -- **Channels/message actions:** [channel boundary](src/channels/AGENTS.md) and [channel responsibilities](docs/plugins/sdk-channel-plugins.md). -- **Agent tools, prompts, admission, or lifecycle:** [agents](src/agents/AGENTS.md) and [Gateway](src/gateway/AGENTS.md). -- **Control UI state, requests, or presentation:** [UI guide](ui/AGENTS.md), including state shared with other Gateway clients. -- **Storage:** [database schemas](docs/reference/database-schemas.md), then its layout, versioning, and storage-changes pages for the affected contract. Read the approval checkpoint before changing schema, transactions, retention, or recovery. -- **Config retirement/migration:** [shared Doctor transforms and startup migration](docs/gateway/doctor/config-migrations.md); reuse this owner instead of new runtime compatibility readers. -- **Audit/identity/receipts:** [audit doctrine](docs/gateway/audit.md). Diagnostic provenance is opt-in and never authorization; changes to collection, reader scope, retained fields, bounds, or contracts require approval. -- **Codex-backed behavior:** personally inspect the exact sibling `../codex` source before implementation or verdict and cite it; wrappers, schemas, and another agent's report do not replace this check. Auth/runtime/catalog routes use `openai`; legacy `openai-codex` input belongs only in migration. Harness upgrades refresh [the harness guide](docs/plugins/codex-harness.md) from `model/list`. -- **Validation commands:** [test suites](docs/help/testing/suites.md) is a command reference; this file and the chosen workflow own check selection. Test authoring also uses [writing tests](docs/help/testing/writing-tests.md) and the owning scoped guide. -- **GitHub:** [contribution rules](CONTRIBUTING.md), the current PR template, and [review feedback](docs/reference/pull-request-review-flow.md). The authorized maintainer workflow owns landing; native `scripts/pr` gates, recovery, and cleanup require [scripts guide](scripts/AGENTS.md). -- **Docs/public links:** [docs guide](docs/AGENTS.md). Update docs with behavior; normal fix notes belong in PRs because `CHANGELOG.md` is release-owned. -- **Releases:** the chosen release workflow and [release contract](docs/reference/RELEASING.md). Preserve the selected release cut and identity through publication and verification. npm-format lock mirrors are verified against `pnpm-lock.yaml`, published in dependency evidence, and kept out of npm tarballs. -- **Secrets/advisories:** [secret semantics](docs/gateway/secrets.md), [auth semantics](docs/auth-credential-semantics.md), and [security reporting](SECURITY.md) for the affected branch. -- **Live channels/native apps:** the owning scoped guide and permitted proof workflow. Telegram claims require Test Server userbot proof with Convex-leased credentials; platform claims require the relevant real device/platform evidence. Mac permission proof needs a stable, properly signed app; see [signing](docs/platforms/mac/signing.md). +- **Plugins/SDK:** [plugins](extensions/AGENTS.md), [loader](src/plugins/AGENTS.md), [SDK](src/plugin-sdk/AGENTS.md) (owns public boundary expansion, including outside callers). +- **Channels/message actions:** [channel boundary](src/channels/AGENTS.md), [channel responsibilities](docs/plugins/sdk-channel-plugins.md). +- **Agent tools, prompts, admission, lifecycle:** [agents](src/agents/AGENTS.md), [Gateway](src/gateway/AGENTS.md). +- **Control UI:** [UI guide](ui/AGENTS.md), including state shared with other Gateway clients. +- **Storage:** [database schemas](docs/reference/database-schemas.md) and its subpages; read its approval checkpoint before schema, transaction, retention, or recovery changes. +- **Config migration:** [Doctor transforms](docs/gateway/doctor/config-migrations.md), never new runtime compatibility readers. +- **Audit/identity:** [audit doctrine](docs/gateway/audit.md); provenance is opt-in, never authorization; changes to collection, reader scope, retained fields, bounds, or contracts need approval. +- **Codex-backed behavior:** personally inspect and cite the exact sibling `../codex` source first; other agents' reports don't substitute. Routes use `openai` (`openai-codex` only in migration); refresh [the harness guide](docs/plugins/codex-harness.md) from `model/list` on upgrades. +- **Validation:** [test suites](docs/help/testing/suites.md) lists commands; [writing tests](docs/help/testing/writing-tests.md) for authoring. +- **GitHub:** [contribution rules](CONTRIBUTING.md), the PR template, [review feedback](docs/reference/pull-request-review-flow.md); `scripts/pr` follows the [scripts guide](scripts/AGENTS.md). +- **Docs:** [docs guide](docs/AGENTS.md); update docs with behavior; fix notes go in PRs (`CHANGELOG.md` is release-owned). +- **Releases:** the release workflow and [release contract](docs/reference/RELEASING.md); preserve the selected cut and identity through publication and verification. npm-format lock mirrors are verified against `pnpm-lock.yaml`, published in dependency evidence, and kept out of npm tarballs. +- **Secrets/advisories:** [secrets](docs/gateway/secrets.md), [auth](docs/auth-credential-semantics.md), [security reporting](SECURITY.md). +- **Live channels/native apps:** the scoped guide; Telegram claims need Test Server userbot proof (Convex-leased credentials), platform claims real device evidence, Mac permission proof a stable, [signed](docs/platforms/mac/signing.md) app. diff --git a/scripts/AGENTS.md b/scripts/AGENTS.md index c84d55170d80..23dc7c6702e5 100644 --- a/scripts/AGENTS.md +++ b/scripts/AGENTS.md @@ -230,3 +230,13 @@ for the evidence fields and supported policy limits. - Before recording intent, `--auto-merge` selects immediate pinned squash for `MERGEABLE/CLEAN` or auto for `MERGEABLE/BEHIND` or `MERGEABLE/BLOCKED`, subject to all admission gates and queue policy. Accepted auto/queue requests are visible pending outcomes, not completion. The request carries `--match-head-commit`; confirmation still requires the exact attempted head. This is a submission-time head precondition, not a server-side freeze against later collaborator pushes. Existing or ambiguous requests are never automatically cancelled, re-armed, or followed by an immediate fallback; an explicitly investigated accepted auto request uses the cancellation recovery above. Ordinary `gh pr merge` can enqueue when queue policy applies. - A confirmed merge receipt precedes audits, comments, and cleanup. A comment POST has one attempt marker and is never blindly repeated. Recovery searches authoritative comments for that marker, reports completion pending, and leaves cleanup to the operator after ownership checks; it works without the original worktree/prep artifacts. Normal uninterrupted completion preserves comments and cleanup, with exact-head leased remote deletion. Authoritative branch absence completes cleanup; inspect warnings for advanced or inaccessible branches. Delayed recovery never deletes a recreated branch by name. - After ownership-checked cleanup, explicitly finish a verified receipt with `scripts/pr merge-complete --confirmed-operator-completion`. It revalidates the exact retained receipt and remote merge, requires native worktree/local branch/remote head branch absence, and never dispatches a merge or deletes resources. A `merged` receipt may post its first completion comment; `commenting`/`commented` require the existing exact attempt marker and never repost. Missing or ambiguous comments preserve pending state. For a retained prior-CI admin receipt, delayed completion reconstructs the historical parent comparison from retained admission main and the verified landed commit before checking cleanup absence. It labels the audit reconstructed after merge, claims no original at-landing audit, and retains the historical CI caveat. Legacy admin receipts without retained prior-CI proof still require owner review; no original audit is fabricated. Read the current outcome OID again after a state transition; stale OIDs are refused. Default `merge-run` remains reconciliation-only. + +## Execution Gotchas + +These commands apply on the host permitted by the task and its workflow; they do +not authorize local execution or a broader test plan. + +- For fs-safe dependency trouble, follow [on-demand vendoring instructions](vendor-fs-safe.md); keep vendor contents local and registry dependencies as the default. +- Restore missing dependencies in a trusted normal checkout with `pnpm install`, then retry once before diagnosing a code defect. Never reconcile a shared/worktree install while other jobs use it. +- Run the CLI through `pnpm openclaw ...` or `pnpm dev`, never `node --import tsx src/index.ts`; the supported wrappers own build freshness and process setup. +- Use installed `oxfmt` for formatting and the repository's `tsgo` lanes for typechecking. Inspect scope with `pnpm changed:lanes --json`; use targeted tests/checks. When avoiding worktree reconciliation, use `node scripts/check-changed.mjs` or `node scripts/run-vitest.mjs` with ready dependencies. Host restrictions still apply. diff --git a/src/plugin-sdk/AGENTS.md b/src/plugin-sdk/AGENTS.md index e6512c110126..c16ccef60f52 100644 --- a/src/plugin-sdk/AGENTS.md +++ b/src/plugin-sdk/AGENTS.md @@ -100,3 +100,16 @@ can affect bundled plugins and third-party plugins. and the most direct provider/plugin tests for the behavior you are centralizing. - Breaking removals or renames are major-version work, not drive-by cleanup. + +## Choose The Capability Surface + +For new capability, use the first path that expresses the actual requirement: + +1. Extend the existing owner or use an existing command, skill, plugin, or supported integration. +2. Use an existing plugin contract. Prefer bundle plugins for skills, MCP servers, and configuration; use code plugins when runtime hooks, providers, channels, or tools are needed. Keep vendor behavior with its vendor plugin and feature behavior with its feature owner. +3. If the contract is missing, define a narrow generic core/SDK capability and move existing bundled implementations and callers onto it together. Repeated independent requests for the same capability trigger this contract review, not another parallel manager or hook. +4. Add universal core surface only when the need is fundamental and existing extension points cannot express it. Explain the gap and ongoing cost; a new hook needs a concrete consumer. + +For example, a new channel action should first use the shared message action +contract. A setup screen needing plugin metadata should use the manifest or +lightweight artifact, not load the plugin's execution runtime.