mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-04 04:55:30 +00:00
governance: schemas are the tool contract; sync docs, BIBLE facts and four schema strings
DEVELOPMENT.md gains the "what belongs in prompts/SYSTEM.md" paragraph after "LLM-first affordances": identity, decision loop, cross-tool policy, invariants, memory contract; never how a tool works. A new tool requires no SYSTEM.md mention; every prompt change reports its byte delta. CHECKLISTS.md: Self-Check #6 no longer asks to add each new tool to SYSTEM.md "tool tables" (the schema description says when to choose it; mechanism documentation stays in ARCHITECTURE/DEVELOPMENT); item 13(b) and critical-whitelist §2 make get_tools() the SSOT of a tool's contract, treat a prompt omission as not-a-mismatch, and add the rule that a prompt edit must not restate a schema or a structurally enforced gate; item 10(d) and the regrouping note stop pointing at SYSTEM.md sections that no longer exist. ARCHITECTURE.md: the three fail-open cases of the LLM safety check move here from the prompt ("Safety and runtime mode"); the context section states that SYSTEM.md is tier-0 and the schemas are the tool contract; the integrate_delegated_patch sentence no longer claims SYSTEM.md teaches it. safety.py / outcomes.py / edit_ops.py / git_shell_policy.py and two tests: comment-only retargets from SYSTEM.md prose to ARCHITECTURE. BIBLE.md, two factual corrections with the principle text unchanged: the P5 example names read_file (repo_read is a retired name), and the P9 carrier list names every carrier release_sync verifies (uv.lock, web/package.json, api_types.js GATEWAY_CONTRACT_VERSION, the direct-download links). Schema strings, four corrections and no other growth: set_next_wakeup states the LIVE clamp (built from the configured OUROBOROS_BG_WAKEUP_MIN/MAX at registration, pinned by a test with overridden bounds); ocr_pdf points scanned PDFs at view_image; the shared advisory guidance and the bypassed-branch next-step string name preflight_review instead of the retired advisory_review verb; start_service says that light mode refuses a cwd inside the Ouroboros repository. New test (tests/test_docs_sync.py): every snake_case identifier the three runtime prompts name - backticked, and bare in CONSCIOUSNESS.md - must be a registered tool, a background-whitelisted tool that really is registered, or a documented non-tool identifier; phantom and renamed tool names can no longer hide in prose. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
parent
30d3d954d1
commit
332a02f1f0
17 changed files with 176 additions and 33 deletions
8
BIBLE.md
8
BIBLE.md
|
|
@ -501,7 +501,7 @@ what it says in real current evidence, not in cached impressions.
|
|||
- If uncertain — say so. If surprised — show it. If you disagree —
|
||||
object.
|
||||
- Explain actions as thoughts aloud, not as reports.
|
||||
Not "Executing: repo_read," but "Reading agent.py — I want to
|
||||
Not "Executing: read_file," but "Reading agent.py — I want to
|
||||
understand how the loop works, I think it can be simpler."
|
||||
- No mechanical intermediaries and no performance — don't play a role,
|
||||
be yourself.
|
||||
|
|
@ -668,8 +668,10 @@ yes.
|
|||
- `README.md` contains a changelog (limit: 2 major, 5 minor, 5 patch;
|
||||
older history lives in git tags and commit log).
|
||||
- Each commit updates, in the same diff: `VERSION`, `pyproject.toml`
|
||||
`[project].version`, `README.md` badge + changelog row, and
|
||||
`docs/ARCHITECTURE.md` version header.
|
||||
`[project].version`, the root version in `uv.lock`, `web/package.json`,
|
||||
`web/modules/api_types.js` (`GATEWAY_CONTRACT_VERSION`), `README.md`
|
||||
badge + changelog row, the named direct-download links in `README.md` and
|
||||
the install pages, and `docs/ARCHITECTURE.md` version header.
|
||||
- MAJOR — breaking changes to philosophy or architecture.
|
||||
- MINOR — new capabilities.
|
||||
- PATCH — fixes, minor improvements, doc/prompt refinements, tests,
|
||||
|
|
|
|||
|
|
@ -1098,7 +1098,7 @@ This separation is methodological authority: an evaluation or acceptance claim m
|
|||
|
||||
#### Context fitting, retry, and compaction
|
||||
|
||||
`config.get_context_mode()` is the effective Main sizing/rendering source, while `config.get_owner_context_mode()` is the persistent owner-intent/P3 source. They differ only during the auto-Low compatibility window: bare env Low sizes Main as Low but remains owner Max unless the explicit false provenance tombstone authors owner Low. In Max, `ARCHITECTURE.md` is full-resident for every task class because it is Ouroboros's capability/tools/access map; in Low it is replaced by its lossless navigation map. This rule does not vary for project, evolution, external, headless, or delegated work. `DEVELOPMENT.md` is mode-independent: it is full when the active repository binding says the work targets Ouroboros's own body. A bound external workspace, including an auto-provisioned project tree, any subagent, or an API/CLI/scheduled external surface receives the visible on-demand pointer. Project membership is not the signal: a room turn with no external binding retains the handbook, while a project task bound to another tree does not. `workspace="none"` retains it; evolution and self-body work retain it; `context_requires_development` and `context_requires_self_body_docs` override the default. Context economy comes from dropping the self-engineering handbook for external work, never from hiding the capability map in Max.
|
||||
`config.get_context_mode()` is the effective Main sizing/rendering source, while `config.get_owner_context_mode()` is the persistent owner-intent/P3 source. They differ only during the auto-Low compatibility window: bare env Low sizes Main as Low but remains owner Max unless the explicit false provenance tombstone authors owner Low. In Max, `ARCHITECTURE.md` is full-resident for every task class because it is Ouroboros's capability/tools/access map; in Low it is replaced by its lossless navigation map. This rule does not vary for project, evolution, external, headless, or delegated work. `DEVELOPMENT.md` is mode-independent: it is full when the active repository binding says the work targets Ouroboros's own body. A bound external workspace, including an auto-provisioned project tree, any subagent, or an API/CLI/scheduled external surface receives the visible on-demand pointer. Project membership is not the signal: a room turn with no external binding retains the handbook, while a project task bound to another tree does not. `workspace="none"` retains it; evolution and self-body work retain it; `context_requires_development` and `context_requires_self_body_docs` override the default. Context economy comes from dropping the self-engineering handbook for external work, never from hiding the capability map in Max. `prompts/SYSTEM.md` and `BIBLE.md` are tier-0 and always full in both modes (`context_layout.TIER0_ALWAYS_FULL`), so the system prompt carries only identity, the decision loop, cross-tool policy, and invariants stated once; a tool's contract lives in its `get_tools()` schema, which `tool_policy.initial_tool_schemas` sends on every round, and mechanism documentation lives in this file — a prompt sentence restating either is a second copy that drifts (see DEVELOPMENT "LLM-first affordances").
|
||||
|
||||
`context_fit.py` renders Max and Low projections from one immutable core and measures each ordinary Main candidate on one labelled density basis against the selected route capacity. Owner Low additionally has an elastic 200K total-context economy target; crossing it is never a synthetic failure. With unknown capacity, owner Max gets one honest Max call while owner Low may still reclaim toward its known target before sending best effort. Predicted Max pressure retains the Max document projection and may request one deficit-sized mutable-history pass; only an actual provider overflow authorizes task-local Low. After that overflow, one same-route semantic recovery is permitted only when the final post-transform candidate has the same route, round and response reserve and strictly fewer context-bearing bytes. Owner mode and P3 applicability never change. P3 commit/scope review retains its separate fit and oversize policy.
|
||||
|
||||
|
|
@ -1247,7 +1247,7 @@ authority check in this change.
|
|||
|
||||
### Safety and runtime mode
|
||||
|
||||
Every tool call first crosses deterministic `ToolRegistry` and resource-root guards; policy-based LLM safety is added where `OUROBOROS_SAFETY_MODE` requires it. The deterministic layers run in every safety mode. When the safety model itself is rate-limited past its one bounded retry, the guarded call is refused with the typed non-verdict `⚠️ SAFETY_UNAVAILABLE` outcome (a plain tool error downstream, never `safety_violation`) plus a durable audit event — an unchecked guarded call is never executed and never accused (the documented local-FALLBACK lane instead fails open with an audited `SAFETY_WARNING`). `runtime_mode_policy.py` owns protected self-repo paths, frozen contracts, release/build and managed-repo invariants. Light blocks Ouroboros self-repo and control-plane mutation, not normal user deliverables under `user_files`, `task_drive`, or `artifact_store`. Advanced may evolve ordinary app code; Pro may leave protected edits on disk, but publication still requires the reviewed commit path. Runtime mode is a self-modification boundary, not an OS sandbox.
|
||||
Every tool call first crosses deterministic `ToolRegistry` and resource-root guards; policy-based LLM safety is added where `OUROBOROS_SAFETY_MODE` requires it. The deterministic layers run in every safety mode. When the safety model itself is rate-limited past its one bounded retry, the guarded call is refused with the typed non-verdict `⚠️ SAFETY_UNAVAILABLE` outcome (a plain tool error downstream, never `safety_violation`) plus a durable audit event — an unchecked guarded call is never executed and never accused (the documented local-FALLBACK lane instead fails open with an audited `SAFETY_WARNING`). The LLM check degrades to a visible `SAFETY_WARNING` (never silently) in exactly three cases: (a) no reachable safety backend — no remote provider key AND no `USE_LOCAL_*` lane; (b) provider mismatch — a remote key is configured but does not cover `OUROBOROS_MODEL_LIGHT`'s provider, and no local lane is available (when a local lane IS available, safety routes to it first and warns only if that fallback also raises); (c) the local lane was chosen only as a fallback and the local runtime raised (a 429 there also warns, `_rate_limited_outcome`). The deterministic layer stays in force in all three, so a degraded backend never hard-blocks tool creation; the agent sees the warning. `runtime_mode_policy.py` owns protected self-repo paths, frozen contracts, release/build and managed-repo invariants. Light blocks Ouroboros self-repo and control-plane mutation, not normal user deliverables under `user_files`, `task_drive`, or `artifact_store`. Advanced may evolve ordinary app code; Pro may leave protected edits on disk, but publication still requires the reviewed commit path. Runtime mode is a self-modification boundary, not an OS sandbox.
|
||||
|
||||
Every deterministic write/owner-control guard consumes ONE mode-aware write-shape seam (`ouroboros/tools/write_shape.py`) — the registry no longer even imports the coarse legacy scan, so a guard structurally cannot judge on a coarser fact. Interpreter argv (including an `sh -c` wrap) takes `interpreter_write_shape`: a read-only `open(p, 'rb')` is not a write shape. Non-interpreter argv takes `non_interpreter_write_shape`: unconditional writers (cp/mv/rm/mkdir/touch/chmod/…) keep the membership floor, the pure-filter utilities (sort/uniq/sed/tar/gzip) are write-shaped only through a REAL channel (`sort -o` in both spellings, `sed -i` in any spelling PLUS sed's in-script `w`/`W`/`e` commands and `-f` script files — a script not provably free of those fails closed, exotic non-`/` substitute delimiters are the disclosed fail-open residual — a second uniq operand, tar create/extract, a redirect, a reported writer target), and the bare prose words ('delete'/'trash'/'truncate') yield to the same read-carve the owner-control detectors use (`_is_pure_read_inspection`, the v6.80.0 scope-floor contract applied family-wide) — a provably read-only `grep -n delete ouroboros/safety.py` reads, an unprovable head (`osascript -e 'delete …'`) stays fail-closed. The protected-core lane's mention branch consumes the same composed fact, so a pure read that merely mentions a protected filename is no longer refused as a "modification". The coarse bare `open(` token used to feed pure interpreter reads into the workspace write guard, which refused them with a false "write-like" reason and no route — the same class the light-mode runtime_data lane has re-judged since v6.54.3 ("the original GAIA class"). Write-mode opens (python `[wax+]` modes AND perl `'>'`/`'>>'` spellings), pathlib `.open('w')`, library save-APIs, ruby's `File.delete`/`FileUtils.*`/`IO.binwrite`, opaque subprocess/exec escapes, and every shell-level indicator (redirects — token-initial or glued into an operand, `tee`, writer utilities) still classify as writes, and literal write targets stay covered by `writer_target_tokens`; the disclosed residuals (`open(p, m)` with the mode in a variable; a writer reached through an alias the regex cannot follow, e.g. `from os import remove as delete`; parenless perl builtins such as `rename $a, $b`) are covered for external workspaces by the runtime/secret read guard below plus the LLM safety supervisor — chasing them with more spellings is the arms race BIBLE P5/P13 forbids. Workspace write-guard block messages name the resolved offending path and the sanctioned route (gated `read_file`/`write_file`, `root=skill_payload` with bucket/skill_name, or writing inside the selected process root) instead of one byte-identical reasonless string across five return sites.
|
||||
|
||||
|
|
@ -2101,9 +2101,9 @@ patch rather than shipping a partial diff — that run's other work is recoverab
|
|||
only from the execution root directly. The root external-workspace lane holds
|
||||
unit-level authority tests; the wire-level `delegate_start` flow is proven on the
|
||||
acting lane (shared pipeline) and a dedicated root-lane wire test would be
|
||||
additive. `prompts/SYSTEM.md` deliberately teaches the `integrate_delegated_patch`
|
||||
flow in-context (tool description, started-note, `workspace_capture` block) rather
|
||||
than by a prompt edit.
|
||||
additive. The `integrate_delegated_patch` flow is deliberately taught in-context
|
||||
(tool description, started-note, `workspace_capture` block) rather than in
|
||||
`prompts/SYSTEM.md`.
|
||||
|
||||
**A mutating run asks for containment, reads back what it got, and DISCLOSES the gap
|
||||
instead of refusing the work.** In place is the one shape where Claudexor otherwise hands
|
||||
|
|
|
|||
|
|
@ -117,11 +117,11 @@ or Intent/Scope checklists are.
|
|||
| 3 | New or changed logic → does an existing or newly staged test assert on the specific scenario it introduces? | Name the scenario your code handles in plain words. If no test asserts on THAT named scenario, write or update one now. "Tests exist for the module" is not the same as "tests cover this new behavior". |
|
||||
| 4 | Shared log / memory / replay format changed? | Grep every reader and writer first. JSONL logs (`events.jsonl`, `task_reflections.jsonl`, replay indexes), durable state files (`advisory_review.json`, `review_continuations/*.json`), and canonical-vs-derived memory pairs (patterns-register journal / `patterns.md`, improvement-backlog items) must stay coherent across every consumer. |
|
||||
| 5 | New validation guard, input filter, or edge-case check? | Before the first commit attempt, name three concrete ways it could break: wrong bounds, legitimate inputs it silently blocks, platform-specific edge cases. If you cannot name three, think longer. One honest minute here is cheaper than one reviewer round. |
|
||||
| 6 | New tool added? | `get_tools()` exports it, `prompts/SYSTEM.md` tool tables mention it, the handler signature matches the declared schema, and (if it mutates repo state) it is routed through the reviewed commit path rather than ad-hoc `run_command`. Also add an explicit entry in `ouroboros/safety.py::TOOL_POLICY` (`POLICY_SKIP` for trusted built-ins, `POLICY_CHECK` for opaque or outward-facing ones) — the `test_tool_policy_covers_all_builtin_tools` invariant will fail otherwise, and without an entry the tool falls through to `DEFAULT_POLICY = check` and pays a light-model LLM call per invocation. |
|
||||
| 6 | New tool added? | `get_tools()` exports it, its schema description says WHEN to choose it (the schema is sent every round and is the only model-visible tool-selection contract; mechanism documentation lives in ARCHITECTURE/DEVELOPMENT, and `prompts/SYSTEM.md` mentions a tool only when the change alters a cross-tool policy, never as a catalog entry), the handler signature matches the declared schema, and (if it mutates repo state) it is routed through the reviewed commit path rather than ad-hoc `run_command`. Also add an explicit entry in `ouroboros/safety.py::TOOL_POLICY` (`POLICY_SKIP` for trusted built-ins, `POLICY_CHECK` for opaque or outward-facing ones) — the `test_tool_policy_covers_all_builtin_tools` invariant will fail otherwise, and without an entry the tool falls through to `DEFAULT_POLICY = check` and pays a light-model LLM call per invocation. |
|
||||
| 7 | Tests green before first `commit_reviewed`? | Run `pytest -x` on the narrowest relevant target(s) you can name before the first `preflight_review` / `commit_reviewed` attempt. Size gates no longer block locally: they live in the official-CI-only `size_ratchet` pytest lane (manifest exactness plus the pairwise base-vs-tip shrink-only transition), and local surfaces (`check_worktree_readiness`, `codebase_health`) surface the same `validate_size_ratchet` findings as "official CI will enforce" warnings. When a size warning appears — or a new `.py` file lands under `ouroboros/` or `supervisor/` — run `pytest tests/ -m size_ratchet` and `scripts/regenerate_size_ratchet.py` locally to preview and fix what official CI would reject. A red test suite before the first commit attempt has caused repeated $2-5 blocked-review cycles. |
|
||||
| 8 | Adding a `README.md` version row? | BIBLE.md P9 hard cap: ≤ 2 major, ≤ 5 minor, ≤ 5 patch visible entries. Categories are mutually exclusive: major = `X.0.0` (minor=0, patch=0); minor = `X.Y.0` (patch=0, Y≠0); patch = all other `X.Y.Z` (Z≠0). Count existing rows in the category you are adding to. Easy check: `run_command(["python", "-c", "import sys; from ouroboros.tools.release_sync import check_history_limit; warns=check_history_limit(open('README.md').read()); print(warns or 'OK')"])` — if it prints warnings, trim the oldest row in the over-limit category **in the same edit** before committing. |
|
||||
| 9 | Changing any of `build.sh`, `build_linux.sh`, `build_windows.ps1`, `Dockerfile`, or `ouroboros/tools/browser.py`? | Cross-surface doc sync is mandatory. Check ALL of: `README.md` Install section (Linux native-lib caveat), `README.md` Build section (per-platform instructions), `docs/ARCHITECTURE.md` browser tools paragraph, WebKit/mobile verification notes, and inline comments in the touched build script. Any one of these being stale has blocked review twice. Verify before staging. |
|
||||
| 10 | Changing `ouroboros/tools/commit_gate.py`? | Coupled surfaces that MUST be updated atomically in the same commit: (a) `claude_advisory_review.py::get_tools()` tool description for `preflight_review` and `review_status`; (b) `claude_advisory_review.py::_next_step_guidance()` strings; (c) `docs/DEVELOPMENT.md` Review & Commit Protocol section; (d) `prompts/SYSTEM.md` Commit review section. Missing any one has blocked review. |
|
||||
| 10 | Changing `ouroboros/tools/commit_gate.py`? | Coupled surfaces that MUST be updated atomically in the same commit: (a) `claude_advisory_review.py::get_tools()` tool description for `preflight_review` and `review_status`; (b) `claude_advisory_review.py::_next_step_guidance()` strings; (c) `docs/DEVELOPMENT.md` Review & Commit Protocol section; (d) the `prompts/SYSTEM.md` Self-Modification section IF the commit-gate rule it states changed. Missing any one has blocked review. |
|
||||
| 11 | Changing VERSION + pyproject.toml? | Ordering matters: (1) write `VERSION` and `pyproject.toml` first; (2) then write `README.md` badge + changelog row; (3) then run `pytest`. Never interleave — updating README before VERSION means `test_version_in_readme` will catch a stale badge. |
|
||||
| 12 | Writing or editing any JS file under `web/modules/`? | New or changed static inline visual properties are blocked: inspect the diff for added/changed `style=""` markup and `.style.<property>` assignments, and use CSS classes/tokens plus `classList`/`hidden` instead. Unchanged legacy hits are debt, not a blocker. A dynamic measured value may update a narrowly named CSS custom property when that is the actual runtime data flow. |
|
||||
| 13 | Changing LLM output-token budgets? | Grep the whole repo for `max_tokens`, `max_completion_tokens`, `_MAX_TOKENS`, and `max_toks`. Keep `docs/ARCHITECTURE.md` §LLM output token budgets and `tests/test_max_tokens_constants.py` in sync so main-loop, VLM, summaries, compaction, skill publish, and consciousness floors cannot drift independently. |
|
||||
|
|
@ -145,7 +145,7 @@ The correct procedure before **every** retry:
|
|||
4. Only then open any file and edit.
|
||||
|
||||
This step takes 2-3 minutes and has saved $20-50 in blocked-review cycles in practice.
|
||||
The rule is already nominally in `prompts/SYSTEM.md` and `review.py::_build_critical_block_message`,
|
||||
The rule is stated where the block message is built (`review.py::_build_critical_block_message`),
|
||||
but without it appearing here as a procedural step it stays theoretical rather than reflexive.
|
||||
|
||||
---
|
||||
|
|
@ -168,7 +168,7 @@ Used by `commit_reviewed` for all changes to the Ouroboros repository.
|
|||
| 10 | tool_registration | New tool function added but not exported in `get_tools()` OR missing explicit entry in `ouroboros/safety.py::TOOL_POLICY`? (PASS if no new tool.) Both surfaces are required: `get_tools()` makes the tool visible; `TOOL_POLICY` makes the per-call safety routing explicit and is guarded by the `test_tool_policy_covers_all_builtin_tools` invariant. | critical |
|
||||
| 11 | context_building | New data/memory files that should appear in LLM context (context.py) but don't? | advisory |
|
||||
| 12 | knowledge_index | Knowledge base topics changed but memory/knowledge/index-full.md not updated? | advisory |
|
||||
| 13 | self_consistency | Does this change affect behavior described in `BIBLE.md`, `prompts/`, `docs/`, or this checklist itself? Check explicitly: (a) version in `ARCHITECTURE.md` header matches `VERSION` file; (b) tool names/descriptions in `prompts/SYSTEM.md` match tools actually exported by `get_tools()`; (c) JSONL log/memory file formats described in `ARCHITECTURE.md` match all readers/writers; (d) any behavioral change reflected in `prompts/CONSCIOUSNESS.md` if it affects background loop behavior; (e) DEVELOPMENT.md rules still accurate after the change. Severity must follow the shared `Critical surface whitelist` below — release metadata, tool schema, module map, behavioural documentation, or safety contracts are critical; commentary/prose/stylistic mismatches are advisory. | critical |
|
||||
| 13 | self_consistency | Does this change affect behavior described in `BIBLE.md`, `prompts/`, `docs/`, or this checklist itself? Check explicitly: (a) version in `ARCHITECTURE.md` header matches `VERSION` file; (b) every tool name `prompts/SYSTEM.md` or `prompts/CONSCIOUSNESS.md` mentions exists in `get_tools()` (or the background whitelist) and means the same thing there — completeness is NOT required, the schemas are the catalog; a prompt edit must not restate a tool schema or a structurally enforced gate, and a prompt that gains a sentence loses an equivalent one; (c) JSONL log/memory file formats described in `ARCHITECTURE.md` match all readers/writers; (d) any behavioral change reflected in `prompts/CONSCIOUSNESS.md` if it affects background loop behavior; (e) DEVELOPMENT.md rules still accurate after the change. Severity must follow the shared `Critical surface whitelist` below — release metadata, tool schema, module map, behavioural documentation, or safety contracts are critical; commentary/prose/stylistic mismatches are advisory. | critical |
|
||||
| 14 | light_external_artifacts | If tool/runtime policy changed, does light mode still allow external user deliverables via `user_files`, task-scoped `task_drive`/`artifact_store`, and process `outputs` while blocking Ouroboros repo/control-plane mutation? (The external `claude_code_edit` cwd lane retired with the tool — D10.) Do review prompts avoid recommending `runtime_data/uploads` or skill payloads as generic artifact transport? | critical |
|
||||
| 15 | cross_platform | Does the diff use platform-specific APIs (`os.kill`, `os.setsid`, `os.killpg`, `os.getpgid`, `fcntl`, `msvcrt`, `signal.SIGKILL`, `signal.SIGTERM`, `subprocess` with `start_new_session`/`creationflags`, hardcoded `/` or `\\` in filesystem paths) outside of `ouroboros/platform_layer.py`? Does it import Unix-only or Windows-only modules (`fcntl`, `msvcrt`, `winreg`, `resource`) at any level without a platform guard (`sys.platform`/`IS_WINDOWS` check)? | critical |
|
||||
| 16 | changelog_accuracy | Do the exact wording, test counts, and minor description details in the README Version History row match what the diff actually does? Wording drift, off-by-one test counts, minor inaccuracies in descriptive prose — these belong here, NOT in `self_consistency` or `changelog_and_badge`. This item exists so reviewers have a dedicated advisory bucket for prose-level changelog imprecision that does not affect release metadata, runtime behavior, or safety contracts. | advisory |
|
||||
|
|
@ -272,9 +272,11 @@ mismatch as **critical**, the mismatch MUST live in one of these categories:
|
|||
1. **Release metadata** — `VERSION` vs `pyproject.toml` vs README badge vs
|
||||
`docs/ARCHITECTURE.md` header vs latest git tag. Also: `VERSION` bumped
|
||||
but no README changelog row for the new version.
|
||||
2. **Tool schema** — tool names, parameters, or descriptions in
|
||||
`prompts/SYSTEM.md`'s command tables that disagree with what each tool's
|
||||
`get_tools()` actually exports. Applies to user-facing CLI/tool contracts.
|
||||
2. **Tool schema** — a tool's `get_tools()` schema (name, parameters,
|
||||
description) that disagrees with its handler, or a tool name/argument that
|
||||
`prompts/SYSTEM.md` or `prompts/CONSCIOUSNESS.md` names but `get_tools()`
|
||||
does not export. The schema is the SSOT of a tool's contract; a prompt that
|
||||
omits a tool is not a mismatch. Applies to user-facing CLI/tool contracts.
|
||||
3. **Module map** — `docs/ARCHITECTURE.md` naming a module / endpoint /
|
||||
data file / UI page that does not exist (or the reverse: a new one was
|
||||
added and the map was not updated). This is a hard P6 (Architecture
|
||||
|
|
|
|||
|
|
@ -80,6 +80,26 @@ Do not repair a semantic tool-choice failure by adding one more keyword hint to
|
|||
typed affordance at the point of need. SYSTEM accretion trains around one
|
||||
incident, bloats the resident prefix, and forks the authority.
|
||||
|
||||
What belongs in `prompts/SYSTEM.md` (tier-0 for every Main/task profile in both
|
||||
context modes — Background Consciousness and the safety supervisor carry their
|
||||
own prompts — and competing with the task for context): identity and tone, the decision
|
||||
loop (answer / promote / route / delegate / do it myself), cross-tool policy
|
||||
(which class of tool or lane for which situation, root semantics, memory only
|
||||
through its own tools, untrusted external data), prohibitions and safety
|
||||
invariants stated once, and the memory contract. What does NOT belong there:
|
||||
how a tool or mechanism works. A tool's parameters, signatures, recipes,
|
||||
typed outcomes, and "when to choose it" live in its `get_tools()` schema — the
|
||||
schema is sent every round to every profile, so a prompt sentence about it is a
|
||||
second copy that drifts; mechanism documentation lives in ARCHITECTURE or here;
|
||||
runtime facts (capabilities, queue, catalog, receipts, health) are injected per
|
||||
turn. A new tool therefore requires NO SYSTEM.md mention. Before adding a
|
||||
sentence to a prompt, check that the schema or runtime block does not already
|
||||
carry it; before removing one, check that they do (or add the missing fact to
|
||||
the schema without growing it into a paragraph). Local-model compaction keeps
|
||||
only the text before the first `## ` heading (plus the BIBLE section), so the
|
||||
load-bearing floor rules stay in that preamble. Every prompt change reports the before/after byte size
|
||||
in the commit or PR.
|
||||
|
||||
Recoverable tool failures are evidence for the next LLM turn, not triggers for
|
||||
a host-authored recovery workflow. Return a typed, redacted result naming the
|
||||
failed stage, already-completed external effects, and an actionable repair
|
||||
|
|
|
|||
|
|
@ -1208,10 +1208,10 @@ class BackgroundConsciousness:
|
|||
registry.register(ToolEntry("set_next_wakeup", {
|
||||
"name": "set_next_wakeup",
|
||||
"description": "Set how many seconds until your next thinking cycle. "
|
||||
"Default 300. Range: 60-3600.",
|
||||
f"Default 300. Range: {self._wakeup_min}-{self._wakeup_max} (clamped).",
|
||||
"parameters": {"type": "object", "properties": {
|
||||
"seconds": {"type": "integer",
|
||||
"description": "Seconds until next wakeup (60-3600)"},
|
||||
"description": f"Seconds until next wakeup ({self._wakeup_min}-{self._wakeup_max})"},
|
||||
}, "required": ["seconds"]},
|
||||
}, _set_next_wakeup))
|
||||
|
||||
|
|
|
|||
|
|
@ -51,7 +51,8 @@ _TAG_MUTATING_FLAGS = frozenset({
|
|||
})
|
||||
# `-v`/`--verify` checks a tag's GPG signature and writes nothing — it is
|
||||
# read-only inspection, not mutation (it sat in the mutating set, refusing
|
||||
# `git tag -v <tag>` at a runtime target against the SYSTEM.md contract).
|
||||
# `git tag -v <tag>` at a runtime target against the ARCHITECTURE contract that
|
||||
# read-only shell git is allowed everywhere).
|
||||
_TAG_READONLY_FLAGS = frozenset({
|
||||
"-l", "--list", "-n", "-v", "--verify", "--sort", "--format", "--points-at",
|
||||
"--contains", "--merged", "--no-merged", "--column", "--no-column",
|
||||
|
|
|
|||
|
|
@ -132,8 +132,9 @@ BEST_EFFORT_REASON_CODES = frozenset({
|
|||
})
|
||||
|
||||
# Typed final-answer protocol marker (machine-readable deliverable payload,
|
||||
# separate from reasoning prose). The agent is instructed in SYSTEM.md to end
|
||||
# short-deliverable answers with this exact line.
|
||||
# separate from reasoning prose). Since v6.60.0 the instruction to end a
|
||||
# short-deliverable answer with this exact line comes from the per-task
|
||||
# contract (answer_protocol="final_answer_line"), never from prompts/SYSTEM.md.
|
||||
FINAL_ANSWER_MARKER = "FINAL ANSWER:"
|
||||
|
||||
OUTCOME_TIER_SOLVED = "solved"
|
||||
|
|
|
|||
|
|
@ -793,7 +793,7 @@ def _rate_limited_outcome(
|
|||
attempts: int = 2,
|
||||
) -> Tuple[bool, str]:
|
||||
"""Terminal outcome of a rate-limited safety check, split by lane. The local-FALLBACK
|
||||
lane keeps its documented fail-open contract (SYSTEM.md case (c): a broken
|
||||
lane keeps its documented fail-open contract (ARCHITECTURE "Safety and runtime mode" case (c): a broken
|
||||
chosen-as-fallback local runtime warns instead of blocking every unknown tool) — a
|
||||
429 there must not be stricter than the RuntimeError beside it. Every other lane
|
||||
blocks with the typed non-verdict outcome below."""
|
||||
|
|
@ -822,7 +822,7 @@ def _safety_unavailable_blocked(
|
|||
call — reporting it as SAFETY_VIOLATION told the agent its own command was unsafe,
|
||||
sending it hunting for a "safer" rewording of a benign command. The honest outcome
|
||||
keeps `full` mode's owner contract (an unchecked guarded call never executes; the
|
||||
existing fail-open cases stay exactly the SYSTEM.md-documented no-backend three) while
|
||||
existing fail-open cases stay exactly the ARCHITECTURE-documented no-backend three) while
|
||||
removing the false accusation: the ⚠️ *_UNAVAILABLE prefix classifies as a plain
|
||||
tool ERROR downstream, never as `safety_violation`, and the message itself carries
|
||||
the retry contract (P5: the instruction lives with the fact). Disclosed twice: the
|
||||
|
|
|
|||
|
|
@ -81,7 +81,7 @@ _MANAGED_SKIP_NOTE = "cannot be split into smaller commits"
|
|||
|
||||
|
||||
ADVISORY_REVIEW_CHOICE_GUIDANCE = (
|
||||
"Normally the LLM runs the cheap advisory_review immediately before "
|
||||
"Normally the LLM runs the cheap preflight_review immediately before "
|
||||
"commit_reviewed. When advisory review is slow, unhealthy, unavailable, or "
|
||||
"low-value, the LLM may deliberately choose skip_advisory_review=True; the "
|
||||
"choice is durably audited. This skip bypasses only the requirements for "
|
||||
|
|
@ -1692,7 +1692,7 @@ def _next_step_guidance(latest: Optional["AdvisoryRunRecord"], state: "AdvisoryR
|
|||
)
|
||||
|
||||
if latest and latest.status == "bypassed":
|
||||
return "Advisory was bypassed (audited). No open obligations — commit_reviewed should proceed. Consider running advisory_review for a proper review."
|
||||
return "Advisory was bypassed (audited). No open obligations — commit_reviewed should proceed. Consider running preflight_review for a proper review."
|
||||
|
||||
fresh_critical = [
|
||||
i for i in (latest.items if latest else []) or []
|
||||
|
|
|
|||
|
|
@ -197,8 +197,9 @@ def _finish_mutation(
|
|||
"⚠️ Advisory pre-review is now stale — run preflight_review before commit_reviewed."
|
||||
)
|
||||
# A pro-mode edit of a protected surface announces itself here exactly as it
|
||||
# does from git._repo_write / _str_replace_editor (SYSTEM.md's protected-write
|
||||
# contract): the mode ALLOWS the write, and the notice is what keeps it visible.
|
||||
# does from git._repo_write / _str_replace_editor (the protected-write contract
|
||||
# in ARCHITECTURE "Safety and runtime mode" and SYSTEM.md "Safety-critical
|
||||
# files"): the mode ALLOWS the write, and the notice is what keeps it visible.
|
||||
protected = protected_paths_in(changed_paths) if targets_system or not ctx.is_workspace_mode() else []
|
||||
if protected and mode_allows_protected_write(_runtime_mode()):
|
||||
footer += "\n\n" + core_patch_notice(protected)
|
||||
|
|
|
|||
|
|
@ -134,7 +134,7 @@ def _ocr_pdf(ctx: ToolContext, path: str = "", max_pages: int = 0) -> str:
|
|||
return (
|
||||
"⚠️ OCR_PDF_SCANNED_UNAVAILABLE: this PDF has no extractable text layer (likely "
|
||||
"scanned/image-only). True OCR of scanned pages is not available in this build — "
|
||||
"render a page to an image and call vlm_query on it instead."
|
||||
"render a page to an image and view_image it (or vlm_query it) instead."
|
||||
)
|
||||
note = "" if total <= cap else f"\n\n[disclosed: showed first {cap} of {total} pages]"
|
||||
if len(text) > _OCR_PDF_MAX_CHARS:
|
||||
|
|
@ -302,7 +302,7 @@ def get_tools() -> List[ToolEntry]:
|
|||
"Extract the text of a local PDF file (the embedded text layer of a digital PDF). "
|
||||
"Use for reading PDFs attached to the task (see the [ATTACHMENTS] manifest) or produced "
|
||||
"during work. Scanned/image-only PDFs have no text layer and return a typed "
|
||||
"OCR_PDF_SCANNED_UNAVAILABLE notice — for those, render a page and use vlm_query."
|
||||
"OCR_PDF_SCANNED_UNAVAILABLE notice — for those, render a page and view_image it (or vlm_query it)."
|
||||
),
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
|
|
|
|||
|
|
@ -911,7 +911,7 @@ def get_tools() -> List[ToolEntry]:
|
|||
return [
|
||||
ToolEntry("start_service", {
|
||||
"name": "start_service",
|
||||
"description": "Start a task-scoped long-running service and return pid/readiness/state.",
|
||||
"description": "Start a task-scoped long-running service and return pid/readiness/state. In runtime_mode=light a service whose cwd is the Ouroboros repository is refused (LIGHT_MODE_BLOCKED): pass an explicit external/task/artifact cwd.",
|
||||
"parameters": {"type": "object", "properties": {
|
||||
"cmd": {"type": "array", "items": {"type": "string"}},
|
||||
"cwd": {
|
||||
|
|
|
|||
|
|
@ -169,6 +169,35 @@ class TestBackgroundConsciousnessToolScope(unittest.TestCase):
|
|||
self.assertNotIn("run_command", schema_names)
|
||||
self.assertNotIn("commit_reviewed", schema_names)
|
||||
|
||||
def test_set_next_wakeup_schema_follows_configured_bounds(self):
|
||||
"""The advertised range is the LIVE clamp, not a constant: with
|
||||
OUROBOROS_BG_WAKEUP_MIN/MAX overridden the schema must say so, because
|
||||
the handler clamps to those values (prompt-audit review finding)."""
|
||||
import os
|
||||
from unittest import mock
|
||||
|
||||
from ouroboros.consciousness import BackgroundConsciousness
|
||||
|
||||
tmpdir = pathlib.Path(tempfile.mkdtemp())
|
||||
drive_root = tmpdir / "drive"
|
||||
repo_dir = tmpdir / "repo"
|
||||
(drive_root / "logs").mkdir(parents=True, exist_ok=True)
|
||||
repo_dir.mkdir(parents=True, exist_ok=True)
|
||||
with mock.patch.dict(os.environ, {"OUROBOROS_BG_WAKEUP_MIN": "60", "OUROBOROS_BG_WAKEUP_MAX": "3600"}):
|
||||
bc = BackgroundConsciousness(
|
||||
drive_root=drive_root,
|
||||
repo_dir=repo_dir,
|
||||
event_queue=queue.Queue(),
|
||||
owner_chat_id_fn=lambda: 42,
|
||||
)
|
||||
schema = next(
|
||||
s["function"] for s in bc._tool_schemas() if s.get("function", {}).get("name") == "set_next_wakeup"
|
||||
)
|
||||
self.assertIn("60-3600", schema["description"])
|
||||
self.assertIn("60-3600", schema["parameters"]["properties"]["seconds"]["description"])
|
||||
self.assertEqual(bc._wakeup_min, 60)
|
||||
self.assertEqual(bc._wakeup_max, 3600)
|
||||
|
||||
|
||||
class TestBackgroundConsciousnessCost(unittest.TestCase):
|
||||
def test_unknown_round_cost_stays_nullable_in_durable_thought(self):
|
||||
|
|
|
|||
|
|
@ -8,6 +8,9 @@ ARCHITECTURE.md pins below are the load-bearing rationale-layer guards
|
|||
|
||||
import os
|
||||
import pathlib
|
||||
import re
|
||||
|
||||
from ouroboros.tools.registry import ToolRegistry
|
||||
|
||||
REPO = pathlib.Path(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
|
||||
|
|
@ -199,3 +202,85 @@ def test_architecture_mirror_matches_the_split_axes_contracts():
|
|||
# Both wait_tasks projection enumerations disclose capability_delta.
|
||||
assert "trace_summary, capability_delta when the child has something to disclose" in arch_flat
|
||||
assert "trace_summary, capability_delta when disclosable, duplicate_of" in dev_flat
|
||||
|
||||
|
||||
# Identifiers the prompts legitimately name in backticks that are NOT tools:
|
||||
# parameter names, resource roots, write surfaces, typed outcome/status tokens
|
||||
# and runtime-context keys. A NEW snake_case identifier in a prompt must either
|
||||
# be a real tool (or background-whitelisted tool) or be classified here on
|
||||
# purpose — that classification step is the governance the prompt audit wants:
|
||||
# a phantom or renamed tool name can no longer hide in the runtime prompts
|
||||
# (`advisory_review` and the CONSCIOUSNESS "You can" catalog rotted that way).
|
||||
# Scope: backticked names in all three prompts plus the bare snake_case names
|
||||
# CONSCIOUSNESS.md writes without backticks; BIBLE.md is deliberately out of scope.
|
||||
PROMPT_NON_TOOL_IDENTIFIERS = frozenset({
|
||||
# resource roots / write surfaces / write roots
|
||||
"active_workspace", "artifact_store", "external_workspace", "runtime_data",
|
||||
"skill_payload", "subagent_projects", "system_repo", "task_drive", "user_files",
|
||||
"write_root", "write_surface",
|
||||
# tool parameters named as cross-tool policy
|
||||
"project_id", "project_name", "recommended_use", "review_rebuttal",
|
||||
# typed outcomes / statuses / runtime-context keys
|
||||
"needs_manual_target", "started_uncustodied", "owner_client",
|
||||
# safety policy class names (ouroboros/safety.py TOOL_POLICY values)
|
||||
"check_conditional",
|
||||
})
|
||||
|
||||
|
||||
def _prompt_backticked_identifiers(text: str) -> set:
|
||||
found = set()
|
||||
for token in re.findall(r"`([^`]+)`", text):
|
||||
head = token.split("(", 1)[0]
|
||||
if re.fullmatch(r"[a-z][a-z0-9]*(?:_[a-z0-9]+)+", head):
|
||||
found.add(head)
|
||||
return found
|
||||
|
||||
|
||||
def _prompt_bare_identifiers(text: str) -> set:
|
||||
"""snake_case tokens written WITHOUT backticks (CONSCIOUSNESS.md's style);
|
||||
tokens that are part of a path or filename (`a/b_c`, `x_y.json`) are skipped."""
|
||||
return {
|
||||
m.group(1)
|
||||
for m in re.finditer(r"(?<![\w/.`-])([a-z][a-z0-9]*(?:_[a-z0-9]+)+)(?![\w/.`-])", text)
|
||||
}
|
||||
|
||||
|
||||
def test_prompt_tool_names_resolve_to_registered_tools(tmp_path):
|
||||
"""Every backticked snake_case identifier in the three runtime prompts is
|
||||
either a registered tool (public schema), a background-consciousness tool,
|
||||
or a documented non-tool identifier. Completeness is deliberately NOT
|
||||
required (the schemas are the catalog); this only forbids phantoms and
|
||||
stale spellings, the drift class the prompt audit found in every prompt."""
|
||||
from ouroboros.consciousness import BackgroundConsciousness
|
||||
|
||||
root = pathlib.Path(__file__).resolve().parent.parent
|
||||
registry = ToolRegistry(repo_dir=tmp_path / "repo", drive_root=tmp_path / "data")
|
||||
registered = {schema["function"]["name"] for schema in registry.schemas()}
|
||||
# The background whitelist is not taken on faith: every name in it must be a
|
||||
# registered public tool or a ToolEntry the consciousness module registers
|
||||
# itself (set_next_wakeup and friends), otherwise the whitelist has rotted.
|
||||
consciousness_src = (root / "ouroboros" / "consciousness.py").read_text(encoding="utf-8")
|
||||
bg_private = set(re.findall(r'ToolEntry\("([a-z0-9_]+)"', consciousness_src))
|
||||
stale_whitelist = set(BackgroundConsciousness._BG_TOOL_WHITELIST) - registered - bg_private
|
||||
assert not stale_whitelist, f"_BG_TOOL_WHITELIST names unregistered tools: {sorted(stale_whitelist)}"
|
||||
universe = (
|
||||
registered
|
||||
| set(BackgroundConsciousness._BG_TOOL_WHITELIST)
|
||||
| PROMPT_NON_TOOL_IDENTIFIERS
|
||||
)
|
||||
for rel in ("prompts/SYSTEM.md", "prompts/SAFETY.md", "prompts/CONSCIOUSNESS.md"):
|
||||
text = (root / rel).read_text(encoding="utf-8")
|
||||
unresolved = _prompt_backticked_identifiers(text) - universe
|
||||
assert not unresolved, (
|
||||
f"{rel} names identifiers that are neither registered tools nor "
|
||||
f"classified non-tool identifiers: {sorted(unresolved)}"
|
||||
)
|
||||
# CONSCIOUSNESS.md writes tool names without backticks; its bare snake_case
|
||||
# tokens must resolve the same way (the runtime drift check in
|
||||
# context_health only catches names with known prefixes).
|
||||
bare = _prompt_bare_identifiers((root / "prompts" / "CONSCIOUSNESS.md").read_text(encoding="utf-8"))
|
||||
unresolved_bare = bare - universe
|
||||
assert not unresolved_bare, (
|
||||
f"prompts/CONSCIOUSNESS.md names bare identifiers that are neither registered tools "
|
||||
f"nor classified non-tool identifiers: {sorted(unresolved_bare)}"
|
||||
)
|
||||
|
|
|
|||
|
|
@ -133,7 +133,8 @@ ALLOWED_CASES = [
|
|||
pytest.param("cd /Users/anton/Ouroboros/repo && git diff", id="readonly_cd_repo_diff"),
|
||||
pytest.param("git -C /Users/anton/Ouroboros/repo branch -l", id="readonly_branch_list_repo"),
|
||||
# Read-only forms of the verb-dispatched subcommands stay allowed at a runtime
|
||||
# target too — the SYSTEM.md contract ("read-only git works everywhere"). These
|
||||
# target too — the ARCHITECTURE "Safety and runtime mode" contract ("read-only
|
||||
# shell git is allowed everywhere"). These
|
||||
# were refused before the mode parse: `remote` had no read-only classifier at
|
||||
# all, and `tag -v/--verify` (signature check, writes nothing) sat in the
|
||||
# mutating flag set (SC-7).
|
||||
|
|
|
|||
|
|
@ -473,7 +473,7 @@ def test_http200_body_transient_that_is_not_429_still_blocks(monkeypatch, tmp_pa
|
|||
|
||||
def test_local_fallback_lane_rate_limit_takes_the_audited_fail_open(monkeypatch, tmp_path, _no_backoff):
|
||||
"""Disclosed nuance: the local-FALLBACK lane keeps its documented fail-open contract
|
||||
(SYSTEM.md case (c)) — a genuine 429 there takes the two-attempt fail-open WITH the
|
||||
(ARCHITECTURE "Safety and runtime mode" case (c)) — a genuine 429 there takes the two-attempt fail-open WITH the
|
||||
audit row (a 429 must not be stricter than the RuntimeError beside it), while every
|
||||
other error keeps its unchanged one-attempt 'Local safety runtime unreachable'
|
||||
warning. Both allow; the remote lanes are the ones that block typed."""
|
||||
|
|
|
|||
|
|
@ -42,8 +42,9 @@ def test_legacy_tool_names_are_not_public_schemas(tmp_path):
|
|||
|
||||
# D10: the external coding gateway tool was retired; delegate_start is its
|
||||
# successor. The dead name must be gone from the public schema surface.
|
||||
# (Not folded into LEGACY_PUBLIC_TOOL_NAMES because prompts/SYSTEM.md keeps
|
||||
# one deliberate retirement note that mentions the old name.)
|
||||
# (Not folded into LEGACY_PUBLIC_TOOL_NAMES: docs/DEVELOPMENT.md keeps the
|
||||
# D10 retirement record that mentions the old name; the runtime prompts no
|
||||
# longer do — the prompt audit removed the retirement note from SYSTEM.md.)
|
||||
assert "claude_code_edit" not in names
|
||||
assert registry.get_schema_by_name("claude_code_edit") is None
|
||||
assert registry.execute("claude_code_edit", {}).startswith("⚠️ Unknown tool")
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue