Split the two reference books into verbatim chapters

The books were one physical file each: ARCHITECTURE 2,386 lines and
DEVELOPMENT 3,886, with 13 and 14 `##` sections. `reference_books.py` had
shipped the chaptered reader — membership, authored introductions, exact
physical source views — with the migration still at zero, so every reader
took the legacy monolith branch and the validator had no production caller.

Each `##` section is now one chapter file under `docs/architecture/` or
`docs/development/`. The only new bytes per chapter are its prologue: the old
section title at H1 (numbering text kept, so every `ARCHITECTURE "8. Git
Branching, CI, and Build"` cross-reference still reads) and one authored
introductory paragraph saying what the chapter owns and why it exists.
Everything after that prologue is the old section body byte for byte, with
`###`/`####` levels untouched — so the residue rules keep reading the exact
subsection headings they exempt, and no section title was renamed.

The move is therefore INVERTIBLE, and `tests/test_reference_book_migration.py`
inverts it: drop each chapter's H1 line and its one introduction, re-prefix
`## `, concatenate in membership order, and require the recorded SHA-256 of
the old body — plus, whenever the base commit is reachable, byte equality with
`git show <base>:<path>`. `docs/reference-books-migration.md` is the operator
transfer table: every row a verbatim move, with its line range at the base, its
destination and an empty rename column. It lives directly under `docs/`, so it
is reviewable without becoming a book member.

`_preamble` now takes the FIRST paragraph under the H1 instead of demanding the
only one before the first H2. Most sections open with prose at the level they
already had, so the old rule could only be satisfied by promoting `###` to `##`
or inventing a sub-heading — either of which would rewrite what this migration
relocates verbatim. A source whose H1 is followed straight by a subsection
still has no introduction and is still refused, which is the property that
keeps an overview from quoting body prose as authored orientation.

`.gitattributes` pins `docs/**/*.md` to LF: chapter line ranges, byte spans and
SHA-256s are physical facts that a Windows checkout must not rewrite, and
`full-test` runs on windows-latest for every PR. The two ARCHITECTURE-derived
generated inventories are regenerated, because a chaptered section now carries
its physical provenance note.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
Ouroboros 2026-09-15 18:23:16 +03:00
parent 1eca97ddaf
commit 35b8151819
37 changed files with 6655 additions and 6281 deletions

7
.gitattributes vendored
View file

@ -16,3 +16,10 @@ tests/fixtures/skill_publish_scanner/*.fixture text eol=lf
# Recorded provider wire (SSE / JSON bodies) is byte-exact evidence: no line-ending
# conversion at commit or checkout, so the replay corpus always sees the recorded bytes.
tests/fixtures/llm_wire/** -text -whitespace
# Reference-book chapters are addressed by physical LF line ranges and byte
# spans (`reference_books` source refs, the generated inventories' chapter
# provenance notes, the migration byte proof). Pin every docs markdown file to
# LF so a Windows checkout reads the same lines and digests as the committed
# bytes.
docs/**/*.md text eol=lf

File diff suppressed because one or more lines are too long

File diff suppressed because it is too large Load diff

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,38 @@
# 2. Startup / Onboarding Flow
Packaged startup is an ordered ownership transaction, and this chapter records that order: what the launcher must prove before the server starts, why onboarding runs after the gateway rather than before it, and how one completion transaction persists settings, next-boot runtime mode and the install-time preset together. It exists because the failure modes here are quiet ones — a read that creates `settings.json` disqualifies the fresh-install latches, and a half-saved completion leaves an install with neither a working provider nor a wizard.
Packaged startup is an ordered ownership transaction. The launcher prepares the platform UI runtime (or runs the Linux browser-mode probe), acquires the single-instance lock, verifies Git, and validates and bootstraps the embedded managed-repo seed — server preconditions, so they precede it. It then removes only identity-proven stale server state and ports, starts the lifecycle thread, and waits on `/api/health` at the authoritative port from `data/state/server_port`; only then is first-run onboarding presented against that live server, before the pywebview shell or browser presentation opens. The server starts the gateway first and the supervisor/worker pool only when provider configuration is structurally sufficient.
Onboarding runs after the gateway because connecting an agent subscription is a live `/api/*` conversation, not a form field — and a gateway without a supervisor is exactly what the readiness predicate produces, so no second server, mode, or onboarding state machine exists. `ouroboros/launcher_onboarding.py` owns the presentation (readiness decision, setup window, window-lifecycle bridge); when completion reports a boot-pinned value changed, the launcher recycles the managed server rather than counting the exit as a crash. Neither launcher nor server boot normalization may CREATE `settings.json` (the install-time latches below are gated on its absence).
`has_startup_ready_provider()` is a structural gate, not a network, credential, entitlement, model, or local-process probe: any non-empty recognized remote configuration (including DeepSeek) or active task-capable local routing flag passes (key list: `server_runtime.has_startup_ready_provider`; `USE_LOCAL_HEAVY` is legacy migration input only, `LOCAL_MODEL_SOURCE` alone insufficient). When the gate is false the server marks startup complete without workers so the web UI serves the blocking onboarding overlay; a later successful settings save hot-starts the supervisor.
Every host renders one served page: `GET /onboarding` returns `onboarding_template.html` with the `settings_setup_contract` bootstrap injected, linking wizard CSS and `web/modules/onboarding_wizard.js` as static assets so steps import the same modules as the rest of the UI — an inlined `srcdoc` string cannot. The desktop setup window opens that URL, the blocking overlay frames it, a browser owner opens it directly; `GET /api/onboarding` is the readiness probe (204 once the gate passes, otherwise the page). Steps: Accounts, Models, Review, Budget, Summary; context mode remains outside this wizard. Accounts includes optional agent connections and mounts the shared login cards in `full` mode because `compact` omits the paste-code entry a Claude login needs when its localhost callback cannot complete; its account facts come from the shared Claudexor status store and become the completion payload's `subscriptionsConnected` declaration — a request to look at the daemon, never an authority.
Completion is one HTTP conversation on every host — `POST /api/onboarding/complete` (`gateway/onboarding.py`; its docstring carries the numbered step order) — so providers and runtime mode cannot be left half-saved. Its single write, re-proved under the settings lock, persists settings, next-boot runtime mode, the fresh-install `OUROBOROS_SAFETY_MODE=light` default no other endpoint may author, the one-shot preset marker, and the durable completion fact; only then does the supervisor start. Install time means three proofs together (`gateway/onboarding.py::preset_eligible`): onboarding never completed here, no preset generation applied, and no `settings.json` yet — three, because "no startup-ready provider" alone is a state an old install reaches whenever its key stops working. `GET /api/onboarding` normalizes what it displays but never persists: a read that created `settings.json` would silently disqualify the install-time latches. A completion failing AFTER bytes reach disk reports that it saved plus the failing stage; a completion whose body outlives the shared writer bound (`settings.py::_run_settings_writer`, the seam every settings writer runs through) answers 503 `settings_save_timeout` with `saved: null` and the wizard offers "Check status" (a re-read of the readiness probe) instead of a blind re-save; a 2xx is not a completion — only the exact success envelope is, because the saved runtime mode and restart receipt live in that body. On a boot-pinned change the framed overlay shows its restart card and a plain browser tab shows the saved-but-restart-required screen (the launcher recycle: above). The overlay sandbox grants popup permission because the agent sign-in link is the step's primary action.
`ouroboros/subscription_install_presets.py` is a pure compiler for sibling task actors, reviewer slots and model assignments, not a reviewer/subagent matrix; its docstring carries the emission rules, the ratified Claude/Codex/Cursor reviewer subset, and the Agy-neutrality contract. Normalized API/local settings alone suffice for API-only and local-only installs; a declared subscription makes `gateway/onboarding.py` read a live account snapshot and the declared raw-model source catalog before the settings transaction — the authority for supported harnesses, exact model ids (Agy's automatic row: `gemini-3.8-flash-high`), and enabled accounts. Durable seat facts (kind/enabled/verified), not the hour's quota reading, decide the once-only preset: a spent Claude window at save time must not yield a Codex-only preset permanently. Task-actor rows stay unpinned and grow linearly, never as a powerset; on the fresh-install path reviewer slots are `subagent_id` references into the roster the preset ships (a seat matching no task actor mints a `review-<harness>` roster row), while an owner-configured roster stays validate-only. Missing required exact discovery is a typed pre-write refusal; no partial preset persists.
`POST /api/onboarding/subagents/preview` runs the same compiler against the open draft without persisting; the editable result becomes `OUROBOROS_SUBAGENTS`, and an owner-edited draft is validated and preserved, not regenerated. Completion persists actors, reviewer disposition, preset receipt, and completion facts through the one-write document-lock/fingerprint boundary; network discovery happens before the lock. A daemon failure keeps the wizard open with an explicit finish-without-agent-defaults path; a read or failed refresh never rewrites saved intent.
Validation is structural: at least one exposed remote configuration, selected managed model source, or local model source; local-only setup routes at least one active lane locally; Main required, while Light, Vision, Consciousness, and Fallback keep inheritance/empty semantics and Heavy is readable only for bounded migration into an explicit API actor; enforcement and runtime mode are closed enums, budgets finite and positive, the MiniMax region closed, a Hugging Face local source needs a filename. Credential length is checked only on fields changed in the payload: rejecting an unchanged short legacy value would discard the whole form, including its own repair.
Provider readiness and provider defaulting are separate. With no OpenRouter, legacy OpenAI base, or OpenAI-compatible endpoint, exactly one registered direct provider receives provider-prefixed defaults and migration of untouched shipped/legacy slot values; OpenAI, Anthropic, Cloud.ru, GigaChat, MiniMax, and DeepSeek each use their own registered defaults. Multiple direct providers stay owner-editable, and OpenRouter keeps router-style routing. An arbitrary OpenAI-compatible endpoint gets no guessed model ids — compatible servers have no universal safe name; the owner selects explicit `openai-compatible::...` routes. A local-source install with no remote provider clears only untouched shipped remote Light/Fallback values that would be unreachable; owner-authored values and explicitly local slots are preserved. This is migration of defaults, not a model allowlist, and never proof the local server is running.
`scripts/build_repo_bundle.py` creates the packaged seed only from a clean named checkout, writes a git bundle of that commit/tags, and records schema, version, source SHA, release tag, bundle hash, and managed branch/remote metadata (the release-tag check itself: §8). The launcher validates the manifest fields and the bundle SHA-256 but no per-file member set: clone-time Git verification proves the manifest source object exists and checked-out HEAD equals it.
`ensure_managed_repo()` owns packaged checkout bootstrap: a first install clones the bundle into a temporary checkout, verifies and configures the pinned source SHA and managed branches/remotes, then moves it into `repo/` (an existing legacy non-git directory is archived first). Once a managed git checkout exists, a changed application manifest does not archive or replace its working tree: bootstrap atomically refreshes managed metadata and the official `managed` remote in place, preserving the local branch tip and owner edits. Ordinary restart performs no network fetch — network movement to an approved official SHA belongs to the pinned managed-update path in `supervisor/git_ops.checkout_and_reset`, not to bootstrap — and `origin` remains optional personal persistence, not the official update authority.
Bootstrap also creates the initial world profile when absent and seeds launcher-owned native skills without resurrecting an intentionally deleted seed. `launcher_bootstrap._per_skill_version_resync` replaces existing marker-owned payloads when their manifest versions differ in either direction. Equal versions retain installed bytes and lifecycle state; the shared content hash reports payload drift or an unavailable comparison without blocking startup. Bundled payload changes increment the skill manifest version independently of the application version. Dependency installation runs only when checkout/bootstrap metadata changed and its boolean result reaches the launcher; after exit code 42 the launcher refreshes bundle metadata and synchronizes dependencies before starting the edited body. A failed install gets one visible retry after five seconds; a second failure stays in logs while startup continues under the five-crashes-in-120-seconds fuse — silently losing a pip failure makes a later ImportError inexplicable, while refusing every offline restart would break a checkout whose requirements are already present.
Managed supervisor bootstrap is the sole owner of destructive dirty-tree recovery. Before any reset/clean, `supervisor.git_ops` writes a merge-aware rescue directory: porcelain status, `changes.diff` captured as raw bytes with a hardened argv/environment (`supervisor/git_ops.py`) because it is the only carrier of a resolution stash cannot capture, a stash-created rescue ref when possible, copied untracked files with completeness metadata, unpushed-commit evidence, `rescue_meta.json`. An incomplete snapshot blocks `rescue_and_reset` — not permission to discard what could not be captured. Normal managed bootstrap then cleans back to the local branch's own HEAD, not to `managed/<branch>`.
Managed update is the second user of that machinery, with the opposite failure policy. Every destructive rollback shares one choke point in `rollback_managed_update`, and both it and boot-resume re-materialization take a FRESH rescue before the first destructive command, because the pre-update snapshot predates the merge and holds none of the resolver's work. This hook is fail-open: a rescue that cannot be taken never blocks the rollback (logged and disclosed), one durable `supervisor.jsonl` line is written at capture time, before destruction, so the record survives a crash mid-teardown, and a `git status` that cannot answer counts as dirty. The rescue is deliberately NOT linked to an active evolution transaction — that would flip the campaign's cycle to abandoned for an unrelated reason. The update transaction persists a pointer to what was rescued: a replayed rollback does not duplicate a snapshot, a retry re-rescues the tree it actually finds, and the resolver's objective names the latest rescue directory — re-materialization re-creates a dirty tree WITHOUT replaying the rescued edits and must never be read as their return. Rescue Git processes are bounded by `OUROBOROS_RESCUE_GIT_TIMEOUT_SEC` with process-tree termination on timeout.
An active Evolution transaction or managed-update merge uses `rescue_and_block`: evidence links to the transaction and the tree is left intact, pausing Evolution rather than erasing partially resolved work; with no such owner, startup uses `rescue_and_reset`. Source/local-development startup skips the managed checkout/reset path (dependency sync plus import test only). Worker startup checks are diagnostic and warning-only: launcher-management environment variables propagate into worker, review, and test subprocesses, so per-constructor auto-rescue would let an incidental child steal or clean another actor's in-progress edits.
`server.py` establishes `OUROBOROS_AGENT_PYTHON` from its actual interpreter immediately after binding the repo import root, before workers or review subprocesses start. Hermetic commit/review preflight uses that handle (then `sys.executable`, then `python3`) so tests run in the environment containing Ouroboros dependencies; plugin verification is part of that preflight, not a launcher claim that every interpreter was live-probed.
User process tools have a separate surface-aware resolver (`ouroboros/process_interpreters.py`, which carries the priority ladder): for exact unversioned `python`/`python3` on the four public process launch surfaces, registry pre-dispatch resolves once before deterministic guards so guard and handler see byte-identical argv. Absolute or versioned interpreters, shell bodies, and non-Python commands stay literal at pre-dispatch (the post-gates Node ladder: §1 `process_interpreters.py` row); resolution emits secret-free provenance, never silently installs dependencies, and fails closed only when a system-owned interpreter cannot be proven. `run_script` accepts an installed executable whose interface takes a script filename followed by literal arguments, without an interpreter-name enum; compiler subcommands use `run_command` or the skill runtime. This is an executable-form contract, not a promise that any language command accepts a script directly.

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,165 @@
# 4. Server API Endpoints
This chapter is the endpoint registry: every mounted browser, CLI and Host Service route beside the handler that owns it, plus the non-loopback authentication gate, the file-root confinement rule and the WebSocket protocol the browser actually speaks. It exists as a test-checked mirror of the executable route collector, so a route added, renamed or removed in code cannot quietly disappear from the map.
If `OUROBOROS_NETWORK_PASSWORD` is configured, non-loopback HTTP and WebSocket access requires authentication; loopback clients bypass the gate, and `/api/health` plus the middleware-owned login/logout paths stay reachable. Browser sessions use a server-keyed, expiring HttpOnly HMAC cookie; `Secure` is set only under TLS so a plain-HTTP LAN session does not enter a login loop. An unauthenticated WebSocket is closed with code 4401 before `ws_endpoint` accepts it. With no configured password, non-loopback access remains open by explicit operator choice.
The executable browser/CLI route SSOT is `ouroboros/gateway/router.py`; file-browser routes are contributed by `gateway/files.py::file_browser_routes()`. `gateway/contracts.py` is the frozen descriptive envelope and endpoint index mirrored by `web/modules/api_types.js` and parity tests; its `TypedDict` classes perform no runtime JSON validation. The loopback Host Service is a separate token-authenticated app assembled by `gateway/host_service.py::create_host_service_app`, not another public owner API.
Every `/api/files/*` operation resolves its requested path and refuses the operation when that resolution leaves the configured file root. In-root symlinks remain usable; out-of-root symlinks may be listed with `is_symlink: true` but cannot be read, written, downloaded, deleted, or traversed. The backend check is authoritative regardless of browser path presentation.
| Method | Path | Handler |
|---|---|---|
| GET | `/` | `server.index_page` |
| GET | `/api/health` | `gateway.state.api_health` |
| GET | `/api/state` | `gateway.state.api_state` |
| GET | `/api/extensions` | `gateway.extensions.api_extensions_index` (unique rows additionally carry `content_hash`, `published` (validated receipt object or null), `published_malformed`; identity-collision rows carry `identity_collision: true` and omit the receipt fields) |
| POST | `/api/skills/{skill}/publish-preflight` | `gateway.skill_publish.api_skill_publish_preflight` |
| GET | `/api/extensions/{skill}/manifest` | `gateway.extensions.api_extension_manifest` |
| GET | `/api/extensions/{skill}/module/{entry:path}` | `gateway.extensions.api_extension_module` (live-registration authorization, reviewed `.js`/`.mjs` siblings from captured texts, `Access-Control-Allow-Origin: *` on every answer) |
| GET | `/api/widgets` | `gateway.widgets.api_widgets` (passive projection of the loader's live UI tabs via `extension_loader.live_widget_projection`; `Cache-Control: no-store`) |
| GET | `/api/extensions/{skill}/settings_section` | `gateway.extensions.api_extension_settings_section` |
| ANY | `/api/extensions/{skill}/{rest:path}` | `gateway.extensions.api_extension_dispatch` |
| GET | `/api/skills/daemons` | `gateway.extensions.api_skill_daemons` |
| POST | `/api/skills/{skill}/toggle` | `gateway.extensions.api_skill_toggle` |
| POST | `/api/skills/{skill}/delete` | `gateway.extensions.api_skill_delete` |
| GET | `/api/skills/lifecycle-queue` | `gateway.extensions.api_skill_lifecycle_queue` |
| POST | `/api/skills/{skill}/review` | `gateway.extensions.api_skill_review` |
| GET | `/api/skills/{skill}/review-history/{job_id}` | `gateway.extensions.api_skill_review_history_detail` (bounded lazy detail from a fixed tail window of `review_history.jsonl`; missing job 404, outside-window honestly unavailable; slot/attempt usage joins from the physical-attempt ledger) |
| POST | `/api/owner/skills/{skill}/attest-review` | `gateway.extensions.api_owner_skill_attest_review` (OWNER-ONLY skip of the LLM review; the deterministic preflight floor still runs, 409 on failure; routes through `run_skill_review_lifecycle` for the post-pass reconcile) |
| POST | `/api/skills/{skill}/grants` | `gateway.extensions.api_skill_grants` |
| POST | `/api/skills/{skill}/reconcile` | `gateway.extensions.api_skill_reconcile` |
| GET | `/api/marketplace/clawhub/search` | `gateway.marketplace.api_marketplace_search` |
| GET | `/api/marketplace/clawhub/installed` | `gateway.marketplace.api_marketplace_installed` |
| GET | `/api/marketplace/clawhub/info/{slug:path}` | `gateway.marketplace.api_marketplace_info` |
| GET | `/api/marketplace/clawhub/preview/{slug:path}` | `gateway.marketplace.api_marketplace_preview` |
| POST | `/api/marketplace/clawhub/install` | `gateway.marketplace.api_marketplace_install` |
| POST | `/api/marketplace/clawhub/update/{name}` | `gateway.marketplace.api_marketplace_update` |
| POST | `/api/marketplace/clawhub/uninstall/{name}` | `gateway.marketplace.api_marketplace_uninstall` |
| GET | `/api/marketplace/ouroboroshub/catalog` | `gateway.marketplace.api_ouroboroshub_catalog` |
| GET | `/api/marketplace/ouroboroshub/installed` | `gateway.marketplace.api_ouroboroshub_installed` |
| GET | `/api/marketplace/ouroboroshub/preview/{slug:path}` | `gateway.marketplace.api_ouroboroshub_preview` |
| POST | `/api/marketplace/ouroboroshub/install` | `gateway.marketplace.api_ouroboroshub_install` (also the adopt transport: `{adopt: true, expected_content_hash}` replaces an external same-name occupant with the sha256-verified catalog payload; adopt forces `auto_review`, conflicts with `overwrite`, typed 400/409/502 codes ride the lifecycle payload) |
| POST | `/api/marketplace/ouroboroshub/update/{name}` | `gateway.marketplace.api_ouroboroshub_update` |
| POST | `/api/marketplace/ouroboroshub/uninstall/{name}` | `gateway.marketplace.api_ouroboroshub_uninstall` |
| POST | `/api/marketplace/ouroboroshub/publication/{name}/clear` | `gateway.marketplace.api_ouroboroshub_clear_publication` (compares the displayed receipt and clears only local waiting state) |
| GET | `/api/files/list` | `gateway.files.api_files_list` |
| GET | `/api/files/read` | `gateway.files.api_files_read` |
| GET | `/api/files/content` | `gateway.files.api_files_content` |
| GET | `/api/files/download` | `gateway.files.api_files_download` |
| POST | `/api/files/upload` | `gateway.files.api_files_upload` |
| POST | `/api/files/mkdir` | `gateway.files.api_files_mkdir` |
| POST | `/api/files/write` | `gateway.files.api_files_write` |
| POST | `/api/files/delete` | `gateway.files.api_files_delete` |
| POST | `/api/files/transfer` | `gateway.files.api_files_transfer` |
| GET | `/onboarding` | `gateway.onboarding_host.onboarding_page` |
| GET | `/api/onboarding` | `gateway.settings.api_onboarding` |
| POST | `/api/onboarding/complete` | `gateway.onboarding.api_onboarding_complete` |
| POST | `/api/onboarding/subagents/preview` | `gateway.onboarding.api_onboarding_subagents_preview` |
| GET | `/api/settings` | `gateway.settings.api_settings_get` |
| POST | `/api/settings` | `gateway.settings.api_settings_post` |
| GET | `/api/reviewer-slots` | `gateway.settings.api_reviewer_slots` |
| GET | `/api/claudexor/status` | `gateway.claudexor_accounts.api_claudexor_status` |
| POST | `/api/claudexor/quota/refresh` | `gateway.claudexor_quota.api_claudexor_quota_refresh` |
| POST | `/api/claudexor/wake` | `gateway.claudexor_accounts.api_claudexor_wake` |
| POST | `/api/claudexor/login` | `gateway.claudexor_accounts.api_claudexor_login` |
| GET | `/api/claudexor/login/{job_id}` | `gateway.claudexor_accounts.api_claudexor_login_job` |
| DELETE | `/api/claudexor/login/{job_id}` | `gateway.claudexor_accounts.api_claudexor_login_job` |
| POST | `/api/claudexor/login/{job_id}/input` | `gateway.claudexor_accounts.api_claudexor_login_job` |
| POST | `/api/claudexor/login/{job_id}/reconcile` | `gateway.claudexor_accounts.api_claudexor_login_job_reconcile` |
| DELETE | `/api/claudexor/credential-profiles/{harness}/{profile_id}` | `gateway.claudexor_accounts.api_claudexor_credential_profile` |
| PATCH | `/api/claudexor/credential-profiles/{harness}/{profile_id}` | `gateway.claudexor_accounts.api_claudexor_credential_profile` |
| POST | `/api/owner/runtime-mode` | `gateway.settings.api_owner_runtime_mode` |
| POST | `/api/owner/auto-grant` | `gateway.settings.api_owner_auto_grant` |
| POST | `/api/owner/context-mode` | `gateway.settings.api_owner_context_mode` |
| POST | `/api/owner/safety-mode` | `gateway.settings.api_owner_safety_mode` |
| POST | `/api/owner/skills/{skill}/presence-runtime` | `gateway.presence_settings.api_owner_skill_presence_runtime` |
| POST | `/api/owner/capability-ack` | `gateway.settings.api_acknowledge_capability` |
| GET | `/api/ui/preferences` | `gateway.ui_preferences.api_ui_preferences_get` |
| POST | `/api/ui/preferences` | `gateway.ui_preferences.api_ui_preferences_post` |
| GET | `/api/model-catalog` | `gateway.models.api_model_catalog` |
| POST | `/api/openai-compatible/models` | `gateway.models.api_openai_compatible_models` |
| POST | `/api/providers/test` | `gateway.models.api_provider_test` |
| POST | `/api/tasks` | `gateway.tasks.api_tasks_create` |
| GET | `/api/tasks` | `gateway.tasks.api_tasks_list` |
| GET | `/api/tasks/{task_id}` | `gateway.tasks.api_task_get` |
| GET | `/api/tasks/{task_id}/events` | `gateway.tasks.api_task_events` (legacy integer rank) |
| POST | `/api/tasks/{task_id}/events` | `gateway.tasks.api_task_events` (read-only v2 cursor) |
| GET | `/api/tasks/{task_id}/artifacts/{name}` | `gateway.tasks.api_task_artifact` |
| POST | `/api/tasks/{task_id}/cancel` | `gateway.tasks.api_task_cancel` |
| POST | `/api/tasks/{task_id}/hurry` | `gateway.tasks.api_task_hurry` |
| POST | `/api/tasks/{task_id}/resume` | `gateway.tasks.api_task_resume` |
| POST | `/api/decisions` | `gateway.tasks.api_decision_answer` |
| GET | `/api/schedules` | `gateway.schedules.api_schedules_list` |
| POST | `/api/schedules` | `gateway.schedules.api_schedules_upsert` |
| DELETE | `/api/schedules/{schedule_id}` | `gateway.schedules.api_schedules_delete` |
| POST | `/api/command` | `gateway.control.api_command` |
| POST | `/api/reset` | `gateway.control.api_reset` |
| GET | `/api/git/log` | `gateway.control.api_git_log` |
| POST | `/api/git/rollback` | `gateway.control.api_git_rollback` |
| POST | `/api/git/promote` | `gateway.control.api_git_promote` |
| GET | `/api/update/status` | `gateway.control.api_update_status` |
| POST | `/api/update/check` | `gateway.control.api_update_check` |
| POST | `/api/update/preflight` | `gateway.control.api_update_preflight` |
| POST | `/api/update/apply` | `gateway.control.api_update_apply` |
| GET | `/api/cost-breakdown` | `gateway.history.make_cost_breakdown_endpoint` (the router's import path; the factory body and `_ACCOUNTING_SUMMARY_FIELDS` live in `gateway.cost_breakdown`) |
| GET | `/api/evolution-data` | `gateway.control.api_evolution_data` |
| GET | `/api/projects` | `gateway.projects.api_projects_list` |
| POST | `/api/projects` | `gateway.projects.api_projects_create` |
| POST | `/api/projects/from-task` | `gateway.projects.api_project_from_task` |
| POST | `/api/projects/{project_id}/update` | `gateway.projects.api_project_update` |
| POST | `/api/projects/{project_id}/delete` | `gateway.projects.api_project_delete` |
| GET | `/api/fs/dirs` | `gateway.projects.api_fs_dirs` |
| GET | `/api/chat/history` | `gateway.history.make_chat_history_endpoint` |
| GET | `/api/logs/{name}` | `gateway.logs.api_logs_tail` |
| POST | `/api/chat/upload` | `gateway.files.api_chat_upload` |
| DELETE | `/api/chat/upload` | `gateway.files.api_chat_upload_delete` |
| POST | `/api/local-model/start` | `gateway.models.api_local_model_start` |
| POST | `/api/local-model/stop` | `gateway.models.api_local_model_stop` |
| GET | `/api/local-model/status` | `gateway.models.api_local_model_status` |
| POST | `/api/local-model/test` | `gateway.models.api_local_model_test` |
| POST | `/api/local-model/install-runtime` | `gateway.models.api_local_model_install_runtime` |
| GET | `/api/mcp/status` | `gateway.mcp.api_mcp_status` |
| POST | `/api/mcp/refresh` | `gateway.mcp.api_mcp_refresh` |
| POST | `/api/mcp/test` | `gateway.mcp.api_mcp_test` |
| WS | `/ws` | `gateway.ws.ws_endpoint` |
| STATIC | `/static/*` | `server.NoCacheStaticFiles` |
| GET | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/identity` | `gateway.host_service._api_identity` |
| GET | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/tools/schemas` | `gateway.host_service._api_tool_schemas` |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/allocate-internal` | `gateway.host_service._api_allocate_internal` |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/inject` | `gateway.host_service._api_chat_inject` |
| GET | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/operations/{operation_ref:path}` | `gateway.host_service._api_chat_operation` (the calling skill's own accepted message: pending, running with its task or turn, the durable answer, or the terminal task status) |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/cancel` | `gateway.host_service._api_chat_cancel` (the existing cancellation owner on work that message started; a typed outcome, never a cancellation that did not happen) |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/decision` | `gateway.host_service._api_chat_decision` (the `task_decision.answer_decision` ingress relayed for a transport skill) |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/presence/turn` | `gateway.host_service._api_presence_turn` |
| GET | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/presence/work/{work_ref}` | `gateway.host_service._api_presence_work` |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/ui/ws-message` | `gateway.host_service._api_ws_message` |
| WS | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/events` | `gateway.host_service._ws_events` |
Rationale: `server.py` owns process startup/lifespan/static mounting, while `gateway/*` owns browser-facing HTTP/WS contracts; this keeps UI and runtime coupling explicit and testable.
### WebSocket protocol
`/ws` is the live browser delivery channel, not a durable state owner: queue, task, Project, review, skill, settings, cost, and update modules persist their own truth, and REST/history endpoints reconstruct it after reload or disconnection. `gateway/contracts.py` describes the frozen envelope shapes and message-type index for Python/JavaScript parity; `gateway/ws.py` performs the actual transport checks — incoming text must decode to a JSON object, extension types must parse as an owned namespace, and built-in `chat` or `command` frames must carry a non-empty payload before entering the message bridge. The authentication middleware above governs socket admission; the public socket never receives the Host Service token or the owned Claudexor daemon token.
The browser constructs one socket for the whole SPA. Feature modules subscribe before connection, and the initial complete Project chat-id set is fetched before the first open so an early Project frame cannot be mistaken for Main traffic. `ws.on(type, listener)` stores listeners in insertion-ordered sets and returns a disposer; emission uses a listener snapshot, so a listener added during dispatch does not receive the current frame and disposing one listener cannot skip its neighbor. Every decoded frame first reaches the generic `message` event and then its type-specific event, which lets Widgets consume reviewed namespaced events without duplicating the socket.
A browser `chat` frame contains the owner text and may add `sender_session_id`, `client_message_id`, `force_plan`, uploaded attachment references, `chat_id`, `project_id`, and `client_surface` — raw sending-surface observables measured at SEND time because the pywebview bridge appears asynchronously after load. The gateway normalizes that payload through `client_surface.normalize_client_surface`, stamps host `received_at`, and persists it on the canonical inbound row — distinct from the `transport` dict (transport is chat-scoped reply routing; the surface fact is per-message provenance). The fact is assembled at its PRODUCER, never inferred at render (the per-producer stamp catalog and closed-key bound: `ouroboros/client_surface.py`); synthetic A2A chats stamp no owner surface (machine traffic never wears one); machine producers stamp nothing (`client_surface` is a reserved schedule-template key rejected at admission); promotion/steering CARRY the originating owner turn's fact. The loop injects a surface note only when sending-surface identity changes within an attempt (viewport excluded — a resize is not a device change). Absence is an honest gap. The client generates a message id when absent and uses it to reconcile its pending bubble, the echoed canonical user row, routing annotations, and mailbox retries; a successful browser `send()` means only that the current socket accepted the frame, not that a task was durably admitted.
Ordinary frames sent while disconnected enter a process-local queue capped at 100 entries (oldest dropped), flushed in order after reconnect and lost on page reload — not a second durable outbox. Attachment messages deliberately set `queue:false`: uploads occur immediately before send, so retaining only the socket frame would leave unowned temporary files; on socket loss Chat refuses the message, cleans uploaded temporaries best-effort, and retains the staged files for explicit retry.
For a chat frame, the gateway validates uploaded filenames as basenames confined under the upload root, exposes the first eligible bounded image as native image content, forwards the complete validated attachment set as task-staging metadata, and calls the local message bridge with the exact thread, Project, sender-session, client-message, and planning facts. The web owner identity is fixed: `chat_id` selects a thread and cannot mint an external owner identity. If the bridge is not initialized, the socket returns a visible assistant warning rather than accepting the message silently.
A built-in `command` frame carries a slash command and enters the same bridge with rebroadcast disabled; runtime command routing, owner authorization, queue authority, and typed outcomes remain outside the socket module. Main header controls therefore reuse the ordinary command contract for Restart, Panic, review, evolution, and background consciousness; Panic is sent only after the shared dialog returns the strict confirmed boolean. The socket does not infer intent from command-looking prose.
Built-in outbound envelopes: `chat`, `photo`, `video`, `document`, `typing`, `log`, `heartbeat`, `extension_lifecycle`, `message_annotation`, `projects_changed`, `task_named`, and `update_status_ready`. Chat progress may carry task lineage, role, requested/effective model lane, delegated route, terminal execution evidence, review projection, cancellation eligibility, outcome axes, artifact references, and nullable cost/finality fields — additive presentation facts; consumers must not infer a missing execution receipt, cost, or task result from the absence of one optional field.
Thread routing is explicit. Project chat, typing, media, and log frames carry `chat_id`; a Project panel consumes its own thread, while Main admits required Project question pointers and the two host-stamped Project lifecycle rows (`project_started`, `project_completion_summary`). `projects_changed` carries a new chat id so every tab can extend its fan-out set before fetching the registry; when even that ordering loses the race, the server-stamped `project_thread` marker on the frame itself keeps Main from adopting it — set once at the message-bus broadcast choke from the registry (a membership lens, never a numeric range, so external transport ids such as Telegram stay unstamped) and enforced by Main's fan-out gate (`chat_activity.mainThreadAccepts`). Task-scoped LOG events acquire their final chat id at supervisor ingress: worker diagnostics carry only their own `task_id`, and `supervisor/log_addressing.py::address_task_event` stamps the audience from host-attested truth (the precedence chain lives in its docstring; an explicit event chat_id of 0 is the hidden partition, `HIDDEN_CHAT_ID`, never "missing"); direct turns carry their chat BY VALUE, stamped at the producer, because the registry entry dies with the turn while queued events drain later. Addressing is honest — an A2A row keeps its true audience, suppressed only at the broadcast choke (`push_log`) so machine traffic never reaches the browser; the same addressing runs in the server-process append sink and at every supervisor handler owning a suppressed type's explicit push, and a genuinely unaddressable event keeps the legacy chat-0 frame. `message_annotation` updates one canonical owner message without creating another bubble; `task_named` updates a card only where that task already exists. Media/document consumers validate MIME, base64, and download-route shapes before building browser URLs.
Extension WebSocket traffic is structurally namespaced by `extension_loader.extension_surface_name()` so an extension cannot shadow a built-in type. On each incoming extension frame the gateway resolves the owning skill and reconciles whether its extension is still desired, reviewed, granted, enabled, and live. A missing or failed handler returns a visible log frame. Out-of-process handlers execute in their extension child off the event loop; in-process handlers first record the required execution/cost disclosure. A non-`None` result returns as `<request-type>.reply`; exceptions become typed error log frames rather than terminating the socket loop.
Server broadcasts snapshot the connected-client list and send to all clients concurrently, so one slow or half-open browser cannot head-of-line-block delivery to every other tab. Failed sends remove only the dead clients and append a durable `broadcast_partial_failure` event; the original domain event stays owned by its durable producer. Restart shutdown closes remaining clients best-effort with code 1012 so they enter the ordinary reconnect path.
The browser reconnects with bounded exponential delay, shows the reconnect overlay, and resets the delay after a successful open; a watchdog closes an apparently open connection after 45 seconds without any inbound frame, so heartbeat traffic proves stream liveness rather than task progress. One served-SHA decision (`ws.js decide()`: keep / reload-changed / reload-unknown) governs both recovery paths so a transient drop cannot destroy in-page state: a changed or no-longer-provable SHA reloads (a restarted server must not keep old JavaScript or CSS alive in PyWebView), an unchanged SHA keeps the page and its queued outbound messages, and an unversioned `/api/state` stays on keep when no non-empty SHA was ever remembered — the owner-selected default under uncertainty, accepting possibly-stale assets as the disclosed tradeoff. While the socket stays down, delayed recovery probes consult `/api/state` without adopting the served SHA; a 200 whose body is not a parseable object counts as a failed probe, probes are single-flight and generation-scoped per disconnect episode, and after several consecutive healthy probes with the socket still down, one forced reload per episode remains as the fuse for a stale browser runtime.
Each Chat instance handles `open` by resynchronizing archive-aware durable history and `close` by withdrawing online/accounting presentation; reconnect deduplication covers overlap between live frames and REST replay, Logs merges the same way, and large history parsing runs off the server event loop. Delivery is live plus replay, not a promise that every transient frame is persisted: durable chat rows, task results, queue snapshots, Project revisions, review ledgers, cost ledgers, and lifecycle state remain the recovery authorities.

View file

@ -0,0 +1,53 @@
# 5. Supervisor Loop
This chapter owns the single scheduler for pooled work: what a healthy tick does, what the queue holds and what its durable snapshot may restore, how a task is addressed and named at admission, how owner waits lend capacity, and the intent-then-custody skeleton every cancellation follows. It exists because these invariants decide whether a stopped task ends honestly or leaves a ghost, and none of them can be reconstructed from any single module's code.
`server.py::_run_supervisor()` is the single scheduler for pooled tasks. A healthy tick publishes liveness, rotates the runtime logs, checks worker health, drains worker, direct-chat, and consciousness events, accepts owner bridge input, enforces deadlines and schedules, runs throttled reconciliation and evolution admission, assigns eligible work, and persists `state/queue_snapshot.json`. Bridge intake precedes timeout, maintenance, evolution, and assignment work so a slow control-plane step cannot make a new owner message invisible. Three consecutive loop failures clear supervisor readiness, stop its watchdog generation, and notify the owner instead of leaving a healthy-looking server that no longer assigns work; a failure raised while a shutdown or restart is already in progress (the lifespan teardown sets a process-local stop event first and joins the loop for a bounded window before workers, bridge and event bus go down) is not a crash — the loop exits quietly, without the counter, the error, or the alarm — and the crash backoff waits on that stop event so a shutdown is never held by it.
`PENDING` and `RUNNING`, guarded by `supervisor.queue._queue_lock`, are the live task-lifecycle authority. Admission reserves identity before project, workspace, attachment, or routing side effects can create a duplicate; refuses a disabled pool, duplicate task, project deletion, accepted or sealed root, or exhausted root budget; attaches the task contract; and preserves stable priority order. Assignment runs against the same locked state and skips reaping slots, budget-paused work, closed project roots, conflicting project writers, and tasks exceeding the root's subagent capacity or depth reservation; evolution tasks are dropped there when `evolution_block_reason()` is set (Light runtime mode, `supervisor/workers.py`). That is the last of three evolution-only runtime-mode fences: owner and post-task entry points refuse a campaign start, `enqueue_evolution_task_if_needed()` independently pauses and disables a carried campaign before queueing it, and assignment drops what still slipped through (`supervisor/evolution_lifecycle.py`). Generic `supervisor.queue.enqueue_task()` has no runtime-mode predicate at all. Configured worker count is therefore not available capacity: the truthful value is the currently assignable idle count after custody, reaping, and admission fences.
A headless task is ADDRESSED when it is admitted, not when it is displayed (`log_addressing.ingress_chat_id`). A registered project's run has exactly ONE destination: an explicit `chat_id` may only agree with that thread, and any other value — the hidden partition included — is refused with a typed 400 rather than honoured or silently overridden, because a run addressed away from its room is the one shape that puts a card in Main whose project holds none of its work. Without a Project, ordinary API tasks default to `HIDDEN_CHAT_ID` (0). The confirmed browser Publish flow explicitly carries `source="web"` and `WEB_UI_CHAT_ID` to request Main; source is caller-declared addressing on the existing owner API, not a new authentication proof. Other non-Project conversation addresses remain refused. A run scoped to a REGISTERED, active project is admitted into that project's thread (dialogue, children, attachments and answer in the room the owner already has; Main still receives the one host-stamped completion row), and Main is told it finished only when its work is actually in that room — addressed there at admission or BOUND to the project; Registration alone does not qualify. Every other run stays in the hidden partition, silent in every chat, read back through the terminal, `--result-json-out`, the chat-blind Logs panel and `GET /api/tasks/<id>`; a reserved but inactive project keeps its chat acceptable so the queue's lifecycle fence refuses with its own typed reason, and a project deleted mid-run keeps its reserved chat. The run is also NAMED at admission and chat promotion, without a new model call: a caller-supplied `title` (`ouroboros run --title` or the top-level contract field; `metadata.title` is refused with a 400 like `metadata.project_id`) is authorship and fills both `title` and `suggested_name`; otherwise the request's first line, stripped of markdown and capped at the project-name length, fills `suggested_name` ALONE, so a truncated prompt never outranks a real name coined later, and a `task_named` frame is broadcast on admission so the live card is never born showing its status phrase as a title — the client buffers a `task_named` that arrives before the card's record exists (`web/modules/chat.js`), so frame order does not matter.
`queue_snapshot.json` is an atomic recovery and diagnostic projection, not a second scheduler. It carries pending and running rows, acceptance and root-budget fences, resident/active/parked worker counts, assignable capacity, and any pool-disabled reason. Startup restores a recent snapshot into an otherwise-empty pending queue and never resurrects ordinary RUNNING work: instead it FENCES every surviving RUNNING row with a durable cancel intent (`reason='server_shutdown'`) and lets the ordinary cancellation custody path terminalize it, expire its open quiz and close the paired owner wait, so a window closed on live work ends honestly instead of leaving a ghost; the restore ledger row names them as `terminalized_running` and the boot notice states the intent, because custody writes each terminal result a watchdog window later. A selected native owner-wait handoff remains eligible beyond snapshot age only through its current waiting source and acknowledged planned-restart transaction. Terminal tasks stay terminal, a task with an active durable cancel intent (or a legacy cancel-requested latch file) is left for cancellation custody, descendants below an accepted or sealed root finalize as cancelled, and malformed durable fence evidence fails closed. Snapshot capture copies the live containers under the queue lock because concurrent HTTP mutation can otherwise crash the supervisor mid-iteration. Assignment mirrors the RUNNING status into the durable task result for EVERY assigned task, not only a subagent, because both orphan healers read the STORED status; a root left unmirrored is a ghost no snapshot-less boot can settle. Orphan reconciliation is therefore a terminal writer for roots too, and closes the same owner quiz and paired owner wait the task-done seam closes.
Pooled completion separates a finished file-save attempt from publishable terminal truth. `worker_process.worker_main` calls `headless.prepare_terminal_task_files` before its own buffered `task_done`, after blocking post-task work; earlier answer/metrics frames retain their order. The worker stays occupied until that first attempt ends. Its private integer `_files_prepared_attempt` identifies the attempt, not save success, and never becomes a durable result or public event field. `events_task_done` re-reads CURRENT and uses `headless.terminal_task_files_ready`: a split drive requires the existing child-bound copyback projection and body, not merely an early terminal post-task checkpoint; pending refs may remain, but workspace artifact finalization cannot still be pending. Queue removal, slot release, accounting and project/evolution hooks stay with the normal event owner.
Legacy or faulted completions use `enqueue_terminal_file_recovery`: `worker_health` owns file-job preparation/recovery, while `task_reaper` owns execution queues and deferred-job replay. Existing lazy facades preserve caller names; this is one D08 custody flow. Missing/unreadable CURRENT publication retains the same RUNNING ownership and retries on the health cadence; it does not emit false Done or replay model work. The helper's transient `terminal_source_present` distinguishes confirmed terminal input, confirmed absence, and an unknown read. Confirmed absence reaches the existing lifecycle-fault owner: an early sticky completed status is preserved while execution becomes `infra_failed`; authored answer, review/objective and cost survive. A CURRENT publication/cancellation that won the race is not faulted. Pooled mailbox cleanup follows file preparation through `settled_mailbox_cleanup_allowed`; accepted pending attachment refs and open post-task work also protect startup cleanup. Direct paths keep their existing ownership.
A required owner wait keeps its task RUNNING and retains the same worker, command queue, live browser and services. The completed-tool checkpoint and task-result `owner_wait` projection reach durable storage before the snapshot confirms the park. The worker lends only its `active_capacity`; the ordinary assignment tick replenishes the configured active capacity through the existing spawn/readiness owner. Addressed input requests a wake, which waits for active capacity and retires an idle replacement or atomically transfers a confirmed-dead exhausted replacement's reservation to the original worker. A failed grant restores both capacity marks; temporary boot/reap reservations still count against the cap, so maintenance cannot buy an extra replacement. Resumption preserves the task attempt and start timestamp and marks its continuation authority consumed before dispatch; no completion, new admission or replay of completed tools occurs. Waiting spares only the idle rail. Stop, deadline and absolute ceiling keep their authority, and cancellation, timeout or crash retires an inactive worker or replaces it when active capacity is missing, without exceeding the configured active limit. A process crash or automatic timeout retry preserves the current-attempt checkpoint as evidence and does not blindly replay the task; cold continuation is limited to confirmed planned restart, which cannot preserve an OS browser session. Every shutdown cleanup recognizes the same restart transaction. An owner-requested MANAGED UPDATE is such a planned restart: its writer fence prepares the owner-wait/planned-restart handoff before stopping the pool and passes those exact task ids as `preserve_running_task_ids`, so a parked wait is requeued instead of interrupted, and a preparation failure blocks the update (repo-writer admission re-opens) rather than terminalizing the wait. At the common direct re-exec seam, a valid managed update in the same restartable phases allowed by `_safe_restart_serialized` arms the already-prepared transaction through the existing one-shot environment handoff. This includes assisted updates whose waiting tasks are already PENDING; the same-PID successor acknowledges it (launcher mode observes exit code 42). No-resume flags suppress this token, and an aborted update's leftover recovery record alone cannot authorize a later manual Restart or rollback. A failed restart callback retains the existing disarm behavior. A manual ROLLBACK returns the tree to an older runtime, so it deliberately parks nothing and keeps the ordinary interrupt semantics. Cold preparation restores the original CostCeiling before constructing either context projection and retains hard clocks, refreshes genuine assignment progress, and rebinds ContextFit to the saved model. Cold continuation completes the saved round's pending budget decision before advancing to another model round or applying queued route overrides. TaskModelWait retains its per-role choices (including Auto), auto-continue choices and completed quota union; both native execution and supervisor timing retain the original execution basis and clock revision. Calendar deadlines and time spent waiting for the owner are unchanged. Successful grant consumes the queue task's resume locator; source evidence remains available through a later crash or timeout.
Ordinary Main/Project roots retain their in-process actor through the same completed-tool owner-wait boundary. `owner_wait.direct_owner_wait` uses that actor's existing mailbox and TaskModelWait controls; no pooled slot is held, lent or synthesized. Pending owner events are flushed before sleeping, and the loop handles controls/deadlines before its post-tool budget decision after wake. Its saved source is evidence, not an automatic cold-restart grant for an ordinary conversation. An addressed text or existing quiz answer resumes the same live stack and browser.
Cancellation is intent-then-custody; intent and outcome are separate fields, because one field carrying both wedges a task forever. The skeleton:
1. **Intent.** Every cancel ingress — the agent `cancel_task` tool, the HTTP single and cascade endpoints, evolution stop, project deletion, the per-descendant mints of a cascade sweep, and the boot migration of legacy `cancel_requested` files — writes one durable row through `ouroboros/cancel_intents.request_cancel` into the locked projection `state/cancel_intents.json` (active intents only; every transition also appends a forensic `cancel_intent` supervisor-ledger row). Every ingress fails closed: a failed intent write refuses the cancel with a typed error (tool `CANCEL_INTENT_WRITE_FAILED`, HTTP 503; a corrupt projection file gets its own refusal, `CANCEL_INTENT_PROJECTION_CORRUPT`, naming the preserved file), so no teardown ever runs without a durable, watchdog-replayable fence. Evolution stop keeps any task whose intent write failed and reports the stop INCOMPLETE with typed per-task outcomes.
2. **Scope.** A cascade ingress mints its intent with `scope: cascade`; recorded scope is widen-only (single→cascade, never narrowed, so Stop-now cannot shrink a cascade). A cascade over an already-settled root with live descendants still mints the durable coordination intent (`allow_settled_target`) — that intent is the watchdog's replay trigger for the subtree — and the sweep mints a per-descendant intent so a crash leaves no live descendant unfenced. Timeout reaping is deliberately NOT a cancel ingress: the reaper keeps its own custody protocol over the `reaping` slot marker and never mints intents, because cancellation is reserved for explicit intent.
3. **Claim.** `supervisor.task_lifecycle.cancel_task_custody` is the ONE settle owner. It claims the intent before any custody mutation (owner + generation, exclusive while alive); a refused claim exits having touched nothing, so two racing custodies can never interleave into a double settle. The one secondary settle site — the pre-assignment pending drop — holds the same claim/generation fence before settling; a queued task whose budget is exhausted now PAUSES instead of terminalizing, so there is no batch terminalizer beside custody. Stopping a LIVE direct-chat turn is addressable through the same custody: there is no worker process to kill, so the chat lane writes the typed `finalize_now` control into the canonical drive's owner mailbox — the one the turn's loop drains at every round boundary, where it ends with zero further model calls (the paid post-task synthesis included: the hard stop records the existing `_skip_post_task_synthesis` marker on the tool context and `emit_task_results` carries it onto the task record, so no summary, reflection or consolidation bills after the stop and no open checkpoint is left for the boot reconciler; a Stop-now that lands AFTER the loop returned, while that paid synthesis is already running on its in-process thread, is still addressable — the pipeline's in-flight key counts as live ownership for `task_has_live_ownership`, custody keeps the durable immediate intent open (typed "still live") instead of settling it, and the synthesis worker consults that intent (or the live marker) before EACH paid stage, skipping the rest and disclosing them as `post_task_stop_reason` `owner_stopped:skipped=<stages>` on a `degraded` checkpoint that rides the result row and `task_cost_finalized`) — arms it exactly once per turn (the turn's own stamp is the latch, atomic against the turn's completion so a turn that already ended gets no control and no owner toast), then waits the short `OUROBOROS_DIRECT_TURN_STOP_WAIT_SEC` bound. The typed outcome is `gone` / `ended` / `live`; `live` releases the claim so the supervisor sweep retries rather than publishing a fabricated `cancelled` row over a turn that is still running.
4. **Kill and re-check.** Custody confirms process death, then re-reads the child's real settled result. Natural completion WINS: a child that finished before the kill keeps its completed result, artifacts, and cost, and the cancel settles as already-settled.
5. **Reconcile and capture.** The task's open delegated runs are reconciled from durable custody rows and always re-audited (open runs and pending invocations are disclosed regardless of the reconcile outcome's shape). Workspace artifacts are captured from the real tree; a failed or owed-but-unrunnable capture is `failed`, never `missing`, and a shared-tree capture carries `attribution: shared_unproven`.
6. **Settle.** The settled result is written with reconstructed-or-honestly-unknown cost, never a fabricated final `$0`. `parent_decision` is stamped only at this outcome.
7. **Owe, then publish.** The owner's terminal answer is registered as OWED in the durable outbox (or a typed no-chat handoff row) BEFORE the intent settles and before `task_done` publishes, so a crash between settle and send replays the answer instead of losing both it and the watchdog trigger. A cascade delivers one root message with a children digest under the deterministic delivery id `cascade:<root_tid>:<request_id>` (replays dedup; a later separate cancel delivers its own); digest membership merges the root's durable descendants with the sweep's outcomes, each child's line rebuilt from its current durable status.
8. **Watchdog.** The supervisor tick dispatches the existing cancel/delivery/ref sweep off drain under one nonblocking in-flight lock (`server_maintenance._run_cancel_delivery_ref_sweep`, unchanged ~20 s cadence). Its `sweep_cancel_intents` re-feeds unclaimed or abandoned-claim intents into custody — replaying a cascade as a cascade, not as a single cancel — so a lost control event or a custody attempt that died mid-teardown cannot wedge a cancellation. The cascade postcondition is judged on physical queue/durable liveness (the intent itself excluded); only that no-live check settles the coordination intent, after the tree's summary is registered as owed.
Readers see the typed projection `cancel_state: "pending"` (with `cancel_reason` beside it) on effective results until the settle; the UI shows an interim "Cancelling…" and restores the Cancel button only when a fetched live non-pending task detail proves the intent is gone. Steering writes — `steer_task`, mailbox follow-ups, `forward_to_worker` — are refused typed while a cancel is pending; the check runs before attachment staging and a refusal removes just-staged inputs. Queue restore and pre-assignment consult the projection under the queue lock, so a cancelled pending task never starts. `task_done` is validated through the DURABLE result unconditionally for every event: a non-settled event status, a settled claim over a non-settled durable row, or a blank status over a non-settled/absent row is refused as a durable lifecycle fault — left to custody when a cancellation is pending, otherwise published through the existing lifecycle-fault owner with a typed infrastructure-failure axis, preserving an existing sticky terminal status; unreadable publication keeps an owned file retry instead of releasing a false terminal; that synthetic terminal rides the normal dispatch seam, including the assisted-update orphan watchdog and the cooperative-checkpoint hooks. Terminal answers ride one durable delivery seam (`supervisor/terminal_delivery.py`): restart-surviving `delivery_id` dedupe shared with the natural final-answer path, a loud UNREVIEWED salvage message (bounded preview, exact omitted count, full-copy receipt) for cancelled and non-retry-reaped tasks, one root message with a children digest for a cascade, nothing for a retryable reap; routing follows the task's lineage chat. The already-settled fast path and the finalize-on-miss lane run the same delegated-run audit as the kill path, so a cancel over a dead task with live delegated runs never reads as a clean completion. An agent-requested cancel (`cancel_task` tool → `cancel_task` control event) publishes nothing of its own for the cancelled, already-settled and not-found outcomes — the custody seam's terminal rows in the task's own thread and the typed tool result are the truth, and a host acknowledgement in the global owner chat was duplicate, untyped speech — while a FAILED settle still speaks as the typed `cancellation_fault` progress incident on the requested task. `_handle_cancel_task` dispatches `_drive_cancel_task_event` off drain with a per-root/task in-flight key; the existing durable claim/generation remains custody authority, and failure releases the local key for watchdog recovery. HTTP cancellation keeps its existing worker dispatch and response contract; unrelated Stop work does not queue behind a new shared file executor. The custody, completion-wins, and owed-before-published invariants are restated in §10.
Stop POLICY is an axis on the same durable intent, independent of cascade scope. An omitted/empty-body cancellation is the synchronous IMMEDIATE teardown, keeping programmatic callers' bounded budgets. An explicit `stop_policy=finalize_then_cancel` answers 202 with the intent OPEN and runs one bounded owner-stop finalization episode (`supervisor/owner_stop.py`): live descendants settle first and feed a bounded child-result projection into the root's final turn; the root receives a deterministic `finalize_now` control whose typed first line (`owner_requested_finalization`) routes to its own loop rail — zero or one tool-less model turn, retained-candidate reuse, terminalizing completed/best-effort under the honest owner reason rather than a false deadline reason. Two anchors are immutable: the grace budget starts at the durable `control_drained_at` (first drain wins, so a task inside a long tool call still gets its final turn when the hard bounds allow), under the request's outer `OWNER_STOP_OUTER_CAP_SEC` cap. Held tasks bypass only generic idle/finalization-grace rails; the task's explicit deadline and absolute ceiling remain independent hard axes and are never widened. Expiry, a hard-bound hit, a pending root, or an already-settled root feeds ordinary custody. Policy transitions are monotonic: an immediate request HARDENS a pending graceful intent (preserving any cascade scope, revoking an unread finalization control, revalidated at drain); graceful can never soften an accepted immediate. A successful graceful root suppresses the redundant cascade summary; Panic bypasses both. The UI projects the pending soft stop through `cancel_state`+`stop_policy` ("Finalizing…"). Beside stopping sits the owner "hurry" control: a typed task-local `kind=hurry` owner-mailbox control (`ouroboros/owner_hurry.py`, `gateway/task_hurry.py`) that skips the next otherwise-eligible acceptance panel with a typed reason, zeroes remaining improvement passes, and makes force-plan projection task-locally advisory — never a chat message, never a settings mutation, never a P3/commit/review-gate weakening. Its effect is attempt-scoped (`task["_attempt"]`); a shared `retry_reset` strips it on every same-id requeue, reaper timeout and crash requeue alike. These invariants hold for every install configuration class.
The event bus is process-lifetime rather than worker-generation-lifetime: full-pool and single-slot respawns reuse one manager-backed queue shared by workers, direct chat, and consciousness. A force-killed producer can corrupt a raw multiprocessing feeder frame, while rebuilding the queue on pool rotation strands surviving producers on the old endpoint; synchronous manager serialization and isolated producer connections avoid both. Live-frame publication of persisted rows is exactly-once and process-symmetric: `ouroboros/utils.py::append_jsonl` streams only runtime `logs/*.jsonl` rows into the process log sink (never `chat.jsonl`, which has its own live channel, and never state/memory/receipt stores), and each process suppresses the types whose live delivery has a dedicated owner — workers via `WORKER_LOG_SINK_SUPPRESSED_TYPES`, the server via the superset `SERVER_LOG_SINK_SUPPRESSED_TYPES` in `make_server_log_sink`. One persisted event produces exactly one live frame, pinned by `tests/test_log_forwarding.py`; an LLM call failure is one durable `llm_api_error` row and nothing else (the live-only `llm_round_error` sibling the same producer once emitted was the same failure twice — Background Consciousness keeps its own live `llm_round_error`).
Heartbeat and progress are different evidence: a heartbeat proves a process or loop is alive; owner-visible progress and model-usage events prove the task advanced. Fresh descendant progress or queued descendants can keep an orchestrator alive, while an explicit deadline, absolute ceiling, cancellation, and budget stop remain hard. After the typed finalization episode (§6), timeout handling freezes its decision under the queue lock, removes the task from ordinary assignment, marks the worker `reaping`, and hands kill, join, salvage, retry, and respawn to the single off-loop reaper. An orchestrator with live descendants is not blindly retried, because a retry would replay its plan and spawn a competing tree.
No retry or new assignment may occupy a timed-out slot until the original process is provably dead. If kill and join cannot establish death, the reaper preserves a low-rank RUNNING result, keeps the slot marked `reaping`, emits a visible wedged receipt and restart hint, and performs no terminal write, `task_done`, retry, or respawn. One slot is sacrificed rather than letting a still-running process race a replacement and overwrite its result; the next supervisor generation reconciles the durable record after old-generation process custody has run.
A spawned or respawned slot is not assignable until its child's PID-bound `worker_ready` row arrives; `supervisor/worker_pool_lifecycle.py` keeps both spawn paths under the existing readiness watcher and lifecycle→queue lock order. Temporary `reaping` reserves active capacity even when not assignable. After `WORKER_READY_MAX_ATTEMPTS` windows of `WORKER_READY_WINDOW_SEC` (`runtime_limits.py`, re-exported by `config.py`), `Worker.readiness_exhausted` records the final outcome for that exact slot; late ready/watcher events cannot reopen it or alter a replacement generation. `worker_pool_lifecycle` owns `_worker_pool_execution_state` and exhaustion disablement behind the existing workers facade; it distinguishes busy/booting/reaping capacity and a valid live owner-wait stack from total exhaustion. Reserve, final enqueue, snapshot and public admission share that fact; repository-writer admission remains a separate outer policy. Total exhaustion closes pooled ingress, including owner `/review`, without blocking direct chat/control or internal boot/update recovery. Once existing RUNNING completion custody is settled, `disable_exhausted_worker_pool` uses the ordinary pool stop to fail unstarted PENDING work honestly with a Restart hint; failed publication retains existing terminalization-retry custody. A new incoming task cannot clear the latch or restart the exhausted attempt budget. Readiness remains separate from liveness and task idle time. Missing events are an empty read; watcher errors release only still-booting, non-exhausted slots to the crash detector with `worker_ready_released`. Linux workers use forkserver; macOS and Windows use spawn.
Unexpected worker death reserves exact worker/meta/task/attempt/root custody under the queue lock and enqueues `confirmed_dead_worker` on the existing reaper. `worker_health.recover_confirmed_dead_worker` performs archive/file work outside that lock and checks strict CURRENT readiness after the shared helper. A saved terminal source wins even after signal death; unknown or incomplete file publication retains the same job via `TerminalFileRecoveryPending` and health-driven deferred replay. Only confirmed absence of a usable terminal source reaches the existing signal/non-signal crash policy: a signal is an infrastructure failure; an otherwise eligible non-signal crash retries within `QUEUE_MAX_RETRIES`, preserving owner-wait replay restrictions and cost. Generation checks and atomic retry rollback prevent stale jobs from affecting replacements. Crash-storm jobs follow discovered dead-worker jobs, with respawn suppressed while their terminal sources settle; the existing storm fence then stops pooled admission. Direct chat remains independently available.
Startup terminal-file recovery runs in `_run_supervisor` after process custody and before `_startup_prune_sweeps`; without providers the lifespan uses the same recovery without spawning a worker. Prior worker/server PIDs are captured before spawn replaces the old PID record. Unknown or still-live ownership defers recovery rather than racing a writer. The existing recovery owner scans only known headless/task-drive roots, protects live/pending work and acknowledged owner-wait continuations, and invokes the shared file helper before orphan-result/post-task reconciliation. Its transient report carries recovered, unresolved, protected and error facts. Any unresolved/protected source or ownership/read error sets `preserve_task_sources` for that startup pass, skipping task-drive/tree/temp-source deletion while unrelated maintenance retains its existing owners. No initial adoption anchor or new durable recovery store is needed.
Startup and throttled maintenance reconcile three residue classes. Process custody checks strict PID, start-time, command fingerprint, owner task, session, and generation evidence before reaping an owned process. Delegated-run reconciliation applies the same owner-gone reasoning to external harness rows; the write-side healing seams (boot-backfill reverse join, cursor pass, sweep refresh) and the counters-as-snapshot rule are specified once under §6 Delegated subagents. Task, review, and project reconciliation repair durable records whose producer no longer exists. None of these are command-line-class kill sweeps, and one development or runtime instance never reaps another. The dedicated watchdog separately observes supervisor-loop liveness and every registered native actor; it alerts and recommends `/restart` but cannot kill an individual in-process thread. Other owner conversations run on independent native actors, without a second scheduler.
Cooperative project checkpointing has two equivalent quiescence triggers: a host-minted genesis or cooperative tree is checked when its root settles with no live descendants, and again when the last child settles beneath an already-terminal root. The second trigger exists because a root-scope budget stop terminalizes the root before its children reach their own dispatch boundaries; a root-only trigger would see a live tree once and never return. Event dispatch detects the condition only after removing the finishing task from RUNNING. The bounded git chain runs on a daemon thread, revalidates quiescence under the queue lock immediately before mutation, and uses a per-root latch that replays a trigger arriving during an in-flight check. Only host-minted project roots are eligible; owner-attached folders are never auto-committed, credential-shaped files stay excluded and disclosed, and every material success, skip, or error receives a durable receipt.
The bridge recognizes `/panic`, `/restart`, `/review`, `/evolve [on|off]`, `/bg [start|stop|status]`, and `/status`; all other text enters ordinary agent routing. External transports may invoke these commands only with positive owner identity and a transport-specific owner-chat binding. The commands reuse runtime-mode, queue, cancellation, and typed-result authority rather than implementing parallel control paths. Runtime logs rotate on the same supervisor tick and archive readers preserve their retained timelines. Only explicitly isolated devtool roots may use the narrow rotation sentinel from §1; normal runtime roots never inherit it.

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,215 @@
# 7. Configuration (ouroboros/config.py)
This chapter owns the settings surface: where the document lives, which functions may persist it, how a document written by an older release is translated into today's vocabulary before any default is merged, the per-surface output-token budgets, and the full registry of shipped defaults with their meanings. It exists so a new key has exactly one home and one default, and so this table stays a test-mirrored projection of `config.SETTINGS_DEFAULTS` rather than a second authority.
`ouroboros/config.py` is the SSOT for paths (HOME, APP_ROOT, REPO_DIR, DATA_DIR, SETTINGS_PATH, PID_FILE, PORT_FILE), process constants (RESTART_EXIT_CODE 42, AGENT_SERVER_PORT 8765), every settings default below, and the load/save/env machinery: `load_settings()`, `save_settings()`, `apply_settings_to_env()` (copies hot-reloadable runtime keys into `os.environ`), `normalize_runtime_mode()` (one clamp shared by the save path, the read coercion, and onboarding validation), `get_runtime_mode()`/`get_skills_repo_path()`, and `acquire_pid_lock()`/`release_pid_lock()`; `ouroboros/update_channels.py` owns `get_update_channel()`/`get_update_branch()`.
Settings file: `data/settings.json` under the data root — `~/Ouroboros/data/settings.json` by default, with `APP_ROOT`, `DATA_DIR`, and `SETTINGS_PATH` independently env-overridable. Access is file-locked. `secret_masking.py` is the wire-placeholder authority for known and owner-defined top-level secrets: `load_settings()` repairs only recognized disk placeholders BEFORE environment precedence is resolved, so a real environment credential is never classified as a mask, and `prepare_settings_for_persist()` applies the same top-level repair at the common writer boundary; nested MCP values are never silently migrated.
`ouroboros/openrouter_attribution.py` is the application-identity SSOT for every first-party paid OpenRouter request (canonical URL + `X-OpenRouter-Title`); a fork must use its own URL rather than competing to rename one app record.
### Reading and writing the settings document
`agent.handle_task` binds the admitted task's `settings_integrity.TaskSettingsSnapshot` for the complete task entry. Its normalized document and exact environment projection are separate private in-memory views: absent and empty values remain distinct, concurrent tasks keep their own models/keys/Supervisor/Review settings, and explicit task route overrides still win. A short `SETTINGS_ENV_LOCK` serializes capture/publication only, never task execution. `runtime_setting`, `runtime_settings` and `runtime_environ` reuse that view; settings writers keep reading the current document, immediate effects stay live, and Access retains its boot pin. `model_wait.copy_wait_context` and task-owned helper threads carry only the existing settings binding alongside their established owners. Out-of-process extensions receive only their permitted typed next-task values through the existing private per-call payload; the child still validates current grants, reads immediate values live and receives no whole settings snapshot or extra credential environment. Failed reload loudly retains the prior environment while disclosing unavailable document-only values.
A settings document on disk was written by whatever release the owner last used, so reading one starts by translating it into today's vocabulary. `normalize_settings_raw()` is that translation and the only copy of it: type coercion against the declared defaults, the deprecated per-subsystem retention keys folded into the unified one, the retired acceptance-pass count consumed into the shared review-cycle cap, the keys a release retired dropped, the renamed model slots promoted, and secret placeholders repaired. Every step preserves an owner customization written under a former key, and the ORDER is load-bearing (the pass count is consumed before the retired purge would drop it; the purge runs before the slot rename, so a retired spelling is never promoted), so every reader applies it BEFORE the shipped defaults merge — `load_settings()`, the owner endpoints' `_owner_read_settings_raw()`, and the Colab re-run's `build_colab_settings()` alike. "Raw" in that name is about the runtime-mode ratchets it deliberately skips, never about the migrations. It is pure and idempotent — it touches no file and no environment — which is what lets a read stay a read and lets a read-modify-write apply it on every save; both in-process readers share one read primitive (`settings_integrity.read_settings_json_verified`), so a pinned snapshot that changed refuses the owner reader exactly as it refuses the loader. This seam carries the VOCABULARY normalization only: the provider normalization (`server_runtime.apply_runtime_provider_defaults`) is a separate, never-persisted derivation every route consumer makes over the effective document, `context_fit.resolve_context_fit_route()` included.
Five functions persist a settings document, and every one calls `serialize_settings()` and commits its output through a byte-exact helper (`utils.write_text_atomic()`, or `Path.write_bytes` on the config saver's rename-less `OSError` fallback) — never a text-mode write, which would turn LF into CRLF on Windows. Three persist THIS process's document through `prepare_settings_for_persist()`, the single point where the disk-authored silence rule and ordinary-mode context/safety ratchets are applied against the value ON DISK, with Cyber configuration authority from effective Access: `config.save_settings()`, `gateway/owner_settings._owner_update_settings()` (which `_owner_write_settings()` is one caller of), and the packaged bootstrap's `packaged_cli._save_settings()`. Two are exempt by design and pinned as such: the one-window raw context-pair migration (`context_mode_compat.normalize_and_persist_context_mode_compat()`, written under the load lock — the raw mapping with only the pair changed, never a defaults-merged document) and `colab_bootstrap.write_colab_settings()`, which writes a generated document for a foreign data root the prologue's on-disk proofs would answer wrongly for. One scan closes the inventory over `ouroboros/**`, `supervisor/**`, `server.py` and the repo-root `launcher.py` (`tests._shared.settings_writers`): a function is a settings writer when it CALLS `serialize_settings()`, or when it names the settings path or file AND does a write-shaped thing; it counts as ROUTED only when it CALLS the prologue, and naming either in prose is neither. The flagged set must equal the five writers plus the scan's declared non-writer matches, so a sixth writer in those roots fails the tripwire whether or not it is routed.
An owner endpoint changes one decision inside a document it does not otherwise own, so it must write the whole document back. `_owner_update_settings(transform, expected_digest)` does that read, change and write inside ONE settings lock: the transform receives the document as it is under the lock and returns what to persist, or nothing at all, which is how a no-change decision avoids rewriting the file. An endpoint acting on an earlier read passes the digest that read saw (`settings_document_digest()`, the same staleness question the onboarding transaction asks), and a mismatch refuses before the transform runs, so a concurrent owner change can never be reverted key by key while the request answers "saved".
### LLM output token budgets
Providers name the same output-token budget differently: OpenRouter/Anthropic-compatible calls send `max_tokens`, while every official direct OpenAI Chat route sends `max_completion_tokens` — a real provider-wire boundary, not naming style. Direct OpenAI also sends the requested `reasoning_effort` provider-wide; model-name prefixes are not capability authority, and only exact-route success-confirmed wire evidence may adapt a request. Runtime floors (numeric SSOT: the constants in code, pinned by `tests/test_max_tokens_constants.py`):
| Surface | Output-token budget |
|---------|---------------------|
| `LLMClient.chat()` / `chat_async()` defaults | 65,536 |
| Main task loop (`loop_llm_call.MAIN_LOOP_MAX_TOKENS`) | 65,536 |
| `LLMClient.vision_query()` and VLM tools (`analyze_screenshot`, `vlm_query`) | 32,768 |
| Review synthesis dedup | 16,384 |
| Chat block consolidation, era compression, scratchpad consolidation | 16,384 |
| Execution reflection and pattern-register update | 16,384 |
| Post-task summary (`agent_task_pipeline`) | 16,384 |
| Improvement-backlog grooming (`improvement_backlog.groom_backlog`) | 8,192 |
| Post-task evolution promotion decision (`post_task_evolution`) | 8,192 |
| Context compaction round summaries | 32,768 |
| Skill publish PR body generation | 8,192 |
| Background consciousness loop | 65,536 |
| Project naming LIGHT one-shot (`project_naming.llm_project_name`) | 256 |
| Update letter LIGHT one-shot (`update_letter.write_letter`) | 1,024 |
| Provider Test (`llm_probe.PROVIDER_TEST_MAX_TOKENS`) | 16 |
### Default settings
A registry of `config.SETTINGS_DEFAULTS` (exact defaults stay canonical in `config.py`; this table is test-mirrored against it). Rows marked env-only are operator environment levers with no settings.json carrier.
| Key | Default | Description |
|-----|---------|-------------|
| OPENROUTER_API_KEY | "" | OpenRouter credential |
| OPENAI_API_KEY | "" | Official direct-OpenAI credential |
| OPENAI_BASE_URL | "" | Legacy OpenAI base-URL override |
| OPENAI_COMPATIBLE_API_KEY | "" | OpenAI-compatible endpoint credential |
| OPENAI_COMPATIBLE_BASE_URL | "" | OpenAI-compatible endpoint base URL |
| CLOUDRU_FOUNDATION_MODELS_API_KEY | "" | Cloud.ru credential |
| CLOUDRU_FOUNDATION_MODELS_BASE_URL | `https://foundation-models.api.cloud.ru/v1` | Cloud.ru base URL |
| GIGACHAT_CREDENTIALS | "" | GigaChat auth key |
| GIGACHAT_USER | "" | GigaChat user login |
| GIGACHAT_PASSWORD | "" | GigaChat password |
| GIGACHAT_SCOPE | `GIGACHAT_API_PERS` | GigaChat API scope |
| GIGACHAT_BASE_URL | `https://api.giga.chat/v1` | GigaChat base URL |
| GIGACHAT_VERIFY_SSL_CERTS | `true` | GigaChat TLS verification |
| GIGACHAT_PROFANITY_CHECK | "" | GigaChat profanity filter passthrough |
| ANTHROPIC_API_KEY | "" | Official direct-Anthropic credential |
| MINIMAX_API_KEY | "" | MiniMax credential |
| MINIMAX_REGION | "" | MiniMax region (empty resolves `global_en`) |
| DEEPSEEK_API_KEY | "" | Optional. DeepSeek direct provider key (`deepseek::...` model values, OpenAI-compatible API at the fixed official endpoint) |
| OUROBOROS_NETWORK_PASSWORD | "" | Non-localhost HTTP gate password (`server_auth.py`; unset only warns — see §8 packaging note) |
| OUROBOROS_SERVER_HOST | 127.0.0.1 | HTTP bind host (`0.0.0.0` for Docker/non-loopback) |
| OUROBOROS_UPDATE_CHANNEL | `stable` | Update channel: stable/qa/development (§8) |
| OUROBOROS_MANAGED_UPDATE_FETCH_TIMEOUT_SEC | 300 | Managed-update fetch ceiling |
| OUROBOROS_RESCUE_GIT_TIMEOUT_SEC | 300 | Per-process ceiling on rescue Git commands |
| OUROBOROS_TRUST_NONLOCAL_BIND_WITHOUT_PASSWORD | unset | Env-only: `1` permits saving a non-loopback bind without a password |
| OUROBOROS_MODEL | google/gemini-3.8-flash | Main model |
| OUROBOROS_MODEL_HEAVY | "" | Legacy slot: readable for migration/history only, out of active routing |
| OUROBOROS_MODEL_LIGHT | openai/gpt-5.6-luna | Light model |
| OUROBOROS_MODEL_ACCOUNTS | "{}" | Role-owned managed account pins; empty means Auto, fallback entries retain order |
| OUROBOROS_PROCESSING_PREFERENCE | "" | One global provider-neutral processing preference: standard/fast/economy; empty preserves adapter default |
| OUROBOROS_MODEL_PROCESSING_PREFERENCES | "{}" | Optional role-owned processing preferences; empty roles inherit the global preference |
| OUROBOROS_MODEL_CONTEXT_WINDOWS | "{}" | Role-owned context sizing assertions; zero means Auto, not a provider limit or scope acknowledgement |
| OUROBOROS_MODEL_VISION | "" | Vision model (empty inherits) |
| OUROBOROS_IMAGE_INPUT_MODE | auto | Send-time image routing (`vision_routing.py`) |
| OUROBOROS_VISION_CAPTION_TIMEOUT_SEC | 90 | Caption-generation ceiling |
| OUROBOROS_MODEL_CONSCIOUSNESS | "" | Background-consciousness model (empty inherits) |
| OUROBOROS_MODEL_FALLBACKS | openai/gpt-5.6-luna | Cross-model fallback chain (`fallback_cooldown.py`) |
| OUROBOROS_MODEL_MAX_CONCURRENCY | 3 | Per-(model,route) concurrent provider-call cap (`model_concurrency.py`) |
| OUROBOROS_MODEL_SLOT_MAX_WAIT_SEC | 180 | Concurrency-slot wait bound |
| OUROBOROS_PROJECT_NAMING_TIMEOUT_SEC | 60 | Project-naming call ceiling |
| OUROBOROS_PROJECT_NAMING_ASYNC_TIMEOUT_SEC | 8 | Proactive-namer async bound |
| OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC | 120 | Update-letter LIGHT one-shot ceiling, slot wait and provider call together (`update_letter.py`) |
| OUROBOROS_FALLBACK_COOLDOWN_ENABLED | true | 429-aware per-process model cooldown |
| OUROBOROS_FALLBACK_COOLDOWN_SEC | 120 | Cooldown window |
| OUROBOROS_FALLBACK_ATTEMPTS_PER_MODEL | 1 | Attempts per model in the fallback walk |
| OUROBOROS_REVIEW_NATIVE_MAX_TRANSCRIPT_CHARS | 900000 | Owner CEILING (chars) on the native review inspection episode transcript; the effective bound is the reviewer window's calibrated capacity, never above this — except that a surface's declared mandatory reading (the advisory's five governance documents) lifts it up to the window, disclosed as `native_mandatory_read_exceeds_bound` when even the window cannot hold it. No round cap exists (`OUROBOROS_REVIEW_NATIVE_MAX_ROUNDS` is retired): exhaustion is a typed fail-closed refusal for verdict shapes and a disclosed incomplete product for the report shape, never a silent truncation |
| OUROBOROS_MODEL_DEEP_SELF_REVIEW | openai/gpt-5.6-sol | Deep self-review model key — the invisible migration source and fallback for the optional `deep_review` reviewer row: with no row saved, `deep_review_slot()` synthesizes the packed api row from it (the historical delivery); a saved row wins and the key is not read — so the provider-default migrations of this key (`server_runtime.py`) reach only installs that still synthesize from it; a row-configured install keeps its row, by design. Not a Settings UI field any more — the row lives in Agents → Review lanes |
| OUROBOROS_MAX_WORKERS | 10 | Active worker dispatch capacity; required owner waits retain additional sleeping processes |
| OUROBOROS_MAX_ACTIVE_SUBAGENTS_PER_ROOT | 6 | Live-subagent cap per root (hard cap 500 ids; depth hard cap 10) |
| OUROBOROS_MAX_SUBAGENT_DEPTH | 3 | Subagent tree depth |
| OUROBOROS_DISABLE_MANAGED_UPDATES | (unset) | Env-only: `1` disables managed updates (`git_ops.py`) |
| OUROBOROS_ALLOW_MUTATIVE_SUBAGENTS | (empty) | Mutative-subagent Auto override |
| OUROBOROS_SUBAGENT_WORKTREE_ROOT | (empty) | Acting-worktree root (empty derives `~/Ouroboros/subagent_worktrees`) |
| OUROBOROS_SUBAGENT_PROJECTS_ROOT | (empty) | Projects root (empty derives `~/Ouroboros/projects`) |
| OUROBOROS_SUBAGENTS | (empty) | Canonical configured-subagent roster (`configured_subagents.py`; §6) |
| OUROBOROS_SUBAGENT_HARNESS | (empty) | Legacy narrow harness input |
| OUROBOROS_SUBAGENT_PROFILE | (empty) | Legacy narrow profile input |
| OUROBOROS_DELEGATE_WAIT_SEC | 120 | Default delegate_wait window |
| OUROBOROS_DELEGATE_WAIT_MAX_SEC | 1800 | delegate_wait ceiling |
| OUROBOROS_DELIVERABLES_ROOT | (empty) | Deliverables root (empty derives the `~/Ouroboros/Deliverables` sibling; `tool_access.py`) |
| OUROBOROS_GC_RETENTION_DAYS | 7 | Unified GC retention (`retention.py`) |
| OUROBOROS_RESTART_DRAIN_MAX_SEC | 120 | Restart drain bound |
| TOTAL_BUDGET | 200.0 | Global budget (USD) |
| OUROBOROS_PER_TASK_COST_USD | 50.0 | Per-task cost cap; also the tree ceiling basis (`task_pacing.py`): the root resolves min(global share, cap minus margin), enabled descendants retain that original threshold while actual monetary admission remains independent, and the wrap-up affordability rail soft-lands under it |
| OUROBOROS_RUB_USD_RATE | (empty) | Manual RUB→USD rate for RUB-priced providers |
| OUROBOROS_PRICING_TTL_SEC | 21600 | Provider-catalog pricing cache TTL |
| OUROBOROS_TOOL_TIMEOUT_SEC | 600 | Default tool timeout |
| OUROBOROS_PER_CALL_TIMEOUT_CEILING_SEC | 1800 | Per-call timeout clamp |
| OUROBOROS_FINALIZATION_GRACE_SEC | 120 | Finalization grace window |
| OUROBOROS_WEBSEARCH_MODEL | gpt-5.2 | web_search backing model |
| OUROBOROS_WEBSEARCH_BACKEND | auto | web_search backend selection |
| OUROBOROS_MAIN_WEB_SEARCH | off | Main-loop inline web search |
| OUROBOROS_MAIN_WEB_SEARCH_ENGINE | auto | Inline-search engine |
| OUROBOROS_MAIN_WEB_SEARCH_MAX_TOTAL_RESULTS | 10 | Inline-search result cap |
| OUROBOROS_OR_PROVIDER | "" | OpenRouter provider-routing preference merged into requests |
| OUROBOROS_SEARCH_CODE_WALL_SEC | 45 | search_code wall-clock budget |
| OUROBOROS_PRESENTATION | (unset) | Env-only: launcher-exported presentation (`desktop_window`/`browser_fallback`; absent renders `web`) |
| OUROBOROS_USER_FILES_ROOT | "" (home) | Env-only: user_files jail root (empty = `$HOME`) |
| OUROBOROS_OBSERVABILITY_KEEP_RAW | unset | Env-only: truthy enables raw observability payload persistence |
| OUROBOROS_GENERATIVE_PROBE | 1 (on) | Generative-write probe toggle |
| OUROBOROS_GENERATIVE_PROBE_CHARS | 5000000 | Generative-probe size companion |
| OUROBOROS_REVIEWER_SLOTS | (empty) | (6.1) Structured reviewer-slot SSOT (`reviewer_slot_config.py`): JSON `{triad[], scope[], advisory, deep_review?}`; each row is EITHER an inline route `{slot_id, route:{kind: api_chat\|agent_session, target_id}, effort}` OR a roster reference `{slot_id, subagent_id, effort}` — mutually exclusive (a row naming both refuses typed; the reference materializes route/effort from the Available-subagents roster at load time, an explicit row `effort` winning over the roster row's) — with a STABLE owner-assigned slot_id (never an array index); an `agent_session` or managed-model `api_chat` route may add the optional `route.profile_id` credential pin (empty = account rotation; direct API-key routes reject account pins); the optional `deep_review` singleton carries the same row keys minus `slot_id` (fixed `deep_review_slot_1`) and, absent, is synthesized as the packed api row from `OUROBOROS_MODEL_DEEP_SELF_REVIEW`. Empty = read the legacy comma keys + phase-5 route envs as the migration source. Malformed value refuses typed at save AND at review time on every surface, task acceptance included (owner R3); env-apply logs and leaves legacy keys unprojected. The save that FIRST gives the triad a retrieving row (agent session or configured-subagent native inspection) returns the one-time R12 migration disclosure in the save response's `warnings` — the rows by id and target, and the measured API packet-panel cost it replaces (≈12 s / ≈$0.07 per model row per task, median of the 2026-09-01 OSWorld traces; ≈75 s / ≈$0.82 for a three-row panel on ProgramBench) against minutes of subscription window per task for a session row; a later save that keeps a retrieving triad is silent, and so — by design — is the reverse transition back to a packet-only triad: R12 is a one-time migration notice, not a routing monitor. The onboarding ladder footnote states the same numbers. |
| OUROBOROS_SUBSCRIPTION_PRESET_VERSION | (empty) | One-shot install-preset marker; endpoint-authored, DISK-ONLY (`ENDPOINT_AUTHORED_SETTINGS`) — its absence authorizes nothing, which is why install time is proved by three facts (§2) |
| OUROBOROS_SUBAGENT_PRESET_RECEIPT | (empty) | Install-preset receipt; endpoint-authored, disk-only |
| OUROBOROS_ONBOARDING_COMPLETED_AT | (empty) | Durable completion fact; endpoint-authored, disk-only |
| OUROBOROS_TASK_REVIEW_MODE | auto | Task acceptance-review mode |
| OUROBOROS_SAFETY_MODE | full | Safety supervision mode (shipped default `full`; a fresh desktop wizard may author `light`); lowering is owner-guarded (the runtime-mode boundary: §6 Safety and runtime mode) |
| OUROBOROS_SAFETY_MAX_TOKENS | 2000 | Safety-check output budget |
| OUROBOROS_SAFETY_CALL_TIMEOUT_SEC | 60 | Safety-check call ceiling |
| OUROBOROS_WEBSEARCH_TIMEOUT_SEC | 480 | web_search ceiling |
| OUROBOROS_DIRECT_TURN_STOP_WAIT_SEC | 2 | Seconds the chat lane waits for a stopped direct-chat turn to reach its next round boundary after the typed `finalize_now` control is armed; past it the outcome is `live` and the supervisor sweep retries custody rather than publishing a terminal (clamped 0-10) |
| OUROBOROS_ONBOARDING_SNAPSHOT_TIMEOUT_SEC | 45 | Bound on the onboarding transaction's settings snapshot read, so a wedged filesystem refuses the step instead of hanging the wizard |
| OUROBOROS_SETTINGS_DOCUMENT_LOCK_TIMEOUT_SEC | 30 | Bound on acquiring the settings-document lock for an owner read-modify-write; a timeout REFUSES before the transform runs — the lock is a precondition of the write, never a hint. The same bound also caps the initiating writer (`_run_settings_writer`: one lock wait plus one held episode; the generic save, the owner endpoints and onboarding completion alike): past it the Save answers 503 `settings_save_timeout` with `saved: null` and the body is left to its thread |
| OUROBOROS_LLM_TRANSPORT_READ_TIMEOUT_SEC | 2700 | LLM transport read timeout |
| OUROBOROS_PLAN_TASK_DEADLINE_MIN_SEC | 300 | plan_task deadline floor |
| OUROBOROS_ACCEPTANCE_REVIEW_EST_SEC | 200 | The floor: the minimum spendable seconds above the finalization reserve required to START an acceptance panel; never below 200 s (a smaller value is raised to it, a larger one wins). The improvement window is this floor ×2 under the adaptive improvement policy and ×1 otherwise. A spendable window at or below the applicable line is refused `review_skipped_deadline_reserve` / `improvement_window_inside_reserve`. The host predicts no review duration (owner R52): once launched, the review is clamped to the owner deadline and the task ceiling (R23) with the per-send money fence, and panel durations are telemetry only. |
| OUROBOROS_REVIEW_MAX_CYCLES | "2" | Shared paid review-cycle cap across the plan/acceptance/commit/skill gates (`unlimited` = no local count cap; per-gate semantics: `review_cycles.py` docstring, §10) |
| OUROBOROS_ACCEPTANCE_MAX_IMPROVEMENT_PASSES | (retired) | Retired alias: a stored value is MIGRATED into `OUROBOROS_REVIEW_MAX_CYCLES` (passes + 1) at settings load; a leftover env value is inert |
| OUROBOROS_ACCEPTANCE_RESERVE_PCT | 5 | Acceptance budget reserve percentage |
| OUROBOROS_OBSERVABILITY_RETENTION_DAYS | (retired) | Retired in 7.0 (`RETIRED_SETTING_KEYS`): observability rows are preserved indefinitely and the reader never deletes, so a retention knob had no reader; a stored value is stripped at settings load and a leftover env value is inert |
| OUROBOROS_REVIEW_MODEL_TIMEOUT_SEC | (unset) | Env-only: logical review timeout (absent = route-owned behavior; late in-flight results stay in custody) |
| OUROBOROS_REVIEW_MAX_TOKENS | 65536 | Env-only: reviewer output budget, clamped to the 8192 floor |
| OUROBOROS_REVIEW_ENFORCEMENT | advisory | Review enforcement: advisory/blocking (closed enum; anything else coerces to the default) |
| OUROBOROS_PREFLIGHT_TIMEOUT_SEC | 1800 | Env-only: TOTAL wall-clock budget for the hermetic pre-commit pytest preflight (node lane + both passes; teardown + containment semantics in `preflight_runner.py`/`process_containment.py`) |
| OUROBOROS_PREFLIGHT_SERIAL | unset | Env-only: `1` selects one serial pytest pass; scrubbed from the candidate environment |
| OUROBOROS_PREFLIGHT_TEST_WORKERS | (unset) | Env-only: xdist worker count for the hermetic parallel pass; floor 2, otherwise `os.cpu_count()`. Read from the OPERATOR environment and scrubbed from the candidate; concurrent-lane sizing rule in `docs/DEVELOPMENT.md` |
| OUROBOROS_AUTO_GRANT_REVIEWED_SKILLS | true | Auto-grant manifest-declared permissions to cleanly reviewed skills (hash-bound; blocking findings never grant) |
| OUROBOROS_TRUST_NATIVE_SEEDED_SKILLS | true | Launcher seed/resync writes hash-pinned `native_seed` verdicts; acts only at seed/resync, no runtime grant endpoint |
| OUROBOROS_CONTEXT_MODE | max | Context mode (`nano`/`low`/`max`), owner-selected outside Cyber Pro; `nano` is the compact owner projection and is recorded as `owner_nano`/`rendered_mode=nano` in physical usage facts. Current scope applicability follows BIBLE P1/P3. Cyber may configure it through the same audited writer. |
| OUROBOROS_CONTEXT_MODE_AUTO_LOW | false | Task-local low-mode overflow retry toggle |
| OUROBOROS_RUNTIME_MODE | advanced | Effective Access light/advanced/pro/cyber_pro; ordinary self-modification boundaries and Cyber agency are defined in §6 Safety and runtime mode. Settings persist the next-boot value; configured review enforcement remains independent evidence. |
| OUROBOROS_SKILLS_REPO_PATH | "" | Extra skills checkout path (expanded at read time, never cloned/pulled) |
| MCP_ENABLED | false | MCP client toggle (§6 MCP) |
| MCP_SERVERS | [] | MCP server list (HTTP/SSE via URL/auth, stdio via command+args and optional cwd/literal/settings-backed env); persisted in settings, never env-exported |
| MCP_TOOL_TIMEOUT_SEC | 60 | Per-MCP-tool timeout |
| OUROBOROS_HUB_CATALOG_URL | `https://raw.githubusercontent.com/razzant/OuroborosHub/main/catalog.json` | OuroborosHub catalog URL (automatic fetch limited to catalog JSON; installs verify SHA-256) |
| OUROBOROS_CLAWHUB_REGISTRY_URL | `https://clawhub.ai/api/v1` | ClawHub registry URL |
| OUROBOROS_PROMPT_CACHE_TTL | 1h | Prompt-cache tier (default/5m/1h). The policy acts at the final send-time wire boundary so it can legalize provider ordering without prompt builders creating provider-specific TTL policy; `review_helpers.cached_prompt_blocks` and `usage_accounting._reservation_cost` also consult it; usage records the applied tier |
| OUROBOROS_EFFORT_TASK | medium | Task reasoning effort (scale none/minimal/low/medium/high/xhigh/max/ultra; Settings exposes all but `minimal`); provider adaptation is exact-route, success-confirmed, disclosed in `request_wire` |
| OUROBOROS_EFFORT_EVOLUTION | high | Evolution effort |
| OUROBOROS_EFFORT_REVIEW | high | Review effort; reaches plan review as every row's default rung unless the envelope declares `reviewer_effort` |
| OUROBOROS_EFFORT_SCOPE_REVIEW | high | Scope-review effort |
| OUROBOROS_EFFORT_DEEP_SELF_REVIEW | high | Deep-self-review effort — the surface default; a saved `deep_review` row's own effort outranks it |
| OUROBOROS_EFFORT_CONSCIOUSNESS | high | Consciousness effort |
| OUROBOROS_RETURN_REASONING | true | Ask OpenRouter to return reasoning; direct/local request copies strip OpenRouter-only fields |
| OUROBOROS_REASONING_SUMMARY | auto | Readable reasoning-summary rendering; presentation-only, never added to history or returned to providers |
| OUROBOROS_TASK_IDLE_TIMEOUT_SEC | 900 | Idle timeout — requires absence of real task/subtree progress; the typed in-flight main-LLM row spares only this rail; a settled child result stamps parent progress, because delivery creates immediate integration work and must not coincide with idle termination |
| OUROBOROS_TASK_ABS_CEILING_SEC | 21600 | Absolute task ceiling, activity-independent; deadline and budget stay separate hard axes |
| OUROBOROS_SUPERVISOR_LIVENESS_DEADLINE_SEC | 90 | Supervisor/direct-turn liveness watchdog — alerts and recommends `/restart` but never frees an in-process lock held by a genuinely wedged turn |
| OUROBOROS_PACING_INTERVAL_SEC | 600 | Pacing reminder interval |
| LOCAL_MODEL_SOURCE | "" | Local-model source (HF repo or path) |
| LOCAL_MODEL_FILENAME | "" | Local-model GGUF filename (split first-shard expanded) |
| LOCAL_MODEL_CONTEXT_LENGTH | 16384 | Local-model context window |
| LOCAL_MODEL_N_GPU_LAYERS | 0 | GPU offload layers |
| USE_LOCAL_MAIN | false | Route Main locally |
| USE_LOCAL_HEAVY | false | Legacy migration input only; excluded from active routing |
| USE_LOCAL_LIGHT | false | Route Light locally |
| USE_LOCAL_CONSCIOUSNESS | false | Route Consciousness locally |
| USE_LOCAL_FALLBACK | false | Route fallback locally |
| OUROBOROS_MAX_ROUNDS | 200 | Max task rounds (hot-reloadable) |
| OUROBOROS_TRANSIENT_RETRY_MAX | 6 | Same-model transient retry budget; pre-dispatch `transport_unavailable` is a separate task-bounded outer wait episode, because no provider attempt was admitted on that route |
| OUROBOROS_SKILL_LIFECYCLE_TIMEOUT_SEC | 1800 | Skill lifecycle-lane timeout |
| OUROBOROS_CLAUDEXOR_HARNESS_INSTALL_TIMEOUT_SEC | 300 | Harness install ceiling (kills the tracked group, typed refusal) |
| OUROBOROS_CLAUDEXOR_QUOTA_REFRESH_TIMEOUT_SEC | 90 | Quota-refresh POST ceiling (clamped 1–90) |
| OUROBOROS_BUNDLE_DIR | (unset) | Env-only: launcher-owned bundle root propagated to embedded children for Node/ripgrep discovery |
| OUROBOROS_BG_MAX_ROUNDS | 10 | Background-consciousness round cap |
| OUROBOROS_BG_WAKEUP_MIN | 30 | Consciousness wakeup floor (s) |
| OUROBOROS_BG_WAKEUP_MAX | 7200 | Consciousness wakeup ceiling (s) |
| OUROBOROS_POST_TASK_EVOLUTION | false | Post-task evolution promotion toggle; agent self-enablement is blocked at the shell/browser/settings/data-write guards; choosing an objective routes through the Main slot because it is a high-leverage decision, while execution stays behind ordinary review and owner gates |
| OUROBOROS_POST_TASK_EVOLUTION_CADENCE | llm | Promotion cadence `llm` or `every_n:k` (malformed normalizes to `llm`) |
| OUROBOROS_POST_TASK_EVOLUTION_BUDGET_USD | 0.0 | Remaining-global-budget start floor, not a cycle cap |
| OUROBOROS_EVOLUTION_PERSISTENT_OBJECTIVE | "" | Owner-only persistent campaign bias; still passes review gates |
| LOCAL_MODEL_PORT | 8766 | Local-model server port |
| OUROBOROS_HOST_SERVICE_PORT | 8767 | Host Service port (loopback-only; §12) |
| OUROBOROS_PRESENCE_MAX_ACTIVE | 2 | Cross-process Presence turn cap (UI-bounded 1–20) |
| LOCAL_MODEL_CHAT_FORMAT | "" | Local-model chat template override |
| GITHUB_TOKEN | "" | GitHub token (push/PR/issues) |
| GITHUB_REPO | "" | Personal `origin` repository |
| OUROBOROS_FILE_BROWSER_DEFAULT | "" | File Browser default root (explicit root required for Docker/non-localhost) |
Direct-provider review fallback (legacy name: OpenAI-only review fallback): when exactly one official direct provider is configured, `config.get_review_models()` compiles that provider's declarative reviewer-role sequence using provider-prefixed model IDs. Current scope covers official OpenAI, Anthropic, MiniMax, DeepSeek, Cloud.ru, and GigaChat; OpenRouter, legacy-base, OpenAI-compatible, and mixed-provider configurations stay outside it. OpenAI, Anthropic, and DeepSeek run three independent Main-model slots; MiniMax mixes Main/Light; Cloud.ru and GigaChat use their one role model for every slot. `_exclusive_direct_remote_provider_env` returns empty when OpenRouter, legacy `OPENAI_BASE_URL`, OpenAI-compatible keys, or multiple official direct providers are present, and the fallback requires `provider_models.migrate_model_value` to make the main model already start with the exclusive provider prefix — exact prefix checking prevents an arbitrary free-text model from silently entering a single-provider route. This is part of the single-provider independence invariant (docs/DEVELOPMENT.md "Provider Independence").
DeepSeek provider specifics (`deepseek::`): the official OpenAI-compatible endpoint is a fixed module constant (`provider_models.DEEPSEEK_BASE_URL`); a proxy or mirror belongs to the generic `openai-compatible::` route. The canonical reasoning scale is projected onto the provider's wire dialect at the send boundary (`minimal`→`low`, `medium`/`xhigh`→`high`, `ultra`→`max`, `none`→`extra_body.thinking.type=disabled`; native tiers pass through), a forced tool choice (`required` or a named tool) is served with thinking disabled because thinking mode accepts only `auto`/`none` (probed 2026-09-03), and every tier-changing projection is disclosed on usage; `reasoning_content` stays on canonical assistant turns for strict v4 tool replay with an explicit empty string for turns produced without provider reasoning, while other lanes strip the field and cross-family switches scrub it. System/assistant/tool content arrays are flattened to strings in the send copy only (the API accepts arrays on user turns alone); canonical block history is untouched. Prompt caching is automatic and cost remains nullable when no exact provider catalog is available. The 1M context claim is admitted only through route-fingerprinted capability evidence or owner acknowledgement. Slash-form `deepseek/...` remains OpenRouter; only `deepseek::...` selects the direct route.
GigaChat provider specifics (`gigachat::`): routed through the native `gigachat` library, not OpenAI-compatible (`llm.py::_chat_gigachat`). OpenAI `tools` map to GigaChat `functions`; at most ONE `function_call` returns per turn, so parallel `tool_calls` collapse to the first; role `tool` results become role `function` and must be valid JSON (plain text wrapped as `{"result": ...}`); the `system` message must be first, so later system-reminders demote to `user`. `reasoning_effort` is deliberately omitted — hidden reasoning can consume the whole output budget and return empty content. GigaChat exposes no automatic live cost source, so cost stays nullable/unknown rather than a hand-maintained tariff. GigaChat models sit below the 1M scope-review floor; when no ≥1M reviewer is configured, the declared alternatives are the owner-selected `low` context mode (whole-repo scope review declaredly not performed; each commit records a typed `skipped_low_context_mode` evidence row) or an owner-selected retrieving scope slot at ≥200K sourced evidence (BIBLE P3); the blocking triad still reviews the full staged diff in both modes.
---

View file

@ -0,0 +1,58 @@
# 8. Git Branching, CI, and Build
This chapter owns release topology: the branch and channel model behind the three official update feeds, the CI job graph with its trust boundaries, the keyless system-E2E suite, the build scripts, the dependency-lock authority and the release proof that accepts only the seven expected assets. It exists because every guarantee here is a rule about what a published artifact may claim, not a build convenience.
`ouroboros` is the local working branch; the runtime setting independently selects one official feed: Stable is the newest plain release tag reachable from both `main` and `ouroboros-stable`, QA is the `ouroboros-stable` tip, Development is the `ouroboros` tip. Promotion and rollback are owner-controlled exact-SHA movements; ordinary restart preserves the local tip, while explicit update owns fetch, target validation, rescue, apply, and rollback. `managed` is the official read/update remote; `origin` is optional personal persistence. Desktop and Colab reject a target without a regular non-empty `BIBLE.md`. External pull requests target `ouroboros`, do not allocate a release version, and receive the collision-free version when maintainers land and re-review them (procedure: CONTRIBUTING.md, docs/DEVELOPMENT.md).
The local `ouroboros-stable` ref is also a recovery fallback maintained by explicit promotion; that local role does not select the official QA feed. Launcher metadata is bootstrap provenance only — runtime status, preflight, Colab bootstrap, and apply resolve the selected channel and exact fetched SHA themselves, so an older frozen launcher cannot silently redirect updates. Stable additionally requires the shared plain release tag; QA and Development do not use version comparison as an admission gate.
### CI topology
`.github/workflows/ci.yml` has 16 jobs, grouped by trigger and secret exposure:
| Tier | Jobs | Trust boundary |
|---|---|---|
| Fork-safe PR validation | `quick-test`, `betterleaks-platform-smoke`, `benchmark-methodology` (also on `ouroboros-stable` pushes and `v*` tags) | `ouroboros` pushes, pull requests into `ouroboros`, manual runs; read-only, no provider secrets, never `pull_request_target` |
| Browser consumer proof | `ui-smoke` | PRs run only the real Publish admission/card/history scenario on Chromium; manual runs and tags retain full host UI and browser-tool Chromium/WebKit coverage. No provider secrets or added default local pytest lane. |
| Scheduled lanes | `system-e2e-mock` (keyless; daily 04:37 UTC, manual runs, `v*` tags — on the release bar), `e2e-live` (`OUROBOROS_E2E_LIVE_OPENROUTER_KEY`, `$30` cap) | `e2e-live` fires only on its own nightly cron (03:17 UTC, seeding the `ouroboros` branch tip rather than the default branch the schedule fires on) or a dispatch that opts in through the `e2e_live` input; without the secret it is one summary line and green, never a pretend run |
| Ordinary desktop tests | `full-test` (no secrets) | Every PR into `ouroboros` runs Windows/macOS alongside Ubuntu quick-test; stable pushes, manual runs and `v*` tags retain the full three-OS matrix |
| Trusted skill matrix | `skill-smoke` (`OPENROUTER_API_KEY`) | `ouroboros-stable` pushes, manual runs, `v*` tags — never pull requests |
| Trusted provider run | `integration-test` | provider secrets; `main`/`ouroboros`/`ouroboros-stable` pushes, manual runs and `v*` tags (the release chain needs it) — never pull requests |
| Tag-only gates and release chain | `marker-guards`, `docker-ui-smoke`, `docker-portable-test` (manual runs or tags, no secrets); `release-preflight` (needs `full-test` + `integration-test`) → `build` (signing secrets) → `release` (needs `build`, `release-preflight`, `skill-smoke`, `ui-smoke` and these three smoke gates); `vendor-package-smoke` needs `build` and stays informational | tag-triggered; a reproducible provider-contract failure blocks tag builds, a typed inconclusive provider outage does not |
Quick and full jobs each run a dedicated blocking `size_ratchet` pytest step — the ONLY enforcing surface for the repository size gates (local runs exclude the marker and warn): manifest exactness on the tip plus the pairwise shrink-only transition against the event base in `OURO_SIZE_RATCHET_BASE_REF`; an unresolvable base degrades to the tip's parent manifest verified against the parent's own tree — never a skip — while a resolvable base without a manifest fails closed. Both jobs also run the browser-module suite (`cd web && node --test tests/*.test.js`), the same node lane the hermetic commit gate executes through `ouroboros/preflight_node.py`. Secret-bearing skill review runs before any step that imports downloaded plugin code — untrusted payload code must never share a process with provider credentials — and a missing required key is red rather than skipped.
Shared browser tests use one `ProcessContainer` per `direct_server_with_data` server incarnation. Teardown always reaps and closes it, including normal parent exit and readiness failure; restart receives a fresh token/Job and a cleanup failure fails teardown. A POSIX pytest process killed before `finally` cannot perform that cleanup.
Tool-schema compatibility has two layers. Fork-safe PR tests build the complete shipped catalog without MCP or extensions, validate every schema as general JSON Schema plus the cross-provider subset (no empty enum, no root `anyOf`/`oneOf`/`allOf`), and require the OpenRouter/function, direct-Anthropic, GigaChat, and direct-OpenAI projections to preserve the complete tool-name set. The trusted `integration-test` lane sends that same full registry in one bounded `delegate_start` canary per physical route without executing the returned call — OpenRouter Gemini/Opus/Fable/GPT/Grok/DeepSeek, the three shipped direct-OpenAI defaults (Main alone among them keeps a second-turn nonce-bearing continuation), direct Anthropic, and optional MiniMax/DeepSeek/Cloud.ru/GigaChat (DeepSeek also keeps the continuation because its tool contract is the reasoning_content replay). Each canary requires positive usage, exact provider/model identity, and a normalized schema-valid tool call; quota/billing, 429, 5xx, and timeout outcomes stay typed inconclusive while contract/auth/model/tool/reasoning 4xx are red; one same-route resend is allowed only for a runtime-classified semantic-empty response (bypassing response caches where supported). `response_finish_reason` is retained in host usage for bounded diagnostics only; malformed raw arguments are reported by type/position and hash without copying the provider payload.
`claudexor-platform-gate.yml` proves the managed Claudexor runtime on three OSes: a fixture lane (fake harness, offline, $0) always, and a live lane only on explicit API keys — subscription auth stays deliberately out of CI, because an interactive machine-bound token must not enter CI secrets. `dependency-graph.yml` reads `ouroboros/claudexor_runtime_pin.json` and submits that direct runtime relationship to GitHub's dependency graph (runs only on pin/workflow changes on `main`/`ouroboros`, plus manual dispatch; `contents: write` only) without presenting Claudexor as a Python or Node package dependency. The Scorecard workflow (`.github/workflows/scorecard.yml`) runs on `main` pushes and weekly, pins every action by full commit SHA, defaults permissions to read-only, and adds only `security-events: write` + `id-token: write` for SARIF upload and OpenSSF publication. `CODE_OF_CONDUCT.md` owns community rules; `CITATION.cff` owns the software and preferred technical-report citations; `site/paper/index.html` owns the canonical paper landing page; `docs/benchmarks/evidence.json` holds the release-bound public benchmark projection; README remains the claim SSOT (guarded by `tests/test_trust_metadata.py` and `tests/test_public_site_metadata.py`).
Dev-facing release/review scripts: `scripts/run_external_review.py` (dual-lane external review; `--contributor` binding; `READY_FOR_INTEGRATION` is evidence, never merge authority), `scripts/contributor_review_evidence.py` (route-neutral contributor-packet binding), `scripts/run_plan_review.py` (the same engine as `plan_task`; review-exempt dev tool), `scripts/validate_scope_receipt.py`, `scripts/claudexor_platform_smoke.py`, `scripts/fetch_claudexor_runtime.py` (pin SSOT: `claudexor_runtime_pin.json`), `scripts/cleanup_test_pollution.py` (dry-run-first). `site/` is the Vite source of the public pages (`site/scripts/sync-assets.mjs` syncs assets); `skills/telegram/` and `skills/unix_computer_use/` are the bundled skills; `packaging/cli/` holds the CLI wrappers and installers; `packaging/appimage/` the AppRun dispatch + desktop metadata; `packaging/systemd/` the opt-in user unit (§1 Runtime topology owns its no-restart rationale).
### System E2E suite (`tests/system_e2e/`)
The deep-integration suite drives a REAL `server.py` per scenario — an isolated repo clone, data root and free loopback port through `KeylessIsolatedServer`, whose child environment can never carry a provider credential — and asserts DURABLE artifacts through `ArtifactOracle` readers (`task_results/`, `logs/*.jsonl`, `state/*`), never an HTTP 200 or a harness exit code. Synchronization is durable-event polling (`wait_until` over those readers), never bare sleeps. A scenario that must act MID-ROUND (the owner stopping a turn while the model is still answering) holds the loopback model on an event-gated `ModelGate` — the request thread blocks until the scenario releases it — instead of racing a latency sleep. Two loopback models share one HTTP base, which answers non-streaming requests with JSON and streaming requests with complete SSE tool indices, usage and terminal framing: `ScriptedStubModel`, an ordered per-scenario tool script whose review-organ calls are classified by prompt markers and answered with canned parse-clean verdicts before the finalization-turn check, and `ReplayModel`, deterministic fixtures bound by `(lineage, slot, attempt)` where a miss or an unconsumed row is red. A scenario may hand the stub a `ReviewScript` — ordered per-kind verdict queues that override the canned all-clean answers for exactly the scripted red rounds and then fall back — and review calls never consume agent script steps in either direction. The prompt markers are VERBATIM literals from the tree under test, greppedback out of their source files by a default-lane pin, so drift is a named failure rather than a silently mute stub. Reviewer routing is pinned through the structured `OUROBOROS_REVIEWER_SLOTS`, because the retired comma keys are dropped at load and would fall back to the shipped paid default panel. `FakeClaudexorDaemon` is a loopback daemon imitation serving the exact `gateways/claudexor.py` client contract — the authenticated handshake reporting the TREE'S OWN runtime-pin identity, the capability/quota answers `route_health` reads, idempotency-keyed project registration and run creation with the engine's replay semantics, run detail and cancel — recording every wire request for scenario assertions.
The scenario inventory is DATA (`SCENARIOS` in `harness.py`) with a two-direction pin scanning every `test_*.py` module of the package: a manifest row without a test is red, and a scenario test without a manifest row is red. It covers boot and identity, the review organs (commit triad plus scope in both enforcement classes, plan review, the acceptance loop), egress hardening, typed tool safety, honest cost names, subagent trees, cancellation, its cascade and the owner stop of an in-process direct-chat turn, managed update (fast-forward, carrier transfer, conflicting refusal, crash mid-apply, rollback), delegated transport including the mutating pull-in and its refusal, and the whole skills lifecycle from payload to delete. Scenario tests carry BOTH the `integration` and `serial` markers plus the `OUROBOROS_E2E_DEEP=mock` env gate, so neither the default local run nor either CI pytest pass executes them; the keyless lane is `OUROBOROS_E2E_DEEP=mock pytest tests/system_e2e/ -o addopts=""`, and in CI it is the daily scheduled `system-e2e-mock` job — schedule or manual dispatch only, secret-free — the one lane those three gates leave open.
### Build scripts
`build.sh` (macOS → .dmg), `build_linux.sh` (→ .AppImage + .tar.gz), `scripts/build_appimage.sh`, `scripts/build_linux_packages.sh` (.deb/.rpm/.red80.rpm), `scripts/smoke_linux_packages.sh`, `build_windows.ps1` (→ .zip), and `scripts/build_repo_bundle.py` (writes `repo.bundle` + `repo_bundle_manifest.json`, the manifest `launcher.py` bootstrap validates) are the release-invariant owners. Linux PyInstaller runs under the same pinned portable Python shipped in the payload, so its bundled libpython keeps the payload's glibc floor instead of inheriting the release runner's newer ABI. The AppImage builder wraps that payload with digest-pinned tool and runtime bytes; the native Linux builder wraps the same x86_64 payload without replacing its runtime. Native package metadata declares the external Git required by bootstrap while bundled Python/Node/browser stay under `/opt/ouroboros`; the packages install the opt-in user unit at `/usr/lib/systemd/user/ouroboros.service` with no activation scriptlet. The release-gating package smoke installs through `apt`/`dnf` and proves Git resolution, desktop files, the installed user unit's launcher/cgroup/no-restart contract, the packaged CLI, and a bounded desktop-launcher start on Ubuntu 22.04/Fedora 42; Astra Linux and RED OS vendor-image runs stay informational evidence, because third-party registry availability cannot block publication. The macOS image keeps the app + Applications symlink + optional CLI installer layout. Release tag prerequisite: `scripts/build_repo_bundle.py` is the release-tag SSOT and verifies the annotated `v$(cat VERSION)` tag points at `HEAD` before packaging.
Betterleaks follows the same release-resource discipline: packaged builds stage the pinned binary and license under `betterleaks-standalone`, source checkouts install the exact managed runtime explicitly under `data/state/betterleaks/`, and the Publish path never downloads it (§13).
Python dependency resolution has one authority: direct requirements and group membership live in `pyproject.toml`, `uv.lock` records the universal solution under the pinned `tool.uv.required-version`, and source/CI sync with `--locked` so metadata drift is an error. The two-interpreter packaging boundary rides projections: build scripts export their PyInstaller input directly from `uv.lock`, the committed `requirements-runtime.lock` supplies embedded `python-standalone` and managed updates (pip, no bundled uv), and a one-line `requirements.txt` pointer serves already-released updaters; CI regenerates the export and requires a clean diff, so neither file becomes a second dependency authority.
Platform builds precompile bundled Python with unchecked-hash bytecode: sealing valid bytecode prevents runtime `__pycache__` writes from invalidating a macOS signature, and runtime children route caches outside the bundle. When signing is enabled, hardened runtime, notarization, xattr hygiene, and strict verification remain part of the stable-release path; prerelease artifacts may be unsigned and their evidence must report the actual signing state.
The release proof begins with the final DMG/AppImage/tarball/ZIP, never its staging directory — validating a staging tree does not establish that the published bytes contain or execute the same payload. Each platform shard checks the embedded repository bundle, packaged CLI, and managed Claudexor seed + Node by starting the owned daemon, completing a fixture task, and verifying an identity-bound stop. The AppImage is extracted for metadata/SBOM inspection, then run FUSE-independently to prove version output, CLI dispatch, browser-fallback readiness, payload lifetime, main-executable libraries, and clean shutdown (browser-fallback evidence, not a claim of a native GTK/Qt backend); the nested cleanup proof follows the live `runtime → AppRun custodian → launcher` chain and requires both the extraction and its private base absent. The proven Linux tarball payload is wrapped into the three native packages, each receiving its own digest-bound package-manager smoke receipt and provenance attestation; a digest-pinned Syft build produces CycloneDX inventories from extracted payload bytes, and the tarball inventory is reused for the identical-byte native wrappers while each wrapper keeps its own digest-bound installation proof. The release job accepts only the seven expected assets, recalculates digests, verifies both predicate types, writes the checksum/evidence capsule, and rechecks the remote annotated tag immediately before publication. Signing credentials stay step-scoped and absent from SBOM/attestation steps.
Public installer naming and links ride the same projection: `release_sync.py::RELEASE_ASSET_TEMPLATES` is the filename SSOT shared by the proof builder, README, and the install pages. A version bump rewrites only named download references and `data-release-download` anchors to immutable `/releases/download/v{VERSION}/...` URLs; the `/releases/latest/download/...` shape is forbidden because GitHub excludes prereleases from `latest`. Generated release notes expose direct links only for the seven proof-accepted assets. The default README and GitHub Pages deployment use the stable `main` boundary (`main:/docs` for Pages); stable promotion advances `main` only after the release is published with all seven proof-bound installers, so an unreleased development VERSION never exposes dead installer links and an omitted promotion leaves the previous working release public.
### Docker
Docker runs the web and server runtime without PyWebView. The image binds `0.0.0.0` and sets no network password by default: a missing password only warns, and `NetworkAuthGate` permits requests when no password is configured — publishing the container port without setting `OUROBOROS_NETWORK_PASSWORD` therefore exposes the owner surface. Set the password (or keep the port unpublished); packaging adds no stronger boundary of its own.
---

View file

@ -0,0 +1,18 @@
# 9. Shutdown & Process Cleanup
This chapter owns the ways this installation stops: ordinary window close, which preserves the shared daemon; the owner's manual Restart, which is a clean stop of everything this server generation owns; and Panic, which is a complete explicit stop. It exists because each has a different custody contract — what is signalled, what is only disclosed as unconfirmed, and what the next generation is expected to reconcile — and confusing them is how a live paid run gets killed or a dead one gets reported as stopped.
Ordinary window close preserves this installation's shared Claudexor daemon and work submitted by other clients, on graceful and forced launcher cleanup paths. It stops the server generation's workers/services, waits for its process group or Job, performs owned cleanup and releases the PID lock. The Windows daemon is outside the permitting launcher Job; retained-subtree cleanup also covers worker/server ancestors and the POSIX stray reaper (§1). An old immutable launcher requires a later package update for this guarantee. Same-pin planned restart retains the daemon. A planned restart after a changed engine pin instead invokes `server_restart._stop_owned_daemon_for_new_pin` in lifespan teardown: compare the landed `load_runtime_pin` with the serving authenticated engine through attach-only `read_owned_gateway`, then use the explicit stop below. Unpublished/unreadable pins and unreachable/foreign/stopped daemons do not alter the handoff. Unconfirmed stop retains custody and diagnostics while restart proceeds; the next generation reconciles old runs without inventing spend. In lifespan teardown the ONE irreversible durable write goes first: the stop flag, the bounded supervisor join and then `kill_workers`, which terminalizes every interrupted task, all ahead of the best-effort extension-reconcile and host-service waits and the kill-all sweeps. That order is what keeps the write reachable inside the launcher's force-exit budget; an external SIGKILL can still cut it short, and the bridge and event-bus teardown that follows it can still lose a late frame. Ordinary server shutdown closes its own Host Service listener. Existing reserved-port sweeps remain separate launcher/recovery/Panic behavior; they are not PID-identity proof and can terminate an unrelated listener on a reserved port.
The owner's manual Restart (`/restart`, the chat Restart button) is a clean stop of everything the current server generation owns, then the re-exec. The bridge handler keeps its checkout-first order: `_safe_restart_serialized` (the update lock and strict managed-update transaction gate, then `safe_restart`'s checkout, dependency sync, import test and stable fallback) runs before anything is stopped, so every refusal — an assisted-update merge mid-resolution ("restart was deferred"), a failed checkout, an unwritable no-resume flag — leaves the server intact and is answered with one "Restart cancelled" line. Past that point the restart always follows: the durable no-resume intent (`state/owner_restart_no_resume.flag` plus the `panic_stop.flag` compatibility pair, consumed by `auto_resume_after_restart`), then `server_restart._stop_owned_work` — one durable cancel intent per RUNNING task, live direct activity and in-flight post-task synthesis (`cancel_intents.request_cancel`, source `owner_restart`, immediate), `kill_workers(force, cancelled)` with Panic's `reconcile_delegate_custody=False`, delegated-run cancellation through the public owner-gone seam `reconcile_orphaned_runs(running_task_ids=set())` over `read_owned_gateway` (attach only: nothing between the cancel intents and the daemon stop may `ensure_owned_gateway`, which would start a dead daemon), and the attested owned-daemon `stop_outcome()` exactly as Panic makes it — then the owner's "Stopping active task" notice and the exit. Nothing there is a veto or a deferral: an unconfirmed or raising worker shutdown and an unconfirmed daemon stop each leave a critical diagnostic (the daemon's own `process_stop_unconfirmed` supervisor row, or the same row when the stop raises), custody stays retained, and the restart proceeds; a stop with nothing to stop is quiet. The remainder is the next generation's existing startup work: owner restart is a no-resume cause (`delegate_recovery.NO_RESUME_CAUSES`), so no handoff is prepared and nothing is adopted; the startup custody sweep reconciles every open delegated run as owner-gone against the daemon it then reaches (a run the stopped daemon took down answers absent → `close_absent_run`, no invented spend; a journal-recovered terminal settles; a pending invocation replays under its own key), the boot backfill heals stored disclosures, and an in-flight owned model operation stays a `dispatched`/`unresolved` physical-attempt row that already counts as spend upper bound. After an unconfirmed stop the re-exec'd generation attaches to the still-live daemon rather than spawning a second one, and the lifespan's one background `warm_owned_daemon()` (provisioned homes only) makes the first delegation after a Restart find the daemon serving. Planned self-restart, managed-update handoff and Panic keep their separate contracts; the planned restarts share only the attested daemon stop, and only when the landed checkout pins another engine (above). The explicit stop is installation-wide: manual Restart ends runs served by the owned daemon, including another client's runs. This differs intentionally from ordinary window close.
Panic is a complete explicit owner stop, not a restart: stop consciousness, record evolution owner-stop state, close the campaign and pending promotion, write `panic_stop.flag`, stop local models and call `OwnedClaudexorDaemon.stop_outcome` before worker cleanup. It then ends tracked foreground commands, executor processes, services, companions and workers before hard server exit. The shared daemon may belong to a previous generation and is outside the Windows launcher Job; its explicit stop, not closing that Job, owns termination. Failure emits a critical `process_stop_unconfirmed` row and retains unresolved custody while the rest of Panic continues. The next launch suppresses automatic work under the panic/no-resume flags.
`stop_outcome` returns `stopped`, `nothing_to_stop`, or disclosed `unconfirmed`. With a valid owned marker and authenticated same-home endpoint it first invokes the already installed managed CLI `daemon stop --json`, using read-only `resolve_cli_command(require_npm=False)` and an owned config root with ambient socket/port overrides removed. Operator shutdown needs the existing exact Node, not the npm installation toolchain; the default installer-facing resolver keeps that separate requirement. It never ensures a runtime, wakes a daemon or probes accounts. Exit0 plus the typed terminal CLI receipt is required; an RPC acknowledgement or clean lease release alone does not prove captured OS processes have exited. Known roots/children are observed through the exit window, and a surviving authenticated endpoint is disclosed without chasing a successor. The pre-request custody snapshot limits forced fallback to those original rows; forced signalling still requires both measured birth and command identity or the manager's own Popen. Legacy Windows empty-birth rows cannot authorize signals or be re-attested from a later PID observation, but healthy attached daemons can stop cooperatively through the existing CLI. Missing CLI or failed/unavailable shutdown retains the previous independently proven fallback. Token refusal, invalid discovery, incompatible/malformed responses and a matching error-code string alone never provide that authority; a typed HTTP transport failure or positively absent descriptor for marked startup retains its existing meaning. Only confirmed death permits signal-stop custody removal; concurrent rows survive.
The short manager lock retires the local startup generation before waits and never spans runtime preparation, network or process-exit work. CLI RPC, termination confirmation and startup slack use `config.CLAUDEXOR_OPERATOR_STOP_TIMEOUT_SEC`; captured-process observation uses `config.CLAUDEXOR_STOP_EXIT_WAIT_SEC`. These existing constants live in `runtime_limits` and are re-exported by config; each signalled ledger root retains its exit window. These phases do not promise an absolute total stop duration. Startup joins purpose-filtered custody before and after preparation; engine writer election remains authoritative. A live startup outlives its caller's window (20 seconds by default, independently narrowable): `daemon_starting` invites joining the same work, whereas `daemon_spawn_failed` identifies the exited child/build/log interval. The authenticated control handshake and the independent normal-admission wait remain distinct. Native Windows Job/identity/stop tests and frozen/embedded psutil imports require Windows/package evidence; portable fixture success alone does not establish that proof.
**Startup failure: typed classification, one durable row, the spawn latch (#844).** When this manager's own child exits without a control endpoint, `claudexor_startup_failure.py` classifies the failure from the child's own recorded `daemon.log` interval (identity-checked, bounded tail read, never an older generation's tail) into `heap_exhausted` (V8's fatal heap-limit banner: the observed spellings and the general form, one line carrying both `FATAL ERROR` and `Allocation failed`), `writer_lease_contended` (the engine's single-writer refusal), `engine_floor` (the root-authority version-floor refusal) or `unclassified`. The classification is diagnostic only: it names the failure in the typed `daemon_spawn_failed` detail, in `last_error` and in one durable `claudexor_daemon_start_failed` supervisor row (pin version/build, exit code or signal, whether a descriptor was written, classification, log path and interval offsets, `latched` = the exit was of the latching class, `ts` = when the exit was first observed), and it never selects behaviour. The latch keys on the typed exit fact alone: the child exited by signal or non-zero and wrote no control descriptor during that spawn (descriptor mtime/size identity recorded at spawn, compared when the exit is first observed); a clean exit or a crash after publishing control is rowed with `latched=false`, and a joined peer startup that vanished is not rowed. While the latch is set, every ordinary caller (delegation, model calls, review, warmup) receives the same immediate `daemon_spawn_failed` refusal, before runtime preparation and without a new spawn — unless a live startup already exists, which it joins without spawning (refusing there would hide a daemon that may come up). Releases: a live authenticated daemon is still attached to, and the attach clears the latch (`claudexor_daemon_start_latch_cleared`, `cleared_by=live_daemon_attached`; an exit the attach itself settles is latched only for the duration of that settle, and both rows are kept); the existing ~600 s periodic supervisor sweep clears it (`cleared_by=supervisor_sweep`) and, only when it actually released one, makes exactly one zero-wait ensure on a short-lived background thread — no admission wait, no startup wait, the supervisor tick never waits on runtime preparation; `daemon_starting` is the expected answer and a refusal re-latches — so a healthy install never spawns or waits in the sweep and a persistently crashing engine costs one spawn per sweep period; the owner's explicit Refresh (`/api/claudexor/wake`) clears it (`cleared_by=owner_wake`) and makes its one ordinary ensure through the same funnel, so each press on a still-crashing engine costs one spawn, an explicit owner action; the owner's Restart or Panic constructs a new manager and therefore clears it. The latch is per-manager state and the manager is a process global: the sweep and the Refresh release the server process's latch; a task worker process holds its own, learned from its own single failed spawn and released only by a live-daemon attach or the worker's respawn. There is no backoff, counter, cooldown, heap policy or new timing constant: the host neither derives the engine's heap from host RAM (V8's default old space is 4096 MiB on any machine with ≥ 8 GiB, so a guessed size only moves the crash; the engine's bounded replay removes the need) nor touches the writer lease. `NODE_OPTIONS` set on the server process is inherited by the spawned daemon unchanged — the operator's documented escape hatch (for example `NODE_OPTIONS=--max-old-space-size=8192`), not a product mechanism. The harvest cadence and the two disclosed residuals are stated once, in DEVELOPMENT.md's Process Custody Rule.
Bounded foreground `run_command`/`run_script` processes ride the in-memory `_active_subprocesses` panic registry; long-lived children — services, executor-backed processes, extension companions, delegated runtimes — enter durable exact-identity process custody via `spawn_supervised` (§1 Runtime topology). Unix process groups and Windows Job Objects provide tree cleanup; durable executor and service records let the host recover after worker death. Normal cleanup may archive logs, while panic skips nonessential finalization — agent-controlled or wedged cleanup must not delay an owner stop. Timeout and signal exits remain distinct in tool results so a killed command never resembles success.

View file

@ -0,0 +1,81 @@
# 10. Key Invariants
This chapter is the short list of properties the rest of the book must not contradict: the constitution persists, release metadata has one projection, the attempt ledger is the monetary authority, cancellation is intent-then-custody, and a dozen more, each naming its owner. It also carries the continuity data-flow map that states, per surface, the canonical source, the bounded projection over it and the rule deciding when a consumer may act. It exists so a change can be checked against a numbered claim instead of an impression.
1. **Constitution and identity persist.** `BIBLE.md` is never deleted; `identity.md` remains a physical file even when its content evolves.
2. **Release metadata has one projection.** `VERSION` is canonical; `ouroboros/tools/release_sync.py::version_carrier_desyncs()` and `sync_release_metadata()` keep the PEP 440 form in `pyproject.toml` and the editable root entry in `uv.lock`, plus the author-facing version in `web/package.json` and both root entries of `web/package-lock.json`, `web/modules/api_types.js::GATEWAY_CONTRACT_VERSION`, the README badge, and this document's header. Changelog prose remains deliberate. Pull requests into `ouroboros` leave these carriers byte-identical to their target; integration assigns the release version.
3. **Configuration and messaging have single owners.** `ouroboros/config.py` is the one IMPORT surface for paths and settings, and the vocabularies live in its leaves — `settings_defaults` (keys, shipped defaults and the retired-key lists), `settings_scales` (the closed clamps), `model_slots` (slot resolution and the frozen model target), `review_model_routes` (the API-pinned reviewer lists), `runtime_limits` (numeric knobs and their clamps) and `settings_integrity` (the verified read primitive) — so a new key belongs to a leaf, never to the facade; messages go through `supervisor/message_bus.py`; concurrent state transitions use the owning file lock.
4. **The attempt ledger is monetary authority.** `state/usage_attempts.jsonl` records every physical model send. State, task, event, and UI totals are projections carrying attempt identity; unknown or unresolved cost never becomes false zero.
5. **Packaged bootstrap is manifest-bound.** A packaged install verifies `repo.bundle` and its manifest once, then runs the managed checkout. Restart preserves its local tip; only explicit update applies an approved exact SHA.
6. **Shutdown is custody-complete** (§9). Normal close verifies child death; panic stops all owned work — the installation's daemon by confirmed ledger identity — without allowing agent code to delay it; the unledgered-daemon and Windows residuals are disclosed.
7. **This document is the present-tense map.** Structural owners, APIs, durable data, UI surfaces, and the rationale for non-obvious guards update here in the same commit as the code (documentation contract: docs/DEVELOPMENT.md; residue ratchet: `tests/test_docs_sync.py`); release chronology lives in git and README.
8. **Skill gates do not collapse.** Discovery, deterministic preflight, content-hash-bound executable review, owner grants, dependency readiness, enablement, and execution remain separate. A PASS does not install dependencies, and `enabled=true` does not prove executable readiness.
9. **Startup rescue has one mutation owner.** Supervisor recovery writes rescue evidence before reset or blocks while preserving the tree. Worker or agent construction remains warning-only and never stages or commits inherited dirt.
10. **Projection over replay.** Interactive status, history, and cost reads are bounded, non-materializing projections; durable owners perform the one authoritative replay or terminal materialization.
11. **UI resources carry a disposer.** Every subscription, listener, observer, timer, stream, and live page instance has explicit teardown; navigation does not leave hidden instances mutating visible or durable state.
12. **Frozen contracts extend explicitly.** `ouroboros/contracts/` is a versioned, backward-compatible ABI — typed shapes together with their parsing/normalization/policy helpers (§11). New capability extends the frozen shape or ships an explicitly versioned successor; existing consumers keep working.
13. **Provider wire adaptation stays exact-route and success-confirmed.** Canonical history remains provider-neutral; typed physical projections may change values, fields, or a registered dialect on one provider/endpoint/API/model only. Failed candidates teach nothing durable, task-local cognition degradation never becomes future dispatch authority, and the physical-attempt ledger remains distinct from terminal request-wire history.
14. **Cancellation is intent-then-custody.** A cancel intent never rides the canonical task status; no teardown begins before the durable, watchdog-replayable intent row exists, and a failed intent write refuses the cancel typed; the settle owner takes an exclusive claim BEFORE any custody mutation and every mutation is fenced by the claim generation (a probed-ALIVE claimant is never abandoned); a settled RESULT does not prove a dead WORKER, so live ownership gates every settled-target cancel while the worker-side snapshot check fails OPEN toward liveness; natural completion wins a late cancel and the stored terminal row survives byte-identical — the kill is about the process, never the result; the deliverable is durably registered as OWED before the intent settles, and a registration failure leaves the intent open for the watchdog; a cascade acts on lineage ROOTS (one cascade, one summary per tree) and settles only on its no-live postcondition. Owners: `task_lifecycle.py`, `cancel_intents.py`, `cancel_publication.py`, `terminal_delivery.py` (§5 for the flow).
15. **Cancellation registry corruption stays visible.** In `state/cancel_intents.json` and `state/terminal_deliveries.json` an absent file reads as empty; enforcement reads disclose a malformed file or row and keep its on-disk bytes, and no mutation rebuilds an existing malformed store from a `{}` collapse. The shared `utils.read_json_dict` helper answers `None` for absence, unreadable or invalid JSON, and non-object JSON alike, and `_read_evolution_campaign()` collapses all of those to `{}`.
16. **Verification receipts reconcile by one typed identity key.** Sameness is equality of the single (kind, value) key — never a match across kinds — so reconciliation is an equivalence that fails SAFE toward strictly fewer reconciliations; disclosed parts are never the comparison (`_outcome_receipts.py`).
17. **Review spend has one ceiling.** Every paid review gate shares `OUROBOROS_REVIEW_MAX_CYCLES`; the per-gate meanings live once in §6 Review stack, the SSOT in `review_cycles.py`.
### 10.1 Continuity data-flow map
The table below is the canonical map for continuity changes. A bounded view is
an interface projection, never a new authority. The actor that makes the
decision must be able to resolve the named source through an existing reader;
otherwise the view is partial and the consumer remains non-final or abstains.
| Surface | Canonical source and owner | Bounded projection | Actor-readable source/ref | Decision and retention rule |
|---|---|---|---|---|
| Owner authority and biography | Canonical `logs/chat.jsonl`, archive generations, and `memory/dialogue_blocks.json` owned by the canonical drive | Main/Project context sections and archive-aware history windows | Existing `chat_history`/archive readers with generation and gap metadata | A known gap is disclosed; summaries/blocks never replace exact current owner directives. Raw generations and durable blocks follow their existing retention owner. |
| Shared understanding and knowledge summaries | `memory/knowledge/overview.md` and each note's authored YAML `summary`, owned by the canonical drive | The resident `## Shared understanding` section and the summary line of every knowledge-index row | `knowledge_read(topic=..., scope='global')`; `knowledge_list` | A missing overview renders as a visible gap line, never a silent omission; summaries stay resident whether or not an overview exists; a body-only rewrite keeps the previous summary |
| Execution evidence | Task results, observability call manifests/blobs, service logs, and process-custody records | Status cards, terminal rows, bounded tails, and compact child summaries | Exact artifact/blob/service-log refs carried by the task result or canonical promotion | A projection cannot certify a missing child/source. Referenced canonical artifacts are promoted before child-drive GC; disposable execution scratch follows unified GC. An omitted-to-artifact verification ledger stub carries only its re-projected `summary`; entries and axes are read from the artifact file it points at. |
| Terminal task/project memory | Root terminal result plus existing task/project summary producers | Cognitive Main terminal summaries and the two Project-root UI lifecycle rows (started + terminal completion) | Task-result ID, project binding, and summary/source refs | Summary is a biography projection, not raw evidence. Terminal outcomes, including failed/cancelled/degraded, remain retained through their canonical result owner. |
| Background Consciousness observations | `data/state/consciousness_observations.jsonl`, append-only enqueue/ACK rows owned by `BackgroundConsciousness` | Pending count/oldest metadata and a bounded recent observation rendering | `read_file(root='runtime_data', path='state/consciousness_observations.jsonl')` | Unacknowledged rows survive restart/overflow/error. Gaps block ACK and the existing direct identity rewrite; only a settled successful cycle appends ACK. |
| Plan/review authority | Exact task-artifact/observability wave bodies, evidence selectors, reviewer route/thread receipts, and the bounded review hot index | Review status, latest wave, obligations, and compact findings; a predecessor's inherited `plan_review_state` is first projected to a compact authority core ordered around the newest wave's identity, acceptance claims, findings, and dispositions, with reviewer transport removed and `need_evidence_seen` last-priority. Every bounded collection names its total and omitted count; the projection discloses `full_chars` plus `source_ref`, and the named `include_authority` source stays complete | Exact artifact/source handle plus SHA/range/thread selectors | Missing or partial evidence is `DEGRADED`/`NOT_RUN`, never PASS. Exact artifacts remain bound to the reviewed candidate SHA; hot indexes may rotate only after the source is retained. |
| Task acceptance (three deliveries) | The FULL host packet (`review_evidence.build_task_acceptance_evidence` under the host ladder, with its `__provenance__` table), the applied host run retained through canonical task source handles, and the paid-identity wallet ledger | The per-delivery work order: the api pack for a packet row; the FULL packet plus absolute pointers and the access disclosure for an agent-session row; the packet without its freely degradable tail plus the real data root for a native inspection row (`loop_acceptance_review.acceptance_retrieving_work_order`) | Exact `evidence_refs` from the packet's enumerable exhibit vocabulary; absolute pointers to the task's active workspace, task result record, artifact directory, verification receipts and tool-trajectory log; `review_projection.panels[].applied_source_ref` for the complete redacted applied review | Refs resolve against the FULL packet only, never the rendered projection; a session's reads are unobserved by the host (disclosed), a native episode's are `host_observed`; the immutable-core overflow refuses every delivery, a partial tool-result projection only packet rows; one strict wallet claim per panel whatever the rows' deliveries (owner R11). |
| Canonical versus execution roots | Canonical budget/data root owns identity, authority, biography, results, and promoted observability; execution drives own tools, workspace, transient trajectory, and per-call manifests while a task runs | Project/fork/task lenses and status projections | Existing canonical-root resolver, task-result pointers, and source handles | A fork is an execution lens, not a second mind. Copy-back/promotion precedes GC for anything referenced by a canonical result; before terminal promotion the canonical reader cannot resolve a ref bound to a child drive (tracked as #805), and missing legacy bytes become an explicit gap. |
---
Acceptance source identity is computed before history-dependent packet budgeting.
The complete receipt/tool sources and work artifacts still invalidate the binding
when their facts change. Applied reviews and completion observations use the existing
write-once `source_handles/context_checkpoints` store and verified `task_source` refs;
source handles stay outside both deliverables and the acceptance artifact manifest.
The existing artifact route selects a published ref through its `source` query and
returns digest-verified bytes. Ordinary artifact downloads retain their existing path.
`api_client.taskSourceDownloadUrl` owns the shared browser URL contract.
The final packet size includes source references and omission notes.
Individual materialized tool records have addresses `tool_trajectory:<corpus-sha>[<source-index>]`,
not positions in the moving tail. `task_acceptance_review(evidence={tool_trajectory_indices: [...]})`
selects earlier records from the retained corpus through the existing source reader.
The full trajectory remains partial when its head is omitted; a selected record
resolves independently only when its actual arguments and result survive final
packet budgeting complete. Selection changes the view, not source revision or
authorship: agent-supplied prose never becomes host evidence, and a readable
source does not prove the reviewer read or understood it.
Child copy-back uses `observability._rewrite_child_ref_tree` for typed refs in
the owned acceptance-checkpoint and trajectory JSON formats. It preserves the
complete source handle and rechecks dependencies even when the outer checkpoint
was copied before. Relative source addresses preserve immutable checkpoint bytes;
rebased observability refs require a newly addressed checkpoint. A failed copy
retains pending custody before cleanup; missing legacy bytes remain unavailable.
If rebasing changes trajectory bytes, the existing handle retains the original
`corpus_sha256` for record citations while `sha256` verifies the transported copy.
Existing row/reviewer addresses therefore survive cleanup without rewriting claims.
Normal copyback selects CURRENT review authority through the existing field reducer, prepares referenced bytes outside the result lock, and publishes only while the selected ref/binding basis still matches. A changed basis repeats preparation outside the lock; unrelated newer fields survive. Pending retry starts from CURRENT, not a stale child body. `child_ref_promotion_scope` memoizes verified work only for one operation; failures are not cached as success. Same-physical-store copies keep the original manifest bytes/digest and return canonical path spelling without adding `promoted_call_manifest`; distinct-root copies keep their existing provenance marker and filename. Missing aliases resolve only to the exact canonical CAS/call address through shared verified readers, including the model-send reverse reader. Existing corrupt or wrong-scope bytes never trigger a convenient fallback. No arbitrary JSON-path crawl, new manifest filename format or persistent copyback ledger is introduced.
`review_projection.publish_acceptance_checkpoint` saves full applied host records
before updating the compact task-result field through `write_task_result` and
emitting the existing `review_reference` invalidation with the terminal task's explicit
chat id (including hidden chat 0); loop/plan callers retain their context default and
Project binding retains addressing precedence. Actual task attempts and
host-only publication revisions order snapshots of each panel; the common merge
also covers effective child reads and copy-back. Neither the projection nor its
ordering stamp grants review authority. Old records without a retained full source
and failed source writes disclose `applied_source_status="unavailable"`.

View file

@ -0,0 +1,93 @@
# 11. Frozen Contracts v1 (`ouroboros/contracts/`)
This chapter owns the ABI promise: which typed shapes and their parsing, normalization and policy helpers are frozen, what is deliberately not frozen, how to extend one of them, and which retirements were taken as deliberate breaks. It exists because skills, extensions and the browser all depend on these shapes, so removing a field is an announced break with a versioned successor rather than an ordinary edit.
`ouroboros/contracts/` is the frozen ABI package for the skill/extension layer: typed protocols and shapes TOGETHER with their parsing, normalization, capability, and policy helpers (`task_contract.py`, `skill_manifest.py`, `plugin_api.py`, `skill_payload_policy.py`, `chat_id_policy.py`, `schema_versions.py`, `tool_context.py`, `tool_abi.py`). The frozen property is backward compatibility, not absence of behaviour: shapes and helper semantics stay stable for existing consumers, and the protocols are verified against the real implementations by `tests/test_contracts.py`. The browser-envelope ABI is additive and lives in `ouroboros/gateway/contracts.py` with the JSDoc mirror `web/modules/api_types.js`, pinned by the contract/parity suites — not a second file in this package.
### 11.1 What is frozen
| Contract | File | Anchored by |
|---|---|---|
| Claudexor login/status envelopes — `ClaudexorLoginJobResponse` (required top-level `job`) and `ClaudexorLoginJobProblem` (required `error`, optional `code`, bounded `required_actions`); the daemon envelope passes through verbatim | `ouroboros/gateway/contracts.py` + `web/modules/api_types.js` | `tests/test_gateway_parity.py`, `tests/test_claudexor_owned_daemon.py` |
| `ToolContextProtocol` — the minimum context surface tools may rely on | `ouroboros/contracts/tool_context.py` | `tests/test_contracts.py` (duck + AST checks) |
| Tool module ABI — `ToolEntryProtocol` + `GetToolsProtocol`; every registry entry satisfies it | `ouroboros/contracts/tool_abi.py` | `tests/test_contracts.py` |
| Gateway browser envelopes (`gateway/contracts.py`; ABI 7.0 removed the `contracts/api_v1` re-export and the `cost_usd*`/`telegram_chat_id`/`project_last_viewed`/`project_hidden` compat aliases) — browser WS/HTTP envelope families and `TaskCreateRequest` (optional caller metadata; `executor_ref` host-owned; costs nullable — whole-object omission over a confident `$0`) | `ouroboros/gateway/contracts.py` + `web/modules/api_types.js` | `tests/test_contracts.py`, `tests/test_gateway_parity.py`, `tests/test_gateway_abi3_removals.py` |
| Provider Test gateway ABI — `ProviderTestRequest` is exactly `{provider_id, overrides?}` and `ProviderTestResponse` exactly `{ok, error?}` (allowlisted request-local overrides; bounded errors) | `ouroboros/gateway/contracts.py` | `tests/test_gateway_parity.py`, `tests/test_provider_key_test.py`, `web/tests/provider_test.test.js` |
| Presence settings card + CAS update (reviewed defaults, local overrides, state fingerprint; CAS touches only presence-profile state) | `ouroboros/gateway/presence_settings.py` | `tests/test_extensions_api.py`, `tests/test_gateway_parity.py` |
| `client_surface` — optional closed-key normalized/bounded client descriptor, host-stamped, propagated to task metadata and chat history; absence is an explicit honest gap | `ouroboros/client_surface.py` | `tests/test_contracts.py`, `tests/test_gateway_parity.py` |
| `cancelable` + cancel-response `cascade` — additive fields with host-attested UI gating semantics | `ouroboros/gateway/contracts.py` | `tests/test_gateway_parity.py`, cancel/history tests |
| Executor-route projection — an opaque dispatch decision distinct from execution evidence; empty means native/no chip; a BLOCKED harness pin projects the route it named onto the live frame from cap_info alone (`executor_blocked_route`; the durable record keeps its empty route so the completion-seam evidence gate stays closed) and its typed `subagent_executor_unavailable` terminal renders `{harness} · blocked`; the sticky renderer is `log_events.js` `executorChip` | `ouroboros/agent.py`, `ouroboros/subagent_dispatch_notes.py`, `ouroboros/subagents.py`, `ouroboros/gateway/history.py`, `web/modules/log_events.js` | `tests/test_claudexor_owned_daemon.py`, `web/tests/review_truth.test.js` |
| `ChatOutbound.executor_observation` — optional event-local task/attempt/run/harness/phase/revision facts; model requires explicit requested/observed provenance, with only requested-model production from the current live timeline. Last activity is neither current liveness nor terminal evidence; foreign/older observations cannot replace their owning facts | `ouroboros/delegate_progress.py`, `ouroboros/subagent_messages.py`, `ouroboros/agent.py`, `supervisor/events_chat_delivery.py`, `ouroboros/gateway/history.py`, `ouroboros/gateway/contracts.py`, `web/modules/api_types.js`, `web/modules/log_events.js` | `tests/test_executor_observation.py`, `web/tests/wire_contract.test.js`, `web/tests/review_truth.test.js` |
| `execution_evidence` — started/settled/succeeded/failed counts, `delegated_run_failure_states`, `evidence_read_failed`, `nanny_nudge_recorded`, `subscription_cost_usd` (None while undisclosed — never 0), `subscription_cost_estimated`, `harness_models`, `applied_access_profiles`; derived from durable custody rows by `delegate_evidence.task_execution_evidence`, attached in `subagents.envelope_from_task` at terminal statuses only (never overwriting `effective_executor`/`executor_route`), and enriched onto the pushed `task_done` frame by `enrich_task_done_event` in `supervisor/subagent_task_truth.py`; `actual_substrate` ∈ harness_used/harness_attempted/native_only from custody evidence only; `substrate_result_fields` = {actual_substrate, delegated_runs_started, delegated_runs_settled, delegated_runs_succeeded, delegated_runs_failed, delegated_runs_source_unresolved, native_contribution}; the `wait_tasks` compact projection carries `dispatch_executor` plus that same set; an unreadable custody log omits the substrate claim and counts everywhere — `evidence_read_failed` means UNKNOWN, never "no run", and absence of evidence on a pre-evidence stored result is never a zero-run receipt; `native_contribution` is the constant "unknown" (no share/ratio is derivable from custody rows); a verifiable `native_only` amends `capability_delta` with `delegated_substrate_unused`; the `log_events.js` executor chip renders layered truth with unverified work-order counts spelled out | `ouroboros/delegate_evidence.py`, `ouroboros/subagents.py`, `supervisor/subagent_task_truth.py`, `web/modules/log_events.js` | `tests/test_execution_evidence.py`, `tests/test_terminal_delegation_receipt.py`, `web/tests/review_truth.test.js`, `tests/test_task_status_flow.py` |
| `TaskCostBreakdown` (root-only, read-time, never persisted; `accounted_upper_bound_usd`; `authority="physical_attempt_ledger"`) + `cancel_state: "pending"` with `cancel_reason` beside it; the browser's one consumer is `log_events.js` `taskCancelPending` | `ouroboros/gateway/contracts.py` | `tests/test_gateway_parity.py`, `web/tests/cancel_run.test.js` |
| Task hurry ABI — `POST /api/tasks/{task_id}/hurry` with exactly `{request_id}` (extra fields refused), `duplicate` = idempotent success; `OwnerHurryProjection` attempt-keyed states; consumers `log_events.js` `taskSoftStopPending`/`ownerHurryProjection` (task-card only) | `ouroboros/gateway/task_hurry.py`, `ouroboros/gateway/contracts.py` | `tests/test_owner_hurry_s3.py`, `tests/test_owner_stop_s3.py`, `web/tests/task_control_menu.test.js` |
| `StateResponse.active_direct_turns`/`active_chat_activities` (phases queued/working/finalizing/budget_paused — one predicate `budget_pause_fact` decides budget pause); `TypingOutbound` activity fields; `ChatOutbound.task_phase`/`task_terminal_status` | `ouroboros/gateway/contracts.py`, `supervisor/active_activity.py`, `web/modules/chat_activity.js` | `tests/test_gateway_parity.py` + the activity test files |
| `project_thread` stamp on all seven outbound frame types, stamped at the message-bus broadcast choke; a stamped frame is never adopted by Main (`chat_activity.mainThreadAccepts`) | `supervisor/message_bus.py`, `ouroboros/projects_registry.py` | `tests/test_message_bus.py`, `web/tests/chat_thread_routing.test.js` |
| Media/link envelopes — media `task_id`/`size_bytes`/`download_url`; `LinkAction {label,url}` with at most twelve absolute HTTP(S) actions; `links` in `WS_MESSAGE_TYPES`; `chat.links` host topic | `ouroboros/gateway/contracts.py`, `ouroboros/tools/core.py`, `ouroboros/event_bus.py` | `tests/test_contracts.py` |
| Owner quiz ABI — `QuizOption {label, detail?, recommended?}` (the asker marks its recommendation on that option; the web card badges it, Telegram stars its button, the durable block keeps `recommended_index`), `QuizOutbound` (quiz_id, question, options, stake, `assumption` (required for optional clarification), additive `wait_for_answer` for a live pooled or ordinary-conversation root that must wait, lifecycle state open/answered/expired_terminal/superseded), separate `QuizStateOutbound` discriminator, `chat.quiz` host topic; the producer is the one escalation verb `escalate(question, options, stake, assumption, wait_for_answer=False)` — a ROOT asks the owner, a SUBAGENT delivers a typed frame to its nearest LIVE ancestor, which answers via `forward_to_worker` or escalates verbatim, so the owner sees only what no ancestor answered; answers arrive through the ONE ingress `POST /api/decisions` (family ids `quiz:{task_id}:{quiz_id}`, `routing:{client_message_id}:{routing_token}`; `interaction:` reserved), request-id idempotent, first answer wins, validated against the STORED options; `option_index` is optional for the quiz family alone — a comment-only answer writes NO `answered_index`, because a stored 0 would replay as "chose the first option"; injected as the typed `KIND_QUIZ_ANSWER` mailbox control and broadcast as `quiz_state` (carrying the recorded `comment` when the owner answered in their own words, so the live card shows `Owner's answer:` exactly as replay does); expiry is structural only (the task-done seam flips open quizzes to `expired_terminal`, and the SAME reconcile closes the paired `owner_wait` so a terminal task never projects `quiz=expired_terminal` beside `owner_wait=waiting`; `owner_wait.set_owner_wait`'s refusal to continue waiting on a terminal result is preserved, not caught) and history replay merges the projection state | `ouroboros/gateway/contracts.py`, `ouroboros/gateway/task_decision.py`, `ouroboros/owner_quiz.py`, `ouroboros/tools/core.py` | `tests/test_gateway_parity.py`, `tests/test_quiz_display.py`, `tests/test_quiz_answer.py`, `web/tests/chat_decision.test.js` |
| Managed update ABI — preflight, `UpdateMergePlan`, pinned apply, `update_status_ready` WS notice | `ouroboros/gateway/contracts.py` | `tests/test_update_apply_routing.py` |
| `ChatOutbound.review_projection` — bounded actor findings via `utils.truncate_review_artifact`, at most `MAX_PROJECTED_ACTOR_FINDINGS` rows (`review_execution_projection.py`) | `ouroboros/gateway/contracts.py` | `tests/test_review_substrate_v2.py`, `web/tests/review_truth.test.js` |
| Skill preflight statuses — `preflight_failed` is fresh-only; a stale failure surfaces as `preflight_failed_stale`; absence means the caller could not know | `ouroboros/skill_review_status.py` | `tests/test_skill_preflight_repair.py`, `web/tests/skill_preflight_repair.test.js` |
| `chat_id_policy` — the SSOT for human-visible vs synthetic chat ids across message bus, history, memory, and consolidation (§12) | `ouroboros/contracts/chat_id_policy.py` | `tests/test_chat_id_policy.py` |
| `task_contract` — normalization + `effective_acceptance_claims(task, closed_plan_wave)`, a pure read-time binder where ingress wins over a closed plan wave (acceptance reads fresh claims without mutating the live contract); an open wave binds nothing and is disclosed as a non-binding exhibit instead; pacing interprets the budget profile via typed `task_pacing.CostCeiling`; Presence promotion and follow-ups copy the ceiling by value (§12) | `ouroboros/contracts/task_contract.py` | `tests/test_contracts.py` |
| `PluginAPI` 2.0 — the full 16-method extension surface, negotiated against the manifest (major strict, minor minimum, closed-set capabilities, typed educational refusals) BEFORE plugin import or out-of-process cataloging; an absent manifest field means legacy `1.3` by construction, `ExtensionRegistrationError`, `FORBIDDEN_EXTENSION_SETTINGS`, `VALID_EXTENSION_PERMISSIONS`, `VALID_EXTENSION_ROUTE_METHODS` (route methods mirrored against server dispatch by `test_extension_route_methods_contract_matches_server_dispatch`); `skill_job_dir(job_id)` creates `jobs/<sanitized>-<hash>/{assets,output,tmp}`; host-mediated permissions (`companion_process`, `supervised_task`, `subscribe_event`, `inject_chat`, `presence`) require review/owner grants; the `ExecutionMode` capability matrix is the SSOT for what a per-call child can proxy | `ouroboros/contracts/plugin_api.py` | `tests/test_contracts.py`, `tests/test_extension_loader.py` |
| `SkillManifest` — unified frontmatter (instruction/script/extension), reviewed `scheduled_tasks` cron metadata, bounded canonical `conflicts`, `presence:` block parsed by `presence_profile.py`; `parse_skill_manifest_text()` tolerates missing optional fields and `validate()` returns warnings without raising | `ouroboros/contracts/skill_manifest.py` | `tests/test_contracts.py` |
| `schema_versions` — opt-in `_schema_version` stamping (`with_schema_version`/`read_schema_version`); wired by extension `health.json`, the projects registry/bindings, and presence bindings | `ouroboros/contracts/schema_versions.py` | `tests/test_contracts.py` |
### 11.2 What is NOT frozen (intentionally)
The full `ToolContext` dataclass (browser state, review history, model overrides, …) stays mutable implementation detail — the protocol pins only the minimum. Raw WebSocket/HTTP *values* are unpinned; only the shape keys are. The SKILL.md body is free-form; only the frontmatter schema is pinned. `state/state.json` and `queue_snapshot.json` carry no `_schema_version` key and read as version 0; `task_results/*.json` is the exception — every write is stamped `_schema_version: 1` and admission is `task_result_schema.py`'s, not this section's.
### 11.3 What to do when extending
Add the field to the active frozen owner — `ouroboros/contracts/` for the package ABI, or `ouroboros/gateway/contracts.py` + `web/modules/api_types.js` for browser envelopes — keeping existing consumers working, and enforce the new surface in the contract/parity tests (CHECKLISTS item 17, `gateway_parity`, owns the review-time criteria). Removing anything from 11.1 is a deliberate ABI break: it requires an explicitly versioned successor and a migration note in the release row — the release ledger, not this map, is the SSOT for retirements.
### 11.4 Recent ABI Retirements
- **ABI 7.0** is one deliberate window, so the breaks land together instead of one per minor:
- *Gateway envelopes.* Five compatibility aliases are gone: `cost_usd` / `cost_usd_with_children` (stripped at
the cost SSOT seams — stored rows are still READ tolerantly), `telegram_chat_id` from the four outbound
frames and the history mapper, and `project_last_viewed` / `project_hidden` from `UiPreferencesResponse` and
its endpoint, which now answers 400 on an unknown key. The `contracts/api_v1.py` re-export went with them,
and the contract became executable (`gateway/schema.py`).
- *Settings keys.* `RETIRED_SETTING_KEYS` in `settings_defaults.py` is the machine-readable list — read the
tuple, not this paragraph, for membership — and `load_settings` strips its members off disk. This window
added `OUROBOROS_SCOPE_REVIEW_FLOOR`, the flat `OUROBOROS_SOFT_TIMEOUT_SEC`/`OUROBOROS_HARD_TIMEOUT_SEC`
pair (superseded by the activity model), `OUROBOROS_REVIEW_NATIVE_MAX_ROUNDS` (the native review episode is
bounded by its transcript ceiling, the owner deadline and the wallet, not by a round count),
`OUROBOROS_OBSERVABILITY_RETENTION_DAYS`, and — classified separately as
`RETIRED_COMMA_LIST_SETTING_KEYS` — the reviewer comma lists and route envs
(`OUROBOROS_REVIEW_MODELS`, `OUROBOROS_SCOPE_REVIEW_MODELS`, `OUROBOROS_SCOPE_REVIEW_MODEL`,
`OUROBOROS_REVIEW_ROUTES`, `OUROBOROS_SCOPE_REVIEW_ROUTES`, `OUROBOROS_ADVISORY_REVIEW_ROUTE`). Their
migration note: move the configuration into the structured `OUROBOROS_REVIEWER_SLOTS` BEFORE upgrading — an
install carrying only comma keys comes up on the shipped default panel. The owner is TOLD: the read seam
logs the dropped keys once per process, and the first supervisor boot with an owner chat bound posts one
system row there (`server_maintenance._startup_retired_settings_notice`, the same sentence —
`settings_defaults.retired_setting_keys_notice` — naming the keys as NOT honored and, from
`reviewer_slot_config.authored_reviewer_slots_state`, what runs now: the authored panel, the shipped default
while the structured key is absent, or NO panel with the parse error while it is malformed and the loader
rejects it), deduplicated durably per retired-key set in `state.json:retired_settings_notified`. A
comma-spelled ENV projection survives as the derived runtime plane, never as configuration.
- *Plugin ABI.* `PLUGIN_API_VERSION` is `"2.0"` with manifest negotiation checked before plugin import or
out-of-process cataloging; an absent field means legacy `1.3` by construction, and a hash-bound PASS is
grandfathered.
- *Durable task rows.* Every `task_results/<id>.json` write is stamped `_schema_version: 1`, with deliberately
NO legacy converter: an inadmissible row is quarantined with log-only visibility and keeps its id occupied.
- *Typed internals.* `ResolvedModelTarget` is the frozen model-target dataclass at the existing resolution
seams, extension registrations publish atomically with a per-publication generation digest, and the dead
`_call_llm_with_retry`, `compute_cost_with_children` and `format_handoff_message` surfaces are gone.
- *The migration note is a program.* `scripts/rc_audit.py` is a READ-ONLY pre-upgrade scan of a third-party
install that emits a machine-readable scope document naming every incompatibility it finds across the five
frozen classes (gateway alias, retired setting, comma list, plugin API, schema stamp), snapping the
retired-key lists at execution time instead of hardcoding them.
- *Deliberately NOT in this window:* the handler ABI — tool handlers returning `ToolResult` instead of `str` —
is backlog, so handler signatures are unchanged. The external-executor family is the one INTERIOR
exception and does not open that window: its producers, decorators and host consumers pass a native
`ToolResult`, while its four REGISTERED entries still publish a `str` projection through
`tool_result._publish_tool_result`. No handler signature and no other family changed.
- `5.25.0-rc.4` retired the native skill upgrade migration banner API (`GET /api/migrations`,
`POST /api/migrations/{key}/dismiss`, and `MigrationsResponse`). The migration note is the release row itself:
dismissed banner state in `data/state/migrations.json` is intentionally ignored by current runtimes.
---

View file

@ -0,0 +1,47 @@
# 12. Host Service, Companion Processes, and Chat IDs
This chapter owns the loopback, token-authenticated callback boundary reviewed skills speak to, the host-supervised companion processes they may own, and the chat-id policy that decides where a message is actually delivered. It exists because a chat id is a value rather than a boolean and the hidden partition is a real destination, and because an owner-bound reviewed transport is a first-class control surface that must be reviewed as one.
Chat uploads have one `gateway.files` storage owner for Host-confined paths and
completed multipart spools. It uses the artifact substrate's streaming hash and
atomic copy; borrowed spools promise descriptor identity, not an original pathname.
The multipart request's worker owns both copying and closing its spool, so HTTP
cancellation waits for both and cannot interrupt cleanup. Host uploads use the
same settled wait before releasing their existing in-flight slot; the skill keeps
ownership of its source. `store_chat_upload` keeps
its Path return. The old 50 MiB chat upload rejection is removed on both ingresses;
Host's existing 25-file request count and the separate Files-browser upload policy
retain their own contracts.
Telegram document mirroring resolves the captured `file_ref` and streams an owned
file handle through its existing multipart client, with legacy inline bytes still
accepted. Its outgoing 50 MiB boundary is separate from the integration's inbound
10 MiB download limit. Oversized files stay saved in the application: a ready,
already-running owner-authenticated Mini App can be opened through its existing
entry button; otherwise the message explicitly says it cannot mirror the file and
directs the owner to the app. Delivery never starts a tunnel or publishes a new
public/token-bearing artifact URL, and this notice does not claim the bytes were
uploaded to Telegram.
The Host Service is a loopback, authenticated callback boundary for reviewed skills (`ouroboros/gateway/host_service.py`, `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}`). Every request authenticates an opaque `x-skill-token` bound to the skill's content hash, executable review, enablement, and grants — secrets never enter the token, a payload edit stales it, and the client wrapper refuses stringification (`skill_token.py`). The frozen route family is exactly: `/identity`, `/tools/schemas`, `/chat/allocate-internal`, `/chat/inject`, `/chat/operations/{operation_ref}`, `/chat/cancel`, `/chat/decision`, `/presence/turn`, `/presence/work/{work_ref}`, `/ui/ws-message`, and WS `/events`; permissions still decide which route works for a given skill. Review of a transport skill evaluates identity binding, attribution, polling bounds, panic cleanup, token confinement, and exfiltration — an owner-bound reviewed transport may be a first-class control surface, not a screen-only integration. External slash commands bind a separate positive-identity external owner slot (`supervisor/state.py`), so an unidentified transport can never bind commands and the local web owner can never lock out a real remote owner. `wait_for_response` remains restricted to A2A-allocated chats. Named messages wait for their exact operation; only the legacy unnamed path uses chat-wide response subscriptions. Admission is one limiter with two policies (`host_service._RateLimiter`): the WS relay lane (`/ui/ws-message`) is a token bucket — a 60-message burst reserve per skill refilling one message per second, so a burst no longer silences a widget for the rest of a minute — while every other lane keeps its 60-per-60-seconds sliding window. A refused relay is visible at the host and aggregated per burst: one warning when the reserve empties, then one warning plus one durable `host_service_ws_relay_dropped` row in `logs/events.jsonl` carrying the dropped count when the lane admits again or the idle bucket is swept; the 429 carries `retry_after_sec` and `dropped_in_burst`, the child's `send_ws_message` stays best-effort `None`, and the in-process broadcast path keeps no bound.
Host attachment copies use one request-local cleanup stack and the shared settled
HTTP-worker wait. Partial copy/refusal/cancellation removes only new unaccepted
copies. Named ingress retains those copies once its existing canonical write is
attempted, through an internal custody callback; a replay never adopts its duplicate copies.
Cancellation waits for that admission worker. Accepted bytes survive disconnects
and unknown write/queue outcomes. Such an error does not claim acceptance and
may retain an unused copy; no new reconciliation or transaction log is introduced.
Operation correlation (#667): a named injected message has `operation_ref=<chat_id>:<client_message_id>` on 202, 200, 504 and disconnect responses. `supervisor.message_bus.accept_local_message` serializes check, canonical inbound-row acceptance and enqueue; its existing `log_chat` writer must succeed before work is queued. The host-minted full origin is carried internally to `record_inbound_message`, which verifies the source and consumes it without writing a second row. Repeated same-id, same-text, same-skill delivery rejoins even before supervisor dequeue; changed content or source is refused with 409. This preserves the existing in-memory queue: a crash after acceptance can lose delivery and reads honestly as `lost`, never authorizing a second enqueue. Reads span retained canonical chat generations. Routing annotations and outbound task ids are discovery hints only: task reads and cancellation require the actual queue/task record's complete `origin_message_ref` to match the authenticated skill's canonical source. `DirectActivityRegistry` carries that same origin. Named response waits poll this exact operation and its retry-aware effective task result; chat ordering alone never proves a reply. Cancel enters the existing durable intent and cascade-custody owner only for work with that origin and the same installation root. A different or unavailable owner root is disclosed as cancel_unsupported before any intent is written. Unaddressable/foreign work remains `cancel_unsupported`; unresolved custody remains explicit and never becomes a false `cancelled`.
Successful extension children also carry a bounded `ws_relay_failures` count map from the existing PluginAPI transport owner through their result envelope and process facts. Missing transport, network errors and HTTP refusals remain visible without changing `send_ws_message -> None` or the producer's success. One host warning reports the aggregate; diagnostics contain no message bodies, URLs or credentials. A child that dies before its final envelope may lose this aggregate; the existing measured death/timeout facts remain authoritative. The separate streaming route implementation carries the same aggregate in completion X after body/background work.
Presence flow: `POST /presence/turn` requires the content-hash-bound `presence` permission, one `binding_id`, one exact transport event, and optionally staged files confined to the skill's state root; the binding resolves from `state/presence_bindings.json` with exact provider/account/conversation/thread origin verification; cross-process locks enforce the installation-wide cap and serialize one `conversation_key`; a stable event-derived task id makes transport retries idempotent; input and output join ordinary dialogue history with full transport/actor provenance. The one typed outcome is message/silent/tool_delivered/deferred — `deferred` only with a correlated `work_ref`, because an unanchored "deferred" would be an unanchored promise — and `GET /presence/work/{work_ref}` polls the bound late result without exposing the general task API. Promotion of Presence work into a managed task clears requested Project/workspace/source widening: a public conversation may promote long work but cannot choose new authority, and the cost ceiling plus return destination follow the promoted root by value. `presence_cancel_work` acts only on a `work_ref` whose stored binding and conversation match the current turn; owner chat and Background Consciousness may `initiate_presence` on an existing enabled binding. Every admitted turn's ceiling also carries the cognitive-memory baseline (knowledge read/write/list, scratchpad, identity, chat history): one mind keeps one memory in every channel, so an external message can prompt an in-turn memory or identity revision on the profile's model slot rather than only a post-task knowledge write; the correspondent gains no tool, and public text still carries no owner-command authority; on the read side, an unselected baseline `chat_history` grant is unbound, so — like `knowledge_read` over the notes kept about people — a reply can surface the owner's own words to the correspondent through the model's judgment. Ceilings frozen before the baseline existed verify their own digest and carry no baseline until their profile is recompiled.
Companion processes are host-supervised: reviewed manifest-declared descriptors enter durable custody, reconcile after lifecycle changes and restart, and stop on disable/unload/panic; `state/extension_generation.json` carries the opposite direction — the server's published live set, which a running task worker adopts at a task's start or at a dispatch miss, so an enable after boot is not invisible until the pool respawns. Worker-side changes write durable reconcile requests (`state/extension_reconcile/`) rather than spawning server-owned children; every reconcile state names the process that answered and whether that marker request was written, and the tool and review receipts pass both facts through. Health observations are process-qualified: aggregate and Skills UI health use the server observation as authority and expose the worker observation with its handoff outcome only as a qualifier, so failed handoff success cannot advance `last_known_good`; restart-budget exhaustion persists a terminal reason in that health state, cleared only by a later successful start. A companion's cwd is the reviewed payload directory, so a payload edit stales review before reload instead of silently mutating a live process. The live projection is `state/extension_companions.json`.
Chat IDs: a chat id is a VALUE and absence is `None`. `HIDDEN_CHAT_ID` (0) is the hidden partition — the Skill Review panel plus every headless task admitted without a registered project — a REAL destination that no browser surface reads: delivery goes through membership routing (`message_bus.notification_chat_route`; every producer that tested `if chat_id:` dropped its notices), a `chat_id=0` history query coerces to Main and the Main filter drops chat-0 rows, so explicit panel rows never become ordinary conversation history, and a chat-0 row reaches a project thread only via a durable lineage binding; negative ids are synthetic A2A traffic and never enter a human stream (the id policy SSOT is the §11.1 `chat_id_policy` row).
---

View file

@ -0,0 +1,17 @@
# 13. External Skills Layer
This chapter owns the skill plane outside the runtime package: where native and marketplace payloads live, the independent gates between installation and execution, and how archive installs, declared dependencies, extension staging and the bundled Telegram transport are admitted. It exists so discovery, review, grants, dependency readiness and enablement stay separate facts, because collapsing any two of them turns a UI toggle into execution authority.
Native and external payloads live in separate data-plane buckets (`data/skills/{native,clawhub,ouroboroshub,external}`), with review, grants, enablement, dependencies, tokens, and health under `data/state/skills/<name>/` (§1 Data layout). Discovery and manifest parsing establish identity, source, hash, provenance, and conflicts — never trust; a conflict declared by either enabled peer is enforced symmetrically without deleting either payload.
Outside Cyber Pro, after real skill feedback, explicit Advisory author finish may accept a revised payload after deterministic preflight without another panel. Cyber may finish without prior feedback while retaining missing, pending or failed review facts. `SkillReviewState.gate_for` is the shared readiness/extension/Host Service admission: `is_stale_for` still reports the original critic hash, `author_disposition.subject_hash` names the author's current bytes, and ordinary Blocking requires fresh critic authority. Repeated finishes preserve the original hash; desired enablement, grants and dependencies remain independent. Skill publication uses the separate mode-aware capture/scanner contract in §6.
The executable sequence is install → deterministic preflight → hash-bound multi-model review → grants → dependency readiness → enablement → execution, and the gates stay independent: a review PASS installs no dependencies, `enabled=true` does not prove readiness, and an extension additionally needs host registration (§10 invariant 8). Mutating lifecycle work flows through one deduplicated queue (`skill_lifecycle_queue.py`); review jobs retain task/source/hash/attempt/actor/terminal evidence in the private full record, with the compact UI history as a projection. Review ordinals are allocated only after a job starts under the lifecycle lock — retry history stays explainable across hash changes without a UI counter becoming review authority; a started failure/cancel/timeout consumes its number with one idempotent terminal row, while pre-start dedupe consumes none. Skill-review history displays its actual review round, snapshot attempt and revision facts rather than renumbering the retained tail. Accepted rebuttals reduce reviewer thrash; a new payload hash still requires fresh evidence.
Skill review combines the deterministic preflight with the authoritative multi-model checklist review; an optional advisory stays fail-open and cannot replace it. Official catalog payloads get their reduced-noise profile only when the sidecar, catalog file set, local file set, and every SHA-256 match exactly. A deterministic preflight failure persists as PENDING, not BLOCKERS. Outside Cyber Pro, PENDING is non-executable even under advisory enforcement; Cyber retains that missing or failed review fact while choosing whether to execute.
`skills/telegram/` is the bundled transport: an in-process extension owns binding/polling/injection/settings while a supervised companion owns the sidecar, tunnel, menu rollback, heartbeat, and singleton; the text bridge works when the Mini App/tunnel is unavailable. The payload is seeded with hash-bound native provenance and stays disabled until token/permission grants; it is never a marketplace install. Its state classes stay separate under `data/state/skills/telegram/` (settings/binding, companion config, menu rollback snapshot, delivery cursors, verified tunnel cache). Colab bootstrap waits for native discovery plus a fresh executable seed projection, grants only missing grantable items under the owner auto-grant policy, then enables. Inbound Telegram files (documents, video, audio, voice; this integration's 10 MiB download cap) ride `/chat/inject` `attachments` — paths confined to the skill's own state root that the host copies into the browser paperclip's `data/uploads` store and stages like any chat attachment, so a caption-only text or a text-less file is one ordinary owner message; owner answers to quiz cards (a tapped option or a reply to the card) ride `/chat/decision`, the same `task_decision.answer_decision` ingress as the web card; the bridge status reports `degraded` (`telegram_startup_deferred`) while a dead network defers token validation, and the task-done push takes its word and icon from the stamped `outcome_phase`.
The optional `model_experience` manifest prose — what the skill adds to the model's context and what it costs in tokens — travels to the model-visible surfaces (the `list_skills` JSON and the installed-skills context section), and a manifest without the section keeps the exact prior rendering; manifest parse refusals teach the repair through the `fix_hint` every `SkillManifestError` carries.
Marketplace installs are bounded archives staged privately and landed atomically with per-file hash checks. Inert resource bytes, including PDF/PPTX, need no suffix allowlist; archive validity, extraction paths, checksum and actual loader support remain separate facts, and Cyber policy findings are warnings. Install metadata drives isolated dependencies (`marketplace/install_specs.py`); manual instructions remain guidance, not execution. Exact download descriptors and per-spec source/postinstall opt-ins share that dependency owner: fresh review and hash-covered declarations authorize literal build/check argv, verified resources and package caches survive under `state/skills/<name>/dependency_cache`, and existing `deps.json` records resolved package metadata, bytes, outputs and diagnostics. Delivery is distinct from a declared executable check; absent checks remain unknown, and declared-output drift invalidates installed readiness. Large resources stay outside the payload/Git patch. Delegated payload capture uses the existing review byte classifier and preserves binary resource descriptors such as PNG/Wasm. Go compiles to a private temporary executable before running literal argv and preserving its exit status; Deno translates current reviewed effects and actual forwarded env names without adding an OS sandbox. Extensions import through staged trees (`_stage_extension_import_tree` under `__extension_imports/`), so concurrent workers cannot remove a peer's live import and stale trees stay reclaimable. Per-call child processes may proxy tools/routes/WS/UI/settings/companion descriptors, while persistent subscriptions and supervised tasks require in-process or companion lifecycle — reported through the generic capability matrix, never inferred from a platform name. Isolated children run the same staged loader in a private base with a scrubbed env, so native crashes cannot kill `server.py`. In-process extensions are more powerful, which is why namespacing, declared permissions, per-skill tracking, and atomic unload are an executable contract rather than convention. Transport metadata records source and session generically; skill repair enqueues an ordinary managed task with an exact selected resource and revision admission (`skill_repair_admission.py`); a legacy `skill_repair` selector remains readable but selects no reduced profile. Shell, browser and delegation remain normal task capabilities; an installed payload needs no mandatory Git copy, and review never forces unload merely because the caller is repairing it.

View file

@ -0,0 +1,32 @@
# Role and authority
This chapter states what the handbook is for and which document owns what: the constitution, the architecture map, the design semantics and the reviewer checklists each keep their own authority, and this book carries only imperative rules for changing the body. It also names the domain manifest that owns module-to-domain assignment and the generated map read beside it, because moving code across a domain boundary is the owner's call rather than a manifest edit.
This is Ouroboros's engineering handbook: imperative rules for changing the
body, grouped by change class, each naming the surface that enforces it (test,
gate, CI lane) or stating that none does. `BIBLE.md` owns constitutional
principles; `docs/ARCHITECTURE.md` owns the current structure, data flow, and
rationale map; `docs/DESIGN.md` owns visual and interaction semantics;
`docs/CHECKLISTS.md` owns reviewer items, severity, and output contracts. This
file does not duplicate their inventories or serve as a changelog.
One more owner to know before moving code: `ouroboros/domains.toml` is the SSOT
of the module-to-domain assignment (1:1 and complete over the tracked runtime
population) and pins the factual cross-domain dependency data as the baseline.
`docs/DOMAIN_MAP.md` is generated from it — read the map, edit the manifest,
then regenerate both with `python scripts/check_domains.py --write`
(`tests/test_domain_manifest.py` makes staleness red, and a new module with no
row is red too). A new cross-domain import direction, a wider cycle group, or a
cross-domain literal copy is a red gate, not a warning: needing one is an owner
decision, not a manifest edit. The witness-level detail behind the baseline —
every module-edge witness, the lazy/guarded/dynamic classification, the cycle
groups — lives in `docs/v7next/DOMAIN_QUOTIENT_REPORT.md` (report only,
regenerated by `python scripts/v7next_domain_report.py`).
Rules here describe current practice or a deliberately enforced standard. When
code and prose disagree, inspect the implementation and history, repair the
authoritative surfaces together, and retain the failure a non-obvious rule
prevents.
---

View file

@ -0,0 +1,496 @@
# Naming and boundaries
This chapter owns the rules that keep the body legible from outside: naming and entity-type conventions, dependency direction, the CLI and headless contract, cognitive quality, LLM-first affordances and the prompt-edit discipline, the documentation contract, the generality question, pricing and admission, and the anti-patterns that have actually cost this system incidents. It exists because each rule here names the test, gate or CI lane that enforces it, or states honestly that only review does.
- Code identifiers, comments, docstrings, commit messages, and user-facing
product UI strings are English.
- Follow PEP 8: modules and variables use `snake_case`, classes use
`PascalCase`, constants use `UPPER_SNAKE_CASE`. Name the observable
responsibility and authority, not the implementation fashion; prefer a clear
function module over a class with no lifecycle.
- Contracts are typed shapes, not service objects. A manager is justified when
it owns lifecycle or mutable state. LLM-callable Tools remain thin
`{verb}_{noun}` functions that validate their public input, call the owning
subsystem, and format a result. No universal `{Domain}Service`,
`{Platform}Gateway`, or class layer is required.
- Dependency direction is the test: UI/CLI → inbound gateway → domain owner;
runtime policy → small host-owned contract → outbound adapter (the
`ouroboros/gateway/` inbound vs `ouroboros/gateways/` outbound map is
ARCHITECTURE "Gateway Boundary v1"). Provider- or transport-specific
decisions do not flow back into core policy.
- Chat authorship is stamped by the producer and preserved through persistence
and replay; never infer it from text or promote host-selected intermediate
output to a model final (`tests/test_terminal_provenance.py`).
Enforcement: naming and entity-type rules are scored in commit review by
CHECKLISTS item 2 `development_compliance` (a)/(b); the transport/core
dependency direction has one CI guard (the "Guard extracted transport imports
stay out of core" step in `.github/workflows/ci.yml`); the remaining boundary
rules have no automated surface — review-only.
### CLI and headless work
- CLI commands parse, call the existing gateway/scheduler, and render text or
typed JSON/JSONL/SSE. They do not create a second task state machine.
- External workspace tasks keep governance bound to the system repository while
contextual tools default through `ToolContext.active_repo_dir()`. Admission
rejects overlap with the system repo/data and records a read-only preflight.
- Physical file resolution must preserve the requested address. Normalize
in-root absolute paths for every resource; refuse an outside address before
relative-path sanitization. Keep repo-prefix and canonical artifact redirects
on the same resolver used by guards and handlers.
- Project focus changes the default target, not the top-level tool surface.
Generic VCS selects active/system explicitly; advisory, reviewed commit,
rollback, promotion, restart, and runtime control keep their intrinsic
system-repository contracts. `executor_ref` selects a process backend, not an
implicit sandbox.
- Project-local installs may run within the workspace policy. Global/system
installs remain safety-reviewed, and `sudo` is non-interactive (`sudo -n`).
- GitHub issue/PR tools resolve the same process binding as shell commands and
carry an explicit `repo` into every dependent call. Project focus overrides
ambient `GH_REPO`; broken room bindings refuse and file-less Projects need an
explicit repository. Presence may override its default repository only through
a host-selected argument binding. Native CLI configuration proves configuration,
not active authentication; discovery never logs in or probes the network.
- Do not add a second scheduler for operator tooling or a generic CLI file
manager. Use the task queue, attachments, logs, and artifact endpoints.
- A process cwd determines relative paths, not the root task's entire write
authority. Reuse the resource binding for other authorized destinations;
preserve child write confinement and actual runtime/credential boundaries.
Process admission preserves original argv and the prepared physical target.
Do not reconstruct semantic permission from command words or guessed effects;
use the Supervisor source and mode contract in ARCHITECTURE "Safety and runtime
mode". Explicit task resources and actual Git targets remain distinct from
semantic judgment; this does not promise an SSH or OS sandbox.
- Do not infer credential authority from ordinary source/config directory
names. Reads and searches share ONE byte masker in every scope: known token
formats and PEM private-key blocks are masked, ordinary long source, hashes
and identifiers are preserved, so search and read deliver the same bytes to
the owner. A runtime data
directory inside a project retains its actual secret/control path rules,
including a forked child's canonical parent (`core_secret_paths.restricted_data_roots`).
Prepare invariant credential/root locations once per list/search/query call
through `make_subagent_secret_target_check`; never retain that predicate across
calls. Target resolution and owner-state/file-identity checks remain per target.
A SUFFIX or a WORD inside a file name never refuses owner input or owner
output on attachment ingest, on export or Deliverables, on a `user_files`
mutation, or in the git lanes, with one surviving tail rule: dotenv
spellings (`.env`, `.env.local`, `prod.env`) still refuse on ingest, on
export, in the git lanes and in the skill/review packs. Path-selected
attachment ingest otherwise checks exact credential leaves and credential/control
directory components; enumerated physical owner stores are mutation-fence
authority, while the git lanes also check content evidence (the bounded PEM
head read, `workspace_patch_capture.pem_private_key_reason`). The capture,
snapshot and cooperative checkpoint consumers use `pem_capture_refusal`:
effective Cyber preserves the original finding as advisory and includes the
requested bytes; ordinary modes retain the existing exclusion.
Restricted file readers mask complete private-key blocks before
selecting a window, preserving character positions and line breaks. The
owner credential fence covers the enumerated locations in
`credential_shapes.owner_credential_locations`, credential leaves and VCS
control directories. The exact SSH config exception does not permit key writes
under `.ssh`; unlisted stores such as `.cargo`, `.terraform.d` and `.kaggle`
are not protected by an arbitrary dotted-directory default-deny.
- An unlaunchable sole cmd element gets an actionable argv/shell hint, never
automatic splitting or an implicit shell. Repo-only edit tools reject their
unsupported roots through their existing argument categories.
Enforcement: `tests/test_headless_cli.py` (task-API admission, typed refusal
terminality, attachment admission), `tests/test_cli_entrypoint.py` (the CLI
surface), and `tests/test_external_workspace_access.py` plus
`tests/test_workspace_authority_binding.py` (workspace policy and system-repo
collision blocking). The no-second-scheduler and no-generic-file-manager
rules have no automated surface — review-only.
### Cognitive quality
Do not lower model quality, reasoning effort, output budget, or context breadth
as an incidental latency/cost optimization (BIBLE P1 owns the principle). An
intentional narrowing is an explicit recorded decision reflected in the plan,
docs, tests, and evidence; outside Cyber Pro it belongs to the owner. Cyber
configuration authority follows BIBLE P0/P3 without rewriting earlier call facts.
No automated surface — review-only: commit review carries
it through CHECKLISTS items 1 (`bible_compliance`) and 21
(`capability_regression`: an accidental narrowing is the named failure class).
### LLM-first affordances
Do not repair a semantic tool-choice failure by adding one more keyword hint to
`prompts/SYSTEM.md`. A model mistake alone does not justify a new host contract:
first establish a real missing capability or information, or an independently
valid requirement. When discoverability is genuinely missing, repair the existing
tool schema or affordance at the point of need. Do not freeze the model's reasoning,
dialogue representation or collaboration strategy to make one incident testable.
SYSTEM accretion trains around one incident, bloats the resident prefix, and forks
the authority. These choices are reviewed through CHECKLISTS' development and
capability-regression items, not a new semantic gate.
What belongs in `prompts/SYSTEM.md` (tier-0 for every Main/task profile in both
context modes — Background Consciousness and the safety supervisor carry their
own prompts — and competing with the task for context): identity and tone, the decision
loop (answer / promote / route / delegate / do it myself), cross-tool policy
(which class of tool or lane for which situation, root semantics, memory only
through its own tools, untrusted external data), prohibitions and safety
invariants stated once, and the memory contract. What does NOT belong there:
how a tool or mechanism works. A tool's parameters, signatures, recipes,
typed outcomes, and "when to choose it" live in its `get_tools()` schema — each
profile receives its own visible schema set on every round (delegated, repair,
credential and contract filters narrow it), so the schema is the SSOT
of the per-tool contract and a prompt sentence about it is a second copy that
drifts, while SYSTEM.md stays the cross-tool selection policy; mechanism
documentation lives in ARCHITECTURE or here;
runtime facts (capabilities, queue, catalog, receipts, health invariants,
memory sections, registry digest, installed skills, review section) are
assembled ONCE per task attempt in `build_llm_messages`, so the Health
Invariants block lists custody obligations as of task start and does not
refresh mid-task (a deliberate frozen-ContextCore / prompt-cache choice; the
integrate schema, the apply receipts and the absorption digest carry the
mid-task fact instead). Tool schemas are re-sent every round.
A new tool therefore requires NO SYSTEM.md mention. Before adding a
sentence to a prompt, check that the schema or runtime block does not already
carry it; before removing one, check that they do (or add the missing fact to
the schema without growing it into a paragraph). Local-model compaction keeps
only the text before the first `## ` heading (plus the BIBLE section), so the
load-bearing floor rules stay in that preamble. Every prompt change reports the before/after byte size
in the commit or PR. The memory contract's resident rule: a note's authored
summary is its resident face in the knowledge index, and a resident memory
carrier that is absent renders as a visible gap, never as silence.
Recoverable tool failures are evidence for the next LLM turn, not triggers for
a host-authored recovery workflow. Return a typed, redacted result naming the
failed stage, already-completed external effects, and an actionable repair
hint; the LLM decides whether to inspect, repair, retry, use another capability,
clean up, or stop. Host code remains responsible only for deterministic
integrity and authority boundaries plus truthful receipts; do not add
task-specific auto-retry, fallback, cleanup, resume, or terminal-flow state
machines.
Explicitly naming a documented default is never a different request. An argument
whose value is what omitting it already means (`directory_strategy="direct"` with
no `scope_paths`) takes the omitted path on a shape that cannot serve the argument
at all; only values that genuinely ask for something are refused there, typed, at
the earliest layer holding the authority to judge them, with the repair named.
A producer that already knows its call failed publishes that fact typed: a
`ToolResult` through `tool_result._publish_tool_result`, or a first-line
`⚠️ IDENTIFIER` marker the legacy adapter maps to a status. Identifier-less
`⚠️ prose` and a bare `ERROR: ...` string are recorded by the registry — in
`tools.jsonl`, the outcome classifier and the acceptance packet — as a successful
call, so the failure the producer saw is lost exactly where the next decision
reads it. A wrapper over an inner producer (`view_image` over the local image
loader, the publication transaction over the GitHub transport) carries the inner
failure and its safe cause forward instead of a stage-only word.
Inside the external-executor family (`delegate_start`, `delegate_wait`,
`delegate_cancel`, `delegate_answer`) the result IS the `ToolResult`: one refusal
author (`delegate_shared._fail`) writes `ok: false` plus `host_code` beside the
domain payload, the producers, decorators and host consumers (bootstrap, recovery,
the unknown-provider hold, the pending-wake replay) pass that value around, and the
four registered entries publish it once, after every decoration, immediately before
returning its text. Publishing earlier is silently discarded: the registry accepts a
published result only when its text IS the string the handler returned. The
acknowledgement of a supervising wake is keyed on the `supervision_wake_id` that
result publishes, not on the tool's name.
Enforcement: the prompt-edit discipline is scored by CHECKLISTS item 13(b)
(a prompt edit never restates a tool schema and is never an incident patch);
the recoverable-failure boundary has no automated surface — review-only; the
typed-refusal rule is ratcheted by `tests/test_typed_tool_refusals.py`, a
source lint over returned literals in `ouroboros/tools/` that flags only
identifier-less heads (`⚠️ prose`, bare `ERROR:`); its per-file allowlist
discloses that identifier-less residual — a file whose count grows fails, a
shrink must be recorded, and a same-file swap of one old site for a new one is
invisible to the count. A marker-shaped refusal the adapter still buckets as a
warning (`⚠️ SOME_IDENTIFIER: …` recorded `ok`) is a separate, larger residual
owned by the adapter vocabulary, not by this lint; a failure text that travels
through a variable, a tuple or a helper is outside its reach and is pinned by
the producer's own tests.
### Documentation contract
`docs/ARCHITECTURE.md` is the map of the body (BIBLE P6), written in the present
tense: what exists, where it lives and how it flows (structure and operation),
and WHY it is so. Every important WHY — cross-module, gate-level or module-local
— stays in the map, at least briefly (a second line under a module row or one
sentence in the owning section); mechanism detail beyond that lives in the
module's docstring, which the map points to by name and never copies.
`docs/DEVELOPMENT.md` is how the body is changed: imperative rules grouped by
change class, each naming the surface that enforces it (test, gate, CI lane) or
stating honestly that none does; mechanisms, reviewer criteria, constitutional
text and defaults belong to their owners (ARCHITECTURE, CHECKLISTS, BIBLE,
`config.py`) and are pointed to, not restated — the ARCHITECTURE settings and
endpoint tables are test-checked registries of those owners, not second
authorities. A change REPLACES the description of the node it touched; release
history lives in git and the README history table, leftovers go to issues.
Residue — parenthesized version stamps, decision codenames, "used to /
previously" narrative — is caught by the shrink-only residue check in
`tests/test_docs_sync.py`, which enforces only the explicit, case-sensitive
matches in `DOC_RESIDUE_PATTERNS`, outside its declared skipped subsections and
language-tagged fences (the untagged module-tree fence in ARCHITECTURE §1 IS
scanned — an owner decision); semantically equivalent historical prose stays
review-only under CHECKLISTS item 7.
### Generality and emergence (P13)
Every non-trivial change picks a level: patch the case in front of you, solve
the class it belongs to, or build a framework for cases that do not exist yet.
The first fossilizes, the third speculates; aim for the second — BIBLE P13's
invariant question and stronger-mind question find it. The proof burden is
symmetric: promoting a case detail into shared structure requires showing it is
an invariant (several real variants, or one already-stable boundary); adding an
abstraction requires a demonstrated class — an imagined consumer is not one. In
doubt, generalize the meaning and the authority, keep the mechanism minimal and
local, and let the next real case pay for the next step. Reviewer findings are
evidence for this judgment, never policy that overrides it. No automated
surface — review-only.
### Pricing and admission
Never add hand-maintained model-price tables, inherited prefix tariffs, or
numeric fallback prices; preserve `cost=None` and `cost_final=false` when no
live source answers the exact route. Unknown price is neither free nor a
model-admission veto; known exhausted budget remains enforceable. Enforced by
`tests/test_pricing.py` (exact-route lookup, no prefix inheritance, unknown
cost stays `None`) and `tests/test_budget_limits.py` (a known exhausted budget
still fences); mechanism: ARCHITECTURE "Budget tracking".
### Anti-pattern: content-derived identity for host-minted records
If the host itself created a record — a chat message, a task, a binding — its
identity is captured at ingress and passed downstream BY VALUE as a typed
reference (e.g. `origin_message_ref` built where `log_chat("in", …)` writes
the canonical row). Never re-derive it later by searching logs/state for a
row whose text hash/equality/prefix matches: in an LLM-first system the text
is routinely rewritten between ingress and use, so content-derived lookup
fails exactly on the normal path. Content hashes are legitimate only as (a)
an INTEGRITY CHECK on an already-known identity, and (b) content-ADDRESSING
where the content IS the identity (artifact stores, observability blobs,
staged-diff review bindings). The enforcement shape is a REQUIRED typed
argument at the consuming seam (`bind_task_to_project(..., *, origin)`: a
valid ref or a closed-enum absence reason; omission raises), so a future call
site cannot silently skip the invariant — `tests/test_projects_v6640.py`
exercises that seam. For fuzzy entities use the LLM-first pattern
(`semantic_dedup`), never string equality.
The same captured reference is also the IDENTITY OF THE WORK, not just its
provenance: a new task id minted from the same owner message (a promoted root,
a mid-run scope call, the timeout retry that replaces a dead attempt) must
INHERIT that origin's project binding (`projects_registry.project_id_for_origin`,
keyed by value on chat id + client message id), never re-derive project
membership from its own id — one convertible unit per message, not one per task
id. A timeout retry is bound at RETRY ADMISSION, inside the same transaction and
under the same claim lock that admits it
(`worker_promotion.bind_retry_to_origin_project`, called from the reaper's
`_run_retry_admission_transaction`): the predecessor's own binding answers first,
then that origin's, and the new row reuses the predecessor's stored origin by
value, so the retried work stays in its room instead of painting a Main card that
offers to turn it into a project. The bind lands ONLY once cancellation can no
longer win the boundary — a binding is immutable, so a bound-but-never-admitted
retry id would answer `project_id_for_task` forever — and a retry suppressed by a
cancelled or already-terminal root is therefore never bound. Enforced by
`tests/test_retry_project_binding.py`.
One named exception inside role (b): a verification RECEIPT with no earlier
ingress point is reconciled by ONE TYPED IDENTITY KEY, matching on the key's
kind AND value, never across kinds — a per-component fallback chain is not an
equivalence relation (the outstanding set came out order-dependent), while
keying makes sameness the kernel of a function and fails safe: a false red
costs a human a second look, a false green costs the thing this surface
exists for. The full mechanism (masked-path rules, `IDENTITY_KINDS`,
projections, bounds, rendering stamps) lives in
`ouroboros/_outcome_receipts.py`, enforced by
`tests/test_v678_receipt_reconciliation.py`. The process-tool lane discloses a
masked exit code in its result envelope only and writes no receipt, so nothing
there participates in masked-pass reconciliation or the masked-verification
nudge. Four rules generalize:
- **Whatever decides must be what is reported**: the reporting path reads the
deciding path through one shared projection, never a re-derivation beside
it — a host-attested artifact that misstates its own basis is worse than
one that says nothing, because a reviewer cannot discount evidence whose
provenance it was told wrongly.
- **A property of a closed set of kinds lives IN the set** (a table row per
kind plus a total lookup that raises on a kind that skipped the table), so
a new kind cannot be added without answering.
- **One canonical identity derivation** — comparison, hashing, counting, and
projection read the same derived object, in the order canonicalize the RAW
values → render → bound; a normalization that discards information the
identity depends on is not a normalization.
- **Changing a stored rendering means versioning it**; reason about a format
migration in BOTH directions — the false-green direction is the one that
gets missed, and unknown must not clear a red.
Disclosed deferred limit: `tools/verify.py` bounds the DURABLE
`artifact_observation` path set at twenty with no omission count — advisory
only (a nudge and a disclosed reviewer flag, never a gate); fixing it means
changing the durable store and deserves its own scope. Beyond the typed seams
and tests above, this section is review-only.
### Anti-pattern: an open default behind a closed exception list
A behaviour that is "on by default, except for these names" keeps its real
rule in a list that can only go stale. The retired `ADDRESSING_ONLY_TOOLS`
(`web/modules/chat_activity.js`) decided whether a chat turn had a card by
subtracting three tool names from the tool count: every new addressing tool
would have minted an empty card, every renamed one would have silently left
the list, and the list restated a fact the host already carried as the typed
routing annotation. Derive presence from the facts the record already holds
(`web/modules/chat.js::blockVisible`) and let the host name the special case
(`typed_routing_action`), never a client-side exception list. The same shape
hides in "hide unless kind ∈ {…}" and "count unless name ∈ {…}": when the
list is the rule, the rule is missing.
### Task-authored messages are never owner text
Who is speaking through a routing act is ONE fact the host mints by value
(`control_routing._routing_issuer`): an owner turn (a direct chat turn — its
ingress stamped `client_message_id` — or a task relaying the owner message it
just drained) or a task speaking for itself (a pooled, Swarm, project or
headless root). Never derive it again from a proxy — a routing contract only
chat turns carry, an empty client id, the chat id of the event — and never give
the model an argument for it. A task's own words are delivered as a task
message (`KIND_TASK_MESSAGE`, provenance `independent_task`), never as
`KIND_OWNER_TEXT`, and the provenance value lands at three seams in one change:
the writer's closed set (`owner_mailbox.TASK_MESSAGE_PROVENANCES`), the render
ladder (`deliver_task_message`: `[Message from independent task <id>]`, never
the ancestor fallback) and the drain mapping (`loop_round_limits`): an
independent task's words are context the receiving model judges, so they enter
no owner corpus — `owner_source_sha256`, the post-drain growth check that
supersedes a paid acceptance panel and the acceptance premises stay the
owner's (owner 4=A). A refusal or delivery to a task issuer is typed in its
tool result ("written", never "read") and in one `task_message_routed` Logs
row naming author and target; no chat is told and no picker options are
attached. Enforcement: `tests/test_task_authored_messages.py`.
### Anti-pattern: a chat id tested for truth
A chat id is a VALUE, not a boolean. `HIDDEN_CHAT_ID` (0) is the hidden
partition — the Skill Review panel plus every headless task admitted without a
registered project — a REAL destination that no browser surface reads; absence
is `None`, and a negative id is synthetic A2A traffic. `if chat_id:` therefore
does two wrong things at once: it drops a partition-bound notice AND re-routes
hidden work to the owner's main chat, which is how a whole `ouroboros run`
went invisible while its children surfaced in Main as a nameless card. Use the
two normalizers instead of a third: `message_bus.notification_chat_route`
answers "where does this notice go" (first DELIVERABLE candidate, `None` when
none is) and `message_bus.coerce_chat_identity` answers "what is this row's
address" (explicit value kept, absence defaulted). Address a task once at admission (`log_addressing.ingress_chat_id`) and pass
the value downstream; a producer that sends to the owner DIRECTLY (nothing
re-addresses it later) resolves the task's durable project binding AT EMISSION
through `log_addressing.resolve_project_chat` and puts it ahead of the row's
chat, because a task bound to a project after admission still carries the chat
it was born in. Explicit browser-source Main addressing is distinct from
the ordinary hidden API default; task type is never source provenance. Enforcement: `tests/test_chat_id_truthiness_guard.py` is the
source lint that keeps the class closed; it also sees the id read straight off a
mapping inside a condition (`if row.get("chat_id") and ...`), the form where no
local exists for the other alternatives to match; its allowlist is where a
deliberate exception states its reason.
### Mutable external-fact inventory
This table is a maintenance inventory, not a second runtime authority. External
facts change independently of Ouroboros releases; prefer live metadata or a
bounded probe where that can answer the exact question, and otherwise keep the
current conservative behavior visible. v6.67.0 documents these facts but does
not migrate their runtime representations. No automated surface checks these
rows — review-only maintenance.
| Location | Fact | Mutability | Current authority | Live/probe option | Risk | Recommendation |
|----------|------|------------|-------------------|-------------------|------|----------------|
| `ouroboros/provider_models.py::_VISION_MODEL_PREFIXES` / `_VISION_OVERLAY` | Which model families accept native image input | High as model families and route capabilities change | Conservative shipped prefixes, overridden by parsed OpenRouter `/models` `architecture.input_modalities` for exact model ids | Exact provider metadata when available; otherwise a bounded image-input capability probe | A stale positive sends unsupported image blocks; a stale negative needlessly captions them | Keep the conservative fallback and exact-model overlay; consider broader provider metadata only in a separately reviewed migration |
| `ouroboros/llm.py::supports_message_cache_control` | Which families support message cache controls | Medium/high as provider routing contracts change | Explicit family rules backed by provider behavior and dated live probes | Provider documentation plus a bounded cache-control send | A false positive can invalidate a request; a false negative loses the prompt cache | Retain the small explicit rules and re-probe when provider behavior changes; do not generalize by model-name resemblance |
| `ouroboros/reasoning_artifacts.py::SIGNED_PORTABLE` and its sealed classifier | Which families' SEALED reasoning artifacts (signed, encrypted, redacted, unrecognized) survive a same-model cross-provider replay; readable artifacts are portable by shape for every family | High; an upstream can bind a reasoning artifact to its endpoint without a routing-contract change | A short vouched family roster plus a shape-first classifier that fails closed on artifacts it cannot read | A same-model cross-provider replay probe of the exact family | A false positive 400s the replayed turn (the reactive strip-and-retry is the net); a false negative pins a portable transcript to one endpoint and forfeits same-model failover | Extend the roster only by a fresh cross-provider replay probe of the exact family, never by model-name resemblance; `openai/` was removed on 2026-07 field evidence despite an earlier passing probe |
| `ouroboros/provider_models.py::_ANTHROPIC_MODEL_ALIASES` / `migrate_model_value` | Direct-provider id spelling compatibility | Medium as providers rename ids and prefixes | Shipped compatibility mapping and current direct-provider id contract | Exact provider catalog/documentation can confirm a current id, but cannot establish whether a saved spelling was intentional | Removing an alias breaks upgrades; guessing aliases can silently reroute | Keep explicit compatibility aliases until a separately documented retirement window closes |
| `ouroboros/server_runtime.py::_RETIRED_MODEL_DEFAULT_REPLACEMENTS` and scope prior/legacy defaults | Which formerly shipped defaults are upgraded automatically | Release-dependent | Release history plus current `SETTINGS_DEFAULTS`; only known former defaults are migrated | A live catalog can show availability, but cannot infer user intent or whether a saved value was a default | Over-broad migration overwrites an explicit owner choice | Keep release-scoped exact replacements and regression tests; review retirement separately |
| `ouroboros/pricing.py::get_pricing` and `ouroboros/llm.py::fetch_openrouter_pricing` / `fetch_cloudru_pricing` | Exact-route model tariffs | High; pricing and FX drift independently | Exact provider catalog with nullable unknowns; provider-settled usage wins | Bounded live catalog fetch and provider-reported settled cost | Static prices look authoritative after becoming wrong and can corrupt admission | Preserve the live nullable design and cover it by regression; do not restore runtime tariff tables |
### Provider Independence
One configured provider must be sufficient for the agent loop, commit review,
scope policy, safety, and context/memory flows. Core capability must not
acquire a hidden OpenRouter or second-provider dependency. (CHECKLISTS item
2(h) and ARCHITECTURE both point here; this is the SSOT sentence.)
Tool-schema changes are provider-contract changes. Every shipped built-in
schema must pass general JSON Schema and the known cross-provider subset over
the complete registry; trusted integration CI sends that same registry in one
bounded tool canary per supported provider family/API surface, in the transport
Main uses (streaming for compatible routes), while pull-request CI remains
secretless. Malformed native arguments and invalid schemas stay red, with
diagnostics limited to structural facts, hashes, and parse position. Do not add
a prose parser, provider hop, or unbounded retry to make that contract green
(canary anatomy: ARCHITECTURE "CI topology").
When adding or changing a provider, update one coherent route contract:
1. credential/readiness detection and exact model-id migration;
2. Main/Light/Fallback and reviewer-slot defaults without overwriting explicit
owner choices;
3. canonical tool/reasoning/image/cache intent at `llm.py`, with provider wire
projection and exact-route recovery delegated to the small transport leaves;
4. nullable pricing/settlement and truthful capability omissions;
5. review and scope routing, including sourced context-window evidence;
6. direct-provider and single-provider regression tests;
7. record the route's real streamed wire into `tests/fixtures/llm_wire/`
(redacted: opaque reasoning payloads and signatures truncated) and replay it;
a hand-written stream fixture is not evidence of a provider dialect. Record
with `curl -N` against the route, or export the private `physical_stream`
blob the runtime already retains for every stream.
Local-only installs keep their local route. Unreachable shipped remote defaults
may be cleared, but explicit owner values are not. Scope authority follows
BIBLE P3: owner-selected Max requires the applicable sourced window evidence;
owner-selected Low records the declared skip rather than pretending a partial
review occurred. Current model ids and defaults belong in code/config, not in
this handbook.
Use `provider_models.ACTIVE_MODEL_SETTING_KEYS` for any new active consumer
(provider detection, model catalog/provenance, credential planning, Provider
Test); `LEGACY_MODEL_SETTING_KEYS` exists only for migration/history. In
particular, `OUROBOROS_MODEL_HEAVY` and the paired `USE_LOCAL_HEAVY` may seed
an explicit configured API actor while the canonical list is absent, but must
never become an active slot, startup-readiness signal, test probe, or fallback.
Do not patch each consumer with its own Heavy exclusion; preserve the shared
split.
The `-pro` suffix is an OpenRouter routing slug, not an official OpenAI model
id; a direct OpenAI Chat slot uses the plain Sol id, because projecting the
slug into Chat Completions would turn an owner route choice into a guaranteed
404. This is a compatibility constraint, not a mutable capability table.
Provider-specific optional features may be unavailable on another single
provider, but the core loop must degrade explicitly rather than crash or
silently reroute.
Canonical assistant history and tool schemas are function-shaped across
providers; do not add a second stored transcript for a provider dialect. Direct
OpenAI tool conversations stay on Chat Completions — custom-first when
non-`none` reasoning is requested, an exact custom rejection may fall back to
function with the same effort, and explicit `none` is a task-local last resort
— and send `reasoning_effort` and `max_completion_tokens` provider-wide;
model-name prefixes are not admission authority. DeepSeek is the second
effort-carrying route (`reasoning_effort` beside the compatible-lane
`max_tokens` carrier): the canonical tiers are projected onto its documented
`low`/`high`/`max` enum at the send boundary (`minimal`→`low`,
`medium`/`xhigh`→`high`, `ultra`→`max`), `none` becomes
`extra_body.thinking.type=disabled`, a forced tool choice (`required`/named) is
served with thinking disabled because thinking mode accepts only `auto`/`none`
(live-probed 2026-09-03), and every projection that changes the tier is
disclosed on usage as `reasoning_effort_clamped`. The carriage is keyed on the
provider id, never a model-name prefix or a target capability field, so a
hand-built target cannot silently drop it.
Apply static, semantics-preserving wire normalization before request-wire
binding; keep learned recovery and its evidence store separate.
All learned request-shape adaptation goes through the one provider-neutral
request-wire driver (`ouroboros/request_wire_contract.py`: exact-route
identity, closed action vocabulary, shared TTL, never executes provider prose
or switches route); do not add a second driver, and explicit `none` is never
durable. Direct Anthropic is the deliberate exception to a purely
reconstructed provider transcript: one private, route-bound byte-for-byte
replay receipt of the unfinished native tool turn
(`ouroboros/anthropic_native_custody.py`), scrubbed on any
provider/endpoint/API/model change and fenced from compaction. Do not
synthesize an effort-to-`budget_tokens` policy.

View file

@ -0,0 +1,325 @@
# Module Size & Complexity
This chapter owns the size discipline — the deterministic line, function and byte gates, the debt manifest they read, and the rule that a cap is paid down by simplifying where the change lives rather than by extracting a passthrough — together with the invariants that keep a growing system readable: projection over replay for hot readers, the source-complete decision pipeline, continuation authority, disposers for UI resources, and the geometry and refresh contract for embedded surfaces. It exists because every rule here is about bounded reading cost, and each names its enforcing surface or discloses that it has none.
P7 makes context fit a maintenance constraint, not a line-count aesthetic.
- Python modules everywhere (including `tests/` and `devtools/`) and
first-party `web/**/*.js` modules (including `web/tests/`) target roughly
1000 lines. The deterministic hard gates: 1600 lines per module (exact-path
debt in `ouroboros/size_ratchet_manifest.py::GIANT_PATHS`; stale or newly
oversized entries fail), 300 lines per non-grandfathered Python function
(`FUNCTION_DEBT`, exact `(path, qualname)` keys), 200,000 UTF-8 bytes per
module (`BYTE_DEBT`, shrink-only), and the exact-current 1001–1500 band in
`BAND_PATHS` (a new or re-entered path requires a nonblank rationale).
Regenerate the manifest with `scripts/regenerate_size_ratchet.py`; it
validates the rendered candidate before writing and refuses an unmerged
index with a typed error. Sources decode as strict UTF-8 and normalize to
POSIX LF before counting, so checkout policy cannot change the inventory;
vendored/minified assets are excluded, and the same production iterator
drives the gates, smoke, health, and census.
- Paying down a size cap: when a change runs into a band, the hard cap, the
function-size gate or the byte debt, pay it down by SIMPLIFYING where the
change lives — simpler control/data flow and interfaces, dead code and
duplicates removed, an existing SSOT reused, prose made compact and legible,
so the module reads BETTER after the change. Extracting a helper, a
passthrough wrapper or a neighbour module is the LAST resort, and a paydown
only when the new unit is a natural boundary that would stay correct with the
parent far below the cap: its own reason-to-change, an explicit contract, a
real caller. A cap-driven bucket, a one-caller passthrough, and bytes bought
by deleting contract-bearing comments, docstrings, messages or tests are
defects, not paydown — report the conflict instead (BIBLE P7 «When adding a
major feature — first simplify what exists»; the scope-review advisory item).
- Methods above 150 lines and more than eight parameters are decomposition
signals (BIBLE P7, CHECKLISTS item 2(c)), not deterministic gates. Existing
baseline debt is not retroactively a failing tree. JavaScript currently has
only the module line-count gate.
- Runtime Python function/method count stays under
`ouroboros/review.py::MAX_TOTAL_FUNCTIONS` (that iterator excludes
tests/devtools; the module gates include them). The ceiling is a high-water
alarm with ample headroom; raising it requires a one-line campaign rationale
in the same commit.
- Enforcement: the OFFICIAL repository's CI runs the dedicated `size_ratchet`
pytest lane as a blocking third step in quick-test and full-test — manifest
exactness against the tip tree plus the pairwise shrink-only transition
against the event base (`OURO_SIZE_RATCHET_BASE_REF`). Local surfaces never
block on size: the default pytest lanes exclude the marker, and
`check_worktree_readiness` plus `codebase_health` report the same
`validate_size_ratchet` findings as "official CI will enforce" warnings.
There is no committed-history replay: the previous manifest resolves
merge-aware from `HEAD` or any of its parents, and a checkout with no
committed manifest anywhere bootstraps from its own tree — so a locally
evolved fork can always take an official update without being trapped by
structural debt it inherited, while the official line keeps ratcheting.
### Pragmatic SOLID
SOLID is a direction for making changes legible to future agents, not a demand
for classes or extra framework surface:
- **SRP — Single Responsibility Principle:** keep one coherent reason and one
clear authority for a unit to change.
- **OCP — Open/Closed Principle:** extend an existing stable seam when it
preserves the contract instead of rewriting unrelated callers.
- **LSP — Liskov Substitution Principle:** an implementation or backend must
preserve the caller-visible behavior of the contract it implements.
- **ISP — Interface Segregation Principle:** consumers should depend only on
the capabilities they actually use, not a broad convenience interface.
- **DIP — Dependency Inversion Principle:** policy should depend on small,
host-owned contracts rather than provider-specific or concrete details.
Apply these principles pragmatically. They do not require a class hierarchy,
DI container, numeric score, AST analyzer, or a new review pass. A SOLID or
minimalism finding must name the exact symbol or authority, the concrete
duplication or coupling, and a smaller alternative that still satisfies the
contract. Diff size, line count, and file count alone are not findings.
Enforcement: review-only — CHECKLISTS item 2(d) scores these rules in commit
review.
### Invariant: Projection over replay (hot readers of growing stores)
A reader that runs per INTERACTION — an HTTP request, a WS/SSE message, a poll
tick, a task turn — must not replay a growing store to produce its answer.
Interactive read cost must be O(response), achieved through a maintained
projection, a cursor, rotation, or a bounded tail — never a full-history scan
filtered down to the answer.
- **Per interaction is the unit.** Work that runs once per boot or per explicit
owner action may scan history; work on a request/message/poll-tick/task-turn
path may not. A scan that is cheap today is not the point — every growing
store crosses the threshold eventually, and the reader degrades exactly when
the system is most used.
- **Storage-agnostic.** A full-table read filtered in code IS a replay (a
`SELECT *` narrowed in Python is the same failure as parsing a whole JSONL
file for its tail), including unbounded collections INSIDE snapshot/state
files.
- **Passive GET.** Read handlers perform no NEW steady-state durable writes.
Exactly two named exceptions exist: (1) substrate-owned integrity repair
under the substrate's own lock (the usage-ledger torn-tail quarantine in
`ouroboros/usage_ledger.py`), and (2) one-time idempotent migrations guarded
by a durable watermark (the legacy usage import). Anything else that "just
materializes a bit of state" on a GET is a mutation hiding on a read path.
- **House precedents — reuse these shapes:** chat log rotation with
archive-aware readers (`supervisor/state.py::rotate_chat_log_if_needed`);
the compact `containment_faults.jsonl` projection maintained beside an
unbounded event log (`ouroboros/delegate_custody.py`); one shared custody replay
per context build and per terminal audit (`delegate_terminal.custody_audit_snapshot`,
consumed by `context_health.build_health_invariants` and `_audit_task_custody`): several
projections of the same growing store share ONE traversal instead of replaying per reader,
which bounds the multiplier, not the scan: the read stays O(history) until a compact
projection replaces it; the fingerprint-keyed
render cache in `ouroboros/_usage_rows_memo.py` — a projection cached while
its input is unchanged, invalidated only by advance/refold, never by TTL.
Interactive result discovery reuses `gateway/task_list_scan.py`'s compact
stat-invalidated facts; decode changed/new files, never cache failed or torn
reads, and fetch selected full rows through the existing schema owner. Live
event followers retain per-file proven lineage across failed reads only; valid
moves, role/schema changes and deletion remove it. Disclose incomplete discovery
with nonfatal history_gap coverage, never client-root trust or a global error
for one unrelated bad file. Diagnostics do not advance legacy ranks or bytes.
Task-event v2 cursors advance after each consumed row; preserve physical byte
positions through rotation and disclose replay/gaps rather than resetting an
unavailable cursor. Keep the legacy GET rank contract separate.
Bound native handle lifetime independently of network backpressure: close
each JSONL read buffer before yielding its per-row cursors. Stamp creation
only where the host allocates a fresh id, before preparation; neither a
supplied id nor a legacy/final result timestamp proves task creation.
Pin each source's logical EOF and helper-built chain metadata until the pass
ends; select subsequent buffers without rescanning the archive prefix. Rebind
the original live inode after rotation; later appends belong to the next pass,
without truncating history or imposing an event-count cap. Whole-chain callers
retain their ordered handle list and share JSONL parsing through borrowed
handles without changing descriptor ownership.
Enforcement: Repo Commit Checklist item 24 (advisory) triggers on diffs that
add or change an endpoint/poller/subscription/timer or read a growing store;
the hot-store growth health invariant
(`agent_startup_checks.py::hot_store_growth_notes`, surfaced by
`context_health.py::build_health_invariants`, thresholds justified in
`ouroboros/context_budget.py`) is the deterministic runtime tripwire. A change
that introduces a new append-only store read on an interactive path must
enroll that store in the `ouroboros/context_budget.py` threshold table (with a
justified constant) in the same commit — an unenrolled hot store is invisible
to the tripwire. Retained execution drives under both `state/headless_tasks`
and `task_drives` are enrolled by direct-child count at
`context_budget.RETAINED_EXECUTION_DRIVES_WARN_COUNT`; startup never recursively
sizes those trees.
### Invariant: Source-complete decision pipeline
Every new or changed continuity surface is reviewed as one narrow chain:
`producer → canonical full source → bounded projection → consumer → decision → retention/GC`.
- The producer records the complete event or artifact before it publishes a
projection or wakeup. The canonical source owns identity, order, bytes, and
integrity state; a cache or hot index is never a second authority.
- A bounded projection names what it omitted and carries a source reference
that the *same actor* can resolve through an existing reader.
`source_complete` is a coverage fact, not a permission to infer missing
material.
- Every over-limit tool result persists its exact source; there is no per-tool
exemption, because the DECIDER must be able to resolve what the actor could
page. A bounded row whose exact source is durable and referenced is an
omission for the acceptance panel, and only `source_unavailable` — no
actor-resolvable source at all — is an unresolved partial that withholds
dispatch. An `api_chat` acceptance reviewer has no tools and cannot resolve
`repo_diff_source_ref`, so a criterion that depends on the unseen part of the
diff is at most `partial`.
- A consumer that can authorize PASS, a destructive rewrite, or replacement of
a full contract must materialize the named source first. A known `partial`
marker and an unverified claim that some host might retrieve more are not
equivalent: the latter is not actor-attested coverage and cannot release the
decision.
- Retention and GC are part of the chain. Anything referenced by a canonical
result, review, identity decision, or project summary is retained or promoted
before its execution root can be collected; an unavailable legacy source is
represented as an explicit gap, never silently treated as complete.
**Control-plane distrust is metadata, not a data-plane operation.** Paid model
output is evidence until a typed validity predicate fails. Control-plane
distrust — profile, route, parser, window — may lower authority to
DEGRADED/SKIPPED/NOT_RUN, but it must not blank, rewrite, or relabel the
artifact or its original cause.
Enforcement: CHECKLISTS item 25 `source_completeness` (critical when
applicable) scores the chain in commit review; the presentation-adapter
contracts below are pinned by the named web tests.
#### Review presentation adapters
`web/modules/review_presentation.js` (grouping, identity, ordering, typed
presentation state), `ouroboros/review_execution_projection.py` (the bounded
cross-domain `executions[]` wire), `web/modules/review_dom_patch.js` (keyed
in-place DOM reconciliation) and `web/modules/harness_presentation.js`
(harness identity marks and labels) are pure read-side presentation — they
never author, mutate, or feed back canonical verdict, lifecycle, routing,
attention, or enforcement authority, and never infer that a requested route
executed. Admission is source-complete: an incomplete row is omitted, never
guessed from chat, repository, timestamps, model, tool name, or activity.
Reuse the existing chat-history, task-detail, and canonical physical-attempt
readers; do not add a review ledger, endpoint, persisted UI state, cost copy,
or enforcement layer. Compact review rows carry no dollars; exact Skill
attempt money appears only in the lazy detail when the history row declares
`physical_attempt_v1`, joining the canonical ledger by exact wave and slot.
Reconnect and folded-group bounds exist so one history rebuild cannot fan out
unbounded task-detail reads. Pin these contracts in
`web/tests/review_presentation.test.js` and
`web/tests/harness_presentation.test.js`; module headers carry the per-module
contracts.
#### Context and growth matrix
| Store / surface | Complete producer and source | Interactive projection / consumer | Growth and retention proof |
|---|---|---|---|
| Background observations | `BackgroundConsciousness.inject_observation` → `state/consciousness_observations.jsonl` enqueue rows | Cached pending/oldest status and bounded `_render_observations` view; identity-update consumer reads the gap marker and source ref | `BG_OBSERVATIONS_WARN_BYTES` in `context_budget.py` / `agent_startup_checks.py`; append-only rows, including unacknowledged rows, are not GC-pruned by the hot-store warning |
| Chat and biography | Canonical `logs/chat.jsonl`, rotated generations, and dialogue blocks | Main/Project context and archive-aware `chat_history` | Rotation/archive readers carry generation/gap coverage; blocks are the compression path, not a deletion of the horizon |
| Plan/review evidence | Exact task-artifact/observability bodies and reviewer route/thread receipts | Bounded review hot index, obligations, and latest-wave status | Exact artifact refs and candidate SHA bind the decision; index rotation cannot certify a missing or partial wave |
| Skill-review root tasks | Per-skill `state/skills/<name>/review_history.jsonl`; `skill_review_runner._append_terminal_history` projects terminal identities to `state/skill_review_root_tasks.jsonl` | `skill_readiness._skill_names_from_review_history` reads a bounded newest-first suffix for acceptance | Derived index is append-only and idempotent by root/task/outcome identity; `SKILL_REVIEW_ROOT_TASKS_WARN_BYTES` warns at 20 MB |
| Task/project execution | Canonical task result plus promoted child artifacts and summaries | Status cards, terminal rows, and Main/Project summary projections | Canonical promotion precedes child-drive GC; disposable task scratch follows the unified retention owner |
### Invariant: Continuation authority and bounded Main projection
Continuation is an explicit relation, not an inferred chat-memory feature. The
router contract requires `predecessor_task_id`: an empty string means a fresh
task, a non-empty value means continuation, and omission or `null` is a typed
refusal before any lookup, enqueue, or provider spend. Queue snapshot/restore
retains the predecessor source, so a restart cannot silently turn the selected
task into a fresh one.
The authored continuation narrative is written at the result owner together
with its exact `get_task_result(include_authority=True)` source. Main's
provider projection is defensive: it deep-copies the authority, removes only
the current task's duplicate nested predecessor, and thresholds only the
closed raw keys `result` and `final_answer` using
`context_budget.PREDECESSOR_RESULT_INLINE_CHARS`; oversized values resolve as
persisted narrative, bounded exact-key legacy lookup, or an explicit
source-resolvable gap — never a raw head/tail slice, an invented summary, or a
mutation of the canonical result.
The startup injection is a bounded continuation ENVELOPE, not a body copy,
minted by the one producer `contracts.task_contract.bounded_continuation_envelope`:
the predecessor's contract core inherits without its nested
`predecessor_authority`, and every field is whole-or-pointer against one strict
serialized budget (previews carry `full_chars` plus a named `source_ref`;
`previous_task_id` keeps the chain walkable). Durable `task_results` bodies are
the untouched SSOT. The bound is per-field, so a pathological row can still
exceed the wire budget — the refusal is typed and loud rather than a silent
$0, and no hop cap exists anywhere: depth belongs to the mind, the floor only
keeps bodies off the wire.
Provider context overflow is a typed recovery fact: after the useful reclaim
and one strictly-smaller same-route retry, a final `context_overflow` skips the
provider-unavailable/forced-provider path, keeps
`execution_status=infra_failed` and `reason_code=llm_api_error`, and records
the typed acceptance bypass and `failure.error_kind`. Ordinary provider
outages keep their existing recovery behavior. Enforcement:
`tests/test_continuation_context_authority.py`.
### Invariant: UI resources carry a disposer
Every long-lived acquisition in `web/` returns or records a disposer, and a UI
instance owns a `destroy()` that releases everything the instance acquired.
The resource kinds this covers: WS subscriptions (`ws.on(...)`),
`document`/`window` event listeners, observers (`ResizeObserver`,
`MutationObserver`, `IntersectionObserver`), timers, `requestAnimationFrame`
loops, and `EventSource`/streaming connections.
An instance that can be closed, hidden, or replaced (project chat panels are
the canonical case) must be destroyable without leaving any acquisition
behind. A UI instance may survive being hidden only under an explicit,
owner-visible retention reason — a project chat with pending work, a widget
card the owner set to Keep running — and even then it owns its disposer, and
Stop / unload / reload / shutdown remain force-destroy boundaries; the reason
is re-evaluated at the instance's next lifecycle point, not continuously. The
untyped shape "hide the DOM node, keep the handlers" remains the leak this
invariant forbids. Late async continuations check a `destroyed` flag before
touching state or re-arming loops. A module widget's disposer is the ordered
dispose with acknowledgement (ARCHITECTURE "Skills and Widgets"): post the
dispose message, keep the bridge answering the child's hooks, then abort,
unlisten and remove the iframe on the acknowledgement or after
`WIDGET_DISPOSE_ACK_TIMEOUT_MS`; a route iframe disposes synchronously; the
masonry's `applyMasonry` returns an idempotent disposer for its observers and
pending frame. That bounded wait is not the forbidden shape: the handlers live
only until the settle promise the page tracks per card key resolves.
Enforcement (honest disclosure): the deterministic leak test runs in the
release-tier `ui_browser` lane, not at commit tier; commit-tier coverage is
the advisory Repo Commit Checklist item 24. The class is closed
deterministically for the instrumented surfaces and advisorily for future
ones.
### Invariant: Embedded surfaces declare geometry and refresh semantics
Every owner-visible embedded or framed surface has an explicit host-owned
geometry/overflow contract, a paired disposer for every long-lived resource,
declared refresh/stream/error semantics, and a named real-consumer visual
verification path. Intentional omissions record why they are safe to defer.
For Widgets, framed `height` values are bounded and module auto-height is
host-controlled. Below its finite ceiling, applying a reported block size must
not change the child's inline-size basis; the host owns vertical scrollbar
mode without disabling the orthogonal horizontal overflow capability, and
content measurement includes the measured document's bottom padding and border
(three separate feedback-loop bugs encoded as one rule).
Feedback-sensitive verification is event-driven on the relevant engine: it
proves temporal convergence to a quiet fixed point with a real consumer or
production-derived fixture that crosses the known wrapping threshold, rather
than comparing two snapshots. A module widget's own faults are declared error semantics, not silence: an
in-frame script error, an unhandled rejection or a CSP refusal reaches the
card's status slot as one bounded, deduplicated line, while the lifecycle
state stays running because the frame really is still mounted. There is no
server-side widget fault ledger, so those faults are visible only while the
Widgets page is open; that is safe to defer because the browser is the
verification path for a widget in the first place.
Module source loading and declarative requests
have a bounded host timeout; declarative job widgets keep their `job_id` and
bounded retry/timeout behavior visible in the refresh contract. Missing or
malformed job status is an immediate protocol error, while unknown non-empty
in-progress labels remain bounded pending states for producer compatibility.
Repo Commit Checklist item 24 points lifecycle changes here instead of
re-deriving a second domain-specific rule; the widget geometry/refresh
contracts are pinned in `tests/test_widgets_ui_static.py` and
`tests/test_extension_surfaces.py`.
---

View file

@ -0,0 +1,274 @@
# Core Governance Artifacts
This chapter owns the availability contract for BIBLE, ARCHITECTURE and DEVELOPMENT in every reasoning flow: the per-flow delivery registry, the one structural fact by which plan review tiers its governance pack, and the truncation invariants that make an omission visible instead of silent. It exists because a reviewer operating without the architecture map is not operating with full context, and the only honest response is to disclose that rather than to shrink the pack quietly.
`BIBLE.md`, `docs/ARCHITECTURE.md`, and `docs/DEVELOPMENT.md` are **core
governance artifacts** — the constitutional, architectural, and procedural
ground truth of the system.
### Invariant: Full availability in reasoning flows
Any flow that requires architectural, constitutional, or procedural reasoning
MUST include these artifacts as **first-class context sections** — not as
optional or opportunistic inclusions via touched-file packs.
Plan review is the one flow whose governance pack is tiered, and by ONE
structural fact — whether a declared `affected_paths` target (a file the work
will CHANGE) resolves under the Ouroboros system repository; `affected_resources`
is prose the host never resolves, and an `evidence` path is a read, not a change —
never by prose and never by a plan-kind
taxonomy, which is what keeps classification un-gameable. This is a tiering,
not an omission: before any work exists the reviewer's subject is the
INTENTION, and every absence is a named pointer or a typed `need_evidence`
finding the host attaches on the next cycle under the same evidence policy —
the locator enters the manifest hash, so the next envelope is a new
fingerprint, never an idempotent replay; nothing is silently omitted (P1).
Two branches follow, and only one of them stops a review. A REQUIRED
governance pack that cannot be assembled is a typed assembly failure and the
review does not run (`PlanPacketError`). Declared or reviewer-requested
evidence the policy cannot attach is a named absence instead — a
`[reviewer-requested]` omission row, or the head attached with the cut named
`truncated_to_<N>` — and the panel still runs and judges with it; a re-asked
locator stays a `need_evidence` request with a `need_evidence_repeat` disclosure,
requiring the same free disposition without expanding request memory or paid cycles.
DEVELOPMENT.md is not resident in a plan-review packet; it is one such request
away. Packet composition, bounds, and wave/replay mechanics: ARCHITECTURE
"Plan construction and review" and `ouroboros/tools/plan_packet.py` /
`plan_spec.py`. Planning room evidence comes from `dialogue_evidence.py` over
`Memory.read_chat_generations` and the shared `project_dialogue.room_membership`;
`plan_dialogue.py` binds its redacted source to author-request identity. Do not
substitute the bounded post-consolidation reader or acceptance directive ledger.
An identical replay keeps the recorded source; real plan/evidence changes capture
current discussion. Follow the per-delivery context/source contract in ARCHITECTURE:
API window fit, native mandatory-read bound, and delegated harness-owned reading
with full immutable files and honest coverage. No independent dialogue byte cap.
Exact-wave custody is fail-closed: the evidence continuation uses a fresh
full-packet dispatch only when no exact artifact reference exists. An unreadable
referenced artifact returns `plan_review_exact_artifact_unavailable`; it never
mints replacement authority.
The context-delivery registry:
| Flow | BIBLE.md | ARCHITECTURE.md | DEVELOPMENT.md |
|------|----------|-----------------|----------------|
| Main task context (`context.py`) | full tier-0 | full in Max for every task class; lossless navigation map in Low | mode-independent: full when the active binding targets Ouroboros's system repo, including evolution/self-body work and a project-room turn without an external binding; visible on-demand pointer for a bound external workspace, subagent, or API/CLI/scheduled external surface. `workspace="none"` and explicit self-body overrides retain full Development. |
| Triad review (`tools/review.py`) | ✅ via preamble | ✅ via `load_governance_doc` | ✅ via `load_governance_doc` |
| ↳ Cold-start density rung | — | — | Shared with scope review and the packed deep self-review (`capability_evidence.cold_start_density_probe` on a `review_helpers.density_probe_sample` slice, wired at the admission seam `review_admission.density_probe_before_size_refusal`): a packet that would be refused or degraded for size while its route has no fresh exact-model density witness gets one bounded probe send on the exact model, then one re-size/rebuild; a budget-refused probe is a typed disclosure and the existing refusal stands. |
| ↳ Anti-thrashing | — | — | Open obligations loaded from `review_state` via `load_state(drive_root)` + `make_repo_key(repo_dir)`, injected unconditionally into `_build_review_history_section` prompt context. Same mechanism in `scope_review.py::_build_scope_prompt` (best-effort when `drive_root` available). |
| Background consciousness (`consciousness.py`) | ✅ full | ✅ full (max) / navigation map (low) | — (not yet required) |
| Advisory pre-review (`tools/claude_advisory_review.py`) | Two delivery classes: an `api_chat` row runs the bounded NATIVE inspection episode (governance docs reached through its read-only tools); an `agent_session` row receives a resolvable pointer marked MANDATORY FULL READ and the session reads the full doc itself — retrieval is disclosed (native reads are host-observed; vendor-session reads are not) | same two delivery classes | same two delivery classes |
| Scope review (`tools/scope_review.py`) | full canonical doc + Atlas accounting; a size terminal or a degradation rung under a cold density cap takes the shared cold-start density rung above (one probe, one rebuild, `density_probe` ladder step) | full canonical doc + Atlas accounting | full canonical doc + Atlas accounting |
| Skill review (`skill_review.py`) | full inline (`api_chat`) / mandatory full source-root read (`agent_session`) | full inline (`api_chat`) / mandatory full source-root read (`agent_session`) | full inline (`api_chat`) / mandatory full source-root read (`agent_session`) |
| Plan review (`tools/plan_review.py`) | full for a SELF-MODIFICATION plan (structural path fact: a declared `affected_paths` target resolves under the system repo, whether or not the file exists yet; an `evidence` read of a repo file does not); otherwise a heading-derived navigation map of BIBLE.md generated at runtime (never a copy) | inline, in full, for a self-modification plan; otherwise the lossless navigation map + a resolvable pointer (W3) | named on-demand pointer; a reviewer that needs it returns `need_evidence` (an exact `::lines=A-B` range for one section) and the host attaches what the evidence policy allows on the next cycle, naming every absence |
| Deep self-review (`deep_self_review.py`) | Three deliveries on the `deep_review` row. Packed api row: full canonical doc + Atlas accounting; a required set unfit under a cold density cap gets one bounded probe send and one rebuild, and a pack still unfit is the typed `deep_self_review_pack_unfit` refusal asking the owner to switch the row (no automatic fallback). Native inspection episode / agent session: MANDATORY full read at the repository root, named with its size in the task; the native episode's `read_file` receipts carry the delivered line extent and the host merges the repository-root intervals for BIBLE.md afterwards — `read` only on full coverage, else `partial`/`missing`/`unobserved` — disclosed in the report header and as a typed `capability_delta` (host-observed reads); a session's reads are `unobserved`. Memory (up to seven whitelisted files) is INLINED byte-exact on every delivery, each entry's disposition disclosed (task text, `deep_review_memory`, header `memory=n/7`) — never receipt-checked | Packed: full (max) / navigation map (low) + Atlas accounting. Retrieving rows: navigation map (`generate_doc_nav_map`), sections read on demand | Packed: full canonical doc + Atlas accounting. Retrieving rows: navigation map, sections read on demand (CHECKLISTS.md likewise) |
Skill Review keeps the full stable governance/host prefix for cache-friendly
API rows; a retrieving session reads those same canonical files from its
source-repository root and receives the byte-exact dynamic tail inline, so the
payload snapshot and per-chunk quorum stay identical without rebilling or
crowding the session window.
Planning has two distinct roots: governance documents are always loaded from
the system repository, while declared targets and evidence locators resolve
against `active_repo_dir_for(ctx)`. Exact user-managed installed-skill payload
paths are the one data-plane exception for CLASSIFICATION only — they never
make a plan a self-modification — and are not attachable as evidence: the
resolver allows only the active workspace and the system repository, so a
payload locator comes back as a named `denied_path` omission. Any declared
path escaping the active subject, a workspace/subject mismatch, or an
unavailable root fails loudly with a named omission. Do not fall back to
reviewing the Ouroboros repo for an external plan.
The SPEC must state the goal, acceptance claims, invariants, in-scope and
non-goals, the load-bearing decisions with their rejected alternatives, and
what is consciously deferred. Plan review publishes exactly `GREEN`,
`REVIEW_REQUIRED`, `REVISE_PLAN`, or the honest `DEGRADED` (no quorum,
or a paid actor still in flight);
findings are inputs the main agent may accept, reject, or defer. A
`need_evidence` finding names a locator the host attaches on the next cycle
or, as a question to the author, the spec id it is about in `breaks`; the
author answers it in the disposition (accept), rejects it, or defers it
openly, and its answer reaches the reviewers only on the next paid cycle.
Optional
`note` findings, including useful premise criticism and simpler alternatives,
remain readable but need no adoption or disposition; a note-only wave closes
immediately under both enforcement modes. The same `review_disposition` call may
voluntarily annotate a current closed note-only wave; it keeps the verdict and
spec unchanged, retains the previous immutable artifact, and never reopens the
review or buys another panel. Outstanding `need_evidence` can close
without a second LLM call through a separate `plan_task` call containing `review_disposition` only —
`{review_fingerprint, items: [{finding_id, decision, rationale}]}` — covering
each required finding exactly once; duplicates, contradictions, unknown, stale, or
incomplete required dispositions fail closed, and mixed or vacuous calls fail before an
attempt is recorded as typed argument errors — an optional field that carries
no meaning beside a disposition (a blank `goal`/`plan`; a `spec` holding only
declared keys whose values are `None`, `""` or `[]`) is ignored, never mistaken
for a second operation, while any non-empty list, undeclared key or non-blank
string is meaning and makes the call mixed. Never replay the plan envelope with
the disposition.
Outside Cyber Pro, a real blocking finding, including one below quorum, stays open pending a paid
delta review of a changed spec or justified rejection, when capacity remains; advice
does not become a blocker through repetition. Blocking `REVISE_PLAN` likewise
requires another panel when another paid cycle is available; advisory
may proceed only under loud host disclosure and the agent's rationale.
Reviewers return findings, including optional alternatives, not a required
competing plan. A blocking finding must name the spec id it breaks, and there
is never a required number of findings.
Force-plan is an LLM-first pre-implementation obligation on the admitted
managed root, not a mechanical permission check. `plan_review_state` owns
durable review facts and `config.get_review_enforcement()` owns the configured
blocking/advisory value. Effective Cyber authority is separate and follows
BIBLE P0/P3; it never rewrites an old wave to GREEN. Every submitted envelope that reaches `plan_task`
supersedes prior authority, so a newer attempt cannot fall back to an older
GREEN. Wave recording, free replays, DEGRADED semantics, structurally dead
slots, and `quorum_unreachable` release live in ARCHITECTURE "Plan
construction and review"; paid cycles are bounded by the shared
`OUROBOROS_REVIEW_MAX_CYCLES`.
**Context mode (Nano / Low / Max).** `OUROBOROS_CONTEXT_MODE` controls the
Architecture projection in the agent's own context: Max keeps `ARCHITECTURE.md`
full for every task class, Low supplies its lossless navigation map, and Nano
uses the compact owner projection.
`DEVELOPMENT.md` is mode-independent and follows the active repository
binding — the signal is the repo binding, a path fact, never a guess from
message text (full for system-repo work including a project room with no
external binding; a bound external workspace, subagent, or API/CLI/scheduled
external surface receives a visible on-demand pointer; explicit structured
overrides win). Tier-0 identity and constitutional context stays full in every
mode. `context_fit.py` renders Max and Low from one immutable captured core on
one labelled density basis; Owner Low has an elastic 200K total-context target
whose miss is non-terminal. Predicted Max pressure never swaps in Low
documents — only actual provider overflow may use a task-local Low projection,
followed by at most one same-route strictly-smaller call, and this never
changes owner mode or P3 commit/scope review.
### Invariant: Exact premises with explicit source ownership
Planning reads the complete retained room through `dialogue_evidence` and
`Memory.read_chat_generations`, with the shared Project membership predicate,
progress and addressed mailbox provenance. Both speakers, options and answers
remain exact. JSONL records and chat line selectors use physical LF boundaries;
valid Unicode inside a message is never a record delimiter. Task acceptance
keeps `review_evidence._accept_owner_directives` over the task-local ledger;
planning does not recreate its retired bounded directive section.
Each consumer redacts at its boundary and discloses missing source or ranges.
A replay or an already-earned paid retry of the same author request keeps its
recorded snapshot and discloses that later messages were not reviewed. A changed
plan or explicit evidence request captures current discussion. Native sizing
measures the complete first request, including its schemas and instruction
wrapper; existing task-local model choices apply before fresh source fitting.
Enforcement: `test_packet_uses_full_dialogue_and_keeps_acceptance_directives`,
`tests/test_plan_dialogue_review_regressions.py`, and the acceptance ledger tests
in `tests/test_loop_misc.py`.
### Invariant: Compaction must earn its rewrite
Context compaction is a deficit-requested materializer, not an independent
threshold, timer, route, or retry policy. It first performs pure selection
over completed atomic units (one assistant tool-call message plus all and only
its contiguous matching results); user turns are hard boundaries; malformed,
missing, delayed, duplicated, visually opaque, or corrupt-capsule units remain
byte-identical. No eligible positive reclaim means no checkpoint, summarizer
call, or transcript mutation.
For a non-empty selection, persist the exact actor-visible checkpoint before
calling the summarizer. Summary input covers complete stable hashed chunks
with gap-free offsets; only typed summarizer context overflow may split a
source recursively. A replacement publishes only after transcript/unit
binding, complete coverage, checkpoint provenance, and a strictly smaller
representation on the caller's ContextFit measurement basis are all proved
(the bounded image proxy and density must match; raw base64 byte count is not
token reclaim). Capsules carry host-only generation, source-hash, part,
checkpoint, and CAS-ref metadata so a later pass can recompact them without
losing the original provenance union. Enforcement: `tests/test_compaction.py`,
`tests/test_loop_compaction.py`, `tests/test_loop_compaction_policy.py`.
### Invariant: No silent truncation
If a core governance artifact cannot fit in the available context budget:
- Do **not** silently omit it or truncate it without a visible marker. Either
adjust the budget/flow to accommodate it, or emit an explicit warning
(`⚠️ OMISSION NOTE: ARCHITECTURE.md omitted due to budget constraints`) so
the operator and the model both know the context is incomplete.
- A reviewer or agent operating without ARCHITECTURE.md MUST NOT be treated as
operating with full context — findings may be incomplete.
- Tools that return multi-model review findings (`commit_reviewed`,
`skill_review`, scope/advisory review helpers) MUST be listed in
`UNTRUNCATED_TOOL_RESULTS` or have an explicit per-tool limit; the default
15KB transport cap is not acceptable for review verdicts.
- A reference-doc **navigation map** (H2-H4 inclusive complete-subtree ranges,
with parent rows overlapping descendants and full sections one `read_file`
away) and a named on-demand pointer are visible, lossless representations —
NOT silent truncation. The low context mode uses these; it never applies
`[:N]` to a doc.
- String bounding goes through the SSOT `utils.truncate_review_artifact`,
never a hand-rolled `text[:cap] + marker`. Besides the marker, that helper
carries an anti-waste FLOOR: a cut saving fewer characters than its own
omission note is pure damage, so below it the text passes through whole. A
local re-implementation loses the floor and can return a value LONGER than
the input it "shortened". The two bounded-string primitives serve different
contracts: `truncate_review_artifact` produces DISPLAY previews (its floor
may return the text whole), while `truncate_within_limit` enforces a STRICT
wire/prompt bound — the omission marker lands INSIDE the limit and the
result never exceeds it.
- Bounding a LIST is subject to the same rule: a `[:N]` slice must be
accompanied by an explicit omitted COUNT, and — where the slice touches an
identity that something downstream compares — a durable hash or reference
for the full set (see `_outcome_receipts.receipt_identity_projection`).
Bounding a set is allowed; hiding that you bounded it is the P1 violation.
Enforcement: `tests/test_tool_capabilities.py` (the `UNTRUNCATED_TOOL_RESULTS`
roster) and the truncation-floor coverage in
`tests/test_owner_facing_honesty.py`.
### Invariant: Owner-facing surfaces show the full text
Disclosed truncation (the `⚠️ OMISSION NOTE` marker) exists to protect **LLM
context budgets** — it is a model-bound mechanism, not a licence to shorten
what the owner reads:
- **Owner/UI-bound surfaces** (chat panels, task_results projections, review
verdicts shown to a person) present the COMPLETE text, or carry a reference
to a durable full copy (e.g. an observability `response_ref`). Reviewer
rationale is a cognitive artifact (BIBLE P1): projecting it truncated while
the full copy sits unreferenced in private blobs is partial memory loss.
Terminal text asserts only recovery facts carried by the round record, never
a route mechanism that the selected transport cannot perform.
- **Model-bound projections** (review packs, context sections, tool-result
transport) keep their disclosed-truncation budgets — those are real context
economics.
- **A cut cheaper than its own marker is forbidden everywhere** (the shared
primitive enforces the floor — see "No silent truncation"). One named
exception: tiny single-line identifier fields (limit < 100, e.g. a
reflection backlog `kind`) keep a plain hard slice — a multi-line omission
marker inside a one-line value is worse damage than the cut it discloses.
Enforcement: `tests/test_owner_facing_honesty.py`.
### Invariant: No "only if touched" gate for core artifacts
Core governance artifacts reach review/reasoning flows unconditionally — NOT
only when they appear in `touched_paths`. `build_touched_file_pack` is for
_changed_ files; core artifacts are a separate concern loaded independently.
No surface of its own — the per-flow presence tests required below are the
mechanical cover; otherwise review-only.
### When adding a new reasoning flow
If you add a new flow that reasons about code structure, system architecture,
or engineering standards, you MUST:
1. Explicitly load `ARCHITECTURE.md` (and BIBLE.md if constitutional reasoning
applies).
2. Log a warning if the file is missing or unavailable — do not silently skip.
3. Add a test asserting the file is present in the assembled context/prompt.
That required presence test is the enforcing surface; CHECKLISTS item 11
(`context_building`, advisory) backstops the review.
---

View file

@ -0,0 +1,232 @@
# Review & Commit Protocol
This chapter owns the three stages of a reviewed commit — prepared preflight, the authoritative gate, and publication binding — together with the shared paid-cycle cap, the free-replay rules, the external-review evidence contract and the release-sync carriers a pull request must leave untouched. It exists because technical failure and commit permission are separate facts, and every rule here keeps a missing review from becoming a PASS.
Keep optional task evidence outside the stable governance prefix; shrink its excerpt before reducing existing review material. A source pointer gives a packet-only model no retrieval capability. Rejoin preserves the original hash and project-local view while any physical reviewer may still read it. Removing an ignored view never deletes the canonical source; no separate notes corpus, blanket ToolResult metadata or mandatory whole-history read belongs to this evidence.
Reviewed commits separate improvement evidence from candidate-bound authority.
Finish the edits and focused tests, then call `commit_reviewed`; standalone
`preflight_review` remains available when an earlier critique is useful.
`docs/CHECKLISTS.md` owns reviewer questions, severity and output contracts;
ARCHITECTURE "Review delivery" owns the dataflow.
1. **Prepared preflight.** Authorization and unresolved-work checks precede
mechanical file preparation, staging/classification/protection and the
fingerprint. Existing free-cycle/budget admission precedes any automatic
preflight. When needed, the same `preflight_review` runs inline with the full
`review_rebuttal` and the independently applicable test preflight. The
candidate must remain unchanged before triad/scope dispatch. Explicit
`skip_advisory_review=True`, disabled and unconfigured paths retain their
audited behavior. A free advisory replay reads freshness but buys neither
another preflight nor another triad/scope wave. Stale coverage still needs
the explicit audited skip; applicable compensating tests run even when the
reviewer backend is available. Explicit test skips are not green proof.
2. **Authoritative gate.** Independently configured deterministic test policy,
staged fingerprinting, triad review, applicable scope review, aggregation,
and pre/post revalidation. The fingerprint binds `git write-tree`, ordered
`HEAD`/`MERGE_HEAD` parents, indexed VERSION, expected `v{VERSION}` tag and
existing target, plus the binary staged-diff hash.
3. **Publication binding.** The created commit/tag is checked against the same
tree, parents, VERSION, tag, and reviewed fingerprint before push. Any
mutation, rebase, conflict resolution, or changed landing parent
invalidates exact-candidate authority and requires the applicable final
gate again.
A technical review failure may permit continuing under owner-selected advisory
enforcement on a known, independently bound candidate. The failure's phase,
reason, received findings and full result stay recorded as failure, never PASS.
Outside Cyber Pro, Blocking enforcement still blocks. Cyber may continue
without prior review, preserving the original findings, missing sources and
pending invocation. It never fabricates candidate bytes, completed custody or
physical effects; Stop, deadline and cost facts remain independently recorded.
Diagnostic
`repo_commit_ready` projects this permission only from an exact repo/hash match;
it does not change the failed review's status or freshness.
Pending triad/scope reconciliation retains the prepared index and never restages
or reconstructs a lost index. Pending delegated preflight binds the existing
`delegate_custody` invocation in `AdvisoryRunRecord.execution` before POST and
replays its canonical request, without another prompt store or physical run.
An explicit audited preflight skip releases only logical admission: retain the
old invocation, source and unknown physical work. It neither cancels that work
nor posts a replacement. A late result updates its original record without
invalidating the newer bypass. Only the existing definite start-failure producer
can automatically discharge a stranded checkpoint; elapsed time cannot.
Both commit and review-only entry points forward the explicit skip. A subsequent
standalone request can rejoin exact historical custody or check new evidence;
released unrelated history is not a logical lock. Pending refusals explain the
audited skip without claiming that it cancels physical work or its cost.
Native preflight uses the existing executor's end/failure events and monetary
ledger, without a new advisory operation checkpoint or exact-rejoin protocol.
An ambiguous returned native outcome remains failed evidence; explicit skip
does not erase its cost or custody. The external review wrapper uses
the same cycle and retains its candidate
checkout/index while custody remains unresolved; its existing outcome records
the checkout, drive root, reason and invocation for reconciliation.
Triad slots review the staged diff against `docs/CHECKLISTS.md`; duplicate
model ids remain independent slots and `config.adaptive_quorum` owns quorum. A
managed-update resolution commit reviews the declared M0→S resolution delta
(`tools/review_subject.py`), bound to the index write-tree the fingerprint
pins. Scope slots inspect touched context plus the repository Atlas; the
assembler reduces optional and unchanged-diff context, records every
degradation, and fails closed when its irreducible pack cannot fit — an
artifact owed in full is reached only after the `-U0` rung and cannot buy fit
by degrading into an invalid review. Owner-selected Low records the distinct
BIBLE P3 scope skip; other route or assembly failure is not a clean verdict.
An agent-session scope slot delivers by retrieval: its verdict is
authoritative once its window is sourced at ≥200K, and "the host did not
observe which files it read" is a provenance disclosure, never a
missing-authority finding. The gate is one logical reviewer interaction per
API slot, with at most one bounded second physical send on a same-route
transport rail for a PACKET api row; a hosted agent-session slot is one
multistep execution whose local extraction reuses its collected transcript.
A native tool-round slot (an api
row bound to a configured subagent) is likewise one multistep episode with no
send count: its bounds are the window-derived transcript bound (measured on
the serialized messages of the next send — each appended element charged as
envelope plus list separator, so the counter equals the wire size), the owner
deadline together with the slot's logical window (each send's transport
timeout is clamped to the remainder and a spent window refuses before
dispatch; each internal recovery send re-reads the same caller deadlines), and the paid ledger; exhaustion is a typed
refusal (`native_transcript_cap_exceeded`) for verdict shapes or a disclosed
`native_incomplete` product for the report shape, and every end leaves its
facts on the actor usage and custody row. A retrieving delivery canonicalizes
its answer by the surface's output SHAPE (`triad_review.review_output_shape`:
`array` | `object` | `report`), never by surface-name branches inside the
canonicalizer: the shape table is form only, and a new object- or
report-shaped surface registers there instead of teaching the extraction rail
another `if`.
Advisory validates its own row enums through the shared canonicalizer's optional
array validator. Unknown verdicts or unknown/missing FAIL severity remain unparsed
unless the existing extraction can faithfully recover them; neither path invents
critical severity or downgrades a finding from identifier presence. Preserve the
full raw result, ordinary PASS rows and genuine empty-clean responses in tests.
Hosted-review model/harness/profile evidence comes from the same final attempt in
`final/telemetry.yaml`, not requested values or cross-attempt summary projections.
Missing observations remain unknown; the exact contributor checker still refuses
unconfirmed model identity, including a display label that cannot prove the pin.
Ordinary delegation requests no extra engine panel; the start receipt names the
serving engine version, and historical runs can retain older review behavior.
Paid review cycles across the gates are bounded by one shared owner knob,
`OUROBOROS_REVIEW_MAX_CYCLES` — a STRING, positive integer or `unlimited`,
default `"2"` (Settings → Behavior → "Max Review Cycles"). Its SSOT is
`ouroboros/review_cycles.py`, whose docstring defines the per-gate meaning
(the retired legacy key is migrated at settings load). `unlimited` removes only the local count —
deadline, budget, and lifecycle rails still bind — and a malformed value fails
closed to the default, logged once.
For task acceptance, the exact-binding tree-wallet claim is a strict
write-ahead stamp bound to every delivery the panel's rows run (owner R11,
2026-09-01: the paid identity is material, not route): a packet row fires it
when the API usage ledger crosses into physical dispatch, a native inspection
row on each paid send of its episode, an agent-session row before its
replayable `START_REQUESTED` — one idempotent claim per panel. Panel assembly,
an unavailable route, or another pre-transport refusal consumes no claim and
leaves the binding retryable. The paid claim itself checks cancellation and
the paid-cycle wallet only (owner R55): the launch floor is evaluated ONCE per
panel, at loop admission, and a RUNNING panel is bounded by the R23 deadline
clamps and the per-send wallet fence. Disclosed residual: a panel whose
evidence build consumed the margin after admission dispatches and may be cut
by the deadline — one ADMITTED panel: admission prices one work-order send per
paid row, but packet rows may use the permitted repair/retry send and native
rows may run several rounds — every send remains deadline- and wallet-fenced
where pricing exists, so the total is NOT bounded to one floor wave; a panel
the deadline actually cuts is DEGRADED, a panel that finishes keeps its normal
verdict; never a free skip. An unavailable claim releases the usage
reservation and blocks every parallel panel slot before reviewer transport
rather than degrading hard authority into fail-open cost telemetry. Disclosed
residual (pre-existing release behaviour, b9bcc2da; issue #588; out of
this change's scope): a compatibility transport
that raises with positive physical capture invokes the paid stamp after the
send; if the tree's last paid cycle is consumed concurrently at that late
stamp, the wallet refusal replaces the captured exception and the substrate
may resend.
Never pay for byte-identical review material (`ouroboros/tools/commit_gate.py` owns
the mechanism): the commit gate refuses a byte-identical staged diff for free
from the FIRST verdict-block (`identical_diff_refused`, quoting the recorded
verdict), and skill review replays a recorded substantive verdict for an
identical snapshot at $0 while the persisted state still covers it. A rebuttal
is identified by CONTENT sha256 — a hash new to the streak buys exactly ONE
paid re-review; a repeated hash is refused free. The two axes stay distinct:
refusal-streak eligibility is about VERDICTS (a rebuttal is spent only by the
substantive verdict it bought), while money is about DISPATCH (every
physically dispatched wave counts whatever its terminal; infra facts refused
at assembly never dispatched and stay outside the count; the paid fact is
recorded write-ahead). A free refusal never wears the form of a verdict: a
refusal that spent nothing is recorded as a typed `not_dispatched` fact plus
its reason, never as a DEGRADED panel, a synthetic actor, or a verdict. That
holds for a plan-review locator the evidence policy cannot attach (a named
omission row the panel is dispatched with), an acceptance packet that
overflows (the ladder, not a verdict), a truncated or self-pageable row (the
cut is named), and a request that was never sent (one
`operation_state='not_dispatched'` seat that stays in the denominator).
Exhaustion is always the typed
`review_cycles_exhausted` event with honest exits — under advisory
enforcement a commit after exhaustion proceeds as a free replay with a loud
typed disclosure; blocking refuses it.
Scope of the review-contract fingerprint (deliberate): it covers the reviewer
roster, routes, enforcement, resolved efforts, and prompt constants —
including the session serialization only when Skill Review actually contains
an agent-session row — while governance-document CONTENTS — `BIBLE.md`,
`docs/CHECKLISTS.md`, `docs/ARCHITECTURE.md`, this handbook and
`docs/DESIGN.md` — are deliberately outside it, so editing those documents
neither lapses recorded verdicts nor frees replays. The accepted
trade-off is that an old verdict can replay under amended governance text;
this keeps routine documentation maintenance from repricing every recorded
review.
### External PR review is not commit authorization
The authoring agent freezes the final committed base-to-head range and gives
it to a separate agent context for read-only review; same-conversation
self-review does not count, and unavailable review is recorded `NOT_RUN`,
never silently presented as clean. `CONTRIBUTING.md` owns the public procedure
and evidence fields. `scripts/run_external_review.py --contributor` is
maintainer-grade large-window tooling: it freezes the configured triad/scope
rows, binds each row to its dispatched prompt receipt and observed response
receipt, and records exact base/head/tree/diff hashes, route/model/profile
facts, terminal settlement, capability deltas, and full redacted
agent-session transcripts; missing, tampered, drifted, unprovable, or
contradictory receipts make the packet `INCOMPLETE`. The lane always executes
the TARGET BASE's own review machinery — invoked from any other checkout it
re-runs itself from a detached worktree of the base commit — so a proposal is
never reviewed by its own copy of the review flow, whatever it touches. This
evidence establishes readiness; it does not authorize commit, push, merge, or
publication — maintainers choose the landing parent and release version,
preserve authorship, and run the normal final exact-candidate gate.
### Release sync
A pull request into `ouroboros` leaves every version carrier byte-identical to
its target: `VERSION`, `pyproject.toml`, the editable root version in
`uv.lock`, `web/package.json`, `web/package-lock.json` (both root entries),
`web/modules/api_types.js::GATEWAY_CONTRACT_VERSION`, the README badge and
latest Version History row, the named direct-download links in README and both
install pages, and the Architecture header. At integration,
`ouroboros/tools/release_sync.py::sync_release_metadata()` projects the chosen
version and `version_carrier_desyncs()` verifies the file carriers (the
history row is pinned by the packaging-sync test); changelog prose remains a
deliberate maintainer edit. The same projection owns the seven public
installer filename templates and rewrites the named direct-download links.
Those links use the immutable exact tag
(`/releases/download/v{VERSION}/...`), never
`/releases/latest/download/...`: prereleases are excluded from GitHub's latest
release, so a latest-style link would fail during an RC. The integration
branch may name installers that are not published yet; public onboarding uses
`main` and `main:/docs`, and stable promotion advances `main` only after the
release and all seven installers are public.
Hermetic preflight uses a disposable worktree, temporary
data/settings/pycache, and scrubbed runtime/secret-class environment. Tests
must rebind imported process-global roots and fail closed on the live data
root; setting only `OUROBOROS_DATA_DIR` is insufficient. A reviewed local
commit is the durability boundary; an `origin` push and CI are follow-up
signals, not prerequisites for local self-modification survival.
---

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,61 @@
# Managed Update Rule
This chapter owns what a managed update may do to the live tree: how one exact official target is chosen and bound, which windows refuse conversation, how dirty local work is stashed rather than merged into history, what the authorized resolver may stage, and why every destructive rollback takes a fresh rescue first. It exists because the pre-update snapshot predates the merge and holds none of the resolver's work, so this ordering is what keeps a failed update recoverable.
- Keep the local work branch and the official update feed separate; the
channel and branch topology live in ARCHITECTURE "8. Git Branching, CI,
and Build" (`ouroboros/update_channels.py`).
- A preflight chooses one exact official target SHA. Apply binds to the
disclosed base/target, closes new writers, drains direct turns,
stops workers and tracked services, then re-plans before mutation. Write
the update transaction before mutation; reopen writers only after a
verified abort/rollback or a healthy restart. Delayed evolution cleanup
acquires the same update lock and honors the same admission owner.
- Conversation admission is separate from repo-writing permission. The chat
lanes refuse turns only during destructive windows (apply/replace/rollback
prologue, materialization). While the ONE authorized assisted resolver holds
the repository (`assisted_resolution` / `committing_assisted`) Main keeps
answering, the registry guard still refuses repo tools to every other task,
and steering reaches the resolver through the ordinary `steer_task` mailbox
(`supervisor/worker_chat_lane.py::conversation_admitted_during_update`).
The server's first-party packages are discovered and imported BEFORE conflict
markers land in the live tree (`preload_owner_control_path`, called after the
resolver readiness proof and before a boot re-materialization). Discovery uses
the imported package paths; the packaged server runs embedded Python from the
materialized Git repository. Tool admission remains owned by the registry.
- Dirty local work never enters merge history: the apply stashes it and
restores it as uncommitted content; a conflicting restore keeps the stash
and discloses the recovery command. The reviewed assisted resolver runs
only when Git reports a real conflict; filenames do not create a second
update policy. Managed materialization and rollback run their internal
`git reset --hard` without interactive confirmation; the
explicit-confirmation rule applies to the owner-facing generic restore
seam.
- The authorized resolver stages the complete merge including tracked binary
files, and review receives their exact staged mode/blob/size plus the
parent object ids; missing exact metadata blocks. This exception does not
weaken the ordinary commit pipeline's binary policy.
- A managed merge commits only with proof that the full suite ran green on
the exact candidate tree ("The commit gate mirrors the CI split" names the
proof authority). Any non-commit terminal of the resolver rolls the live
tree back and best-effort preserves the attempt on the deterministic
failed-update branch; the fresh rescue snapshot, not that branch, is the
carrier rollback itself depends on.
- Take a fresh rescue before every destructive rollback and before
boot-resume re-materialization — the pre-update snapshot predates the
merge and holds none of the resolution. The hook is fail-open (never block
a rollback on it), but its outcome is disclosed durably at capture time,
before the destruction; record the pointer in the update transaction so a
replayed rollback does not re-snapshot and a retry rescues what appeared
since.
- Manual Restore reuses the same writer fence and pins the previous HEAD on
a local recovery branch before reset. Promotion resolves the development
SHA once and uses that exact SHA for both the local QA ref and any remote
push.
Enforcement: `tests/test_update_merge_policy.py` (what the merge policy
refuses), `tests/test_update_dirty_stash.py` (the dirty-tree path),
`tests/test_update_hardening.py` and
`tests/test_update_tx_corrupt_quarantine.py` (transaction integrity and
quarantine).

View file

@ -0,0 +1,39 @@
# Mutation Attribution Rule
This chapter owns the rule that attribution is evidence rather than exclusion: the host captures a baseline when a queued root starts, and blockers ride into review and acceptance evidence instead of becoming structural outcome vetoes. It also fixes what a reviewed commit may stage and where an unversioned interpreter may be resolved, because both decide whose changes a commit actually carries.
- Attribution is evidence, not exclusion: the host captures a `system_repo`
baseline when a queued root task starts and a terminal candidate snapshot
at outcome derivation; blockers (pre-existing dirt, stale/missing
baseline, failed scan) ride into review and acceptance evidence for the
LLM panels to weigh — pre-existing owner work creates ambiguity a
reviewing actor must see without the host inventing a semantic outcome. Do
not turn blockers into structural outcome vetoes, and do not add a
lease/holder service, a second ledger, or runtime writer keyword scanners.
- The acceptance packet reads mutation evidence from the canonical results
root (`budget_drive_root` first), the same root the writer and the outcome
consumer use.
- Git staging is attribution-based: `paths=None` means the clean-at-baseline
candidate set, an explicit list must be its subset, and empty never means
`git add -A`. Preserve pre-existing user dirt as excluded evidence.
Whole-tree staging belongs only to typed managed update/release
transactions and the typed external patch-capture transaction; contexts
without a captured baseline keep the legacy staging contract.
- Resolve unversioned Python only for `run_command`, `run_script`,
`start_service`, and run-kind `verify_and_record`, once BEFORE the shell
guard; guard and handler receive identical argv. Resolve bare Node for the
same four surfaces but once AFTER the dispatch gates — the node health
check executes an argv[0]-steered candidate, and probing before the gates
would run a planted PATH shim for a request the fences would refuse.
Never rewrite explicit paths, versioned interpreters, shell bodies, or
remote execution; never install a dependency in response to
`ModuleNotFoundError`; with no usable runtime the argv runs as written and
fails honestly (a rewritten absolute shebang is a disclosed residual).
- Skill Review ordinals and provenance stay in `review_job.json` and the
append-only `review_history.jsonl`: allocate under the lifecycle lock,
consume a round only after actual start, write one terminal row per
`job_id`, and compute legacy ordinals at read time without rewriting
history.
Enforcement: `tests/test_mutation_attribution.py`.

View file

@ -0,0 +1,146 @@
# Process Custody Rule
This chapter owns the durable ledger every long-lived child process must enter, the strict fingerprint the reaper kills by, the single worker tree-kill seam, and the separate contracts for stopping an installation-owned daemon and for latching a failed start. It exists because an unledgered process survives the death of its owner invisibly, and because a command-line-class match would let one instance reap another's children.
Long-lived OS processes (anything `subprocess.Popen`-ed or `mp.Process`-ed
without a bounded wait in the same call) MUST be spawned through
`ouroboros.process_custody.spawn_supervised(cmd, drive_root=..., purpose=...,
scope=...)` — or, when an existing manager owns the Popen call, registered
via `record_process(...)` write-through immediately after spawn. The custody
ledger (`data/state/process_ledger.jsonl`) is what lets the orphan reaper
find children after an abrupt worker/server death; an unledgered process may
evade durable generation-aware reaping (custody complements the in-process
tracking, port sweeps, and Windows Job Objects — it does not replace them).
Scopes are `task`, `session`, and `daemon`; skill companions are the
documented daemon-scope exception, reaped only when the owning skill is
uninstalled or the entry is from a foreign dead server generation — log-only
by default and fail-safe: an unknown install set means keep-all, never a
mass-kill, and same-session companions of installed skills are always kept.
The reaper kills strictly by (pid, start_time, cmd_sha256) fingerprint —
never add command-line-class matching, which would let a dev instance reap a
packaged instance's processes. `tests/test_process_custody.py` enforces the
chokepoint with an explicit allowlist for bounded synchronous helpers.
An installation-owned daemon uses `daemon` scope, with legacy purpose retention
provided by its lifecycle owner and checked against the existing ledger identity;
server-generation changes alone must not kill it. A worker's process tree is
killed only through `supervisor.worker_pool_lifecycle.kill_worker_tree` — pool
shutdown and restart, the managed-update fence, unready-slot replacement, cancel
and timeout custody alike — which spares the ledger's live `daemon`-scope roots
(`process_custody.live_daemon_root_pids`, including the owner's retained legacy
purposes) and, for one task's cancel or timeout only, the kept services; a direct `kill_pid_tree` on a worker anywhere else in
`supervisor/` is a defect. The explicit stop (`OwnedClaudexorDaemon.stop_outcome`, used by Panic and
Restart) separates authenticated cooperative shutdown from forced signalling.
A valid owned marker and authenticated same-home endpoint permit the already
installed managed CLI `daemon stop --json`; read-only resolution must not ensure
or install a runtime, start a daemon, probe accounts or inherit another home's
socket override. Pure operator commands select the existing exact Node with
`resolve_cli_command(require_npm=False)`; npm-dependent installation retains its
separate toolchain requirement, so ordinary stop works on Windows and bundled-only installs. Require exit0 and a terminal stop receipt, then observe known
root/child exits. RPC acknowledgement and lease release alone are not physical
exit proof. A surviving endpoint is disclosed, never chased as the old target.
Capture existing custody before the request; forced fallback may signal only
those unchanged rows with both measured birth and command, or the manager's own
Popen. Legacy Windows empty-birth rows retain their old permissive keep and cannot
be re-attested from a later observation, but a healthy daemon can stop through CLI.
Preserve received HTTP refusals before reading the body: token/protocol/decoded
response failures and invalid discovery are not transport-unavailability evidence.
A typed network failure or positively absent descriptor for marked startup keeps
its existing fallback meaning; a matching error-code string alone gives no right
to signal. Unknown/missing CLI does not relax that rule. Confirmed signal-stop
exit permits `process_stopped` and pruning; concurrent or unreadable ledger bytes
survive the append lock and exact-prefix compaction. Partial success stays
unconfirmed, with the existing critical supervisor row and retained custody.
Manager-lock acquisition is bounded; preparation/network/exit waits remain outside
it, and Stop retires delayed spawns. The existing operator CLI and captured-exit
bounds are `config.CLAUDEXOR_OPERATOR_STOP_TIMEOUT_SEC` and
`config.CLAUDEXOR_STOP_EXIT_WAIT_SEC`, defined once in runtime_limits. Join existing
purpose-filtered startup custody before and after runtime preparation; caller wait expiry never kills it or replaces
engine writer election. Keep startup and normal admission waits independent,
identify current PID/build/log interval rather than an old log tail, and preserve
existing malformed/foreign ownership markers. Publish a missing marker atomically
only after revalidating the home under the shared JSON lock. A failed
owned-daemon start latches on the TYPED exit fact only
(`ExitFact.failed_without_control`: non-zero or signal exit with no control
descriptor written during that spawn); `claudexor_startup_failure.py`
classifies the child's own log interval for the diagnostic label and the one
`claudexor_daemon_start_failed` supervisor row, never for behaviour (BIBLE
P5). Harvest the exit fact at every spawn decision, at attach and at stop
(`_settle_exited_child`), never only on a caller's wait expiry: with the real
crash cadence (V8 dies after the 20 s startup window) the waiting caller gets
`daemon_starting` and the child dies with nobody waiting, and the next caller,
attach or owner stop must still record the row and the latch. Take the latch
under the same lock as the reap, before reading the log, so a concurrent
same-process caller meets the latch, not a free spawn slot; `_spawn` re-checks
the latch under its own lock and never replaces an exited, unsettled child (the
next settle owns that exit fact). Two residuals are
disclosed, not closed: the descriptor identity is sampled at the first
observation of the exit, not at the exit itself, so a foreign publisher in
that window reads as written and costs at most one extra spawn; and between
the sweep's release and its retry an ordinary caller can pass the refusal and
become the spawner, in which case the retry joins that same live child (one
spawn either way, only who pays the startup wait differs). Do not add a
backoff machine, a retry counter, a cooldown constant, host-side heap sizing
or writer-lease handling, and do not add a third retrier: the periodic
supervisor sweep (`clear_start_failure_latch`, then its own single zero-wait
ensure on a short-lived thread, only when it released a latch), the owner's
explicit Refresh (`/api/claudexor/wake`: clears, then one ordinary ensure), a
live attach, or a new manager (Restart/Panic; a task worker's manager is its
own instance with its own latch) are the only releases, and ordinary callers
never make the retry; `NODE_OPTIONS` passthrough is the operator escape
hatch (ARCHITECTURE §9).
Ordinary close preserves the shared daemon on every platform, including forced
worker/server/stray cleanup. Exclusions protect the whole subtree, not merely a
ledger row. The launcher Job allows explicit breakaway only when the daemon asks;
ordinary generation children remain covered. Do not apply this lifetime to all
skill companions or broaden reserved-port/service ancestry rules. Windows birth,
command, Job and explicit prior-generation stop require native Windows evidence;
psutil is a main Windows-only dependency, and both embedded and frozen imports
need packaging proof. A portable fixture does not establish those claims. An old
immutable launcher keeps its disclosed limitation until its package is updated.
Platform-lock/ABI/fingerprint rationale is specified once in ARCHITECTURE §1
"Platform substrate"; code keeps the local invariant and that pointer.
Explicit stop and next-start runtime selection remain separate contracts: a
newer engine pin is never hot-swapped and the daemon's next start selects it; a
planned restart whose landed checkout pins another engine version or build ends
the serving daemon in the lifespan teardown through
`server_restart._stop_owned_daemon_for_new_pin` (`load_runtime_pin` from the
checkout against `read_owned_gateway`, then the shared attested stop — never a
veto), while an unchanged, unreadable or unpublished pin and an unreachable daemon
leave the handoff untouched (`tests/test_planned_restart_engine_pin.py`). Service quiescence excludes
zombie-only groups, but checks every member before releasing a writer fence
(`tests/test_claudexor_custody_lifetime.py`, `tests/test_process_custody_liveness.py`).
The owner's manual Restart keeps its checkout-first order: the update gate and
`safe_restart` refuse before anything is stopped, and `server_restart._stop_owned_work`
runs only after the durable no-resume flags. Reuse the `request_cancel` ingress,
`kill_workers` with `reconcile_delegate_custody=False`, `reconcile_orphaned_runs` over
`read_owned_gateway`, and the typed `OwnedClaudexorDaemon.stop_outcome`; never call
`ensure_owned_gateway` between the cancel intents and the daemon stop, and never read
the manager's private error state to tell "nothing to stop" from "unconfirmed".
Past the gate an unconfirmed step is a critical diagnostic with custody retained,
never a deferral or a veto — do not add a "deferred" state, a drain, a lease, a
census or a new custody event kind for it; the next generation's startup sweep is
the recovery. The lifespan warmup is one background `ensure_owned_gateway` for a
provisioned home and needs no invalidation of its own: a Stop retires it through the
start generation (`tests/test_manual_restart_execution.py`).
Out-of-process extension HTTP responses execute their standard Starlette ASGI
response in the child. The runner owns staging, Popen registration and cleanup;
`extension_route_stream` owns only portable pipe frames and ASGI delivery. Preserve
ordered headers, HEAD/Range and background actions. Consumer backpressure is not
an idle failure and no total/pre-header response deadline applies. Bind cancellation
to the existing loaded bundle, before spawning, and detach on completion. A
cancellation during startup retains the worker future and process context until
it exits; move context entry/exit, spawn registration, pipe shutdown and termination
off the ASGI event loop. Static captured module sources spawn no child. Chunk size
and post-response cleanup grace come from the runtime-limits owner via config.py
and are not stream deadlines. Bound each frame before reading its payload, not
the cumulative response; retain sanitized bounded stderr and the actual exit
code after draining a child that dies abnormally.
Failed final sends remain delivery failures; background failure after a successful final
body is a separate diagnostic. Widget pull credits bound transport buffering;
large URL downloads use the existing native/browser file owner, never an automatic
HTTP-stream-to-Blob conversion.

View file

@ -0,0 +1,46 @@
# Platform Abstraction Rule
This chapter owns the platform seam: which modules, calls and signals must go through `platform_layer.py`, the one documented function-local exception, and the shared atomic state-file helpers whose byte-exactness keeps a manifest or a hashed receipt identical across operating systems. It exists because the alternative is platform-conditional code scattered through the runtime, where an AST guard cannot see it.
Platform-specific code goes through `ouroboros/platform_layer.py`: platform
modules (`fcntl`, `resource`, `grp`, `pwd`, `msvcrt`, `winreg`,
`ctypes.windll`), direct `os.kill`/`os.killpg`/`os.setsid`/`os.getpgid`,
`signal.SIGKILL`/`signal.SIGTERM`, and platform-conditional subprocess flags
(use `subprocess_new_group_kwargs()` / `subprocess_hidden_kwargs()`). Use
`pathlib.Path` for filesystem paths, never string concatenation with
hardcoded separators. One documented exception: a function-local import of a
platform module under an explicit `sys.platform` guard is allowed outside
the layer (e.g. the guarded `resource` import in
`ouroboros/extension_process_runner.py`) — the AST guard inspects only
top-level imports and does not see function-local ones, so review must.
Enforcement: `tests/test_platform_guard.py` scans `ouroboros/`,
`supervisor/`, and `server.py` for top-level platform imports, the four `os`
calls, and the two signals; `launcher.py` and subprocess flag patterns are
deliberately not scanned — code review and the `cross_platform` checklist
item cover them — and the CI matrix runs tests on Ubuntu, Windows, and
macOS. New platform behavior: add the cross-platform wrapper to
`platform_layer.py`, use it in callers, and add platform-conditional tests
when behavior differs across OSes.
### Shared state-file helpers
Durable state files use the SSOT helpers in `ouroboros/utils.py`:
`atomic_write_json` / `read_json_dict`, and `write_text_atomic` /
`write_bytes_atomic` sharing one atomic full-overwrite seam (temp-sibling +
`os.replace`, permission bits preserved) — a crash leaves the old complete
file intact, and appends are intentionally NOT atomic (a separate contract).
Both are BYTE-EXACT on every platform: the text variant is the bytes variant
plus a UTF-8 encode, so the newlines the caller wrote are the newlines on disk
and no platform translation rewrites a manifest, a hashed receipt or an LF
source file.
Prefer these over bare `Path.write_text`/`write_bytes` for full-file
overwrites; lockfiles go through
`platform_layer.acquire_exclusive_file_lock` /
`release_exclusive_file_lock`. Narrow exceptions: `supervisor/state.py`
keeps `atomic_write_text` for its mirrored state writes, and
`ouroboros/config.py` keeps its settings-file lock because the settings path
is bootstrapped before broader runtime helpers may depend on settings state.
Enforcement: `tests/test_atomic_write_v639.py`; the prefer-the-helper rule is
review-only.

View file

@ -0,0 +1,312 @@
# Design System
This chapter owns the engineering rules that preserve the visual and interaction semantics `docs/DESIGN.md` defines: where values may live, which component is the single source of truth for a control, what counts as review debt, how history and chat viewport transactions behave, and how a visible change is actually verified. It exists because the failures it prevents are silent ones — a size token declared without a colour token, a control that widens its column, a dialog that bypasses the design system entirely.
`docs/DESIGN.md` owns visual and interaction semantics; this section owns
the engineering rules that preserve them — where values may live, which
component is the SSOT, what counts as review debt, how a visual change is
verified. `web/ui.css` owns shared values and field/button/status/popup recipes;
both `web/index.html` and `web/onboarding_template.html` load it before their
page styles. `web/style.css` and the page sheets keep shell/page composition.
Documentation keeps semantic roles and failure-prevention rules, not a copied
color/radius/dimension inventory. These are reusable components in the existing
SPA, not a relocatable-page or multi-instance panel framework.
- Select shared control classes on the controls themselves (`.ui-control`,
`.ui-checkbox`, `.ui-field`, `.ui-field-help`), not an expanding list of
page ancestors. A family migration removes replaced chrome in the same
change without claiming unrelated page typography migrated. Domain drafts,
validation and serializers stay with their current owners.
- `web/modules/ui_primitives.js` is the self-contained safe-field renderer,
collector, attribute escaping and tone/status source; `ui_helpers.js` and
`utils.js` retain the corresponding re-exports. Keep this leaf independent
of shell, network and document-global initialization so real first-party
and author consumers share one implementation. `web/tests/ui_primitives.test.js`
pins that portability and the escaping/password contract.
- A text declaration on a migrated surface names a `--type-*` size token AND
a named foreground token: a rule that declares a size and no colour is the
exact defect that made secondary text inherit near-white primary ink.
`tests/test_web_typography_static.py` keeps the class closed on the
migrated families/regions; extending that guard and migrating its subject
are the same commit (DESIGN §8 names the boundary).
- The variable contract is checked in BOTH directions across the whole
stylesheet by the same test file: a `var(--x)` must resolve — an
undeclared one silently renders its hardcoded fallback, which becomes the
real value nobody can find — and a `:root` token must have a reader,
because a token that resolves nowhere is what makes surfaces reach for
literals. Each document resolves against the sheets it actually loads;
neither page may shadow shared palette names. Fix a dangling name by
pointing it at an existing token, not by declaring a new one.
- Layout and controls: top-level pages use a fixed `renderPageHeader`
outside an independently scrolling body; page icons come from
`web/modules/page_icons.js`; primary actions (including Refresh) live in
the `renderPageHeader({ actionsHtml })` slot; tab strips are one
design-system control (`renderTabStrip` + `bindTabStrip` in `page_header.js`,
`.app-tab-strip`/`.app-tab` and the `--pill-*` tokens). The binder owns
selected state, ARIA, roving focus and strip-only reveal; callbacks own
loading/panels, and programmatic `select()` never calls them. Dispose the
binder and its resize observer with the owning page. The navigation column's
height must not grow with the number of items in a collection inside it; a
variable-length collection owns a bounded window with its own scroll
(`--nav-projects-list-max-height`). Scroll bodies share `.scroll-fade-y`,
except a dense list of rows shorter than its 32px edge, where the fade would
cover a whole row; `scroll_fade.js::bindScrollFade` enables an edge only when
content is actually hidden there and returns its observer/listener disposer;
masonry packing uses `web/modules/masonry.js::applyMasonry` (CSS Grid row
packing leaves row gaps under shorter cards): it packs in the page's key
order and writes only `--masonry-*` custom properties — never move
`<article>` nodes to reorder, a moved `<iframe>` reloads; widget order and
the per-card start-mode override (`widget_start_mode`, values from
`extension_ui_validation.WIDGET_START_MODES`) persist through
`/api/ui/preferences` + `data/state/ui_preferences.json`, never in
extension manifests. New visual dimensions become CSS variables first
and are consumed by shared classes; new inline `style=""` markup and
`.style.<property>` assignments are review debt (a dynamic measured value
may update a narrowly named custom property when that is the real runtime
data flow).
- Containment: a control never widens its column, and horizontal overflow
lives in the wrapper that owns the wide content and declares
`overflow-x: auto` (code block, `.md-table-wrap`, tab strip, Costs table
cells) — never in a page scroll body, whose `overflow-y: auto` alone
already makes `overflow-x` compute to `auto`. The shared
`select.ui-control` recipe therefore clips its own value
(`overflow: hidden`): WebKit computes `overflow: visible` on a native
select, so an unclipped option label becomes scrollable overflow of that
page scroller. A grid track holding controls takes a minimum that yields
to its container — `minmax(0, …)`, or
`repeat(auto-fit, minmax(min(100%, Npx), 1fr))`; a fixed px minimum
rescued only by a viewport media query is review debt, because the
viewport does not know how wide the content column is. The global webkit
scrollbar recipe sizes both axes. Enforced by
`tests/test_web_typography_static.py::test_select_control_clips_its_value`,
its `::test_webkit_scrollbar_recipe_covers_both_axes` neighbour,
`tests/test_ui_settings_overflow_browser.py` (WebKit, the native-select clip)
and `tests/test_ui_settings_grid_tracks_browser.py` (Chromium, yielding
tracks); two gaps stay open — the
wizard document loads `ui.css` without `style.css` and keeps native
scrollbars, and an element setting the standard
`scrollbar-width`/`scrollbar-color` opts out of the webkit recipe on Blink.
- One semantic button variant expresses one action role: neutral Settings
and onboarding controls use the existing `.btn.btn-default`; a one-action
result row uses the named `.settings-action-row` contract (status first,
action docked right); notifications use the shared toast host. Working,
warning, error, and destructive states keep consistent meaning across
Chat, Logs, Settings, and Skills.
- A list editor reveals the entry it just added through
`ui_helpers.revealNewRow(row, field)` — the one seam for "scrolled into
view, caret in the first field" — and a freshly added entry shows no
error before the owner tries to save.
`tests/test_available_subagents_ui_static.py` pins the seam; the
`ui_browser` acceptance in `tests/test_ui_smoke_agents_panel.py` pins the
behaviour.
- Task outcome truth stays in `log_events.js::taskOutcomeSeverity` and
`taskTerminalPhase`; `taskPresentation` is the one compact factual
projection consumed by chips, live completion, history replay, and child
terminal presentation. Its host mirror is
`project_dialogue.outcome_phase`, pinned to the browser by one shared
fixture (`web/tests/fixtures/outcome_phase_parity.json`): a new axis, reason
or acceptance status is added to both sides in the same commit, with a row
in that fixture. The detail line under the headline comes from
`taskReasonDetail` in the order soft stop, hard failure or cancellation
reason, host acceptance decision (status plus stored rationale), typed
reason phrase or raw code — never from a second producer. A non-terminal diagnostic may add a timeline fact
but must not promote the whole task; unknown event names never acquire
Chat severity from `error`/`crash`/`fail` keyword matching. The Chat
header reports connection, the `/api/state` activity census, the owner's
own unconfirmed sends and live task cards only; failed task status does
not synthesize header attention, a toast, unread state, or an owner
action. Never derive header liveness from a WS frame: a typing frame is
a submission receipt, and the `/api/state` census is the only inserter
into the client live-activity set (contract and residuals: the
`DirectActivityRegistry` / `active_chat_activities` paragraph of
ARCHITECTURE.md; enforced by
`web/tests/chat_header_census.test.js`).
- Executor presentation consumes the existing task/run attempt facts. Keep
`executor_observation` event-local through Agent, supervisor delivery, progress
history and both Chat metadata paths; ordinary coordinator notes inherit none.
Its current producer reads the already-polled typed timeline, with a requested
model only for the matching harness. Do not borrow a final-attempt model or
parse progress prose to fill an absent live observation. Label last activity
separately from current computation, configured/coordinator model and settled
observed-model history. Preserve terminal-only `execution_evidence` and
`actual_substrate`, with no new poller or execution-state store. Tests:
`tests/test_executor_observation.py` and `web/tests/wire_contract.test.js`.
Render the projected chip/model facts through
`harness_presentation.js::executorIdentityMarkup`; keep execution-evidence
selection in `log_events.js` and avoid a second label builder in Chat.
- Preserve task-owned model-call provenance through result storage, copy-back
and terminal/history rendering. Show the last usable solve response separately
from initial routing, executor observations and final-answer authorship;
post-task or cost-only updates cannot erase it. Fan-out counters describe
emissions and wall-clock intervals, never inferred execution waves.
- Chat viewport invariant: sample live-edge intent before an ordinary
transcript mutation — native scroll anchoring is not proof the owner's
visible message stays stable, so focused regressions disable it. Follow
only inside the 48 CSS-pixel zone, otherwise preserve the visible keyed
message, nested-card, or Reviews anchor; route late
application-controlled DOM writes through the existing stable-viewport
seam, keeping awaited Load-older, reconnect reconciliation, and
cross-instance restoration as explicit lifecycle transactions. Browser
coverage is chosen by risk; this WebKit-sensitive contract requires the
engines exercised by its marker-gated UI smoke.
History pages and reconnect merge into the existing keyed card/row owners. Preserve
actual selected/focused/expanded nodes; rebuilding an equivalent node is not preservation.
Timeline patches compare generated markup so unchanged enhanced markdown retains its controls;
the Reviews reconciler separately owns lazy attempt-detail state and cannot replace that behavior.
Physical source identity orders equal-time archive rows without rewriting JSONL.
A historical frame never grants current activity or replaces newer terminal evidence.
Eviction releases only its own page's media, markdown and decision views, protecting
visible reading, focus and selection. Exact page handles retain return navigation;
read gaps and sparse empty pages never become false EOF. Readable recent rows survive an
unavailable archive with explicit gap/retry and no fabricated physical cursor. Flush historical
timeline changes once per card, and skip idle scroll cleanup when no work is pending.
History chrome describes the rendered transcript, not the pager cache: `canNewer` (contiguous
cached descriptors around `focus`) is never permission to tell the reader that newer messages
exist. A server page with zero rows for the room is a bounded scan: it advances the cursor and
keeps `has_more` honest, but it is not a reading position, not `focus`, and not newer.
Automatic continuation loads pages only at the older edge; at the live edge it only fills the
gap toward already-mounted rows through exact page handles, never a rebuild. Return to the
present is the explicit floating button, which rebuilds the chain with `latest()` only while
non-empty pages above the window are still missing. Distant evicted pages plus retained live
rows may leave a mid-transcript hole; that residual is disclosed rather than covered by a
second load-newer control. Tests:
`tests/test_chat_history_paging.py`, the pass-through cases in
`web/tests/chat_history_pager.test.js`, the sparse-walk, anti-loop and dense-return cases in
`web/tests/chat_history_integration.test.js`, and
`tests/test_chat_history_paging_browser.py`.
The Project work pointer is a navigation component over the existing Chat card
registry (`project_work_pointer.js`), updated inside the same viewport mutation
transaction. Its label names the card on one line (`projectWorkLabel`: coined
name, else title, capped; the `.project-work-pointer-label` CSS ellipsizes) and
never restates the card's full status headline; without a represented root card it is
hidden, not shown disabled. Preserve its loaded-window coverage disclosure; a
represented unfinished card is not independent proof of current execution. Its click changes
only the messages container's scroll position and existing reading intent, never
message routing. Dispose it with the chat; do not add a second card tree, poller
or task-state store for this navigation affordance.
### Responsive and accessible behavior
Navigation, headers, controls, and dialogs stay operable by pointer and
keyboard, preserve focus order, and fit the relevant narrow viewport without
stealing usable text space; use the shared responsive component before
adding a page-specific layout. A visible change is inspected with vision in
at least one relevant real consumer flow. A stored screenshot alone is not
verification; mobile or WebKit is not a universal requirement and is
selected from risk. Containment is the WebKit-sensitive exception — a native
select is not clipped there — so a change to a control recipe or a page
scroll body is verified on the engine that shows the class (Playwright
WebKit for native-control clipping, Chromium for engine-independent track
geometry), measuring overflow on the scroll body's `scrollWidth` rather than
on `documentElement`. Review-only: scored by CHECKLISTS items 2(i) and 30
(`web_design_system`).
### Browser dialogs
For agent page readiness, use `browser_action(action="wait", selector=..., state=...)` on the current page, or `browse_page(wait_for=..., state=...)` after navigation. States are `attached`, `visible`, `hidden`, and `detached`; hidden also accepts an absent element. A timeout returns the requested state, URL, current match count and first-element visibility rather than navigating again. These observations do not decide whether the task should continue; bounded `evaluate` remains available to its existing profiles.
Image-reader regressions must include a real PNG under a synthetic user home with distinct task and canonical skill roots. Exercise same-round auto-attachment, the durable local copy and the actual send-time image block; a placeholder PNG, a flat drive or a mocked attachment helper cannot prove that path. Keep secret/owner-state and protected-artifact denial controls.
Browser-boundary regressions run the installed Chromium and WebKit (`PLAYWRIGHT_BROWSERS_PATH`; a test never installs a browser, and `OUROBOROS_EXPECT_BROWSER_ENGINES` turns a missing engine from a skip into a failure) against real loopback servers bound through `server_entrypoint.bound_service_socket`, so the control endpoint under test is an actual recorded binding rather than a fixed port. The redirect residual is one strict xfail (`tests/test_browser_private_service.py`, the server-side dispatch counter) beside the passing content-refusal proof (`tests/test_browser_redirect_chain.py`); do not turn either into the other. The private-service proof takes its LAN target from `OUROBOROS_TEST_PRIVATE_BROWSER_HOST`/`_ADDRESS` (`OUROBOROS_EXPECT_PRIVATE_BROWSER=1` fails instead of skipping) and writes the images it viewed to `OUROBOROS_BROWSER_EVIDENCE_OUT`.
`window.prompt`, `window.confirm`, and `window.alert` are forbidden in
`web/modules`: PyWebView shells implement them inconsistently, native
dialogs bypass the design system and browser tests, and the macOS shell has
no prompt delegate, so `window.prompt` silently returns `null`. Use
`confirm_dialog.js::openConfirmDialog` — confirm mode returns a strict
boolean, input mode returns `{confirmed, value}`, alert mode renders one
acknowledgement action; Close, Cancel, backdrop, Escape, and supersession
are always non-confirming. Critical actions test the exact confirmed result
and keep the confirmation plus side effect in one injectable flow.
`tests/test_web_dialogs_static.py` keeps the native-dialog class closed.
`ui_interactions.js::bindDialogFocus` owns the modal keyboard boundary and
conditional restoration; callers mount first and dispose before removal,
keeping their own result/cancel contracts. `bindMenu` adds action-menu keyboard
and dismissal behavior over `bindPopoverPosition`; editable suggestion lists
use positioning alone, preserving native input and domain route serialization.
Mount `.ui-popup` outside clipping ancestors and
consume its measured `--ui-popup-*` properties in shared CSS. The owner retains
markup, portal removal and action dispatch; no overlay registry is needed.
Dispose before removing a popup, and close/restore a menu before opening a
dialog from its action. `web/tests/ui_interactions.test.js` pins callbacks,
focus, geometry and cleanup; actual menu/chooser/dialog browser consumers
remain necessary for viewport and engine-sensitive behavior.
Files keeps one current editable document through cancelled navigation, ordinary
folder refresh, failed Save and clipboard feedback. Pointer/keyboard submission
shares one in-flight write; newer text remains dirty after an earlier save.
The New Project adapter shares dialog focus and menu behavior while retaining
all source modes and its selected target independently of browser navigation.
`tests/test_ui_smoke_files_project_drafts.py` verifies these real consumers.
### Declarative widgets
A module handler that calls `OuroborosWidget.openExternal(url)` or `window.open(url)` from an anchor click also calls `event.preventDefault()`. Invoke the helper directly during the gesture, before awaiting other work; automatic relaying respects an already-handled click.
`web/modules/widgets.js` is the host for reviewed widget declarations:
forms/actions, text/data/media, tabs/charts, async jobs, files,
map/calendar/kanban, and composition through `group`, `metric`, and
`callout`. Nested interactive components use stable identity and one
disposer; `subscription.render` is transitively passive. Data updates patch
the existing component/field nodes at the mount's identity seam, preserving
selection, composition, native popup state and password input. Passwords stay
only in their mounted control, never in the retained form-value snapshot.
Forms and actions own visible pending/result/error feedback by component id;
an optional status component or a sibling's shared data target is not that
action's result. Preserve the existing job identity, bounded retries and
disposal contracts. Escape text and attributes for their actual HTML contexts,
constrain media to extension
routes or safe data URLs, and keep charts accessible through a semantic
table. Rare `kind: "module"` UI runs only in a sandboxed opaque-origin
iframe with no `allow-same-origin`; its document policy admits scripts,
images, media and fonts only from the skill's own prefix (plus
`data:`/`blob:`; `connect-src` closed) and its parent bridge proxies only the
owning extension route — never load skill JavaScript into the SPA origin.
Both framed mounts live in `web/modules/widget_module.js` (the child-side
bootstrap is `widget_frame.js`) and return their disposer to the `mountTab`
dispatcher in `widgets.js`; the framed card chrome — launch policy (owner
override > author `render.start` > kind default), `retain`, Start/Stop, the
policy menu, the facade — lives in `widget_card.js`, reorder handles in
`widget_reorder.js`, chart/table helpers in `widget_chart.js`, the pure
list-signature and keyed-patch helpers in `widget_list.js`; the page
compares the list signature after every `GET /api/widgets` and touches no
card node when it is unchanged. A failed list read exposes contextual Retry
through that same reconciliation; it preserves unchanged frames and the
owner's Stop choices. Do not turn Retry into a global refresh/remount.
Long-running actions use a durable job id and resumable status polling.
Every timer, listener, observer, stream, abort controller, chart, and
mounted widget has a paired disposer. Enforcement:
`tests/test_widgets_ui_static.py` at commit tier; in the release-tier
`ui_browser` lane `tests/test_widgets_ui_browser.py` (geometry, job retry),
`tests/test_widgets_ui_browser_lifecycle.py` (launch policy, ordered stop,
`retain`, the streaming bridge), `tests/test_widgets_ui_browser_patch.py`
(keyed patch of a running card, reconnect reconcile) and
`tests/test_widgets_ui_browser_capabilities.py` (the frame CSP, sandbox and
permissions boundary on Chromium and WebKit) — run all four before a release
that touched Widgets. `tests/test_widgets_ui_browser_identity.py` additionally
pins retained interactive nodes, composition/password lifetime, local action
feedback and non-destructive list Retry through real declarative consumers.
### Optional author controls
Use `ouroboros.server_web.read_author_kit_assets(request.app.state.repo_dir)`
to read the fixed installed `web/ui.css` and `web/modules/ui_primitives.js`
sources for an author-owned page. Resolve at the page/kit GET that serves a
new mount, not at extension registration; the request root is propagated by
both in-process and out-of-process dispatch. No bundle cache, new endpoint or
auth exception belongs in the helper.
`docs/examples/author_ui_kit/` contains the two ordinary extension recipes:
a module gets source text from its own route through `OuroborosWidget.fetch`,
adds CSS and imports the self-contained module from a frame-local Blob URL;
a route iframe embeds the same source safely in its initial HTML under its
own CSP. Revoke temporary Blob URLs. These paths need no opaque `/static`
request, new bridge message or widget schema flag. Keep the kit optional and
author-overridable; retained mounts keep their loaded source, with no theme
poller or forced remount. Tests `test_author_ui_kit.py` and
`test_author_ui_kit_browser.py` cover source-root delivery and actual framed
consumers; they do not certify an arbitrary author's CSP or application.

View file

@ -0,0 +1,50 @@
# MCP Client Integration
This chapter owns the boundary for configured MCP servers: the base runtime is a client and never a server, descriptions and results are untrusted data rather than policy, and a stdio entry passes one executable and an exact argument list without a shell. It also fixes how referenced settings values reach a process environment, because a value's secret classification decides whether it is masked in diagnostics and logs.
The base runtime is an optional CLIENT for trusted HTTP/SSE and local stdio
MCP servers — never an MCP server (structure and module ownership:
ARCHITECTURE "MCP and browser-facing external tools";
`ouroboros/mcp_client.py`). MCP descriptions and results are untrusted data,
not policy: configuration trust must not turn remote prose into policy.
Enabled tools join the initial capability envelope, still pass runtime
safety and the caller's ordinary capability ceiling; discovery failure
becomes a visible capability omission. Stdio accepts one executable command
and an exact string argument list without a shell. Optional `cwd`, literal
`env`, and `env_from_settings` (environment names mapped to existing string-valued
setting keys) apply to listing and calling through the same manager configuration.
References override matching literal names. The owner chooses MCP references in
Settings; a caller's existing MCP tool grant does not authorize new references.
Unknown fields remain saved and show a not-applied warning without disabling a
valid server. Invalid known fields, missing references and incompatible transport
fields produce `MCP_CONFIG_ERROR`. Response-only `auth_configured` is discarded at
the configuration boundary. Omission preserves SDK environment/cwd defaults.
Settings' existing built-in/custom-secret classification determines masking:
ordinary selected values remain readable, while exact secret values (including
short values and JSON-escaped echoes) are masked in descriptions/results/stderr.
Executable schema properties, required fields, enum and default values stay intact.
Resources, prompts, and MCP server behavior remain separate architecture changes.
Enforcement: `tests/test_mcp_client.py`, `tests/test_process_environment.py`.
`start_service` overlays ordinary literal `env` on the existing minimal host
baseline; root tasks can additionally choose `env_from_settings` through their
host-resolved process authority. Restricted and Presence tasks receive no new
Settings-selection authority, while their previous literal env and configured MCP
access remain available. Skill grants do not authorize an unrelated service.
Both service backends and MCP reuse `workspace_executor.resolve_process_env`;
referenced Settings fields retain the same secret/ordinary classification.
Cwd still uses the host-owned resource binding; local import scrubbing and
interpreter defaults remain. Docker forwards values through inert CLI environment
aliases and restores the selected names inside the container, preserving host CLI
configuration. Values do not enter host argv or generated shell source.
Service diagnostics and finalized log blobs mask referenced secrets; ordinary
PORT/PATH/DEBUG values, protocol identity/state and executable schema stay intact.
The executor's record-stop owner finalizes local logs through the existing service
log owner before forgetting the in-memory selections, including task/global cleanup
and replacement after exit. Unconfirmed termination retains the existing record;
a later cleanup can settle it. Live child logs and logs surviving worker loss keep
the existing private raw-log contract until successful finalization; oversized or
uncapturable logs retain their existing explicit omission/error report. No secret
values are added to the durable process ledger. Enforcement:
`tests/test_process_environment.py`, `tests/test_workspace_executor_services.py`.

View file

@ -0,0 +1,32 @@
# Gateway Boundary Pattern
This chapter owns the direction of dependency at the browser boundary: inbound routes enter through `ouroboros/gateway/`, outbound provider and harness adapters live in `ouroboros/gateways/` and carry no domain policy, and the frontend calls the one typed client. It exists so the UI can evolve without importing the agent body and the runtime can evolve without ad-hoc browser contracts.
Browser-facing backend work enters through `ouroboros/gateway/` and frontend
calls go through `web/modules/api_client.js` (structure: ARCHITECTURE
"Gateway Boundary v1"; the endpoint index lives in
`ouroboros/gateway/endpoint_index.py`, re-exported by `contracts.py`).
Outbound provider/harness adapters belong in `ouroboros/gateways/` and carry
no domain policy — do not copy policy into an adapter, promote the
`gateway/host_service.py` callback surface into a general owner/task API,
or require a class where established function owners already preserve the
boundary. Enforcement: CHECKLISTS item 17 (`gateway_parity`) and
`tests/test_gateway_parity.py`.
Named skill chat ingress uses the existing message-bus canonical writer before
enqueue. Its internal callback retains request-owned upload copies once the
canonical write is attempted, even if that write fails with an unknown outcome;
replays do not adopt new copies. The existing settled HTTP-worker wait retains
copy/admission until completion under cancellation, and a local cleanup stack
removes only unaccepted destinations. Accepted attachments remain through an
unknown write/queue outcome or disconnect. Such failure may retain an unused
copy but never claims acceptance. Cancellation support and intent writes must
address the same installation root as the operation being read. The server consumes that internal accepted source without logging a
second row. Operation reads/cancel verify the complete source against actual
task or direct-turn ownership; presentation annotations are discovery hints.
Named waits use that operation's state, while unnamed legacy waits retain their
chat callback. Tests: `test_host_service_operation_identity.py` and
`test_host_service_operations.py`. Successful child WS relay failures travel as
bounded counters through the existing process-facts channel, preserving the
producer result and best-effort `None` API (`test_extension_ws_diagnostics.py`).

View file

@ -0,0 +1,253 @@
# Build & CI
This chapter owns the build and test topology: the one dependency-lock authority and its packaging projections, the seven opt-in pytest marker lanes, the parallel and serial split CI actually runs, the hermetic commit gate that mirrors that split in a disposable checkout, and the GitHub Actions secret-gating shape. It exists because the gate's verdict has to be reproducible and un-weakenable by the candidate it is judging.
### Python dependency locks
`pyproject.toml` is the direct-dependency SSOT and `uv.lock` the reviewed
cross-platform resolution; local and CI use `uv sync --locked`, and no
independent hand-written requirements authority exists. Release packaging
exports build requirements ephemerally and commits
`requirements-runtime.lock` for embedded pip; `requirements.txt` is a
generated pointer for older managed updaters, never an authority. A
dependency change updates the metadata, runs `uv lock`, regenerates the
runtime export with the exact README command, and leaves the CI clean-diff
check green. The pinned `tool.uv.required-version` and digest-pinned
`setup-uv` action make resolver changes deliberate rather than an ambient
CI upgrade. Documentation may pair the checkout-free `uv tool install` form
with a full commit SHA to pin the source revision, but must not claim it
locks dependencies or describe it as a release-artifact install or
contributor development environment.
### Pytest marker lanes
Default local pytest excludes seven costly or environment-dependent lanes —
`integration`, `browser`, `ui_browser`, `ui_browser_docker`,
`portable_detail`, `skill_smoke`, and `size_ratchet` — and CI opts into
them explicitly (job topology and provider matrix: ARCHITECTURE "CI
topology"):
- `integration` runs real provider checks, including the trusted
direct-OpenAI canary rows derived from `OPENAI_DIRECT_DEFAULTS`. Missing
core credentials are red in the official repository job; explicit
quota/429/5xx/timeout may be typed inconclusive, while
contract/auth/model/reasoning/tool 4xx stay red. Secretless request-wire
and Anthropic-custody contracts remain in ordinary pull-request tests: do
not move provider secrets into PR jobs or duplicate the trusted lane. The
KEYLESS `tests/system_e2e/` scenario lane rides the same marker plus `serial`
and the `OUROBOROS_E2E_DEEP=mock` env gate — real isolated servers, no
provider keys (ARCHITECTURE "System E2E suite"). Because those three gates
shut it out of every other pass, the only thing that runs it is the dedicated
`system-e2e-mock` CI job on its daily schedule, manual dispatch or release
tag: never an ordinary branch push or pull_request, and carrying no secret.
- `browser` / `ui_browser` / `ui_browser_docker` launch real Playwright
engines (agent browser tools / the host UI / the `ouroboros-web:test`
container; the docker lane skips cleanly when Docker is unavailable
locally). The marker is the source of truth for what the lane collects;
the four Widgets lifecycle suites listed under "Declarative widgets" run
in it. `portable_detail` covers build/portable artifact invariants.
Existing `ui-smoke` runs only `tests/test_skill_publish_browser.py` on PRs,
using Chromium and real task admission, Main card and history. Manual/tag
runs keep the full host UI and Chromium/WebKit browser-tool suites. These
checks do not join the default local pytest run or require a paid provider.
- `skill_smoke` installs the nine pinned official skills from the LIVE
catalog (list in `tests/test_skill_smoke_official.py`) and runs as the
dedicated 3-OS CI job in serial pytest invocations with real network and
real pip; red means investigate — there is deliberately no fallback-skip.
Its paid review tier runs as a SEPARATE pytest step (fresh process) that
alone carries the provider key, ORDERED FIRST and ubuntu-only: the other
tiers import downloaded plugin code in-process, and running the secret
step first means the runner has never executed payload code while the
secret was present. A missing key is a hard red, not a skip.
- `size_ratchet` carries the live-repo size gates and is the ONLY blocking
surface for repository size (rules under "Module Size & Complexity"; only
checks against the live repo carry the marker). The base fallback
verifies the parent manifest against the parent's own tree — accepting a
copied manifest would allow debt laundering, so a resolvable base that
lost its manifest fails closed.
`skill_smoke` and `size_ratchet` tests must NOT also carry the `serial`
marker or join `_SERIAL_TEST_FILES`: the `and not <lane>` markexprs in
quick/full-test are the lane barrier, and single-lane assignment keeps each
test's placement unambiguous. When adding a new opt-in lane, register the
marker in `pyproject.toml`, add a collect-only zero-test guard in CI, and
keep the default local addopts free of network and Docker requirements.
### Parallel CI and the `serial` marker
CI runs the default suite in parallel — `python -m pytest tests/` with
`-m "not serial and <the seven lane exclusions>"`, `-n auto --dist
loadscope --max-worker-restart=0 --timeout=300 --timeout-method=thread` —
followed by a serial pass for `-m "serial and <the same exclusions>"`
(`.github/workflows/ci.yml`, jobs `quick-test` / `full-test`). Two rules
keep new tests from breaking that:
- Mark real-process / real-port tests, and tests that mutate process-global
state WITHOUT reliable fixture isolation, `@pytest.mark.serial` (or add
the file to `_SERIAL_TEST_FILES` in `tests/conftest.py`). Under `-n` such
a test flakes on kill/reap or port-reclaim timing, or crashes its
worker — and with `--max-worker-restart=0` a dead worker fails its WHOLE
co-located batch, showing up as spurious failures in unrelated files.
- Keep every other test parallel-safe so it stays in the fast pass: use
`tmp_path` (never a fixed path), `monkeypatch.setenv`/`delenv`/`setattr` for
environment and attribute changes, and no execution-order assumptions. The
autouse `tests/conftest.py::_os_environ_isolation` snapshot restores
`os.environ` at every test boundary, so a bare assignment no longer leaks;
monkeypatch stays the rule because it reverses exactly the named change
inside the test, before the snapshot runs. A
module-global mutation that is reliably snapshot-and-restored by a
fixture may stay in the parallel pass — the pattern is
`tests/conftest.py::_isolate_workspace_executor_globals`.
### The commit gate mirrors the CI split
`ouroboros/preflight_runner.py::run_hermetic_pytest` mirrors CI in one
disposable checkout and scrubbed temporary data root: the node test lane
(`cd web && node --test tests/*.test.js`, content-keyed — a candidate
without web tests never requires node, while an active web suite cannot
silently disappear when node is missing), then the same two logical pytest
passes (parallel `not serial`, then flag-free `serial`).
The browser no-undef check has two layers: the dependency-free acorn walker
in that suite (`web/tests/no_undef.test.js`) is the hermetic gate's, and
both CI jobs additionally run ESLint's `no-undef` (`web/eslint.config.js`,
exact-pinned, installed with `npm ci` from `web/package-lock.json`) as an
independent second opinion — CI-only, never part of the gate.
`LANE_EXCLUSION_EXPR` and `PARALLEL_PASS_FLAGS` are executable SSOTs pinned
against both CI jobs; the candidate is captured as one hardened
worktree-vs-`HEAD` binary diff, and a capture or apply failure is the typed
`PREFLIGHT_CANDIDATE_ASSEMBLY` hard block, never a test failure.
The `pyproject.toml` `addopts` line is the single home of the per-test timing
report (`--durations=25 --durations-min=1.0`): it is prepended to every argv, so
the same slowest-test evidence appears in a plain local run, in both CI jobs and
in both gate passes without any surface pinning its own copy.
Contributor rules:
- The candidate cannot weaken the pass: `PYTEST_*`/`NODE_OPTIONS` are
scrubbed, and so is owner runtime state — `OUROBOROS_*`, secret-suffixed
keys and every settings key `config.apply_settings_to_env` projects
(derived from `settings_env_keys()`), so the verdict cannot depend on the
operator's install profile; required plugins are probed outside candidate control and
forced on with host-owned worker evidence, post-commit checks also
inspect `HEAD~1` so suite deletion cannot hide after the commit exists,
and exit status owns the verdict — rendered diagnostics do not.
`OUROBOROS_PREFLIGHT_SERIAL=1` is the explicit temporary rollback lever,
never a silent fallback.
- A red post-commit gate is warning-only for an ordinary commit (the local
commit is preserved for forensics); evolution publication refuses to
auto-push while the warning stands, and inside a managed update the gate
blocks boot promotion and routes through rollback — an incomplete
rollback leaves `gate_blocked` so boot retries recovery instead of
promoting the rejected merge.
- The managed mandate is "the full suite provably ran green on the exact
committed tree", not "run it twice": the reuse authority is the
process-held runner proof (`ctx._preflight_test_proof`) covering the tested
tree, installed index, effective passes and execution environment. Equivalent
ordinary preflights and post-commit runs reuse it too, after distinct baseline
checks. Every workload binds HEAD: equal file trees do not establish
equivalence for history-sensitive tests, including ordinary unmarked tests.
A newly created commit therefore requires a new run; repeated preflights with
unchanged HEAD still reuse the existing proof. A skip, no applicable
suite or a mocked `None` return cannot mint a proof. The
durable `tests_evidence` record is forensic telemetry the gate never
consults, so a restart forces a rerun. Review-binding and tag-binding
mismatches use the same managed failure route.
Creation/reuse observations use the existing event log with the exact subject
and workload fingerprint; logs never become reuse authority. Preserve the
executable's invocation path as well as its resolved binary identity, because
separate Python environments may symlink the same binary.
- Process containment is unconditional, including after a green pass:
Windows uses a kill-on-close Job Object; POSIX uses an environment
membership token plus a process-group enumeration backstop and promises
honest detection with a fail-closed verdict for attributed members, not
guaranteed teardown of an arbitrary detached process. A same-uid unreadable
stranger is a warning, never membership proof; an unobserved descendant
that detached and hid its token remains a disclosed detection gap. Known
roots, groups and the retained set of observed members still fail closed when unreadable
(`tests/test_preflight_process_containment.py`). A crashed worker, a timeout-killed
worker, a missing plugin, containment failure, and ordinary test failure
keep distinct diagnostics.
- Mark process/port/global-state tests `serial`; make a merely slow test
faster or split it — marking it serial removes the 300s per-test timeout
and lets it consume the remaining total gate budget.
### GitHub Actions: secrets in step-level `if:` conditions
GitHub Actions rejects `secrets.*` inside step-level `if:` expressions, and
a step's own `env:` block is not visible to that same step's `if:`. Derive
a non-secret boolean in the job-level `env:` block, gate steps with that
boolean, and map the actual credentials only inside the first-party steps
that need them — later SBOM and attestation steps then inherit none of
them.
```yaml
jobs:
build:
strategy:
matrix:
os: [ubuntu-latest, macos-latest]
env:
HAS_APPLE_SIGNING: ${{ matrix.os == 'macos-latest' && secrets.BUILD_CERTIFICATE_BASE64 != '' && secrets.P12_PASSWORD != '' && secrets.KEYCHAIN_PASSWORD != '' && secrets.APPLE_TEAM_ID != '' && 'true' || 'false' }}
steps:
- name: Import Apple signing certificate
if: env.HAS_APPLE_SIGNING == 'true'
env:
BUILD_CERTIFICATE_BASE64: ${{ secrets.BUILD_CERTIFICATE_BASE64 }}
P12_PASSWORD: ${{ secrets.P12_PASSWORD }}
KEYCHAIN_PASSWORD: ${{ secrets.KEYCHAIN_PASSWORD }}
run: |
echo "${BUILD_CERTIFICATE_BASE64}" | base64 -d > cert.p12
security create-keychain -p "${KEYCHAIN_PASSWORD}" build.keychain
security import cert.p12 -k build.keychain -P "${P12_PASSWORD}"
- name: Cleanup keychain
if: always() && matrix.os == 'macos-latest' && env.HAS_APPLE_SIGNING == 'true'
run: security delete-keychain build.keychain
```
```yaml
# ❌ WRONG — workflow fails to parse
- name: Bad
if: secrets.BUILD_CERTIFICATE_BASE64 != '' # parse error
env: # not visible to this step's if:
P12_PASSWORD: ${{ secrets.P12_PASSWORD }}
```
`tests/test_build_scripts.py::TestMacOSSigning::test_ci_uses_env_context_for_condition`
enforces this for `.github/workflows/ci.yml` only; other workflow files are
not scanned by it.
### Apple signing & notarization (macOS Build job)
Prerelease artifacts may intentionally be unsigned and must report that
state; stable publication applies the configured signing and notarization
policy rather than implying credentials or success that were absent. Only
the non-secret `HAS_APPLE_SIGNING` gate is job-wide; certificate/keychain
values exist only in the import step and Apple ID notarization values only
in the first-party build step. Notary/stapler failures are soft outcomes
recorded through `NOTARIZE_OUTCOME`, so a transient Apple service problem
does not silently drop an otherwise valid signed artifact; cleanup uses the
`always()` plus matrix/env guards, and signing material never persists
across runs.
### Release proof capsule
The artifact pipeline — per-platform archive smokes, native Linux packages,
the AppImage custody chain, SBOM and attestation binding, and the
seven-asset release job — lives in ARCHITECTURE "8. Git Branching, CI, and
Build" and `.github/workflows/ci.yml`. The honesty invariants a change must
preserve:
- Publication is draft-first with a per-tag concurrency group; the remote
annotated tag is revalidated against the event SHA immediately before
draft creation AND again before publication, and a published release is
never overwritten by a rerun.
- Vendor-distribution smokes (Astra, RED OS) are reported evidence, never
release authority — third-party registry reachability is outside the
publication pipeline's control.
- The AppImage smoke deliberately makes no native GTK/Qt claim; packaged
native webview coverage remains a separate Linux distribution contract.
- `OUROBOROS_SKIP_PLAYWRIGHT_INSTALL_DEPS=1` is only a local-builder escape
hatch — it skips Playwright's host-library installation, not
browser-binary bundling — and a build using it must disclose that browser
host compatibility was not locally proven.
- Never represent a later checksum inventory as build-time provenance, an
SBOM, or packaged smoke evidence that the original build did not create.

View file

@ -0,0 +1,98 @@
# Reference books: chapter migration transfer table
This is the operator record of the physical split of the two reference books
(`docs/ARCHITECTURE.md`, `docs/DEVELOPMENT.md`) at base commit `5585133db86419c1a28673e498de4fb13c6b2d1e`.
It is a docs file and deliberately NOT a book member: it lives directly under
`docs/`, so `validate_reference_books` never sees it in either book's chapter
population.
Every row below is a **verbatim move**. The moved bytes are every byte of the
old `##` section AFTER its heading line, including its `###`/`####`
sub-headings at the levels they already had. Nothing was merged, rewritten,
summarized, reordered or deleted, and no section title was renamed — the
rename column is empty for every row. Semantic work (merging duplicate
explanations, retiring obsolete history) is a separate later package and will
add its own dispositions here.
## What changed, exactly
Per chapter file, the only new bytes are a two-line prologue:
1. `# <the old section title>` — the old `## N. Title` text at H1, numbering
text kept so every cross-reference in the corpus still reads;
2. one authored introductory paragraph (2–4 sentences: what the chapter owns
and why it exists), the only new prose in this migration.
Everything after that prologue is the old section body, byte for byte.
Per entrypoint, the body is replaced by an ordered `## Chapters` membership
list; the H1 line stays byte-identical in both books.
## Byte proof
`tests/test_reference_book_migration.py` reverses the prologue of every chapter
(drop the H1 line, drop the one introductory paragraph, re-prefix `## ` to the
H1 text), concatenates the results in membership order, and requires the
SHA-256 below. It also compares against `git show <base>:<path>` whenever the
base commit is reachable, so the recorded digest cannot be a fiction.
| Book | Old file | Old bytes | Old SHA-256 | Preamble bytes (replaced) | Moved-body bytes | Moved-body SHA-256 |
|---|---|---|---|---|---|---|
| architecture | `docs/ARCHITECTURE.md` | 724691 | `5db278f8ef5060c4aff5ee1e8743c279661ddd975a311a9bafb5858f32b080de` | 610 | 724081 | `f1f054c700a15533687e0cf81cf19ac53ffcf022eb179f84c0cbccacbb1d305e` |
| development | `docs/DEVELOPMENT.md` | 275551 | `50eb460602f1501915195e3ad1918366312e6292b0dccec6f7ef53f5a5302f3b` | 60 | 275491 | `bffc00227bc5e91f054b38eaed63acd256c2ddf111e231754bfbe92c7dd122e3` |
## `docs/ARCHITECTURE.md`
Entrypoint preamble before: `# Ouroboros v7.0.0 — Architecture & Reference` (byte-identical, the release version carrier `release_sync.VERSION_CARRIER_SPANS` writes), then two paragraphs (`This file is NOT a changelog…` and `This is the present-tense operational map…`) and a `---` rule.
Entrypoint preamble after: the same H1, then ONE merged paragraph carrying both original paragraphs' claims (present-tense map in three layers, not a changelog, WHY stays in the book, rationale self-contained), then `## Chapters`. The `---` rule is dropped.
| Old `##` section | Lines at base | Section bytes | Destination chapter | Disposition | Title rename |
|---|---|---|---|---|---|
| `1. High-Level Architecture` | 9–659 | 188782 | `docs/architecture/01-high-level-architecture.md` | verbatim move | — |
| `2. Startup / Onboarding Flow` | 660–695 | 14555 | `docs/architecture/02-startup-onboarding-flow.md` | verbatim move | — |
| `3. Web UI Pages & Buttons` | 696–868 | 88631 | `docs/architecture/03-web-ui-pages-and-buttons.md` | verbatim move | — |
| `4. Server API Endpoints` | 869–1031 | 23636 | `docs/architecture/04-server-api-endpoints.md` | verbatim move | — |
| `5. Supervisor Loop` | 1032–1082 | 32383 | `docs/architecture/05-supervisor-loop.md` | verbatim move | — |
| `6. Agent Core` | 1083–1871 | 258915 | `docs/architecture/06-agent-core.md` | verbatim move | — |
| `7. Configuration (ouroboros/config.py)` | 1872–2084 | 33538 | `docs/architecture/07-configuration.md` | verbatim move | — |
| `8. Git Branching, CI, and Build` | 2085–2140 | 17850 | `docs/architecture/08-git-branching-ci-and-build.md` | verbatim move | — |
| `9. Shutdown & Process Cleanup` | 2141–2156 | 12956 | `docs/architecture/09-shutdown-and-process-cleanup.md` | verbatim move | — |
| `10. Key Invariants` | 2157–2235 | 15834 | `docs/architecture/10-key-invariants.md` | verbatim move | — |
| `11. Frozen Contracts v1 (`ouroboros/contracts/`)` | 2236–2326 | 19130 | `docs/architecture/11-frozen-contracts-v1.md` | verbatim move | — |
| `12. Host Service, Companion Processes, and Chat IDs` | 2327–2371 | 10594 | `docs/architecture/12-host-service-companions-and-chat-ids.md` | verbatim move | — |
| `13. External Skills Layer` | 2372–2386 | 7277 | `docs/architecture/13-external-skills-layer.md` | verbatim move | — |
## `docs/DEVELOPMENT.md`
Entrypoint preamble before: `# DEVELOPMENT.md — Development Principles & Module Guide` (byte-identical), then `## Role and authority` directly, with no introductory paragraph of its own.
Entrypoint preamble after: the same H1, then ONE newly authored orientation paragraph (what the handbook is, how the chapters are ordered, read the chapter for the class of change in hand), then `## Chapters`. No moved prose.
| Old `##` section | Lines at base | Section bytes | Destination chapter | Disposition | Title rename |
|---|---|---|---|---|---|
| `Role and authority` | 3–32 | 1707 | `docs/development/01-role-and-authority.md` | verbatim move | — |
| `Naming and boundaries` | 33–526 | 35291 | `docs/development/02-naming-and-boundaries.md` | verbatim move | — |
| `Module Size & Complexity` | 527–849 | 21621 | `docs/development/03-module-size-and-complexity.md` | verbatim move | — |
| `Core Governance Artifacts` | 850–1121 | 20584 | `docs/development/04-core-governance-artifacts.md` | verbatim move | — |
| `Review & Commit Protocol` | 1122–1351 | 15585 | `docs/development/05-review-and-commit-protocol.md` | verbatim move | — |
| `Rules by change class` | 1352–2963 | 118624 | `docs/development/06-rules-by-change-class.md` | verbatim move | — |
| `Managed Update Rule` | 2964–3022 | 3732 | `docs/development/07-managed-update-rule.md` | verbatim move | — |
| `Mutation Attribution Rule` | 3023–3059 | 2289 | `docs/development/08-mutation-attribution-rule.md` | verbatim move | — |
| `Process Custody Rule` | 3060–3203 | 10776 | `docs/development/09-process-custody-rule.md` | verbatim move | — |
| `Platform Abstraction Rule` | 3204–3247 | 2514 | `docs/development/10-platform-abstraction-rule.md` | verbatim move | — |
| `Design System` | 3248–3557 | 22543 | `docs/development/11-design-system.md` | verbatim move | — |
| `MCP Client Integration` | 3558–3605 | 3454 | `docs/development/12-mcp-client-integration.md` | verbatim move | — |
| `Gateway Boundary Pattern` | 3606–3635 | 1986 | `docs/development/13-gateway-boundary-pattern.md` | verbatim move | — |
| `Build & CI` | 3636–3886 | 14785 | `docs/development/14-build-and-ci.md` | verbatim move | — |
## Chapter granularity
One chapter per old `##` section, with no merges. Two adjacent pairs were
under the ~60-line merge threshold on both sides — Architecture §12/§13 and
Development "MCP Client Integration"/"Gateway Boundary Pattern" — but neither
pair shares a subject (a host callback boundary is not the external skills
plane; an outbound MCP client is not the inbound browser boundary), so the
default 1:1 mapping was kept. It also keeps every existing cross-reference of
the form `ARCHITECTURE "8. Git Branching, CI, and Build"` resolving to exactly
one chapter.

View file

@ -2,6 +2,8 @@
Machine extraction of the `docs/ARCHITECTURE.md` "Data layout (`~/Ouroboros/`)" tree — the durable-file orientation carrier (this tree's counterpart of the reference PERSISTENCE_OWNERS derivation checklist) — regenerated by `python scripts/regenerate_inventories.py`. Do not edit. Every entry is probed against reality: repo entries must exist as tracked paths; data-plane entries must appear as a literal in the runtime sources that construct them. A durable file renamed or removed in code while its tree row survives = red (`tests/test_generated_inventories.py`).
Source: `docs/architecture/01-high-level-architecture.md`, physical LF lines 563-653; UTF-8 SHA-256 `627473114d9e305312364d55f86a9beb2531c86fc65b983d4bd6a1c5e1d9cc3d`.
- entries: **79** (code-ref: 72, repo-dir: 6, repo-path: 1)
| entry | probe | resolution |

View file

@ -2,6 +2,8 @@
Machine extraction of `docs/ARCHITECTURE.md` §11.1 (the frozen-ABI SSOT), regenerated by `python scripts/regenerate_inventories.py`. Do not edit — edit §11.1 and regenerate; `tests/test_generated_inventories.py` pins byte-identity and the resolution invariants (a §11.1 row whose owner or anchor file disappeared from the tree = red).
Source: `docs/architecture/11-frozen-contracts-v1.md`, physical LF lines 7-36; UTF-8 SHA-256 `9ecac7008228a91d615e9ad6cd7ecc80934a2af731ea647e9f2a174069a34c7c`.
- table rows: **25**
- browser-envelope prose owners:

View file

@ -56,14 +56,27 @@ def _whole(source: MarkdownSource) -> SourceRange:
def _preamble(source: MarkdownSource) -> SourceRange:
"""The authored introduction: the FIRST paragraph after a source's H1.
It must open the chapter — a source whose H1 is followed straight by a
subsection has no introduction and is refused, which is the property that
keeps an overview from quoting body prose as if someone had written it for
that purpose. What it deliberately does NOT require is that the
introduction be the only paragraph before the first subsection: a chapter
carries its relocated section body at the heading level that body already
had, and most sections open with prose, so demanding a single paragraph
would force either a rewritten heading level or an invented sub-heading.
The overview says in its own words that it holds introductions, not
complete chapters, and carries each chapter's physical path beside them.
"""
first = source.headings[0] if source.headings else None
if first is None or first.level != 1 or not first.title:
raise ValueError(f"{source.source_path}: chapter needs a nonempty H1")
stop = next((h.span.start_byte for h in source.headings[1:] if h.level <= 2), len(source.raw))
paragraphs = [p for p in source.paragraphs if first.span.end_byte <= p.start_byte < stop]
if len(paragraphs) != 1 or not source.text_at(paragraphs[0]).strip():
raise ValueError(f"{source.source_path}: chapter needs one authored introductory paragraph before its first H2")
return paragraphs[0]
stop = next((h.span.start_byte for h in source.headings[1:]), len(source.raw))
intro = next((p for p in source.paragraphs if first.span.end_byte <= p.start_byte < stop), None)
if intro is None or not source.text_at(intro).strip():
raise ValueError(f"{source.source_path}: chapter needs an authored introductory paragraph under its H1")
return intro
def _member_paths(entrypoint: MarkdownSource, book_id: str) -> tuple[str, ...] | None:

View file

@ -0,0 +1,144 @@
"""The chapter split moved bytes, not meaning: a reversible byte proof.
Each chapter file is exactly a two-line prologue plus the old `##` section body:
# <the old section title>
<blank>
<one authored introductory paragraph>
<the old section body, byte for byte, starting with its own newline>
So the move is INVERTIBLE, and the inverse is this test's whole method:
first = raw.index(b"\\n\\n") -> the H1 line is raw[:first]
second = raw.index(b"\\n\\n", first + 2) -> the introduction is between them
section = b"## " + raw[2:first] + raw[second:]
Concatenating those sections in membership order must reproduce the old
monolith from its first `## ` heading to EOF, byte for byte. The recorded
digests below make that a complete proof without Git; when the base commit is
reachable the same reconstruction is compared against `git show` as well, so a
recorded digest can never stand in for bytes nobody checked.
The entrypoint preamble is deliberately NOT part of the byte proof: it was
replaced by one merged/authored paragraph plus the `## Chapters` membership
list, and `docs/reference-books-migration.md` records what it said before. Its
H1 line IS pinned here, because that line is the release version carrier.
"""
import hashlib
import pathlib
import subprocess
import pytest
from ouroboros.reference_books import BOOK_ENTRYPOINTS, load_reference_book
REPO = pathlib.Path(__file__).resolve().parents[1]
# The integration base this migration was cut from.
MIGRATION_BASE = "5585133db86419c1a28673e498de4fb13c6b2d1e"
# Recorded at the base commit by the split itself (and mirrored in
# docs/reference-books-migration.md): the whole old file, and the part of it
# that moved -- everything from the first `## ` heading to EOF.
OLD_MONOLITHS = {
"architecture": {
"old_bytes": 724691,
"old_sha256": "5db278f8ef5060c4aff5ee1e8743c279661ddd975a311a9bafb5858f32b080de",
"preamble_bytes": 610,
"moved_bytes": 724081,
"moved_sha256": "f1f054c700a15533687e0cf81cf19ac53ffcf022eb179f84c0cbccacbb1d305e",
"h1": "# Ouroboros v7.0.0 — Architecture & Reference",
},
"development": {
"old_bytes": 275551,
"old_sha256": "50eb460602f1501915195e3ad1918366312e6292b0dccec6f7ef53f5a5302f3b",
"preamble_bytes": 60,
"moved_bytes": 275491,
"moved_sha256": "bffc00227bc5e91f054b38eaed63acd256c2ddf111e231754bfbe92c7dd122e3",
"h1": "# DEVELOPMENT.md — Development Principles & Module Guide",
},
}
def _restore_section(raw: bytes) -> bytes:
"""Undo one chapter's prologue and return the old `## ` section bytes."""
first = raw.index(b"\n\n")
second = raw.index(b"\n\n", first + 2)
h1 = raw[:first]
assert h1.startswith(b"# "), h1[:40]
introduction = raw[first + 2:second]
assert introduction.strip() and b"\n\n" not in introduction, h1
return b"## " + h1[2:] + raw[second:]
def _reconstruct(book_id: str) -> bytes:
book = load_reference_book(REPO, book_id)
assert not book.legacy, f"{book_id} is not chaptered"
return b"".join(_restore_section(chapter.raw) for chapter in book.chapters)
def _base_bytes(rel: str) -> bytes | None:
try:
return subprocess.run(
["git", "show", f"{MIGRATION_BASE}:{rel}"],
cwd=REPO, check=True, capture_output=True,
).stdout
except (OSError, subprocess.CalledProcessError):
return None # shallow clone or exported tree: the digests still bind
@pytest.mark.parametrize("book_id", sorted(BOOK_ENTRYPOINTS))
def test_chapters_reconstruct_the_old_monolith_body_byte_for_byte(book_id):
recorded = OLD_MONOLITHS[book_id]
moved = _reconstruct(book_id)
assert len(moved) == recorded["moved_bytes"]
assert hashlib.sha256(moved).hexdigest() == recorded["moved_sha256"]
old = _base_bytes(BOOK_ENTRYPOINTS[book_id])
if old is None:
pytest.skip(f"base commit {MIGRATION_BASE} is unreachable in this checkout")
assert len(old) == recorded["old_bytes"]
assert hashlib.sha256(old).hexdigest() == recorded["old_sha256"]
assert old[recorded["preamble_bytes"]:] == moved
assert old[:recorded["preamble_bytes"]].startswith(recorded["h1"].encode("utf-8"))
@pytest.mark.parametrize("book_id", sorted(BOOK_ENTRYPOINTS))
def test_entrypoint_keeps_its_h1_line_and_carries_only_the_membership(book_id):
rel = BOOK_ENTRYPOINTS[book_id]
raw = (REPO / rel).read_bytes()
lines = raw.decode("utf-8").split("\n")
assert lines[0] == OLD_MONOLITHS[book_id]["h1"]
book = load_reference_book(REPO, book_id)
# One `## Chapters` heading and nothing else: the entrypoint orients, the
# chapters carry the book.
assert [h.title for h in book.entrypoint.headings] == [lines[0][2:], "Chapters"]
members = [chapter.source_path for chapter in book.chapters]
assert members == sorted(members), "membership order must be the reading order"
for chapter in book.chapters:
assert chapter.source_path.startswith(f"docs/{book_id}/")
assert f"]({chapter.source_path[len('docs/'):]})" in book.entrypoint.text
def test_no_chapter_body_was_reheaded_into_a_duplicate_title():
"""Relocation kept `###`/`####` levels, so a section title still names one
physical place across the whole book -- what `read_book_section` needs."""
for book_id in BOOK_ENTRYPOINTS:
book = load_reference_book(REPO, book_id)
titles = [h.title for source in (book.entrypoint, *book.chapters) for h in source.headings]
duplicates = sorted({t for t in titles if titles.count(t) > 1})
assert not duplicates, f"{book_id}: ambiguous section titles {duplicates}"
def test_the_transfer_table_is_committed_and_is_not_a_book_member():
table = REPO / "docs" / "reference-books-migration.md"
assert table.is_file(), "the operator transfer table must be reviewable"
text = table.read_text(encoding="utf-8")
assert MIGRATION_BASE in text
for book_id in BOOK_ENTRYPOINTS:
book = load_reference_book(REPO, book_id)
assert str(table.relative_to(REPO)) not in [c.source_path for c in book.chapters]
for chapter in book.chapters:
assert f"`{chapter.source_path}`" in text, chapter.source_path
assert OLD_MONOLITHS[book_id]["moved_sha256"] in text

View file

@ -104,12 +104,16 @@ def test_duplicate_named_sections_are_ambiguous_instead_of_first_match():
read_book_section(book, "Exact section")
def test_current_reference_books_pass_the_explicit_transition_policy():
def test_current_reference_books_pass_the_final_chaptered_admission():
"""The production caller of the validator: the tracked tree, chaptered.
This is the docs-lane check item 8 asks for -- a missing chapter, an
unlisted one, or a chapter whose authored introduction was lost fails here
rather than at the next review that assembles a book.
"""
root = pathlib.Path(__file__).resolve().parents[1]
tracked = subprocess.check_output(["git", "ls-files", "-z"], cwd=root).decode().split("\0")
# The physical split switches this same CI/preflight requirement to True;
# current monolith compatibility must not select that policy from filenames.
assert validate_reference_books(root, tracked_paths=tracked, require_chaptered=False) == ()
assert validate_reference_books(root, tracked_paths=tracked, require_chaptered=True) == ()
def test_docs_sync_reads_full_chapters_and_never_loses_residue_between_files(monkeypatch):

View file

@ -74,16 +74,37 @@ def test_missing_authored_preamble_is_not_generated_from_body():
load_reference_book(Path("unused"), "architecture", corpus.__getitem__)
def test_a_historical_monolith_revision_still_reads_as_one_complete_legacy_source():
"""The migration is forward-only; an exact older revision must still compose."""
monolith = b"# Book\n\nOrientation.\n\n## Runtime\n\nProcesses carry the work.\n"
book = load_reference_book(Path("unused"), "architecture", lambda _: monolith)
assert book.legacy and not book.chapters
assert compose_book(book).encode() == monolith
view = overview_book(book)
assert "docs/ARCHITECTURE.md" in view.text
assert "Runtime" in view.text and not view.source_complete
@pytest.mark.parametrize("book_id,path", [
("architecture", "docs/ARCHITECTURE.md"),
("development", "docs/DEVELOPMENT.md"),
])
def test_current_production_monoliths_remain_byte_identical(book_id, path):
def test_current_production_books_are_chaptered_and_composition_covers_the_closure(book_id, path):
root = Path(__file__).resolve().parents[1]
book = load_reference_book(root, book_id)
assert book.legacy
assert compose_book(book).encode() == (root / path).read_bytes()
assert not book.legacy and book.chapters
closure = (book.entrypoint, *book.chapters)
composed = compose_book(book)
for source in closure:
# Every declared source, whole, exactly once: a composed book is the
# complete book or it is a lie about coverage.
assert composed.count(source.text) == 1, source.source_path
assert source.source_path == path or source.source_path.startswith(f"docs/{book_id}/")
assert len(composed) >= sum(len(source.text) for source in closure)
view = overview_book(book)
assert path in view.text
assert view.sources[0].path == path
for chapter in book.chapters:
# The compact view orients by authored introduction and addresses the
# PHYSICAL chapter, never a line of the composed book.
assert f"Source: `{chapter.source_path}`" in view.text
assert not view.source_complete