Fourth authoritative review (two reviewers clean): the index, toggle and
reconcile responses dropped the runtime `process` and `server_reconcile` facts
that the skill tools already return — a small leaf (gateway/extension_receipts.py,
mapped in ARCHITECTURE) now projects them into every response incl. load errors,
the typedefs declare them, and the skill card shows a compact live-state
qualifier; ARCHITECTURE says "omitted route" (not render) and CREATING_SKILLS
discloses that deterministic preflight reports the route-less iframe while
runtime registration preserves it; a module entry must match the browser's
filename grammar ([A-Za-z0-9._-]+) so a preflight-clean name can no longer be
silently rewritten by the loader.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Re-review of round 3: a worker-local reconcile result was written into the one
unqualified health.json even when the worker→server marker handoff failed, so a
worker's LIVE could erase the server's BROKEN and advance last_known_good.
Server observations now exclusively own status / regressed / last_observed /
last_known_good; worker observations persist under observations.worker with
their process and server_reconcile facts (locked read-modify-write); regression
aggregation never lets a worker-local state override the durable server
authority; the Skills API exposes health_observations as a qualifier. The
skill-review response gets its own SkillReviewResponse typedef carrying the two
optional reconcile keys (removed from SkillGrantResponse); no version carrier
touched.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
A module widget was one file inlined into the frame; a game or emulator
needs more. When a module tab registers — in-process `register_ui_tab`
and the host-side out-of-process catalog alike — the loader now captures
every reviewed `.js`/`.mjs` under the skill directory into the bundle's
`module_sources`, keyed by POSIX relative path. The walk is the
review-hash surface (`skill_loader._iter_payload_files`: dependency and
cache directories and escaping symlinks are not reviewed, so they are
not captured) minus dot-prefixed segments; a non-UTF-8 JavaScript file
or a missing/escaping entry fails the registration loudly, as the entry
did before. The validator stores the declared entry stripped once, so
the capture key and the URL the page builds agree (closes the A9/A11
double-normalization class).
The route becomes `/api/extensions/{skill}/module/{entry:path}`. The
endpoint reads the new one-lock accessor `live_module_sources`: 409
when the skill has no live bundle, 404 when the path was not captured,
400 for a backslash, NUL, an empty/`.`/`..` segment (percent-escapes
arrive decoded) or a non-JavaScript suffix; the body is served from
memory with `application/javascript; charset=utf-8`, `Cache-Control:
no-store` and `Access-Control-Allow-Origin: *`, with no per-request
disk read. `live_widget_projection` drops its per-tab source column,
which no caller consumed. The existing `other.js -> 404` pin is flipped
deliberately: with `other.js` in the reviewed payload it is now served.
Docs: endpoint index, ARCHITECTURE table row and Skills-and-Widgets
doctrine, CREATING_SKILLS "Loading more than one file" (classic
`<script src>` or `import()` with the CORS header the opaque frame
needs; the declared entry still executes as a classic script; the
`script-src` prefix lands with the frontend slice).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
- New ouroboros/gateway/widgets.py: GET /api/widgets projects live UI tabs
from the in-memory extension snapshot only (no discover_skills, no
stale-review reconcile, no schedule sync, no hashing, no writes). Each
card carries the owning skill's live payload content_hash as `revision`
-- a revision fact for the page's change signature, not an ETag. Wired
through router, endpoint_index, api_client (`apiClient.widgets`),
api_types (WidgetTab, WidgetsResponse) and the parity/smoke suites. The
Widgets page itself still reads GET /api/extensions; the frontend switch
is a later phase.
- api_extension_module: authorization is the live loader state alone.
`extension_loader.live_bundle_facts` resolves the reviewed payload
directory from the loaded bundle and the live tab registration must still
declare exactly this entry, so both discover_skills walks leave this hot
read path. Adds `Access-Control-Allow-Origin: *` (the opaque-origin module
frame fetches anonymously cross-origin); `Cache-Control: no-store` stays.
An unloaded skill now answers 409 "not live" before the 404 entry check.
- Remove the dead two-phase fields `ui_host_pending` / `ui_tabs_pending`
(always True / always []; no frontend, CLI, or colab reader) and their
test pins; type the `live` block of GET /api/extensions as
ExtensionLiveSnapshot (homed in gateway/widgets.py to keep contracts.py
in its size band).
- ARCHITECTURE.md: module tree entry, endpoint rows, Skills and Widgets
rationale for the passive read and the CORS header.
- tests/test_gateway_widgets.py pins zero discovery/reconcile/sync/hash
calls on both read paths and the exact WidgetTab shape.
No version bump: contribution branch, release carriers byte-identical to
the base.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
- skill_review_gate carries a typed preflight_failed fact when the caller
has the persisted findings (absence is a fact: status-only callers never
fabricate False); the pending aggregation reuses the same predicate.
- Card feeds pass findings: /api/extensions (_review_fields and the index),
marketplace installed payload, skill summaries, and review outcomes.
- The card renderer offers Repair as the primary action for a repairable
payload whose preflight failed (Review would deterministically fail the
same way), admits self_authored payloads to the repair/make-runnable
paths (their payload lives under skills/external/), hides the Skip
review dead end (attestation reruns the preflight and 409s), and renders
the preflight finding as a human-readable diagnosis instead of raw JSON.
- The marketplace lifecycle shows Repair with the same diagnosis for a
preflight-failed install.
- Tests: gate contract + predicate + attestation-never-bypasses (py),
/api/extensions end-to-end fact, and 7 renderer/diagnosis cases (node).
Every unique /api/extensions row now carries published (the validated
record published object or null) and published_malformed (true when
the record file exists but fails schema-v1 validation). Collision rows
keep their existing pin of never reading per-skill state, so both
fields are absent there.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Prevent due manifest schedules and stale review reconciliation from acting while one canonical skill name maps to multiple payloads, while preserving schedule recovery state.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Establish the shared top-level root policy and immutable selected-target foundation, including pure skill identity discovery and collision-safe lifecycle entry points.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Seed the combined Telegram bridge and Mini App as a native skill, add symmetric skill conflicts and managed repair routing, preserve the standalone PoC process boundary, and wire native delivery through Colab and packaged bootstrap.
Owner-approved big release (plan: большой_релиз_v6.26.0). Workstreams:
- WS-A lint gate: ruff F-rules step in CI quick-test; fixed the real F823
NameError class (supervisor/events.py utc_now_iso shadowing); make lint/health.
- WS-B memory integrity: atomic dialogue_blocks with corrupt-quarantine
(.corrupt-<ts>.bak + memory_store_corrupt event); honest scratchpad journal
(block_append_failed, corrupt storage renders as corruption not "(empty)");
full Pattern Register window (16K cap vs 3K cut); merge-aware scratchpad
consolidation under the sidecar lock; backlog writes through locked helpers;
chat omission notes ("[N older unconsolidated messages omitted]").
- WS-C provider SSOT: provider registry (prefixes/credentials/resolution) in
provider_models.py replacing 5 duplicated knowledge sites; credential-aware
consolidation model; pricing empty-fetch retry; direct-route correctness
(o-series max_completion_tokens, reasoning_effort, Anthropic error bodies,
per-request timeouts on cached clients); deep/plan review budgets reserve
output headroom inside the 1M window (min(SSOT, window − output − margin)).
- WS-D races/state: utils.update_json_locked (locked RMW, loud TimeoutError);
task_results merge under per-file lock (cancel-latch safe); update_state
migration for owner binding, evolution counters, budget_line, post-task
activation; queue-lock coverage (enqueue, timeouts, worker health, task_done,
snapshot); visible supervisor death (supervisor_error + owner notice + ghost
consciousness/chat-agent cleanup on re-init); chat.jsonl rotation via
os.replace under the append lock; unbounded _outbox removed; host-service
port probe + actual-port panic sweep; skill lifecycle lane deadline
(OUROBOROS_SKILL_LIFECYCLE_TIMEOUT_SEC, dedupe-leak-safe); consilium
force-plan structured flag (plan_review_aggregate from the FULL result);
atomic writers (update-intent, repo-manifest, metrics cache, post-task pair);
review_state.save_state raises on lock timeout (advisory ledger honesty);
apply_pending_request no-ops while a campaign is active;
api_update_apply kills workers only AFTER update validation, with respawn
on aborted checkout and "interrupted" terminal status.
- WS-E immune hardening: triad anti-refusal coverage contract (empty array
needs NO_FINDINGS sentinel or bare-[] body; refusal prose with [] is
parse_failure and never enters quorum); BIBLE P3 "Owner-chosen enforcement,
loud advisory" bound + CHECKLISTS sync; every advisory pass-through of a
blocking signal writes review_advisory_override + persistent
advisory_overrides (surfaced by review_status); ToolEntry.mutates_worktree
with dispatcher-level before/after worktree diff invalidation (covers error
paths; read-only runs no longer invalidate; redundant manual calls removed);
scope review fails closed without its checklist; advisory final write is a
locked re-read merge; synthesis severity defaults to critical with WARNING
fallbacks; find removed from SAFE_SHELL_COMMANDS; anti-thrashing state
survives advisory criticals; preflight env scrubs secret-class variables;
.git/index.lock age gate in worker startup checks.
- WS-F security: file-browser symlink containment on the RESOLVED path across
all endpoints (out-of-root symlink targets are listed but inert; tests
rewritten to the new contract deliberately); HMAC-signed session cookies
(server-side persisted key, 30-day TTL, Secure on TLS) replacing the
permanent password-derived cookie; password-class settings mask to a
constant placeholder; conservative SSRF guard for the MAIN agent (link-local
/cloud-metadata only, LAN stays reachable, per-request route re-validation);
ClawHub zip-slip hardening (":"/backslash segments rejected + post-join
containment) and lazy no-proxy OuroborosHub opener; single-execution
signature dispatch for extension handlers (no TypeError re-run after side
effects); onboarding postMessage origin checks; SHA256-pinned
python-standalone download with pipefail; payload-resident dependency
fingerprints only corroborate durable deps.json; skill payload re-hash
immediately before spawn (TOCTOU narrowing).
- WS-G process custody: ouroboros/process_custody.py — spawn_supervised
chokepoint + durable data/state/process_ledger.jsonl (pid, pgid,
fingerprint{start_time, cmd_sha256}, purpose, scope task|session|daemon,
owner_task, session_id); platform_layer.process_start_time primitive;
startup + periodic reaper killing ONLY strict-fingerprint matches from dead
generations/tasks (never by command-line class); migrations: services
(the orphan hole), workspace executor + local model + extension companions
(ledger write-through), worker_pids (write-through; legacy path retained);
parent lifelines (ppid watchdog, group-suicide only as group leader) in
worker_main, extension runner, Claude readonly child; conformance test
pinning the Popen allowlist; ARCHITECTURE/DEVELOPMENT/CHECKLISTS entries.
- WS-H native multimodal chat: supports_vision capability map (static
prefixes + OpenRouter /models input_modalities overlay); web chat uploads
ride the WS frame as structured attachments (additive ChatInbound field)
and image uploads become NATIVE image blocks via the existing Path B;
browser screenshots inject natively for vision models via the multipart
user-merge (tool result stays a string; file persisted under
data/uploads/screenshots for re-view); K=3 newest-image eviction with
caption placeholders carrying the vlm_query re-view path; image-aware token
estimates (fixed ~1.1K-token equivalent instead of base64 length, fixing
permanent emergency-compaction wedges); compaction renders images as
captions (no base64 into the summarizer); GigaChat/local lanes emit explicit
"[image omitted: model has no vision]"; internal _caption/_source_path
metadata stripped from provider payloads.
- WS-I housekeeping (partial): SETTLED_STATUSES SSOT (+ cycle-safe mirror pin);
owner_inject.py renamed to owner_mailbox.py; version-neutral envelope
wording; files.py import-block cleanup. Remaining WS-I/WS-J/WS-K items are
deferred with the owner's context-budget priority on review+release.
Review notes: triad+scope ran via scripts/run_external_review.py on the core
pack across 3 rounds to convergence (scope responded=PASS each round; round-2
criticals fixed: update_apply kill-order + respawn, lifecycle dedupe leak on
lane timeout, OUROBOROS_MAX_ROUNDS hot-reload + docs, toggle_evolution
NameError, async Anthropic timeout forwarding, supports_vision local check,
budget-update lock visibility, ChatInbound additive attachment contract,
file-browser doc sync). The FULL combined diff exceeds every triad model's
context window (~1.59M tokens > 1.05M) — reviewed in packs; remaining
cross-pack findings were verified as slicing artifacts. Adversarial critics
(GPT; Gemini/Opus rounds) ran on the working tree. Deliberate tradeoffs:
metadata-based eviction captions (no light-LLM call in the hot path);
worker ledger records use live-cmdline fingerprints with a synthetic-arg
fallback only where the OS offers no cmdline.
First function-count paydown in gate history: remove 14 dead functions
(incl. 3 owner-acked deletions in protected registry.py/git_ops.py), fold
10 duplicate-body clusters into SSOT survivors, and inline 30 trivial
wrappers (3009 -> 2953 gated functions; behavior preserved; 11 manifest
positions rejected per the manifest's own criteria with recorded reasons).
Raise MAX_TOTAL_FUNCTIONS to 3500 per owner decision to stop gate churn
and compress the gate-comment archaeology; drop the dormant
file-size-budget parser whose DEVELOPMENT.md section no longer exists;
sync DEVELOPMENT.md (function-gate precedent, governance-loader table)
and ARCHITECTURE.md (structural smoke-gate rationale, stale post-edit
revert references); sync all release carriers to 6.25.0-rc.3.
Reviewed: real triad+scope dry-runs x3 (converged; gemini round-3 finding
verified false positive) + three-critic adversarial review (convergent
doc-staleness finding fixed in-diff). claude-fable-5 triad slot exceeded
its 1M window on this diff size (deterministic); 2/3 quorum per substrate
rule.
Consolidate repeated gateway, review, skill lifecycle, marketplace, memory/context, release, and web UI paths while preserving public API shapes, review gates, and runtime contracts.
Verification: python3 -m pytest tests -q --tb=short; external triad review PASS with scope review skipped only for context budget advisory.
Move browser-facing HTTP and WebSocket ownership into a dedicated gateway package, add the frontend API-client boundary, and keep release metadata aligned for the pre-release.
Make the executable review gate explicit across skill review results, catalog/API payloads, lifecycle messages, runtime gates, and the Skills UI so agents stop treating advisory findings as passed reviews under blocking enforcement.
This follow-up also makes persisted advisory_pass verdicts revalidate against the current enforcement mode before execution, preserving the skill review trust boundary when operators switch from advisory to blocking.
Layered on top of rc.5..rc.7 catch-up reduction line. Preserves the
rc.7 prompt-injection fix in renderSkillRepairPrompt (web/modules/utils.js)
and the rc.6 docs/accounting cleanup. Adds the Quality, DX & Closed-loop
Skills wave from the user-approved finish plan:
- runtime_mode=light reframed as a minimal compatibility/self-modification
guard. Path-aware shell filter (_light_shell_repo_mutation): blocks
simple writer commands (cp, mv, rm, sed, sort -o, uniq, ...) only when
their target resolves inside the Ouroboros checkout, blocks shell-
wrapped writes via 'sh -c' / 'bash -c', blocks mutative direct git
through run_shell, and otherwise lets ordinary python/node/bash
diagnostics run. Removed the heavy Python AST scanner + script-content
scan that was over-blocking legitimate work.
- Chat scroll fixes: ResizeObserver re-attach, near-bottom threshold,
Playwright smoke now asserts scrollTop ~ scrollHeight after send.
- Recent-chat dedup: read_jsonl_tail_after_offset honours
dialogue_meta.last_consolidated_offset; provenance-aware via a chat
log generation signature (first_line_sha256 + size) so log rotation
cannot silently drop entries.
- Stable WORLD.md injected into context.build_memory_sections.
- plan_task and the commit triad accept duplicate model IDs as valid
reviewer slots (single-provider stochastic sampling); _get_review_models
no longer pads 2-slot configs to 3.
- Settings: OUROBOROS_AUTO_GRANT_REVIEWED_SKILLS owner-confirmed via the
desktop launcher bridge (request_auto_grant_reviewed_skills_change),
hot-read from settings.json so toggle changes take effect without
restart, /api/settings POST drops the key. Consistent truthy parsing
on JS + launcher sides.
- skill_review.py runs an optional fail-open Claude Code advisory over
the skill payload only (include_repo_diff=False), and injects its
output as inert evidence BEFORE the authoritative output contract.
- skill_exec emits skill_exec_finished / skill_exec_failed events;
worker enqueues them on ctx.event_queue, supervisor.events dispatches
them to logs/events.jsonl + the live log + the in-process event bus
for skill.lifecycle subscriptions. host_service_api allows manifest-
declared skill.lifecycle subscriptions without an extra grant.
- run_shell auto-rewrites grep "A\|B" argv-mode to grep -E "A|B" with
SHELL_REGEX_AUTO_CORRECTED prefix; explicit -E/-G/-P/-F still pass
through. SAFE_SHELL_COMMANDS no longer includes sort/uniq.
- Extension loader hardening: load_extension(drive_root=...) is
mandatory (no silent ~/Ouroboros/data fallback), TestClient lifespan
+ settings hot-reload pin to app.state.drive_root, _sweep_stale_-
extension_imports preserves the live import root via a keep-list,
fixture cleanup uses unload_extension. Adds clean_extension_runtime_-
state superset helper and tests for cleanup, drive_root requirement,
and live-root preservation.
- Test pollution: scripts/cleanup_test_pollution.py (dry-run-first
utility), tests/_shared._make_safe_mock_ctx for advisory-workflow
ctxs, _make_safe_mock_ctx adoption in test_advisory_workflow*.
- PluginAPI v1.2: PLUGIN_API_VERSION bumped, skill_job_dir added to
the frozen Protocol + the contract test.
- docs/ARCHITECTURE.md, docs/CHECKLISTS.md, docs/CREATING_SKILLS.md,
prompts/SYSTEM.md updated for new behavior (light-mode wording,
cleanup script, skill.lifecycle topic, advisory pre-review path,
duplicate reviewer slots, auto-grant settings + permissions wording).
VERSION 5.15.0-rc.8 / pyproject.toml 5.15.0rc8 / README badge + Version
History row + ARCHITECTURE.md header all in sync.
First catch-up RC of the v5.15.0 reduction line. The v5.14.0 release missed its
-8000 LOC target by an order of magnitude (delivered -539 LOC vs goal); this RC
starts the honest catch-up by attacking the test layer where audit confirmed
real consolidation room exists.
Shared helpers (tests/_shared.py + tests/conftest.py):
- New `clean_extension_runtime_state()` superset cleanup helper. Three
duplicated copies in test_skill_exec / test_extensions_api / test_extension_loader
collapse to one shared helper with three thin autouse fixtures.
- New `ensure_claude_agent_sdk_mock()` SDK-mock installer. Three duplicated
copies in test_commit_gate / test_advisory_observability / test_claude_code_gateway
share one implementation.
- `tests/conftest.py` drops four unused fixtures (`make_git_repo`, `tool_context`,
`make_chat_mock`, `make_extension_skill`) — verified by grep that no test
parameter ever requested them. -73 LOC.
Per-file parametrization:
- test_extension_loader.py: 13 UI-tab rejection tests + 3 route-rejection tests
collapse into two parametrize tables.
- test_phase7_pipeline.py: 11-case preflight cluster + 4-case JSON parser cluster
parametrized; new `git_ctx` and `review_ctx` fixtures replace 12 paired
`_get_git_module() + _make_ctx(tmp_path)` setups.
- test_review_fidelity.py: 3-case truncation suite parametrized.
- test_advisory_observability.py: resolve-claude-code-model trio parametrized;
budget-gate trio (`_advisory_budget_gate_returns_skipped`,
`_handle_advisory_pre_review_returns_skipped_status`, `_budget_gate_skip_persists`)
share one new `budget_gate_env` fixture + `_budget_ctx` helper.
- test_commit_gate.py: registered-tools triplet parametrized.
- test_pr_tools.py: 37 paired `_make_temp_git_repo + _make_ctx` setups
collapse to a single `_setup` helper call.
Net: -317 LOC across 13 files, 3376 tests still green (full suite).
Note on plan-vs-actual: RC.1 budget was -3500 LOC. Audit-confirmed safely-
achievable ceiling for tests-only consolidation is ~1200-1350 LOC across the
12 large test files; the remaining ~700-1000 LOC will come from continued
parametrize work in test_advisory_workflow / test_skill_loader / test_runtime_mode_elevation
during stretch RC.5 if the diff-size watchdog flags a cumulative shortfall.
The bulk of the missing -8000 LOC will come from RC.2 (production dedup) and
RC.3 (legacy/compat removal) which have more deterministic surface.
Adds reviewed skill authoring/runtime capabilities, hardens skill lifecycle/provenance controls, and repairs the mobile/dashboard/marketplace/widget UI surfaces as one coordinated pre-release update.
Restores the visible grant-access affordance and adds a real on/off
toggle on the Skills tab. Three coupled changes:
1. Dual-track core-key grants. Extensions and scripts now share the
same content-hash-bound owner-grant model. extension_loader.py
replaces the hard "extensions cannot have core keys" rejection
with a per-key grant check that reads grants.json and refuses
load only when the persisted grant is missing or stale.
PluginAPIImpl.get_settings forwards forbidden keys only when the
loader passed them in via the new granted_keys ctor arg. The
contract docstring in plugin_api.py is updated to reflect the
new behavior. grant_status_for_skill now treats both script and
extension types as eligible (instruction stays unsupported).
launcher.py request_skill_key_grant accepts both types and posts
to the new /api/skills/<skill>/reconcile endpoint after writing
the grant so the running server picks up the change without a
manual disable/enable cycle. The reconcile endpoint clears
_load_failures and re-runs load_extension in the server process
(the launcher and server are independent OS processes).
2. video_gen demos the grant flow on a default install. Its manifest
switches from VIDEO_GEN_KEY (which was never in
FORBIDDEN_SKILL_SETTINGS and could not surface a grant button) to
the actual OPENROUTER_API_KEY it uses. Version bumped 0.6.0 to
0.7.0 so the per-skill version-aware re-seed fires on next launch.
3. Skills tab UX rewrite. The previous wall of badges (NATIVE / PASS
/ LIVE / ENABLED / GRANT MISSING) is replaced by:
- Single human status chip per card (Active / Off / Needs review
/ Needs access grant / Loaded — UI tab pending / Failed to load).
- iOS-style on/off toggle switch (input type=checkbox role=switch)
instead of the old Enable/Disable button.
- Friendly page header copy (no data/skills/* / ext.<skill>.* /
skill_exec jargon).
- Optional source chip (External / ClawHub / User repo) next to
the title for non-built-in skills (P1 provenance signal).
- Lock-hint row under the toggle when review or grant is missing.
- Promoted review-findings disclosure to the front face so
fail/advisory verdicts are one click away, not two.
- Friendly Access panel with Grant access button when core keys
are required.
- Show details disclosure for power-user metadata (permissions,
type, version, runtime info, version drift, provenance).
- WCAG-compliant 44x44 hit target around the 40x22 toggle track,
emoji-font fallback for the lock glyph, role=switch + aria-checked
so AT users hear "weather, on, switch" not the awkward
"Disable weather, checked, checkbox".
widgets.js drops the weather:widget debug label leaked into the
muted subtitle; cards now show "from <skill>" only when the title
differs.
Tests: 2726 passed, 0 failed. Five new test_extensions_api,
test_extension_loader, test_skill_loader, and
test_runtime_mode_elevation cases cover the dual-track grant flow,
the cross-process reconcile endpoint, partial-grant merging, and
stale-content-hash defense in depth. Three rounds of
adversarial-multimodel-review (gemini + gpt + opus) signed off as
SAFE TO COMMIT after seven medium/low findings were addressed.
No release: per request, version stays 5.2.1 and no tag/artifact
build is triggered.
Restore ClawHub marketplace visibility and reviewed skill key grants under the accepted local-trust model while keeping owner-facing runtime-mode changes behind the launcher bridge.
Two related fixes that close a same-process privilege-escalation path
demonstrated by the agent in a video-gen skill task and unblock the
natural light-mode workflow at the same time.
Q1 — light allows skills (Frame A): removes runtime_mode == "light"
blocks at ouroboros/tools/skill_exec.py::_handle_skill_exec, the
extension-tool dispatch in ouroboros/tools/registry.py::ToolRegistry.execute
and the runtime_mode_light short-circuit in
ouroboros/extension_loader.py::_extension_runtime_state /
load_extension. Skills already pass tri-model review_skill PASS +
enabled.json toggle + content-hash freshness + sandboxed subprocess
+ FORBIDDEN_SKILL_SETTINGS denylist; gating execution by runtime_mode
only denied owner-approved capabilities.
Q2 — elevation ratchet (multi-layer):
1. ouroboros/config.py::save_settings(allow_elevation=False) chokepoint
compares incoming OUROBOROS_RUNTIME_MODE against a boot-time
baseline (_BOOT_RUNTIME_MODE pinned by initialize_runtime_mode_baseline
from server.py::lifespan). The baseline is exported to
OUROBOROS_BOOT_RUNTIME_MODE env var so subprocesses (run_shell,
skill_exec, mp.spawn workers) inherit the pin via env. Once any
baseline source is available, allow_elevation=True becomes inert.
2. ouroboros/tools/core.py::_data_write blocks writes whose
Path.samefile resolves to SETTINGS_PATH (handles macOS APFS /
Windows NTFS case-insensitive bypass via parent-resolve +
name.lower() fallback).
3. server.py::_merge_settings_payload skips OUROBOROS_RUNTIME_MODE
from /api/settings POST body. api_settings_post adds belt-and-braces
revert. web/modules/settings.js no longer sends the key; the
Behavior tab segmented control becomes display-only with disabled
buttons + tooltip explaining the manual procedure.
4. ouroboros/tools/registry.py::_run_shell_safety_check runs two
substring filters on every run_shell invocation:
- Argv-level _detect_runtime_mode_elevation (conjunctive: BOTH
save_settings AND OUROBOROS_RUNTIME_MODE, OR the dotted path
ouroboros.config.save_settings) fires in ALL modes; the
module-level _LIGHT_MUTATION_INDICATORS (extended with
.write_text(, .write_bytes(, os.replace(, os.rename() fires only
in light.
- File-content scan (iter-3 fix for the file-based subprocess
bypass): _extract_script_file_args walks past interpreter flags
and resolves the script file argument; if it lives inside
repo_dir or drive_root, _scan_script_files reads up to 256 KB
and re-runs the same indicators on the content. Closes the
deterministic 2-step chain data_write helper.py + run_shell
python helper.py that previously poisoned settings.json directly.
5. ouroboros/file_browser_api.py guards api_files_write / delete /
transfer / upload via _is_owner_only_settings_file (samefile-aware).
6. supervisor/workers.py::worker_main calls
initialize_runtime_mode_baseline so spawn-method workers re-pin
from inherited env.
7. ouroboros/tools/registry.py extension dispatch now goes through
ouroboros.safety.check_safety (extracted to _dispatch_extension_tool
helper).
8. prompts/SAFETY.md gains a DANGEROUS clause naming the runtime_mode
self-elevation channels.
Tests: 2705 passed, 0 failed (added 47 new tests in
tests/test_runtime_mode_elevation.py covering chokepoint, downgrade,
data_write block with case-mismatch + symlink + env override, API
drop, set_tool_timeout regression, onboarding flow, boot-baseline
post-pin inert, env-var subprocess propagation, file-content scan
across all modes + light pathlib.write_text + benign-script
false-positive negative + interpreter parser unit tests; conftest.py
autouse fixture resets baseline between tests).
Pre-release went through two cycles of multi-model adversarial review
(Gemini + GPT + Opus critics) plus two cycles of the actual triad+scope
reviewer pipeline. Phase B (move owner settings outside drive_root
for OS-level containment) and the Privilege Boundary section in
prompts/SYSTEM.md are deferred per the minimal_safety profile and
recorded in data/memory/knowledge/improvement-backlog.md.
External review tool: scripts/run_external_review.py invokes
ouroboros.tools.parallel_review.run_parallel_review on the staged
diff with FULL raw triad+scope output (no truncation).
Note on changelog rolloff: the v4.42.1 patch entry was rolled off in
this release to respect the P7 5-patch-row cap. Its full body remains
at git tag v4.42.1.
Complete the three-layer runtime-mode boundary:
- light keeps the existing repo-mutation / skill_exec / extension dispatch block.
- advanced can still evolve the application layer, but now blocks protected core,
frozen-contract, and release/managed-repo invariant paths.
- pro can edit protected paths on disk, but protected commits must pass the extra
core-patch review gate before they land.
Adds a shared runtime-mode policy module and wires it through registry, git commit
paths, and the Claude gateway. Adds the pro core-patch gate, extracted review
revalidation helpers, protected-path rename/backslash coverage, pro CORE_PATCH_NOTICE
coverage, and fail-closed handling when core-patch scope review does not respond.
Moves shipped review defaults to the GPT-5.5 family: commit triad starts with
openai/gpt-5.5, scope review defaults to openai/gpt-5.5, and deep self-review
uses openai/gpt-5.5-pro. Direct-provider normalization now also maps scope review
models for OpenAI-only and Anthropic-only installs.
Fixes audit debts found during the re-review: full-repo prompt secret redaction,
canonical-doc de-duplication for scope/plan review, no accidental .review-drive
telemetry, managed-repo rescue reset fail-close on incomplete rescue, CI quick-test
coverage for ouroboros-three-layer, and updated SYSTEM/SAFETY/ARCHITECTURE docs.
Validation:
- python3 -m py_compile on changed runtime/review modules
- pytest tests/ -q (full suite green; 1 skipped; warnings are pre-existing local env/deprecation warnings)
- gpt-5.5 scope + triad review cycles until no unrebutted critical findings remained within the $100 review budget
No GitHub prerelease is published in this commit by user request; no tag is created here.
Closes the re-audit cycle. Six non-committing advisory+triad+scope
cycles converged with 0 open findings before this commit.
Extension runtime:
- extension_loader.py stages each load under data/state/skills/<name>/__extension_imports/<uuid>/skill/
to defeat Python module caching on rapid in-place edits; fail-closed if
any symlink in the skill tree resolves outside the reviewed checkout.
- extensions_api.api_extension_manifest reports live-SSOT load_error,
not the stale discovery copy, so index and manifest agree.
- tools/registry.list_non_core_tools() now merges live ext.* tools.
Release / packaging guards (BIBLE.md P7):
- scripts/build_repo_bundle.py fail-closed if no release tag is resolved,
verifies the tag is annotated (git cat-file -t == tag), and that it
points at HEAD.
- build.sh / build_linux.sh / build_windows.ps1 enforce the same three
checks before invoking PyInstaller. Synthetic/fake tags no longer
accepted.
- _validate_source_branch no longer treats an ls-remote-only branch as
usable; requires local ref availability, symmetric with
_ensure_source_sha_tracks_branch.
- release_sync._normalize_pep440 now collapses alpha->a and beta->b per
PEP 440 canonical pre-release spelling.
Onboarding preservation:
- Wizard no longer silently wipes OPENAI_BASE_URL, OPENAI_COMPATIBLE_*,
CLOUDRU_FOUNDATION_MODELS_BASE_URL on re-run. The wizard only resets
keys it actually exposes.
Docs and prompts sync:
- ARCHITECTURE.md describes the bundle-bootstrap model, the
__extension_imports/ staging surface, the build-script release tag
prerequisite, the extension HTTP review endpoint, and
VALID_EXTENSION_ROUTE_METHODS as part of the frozen contract.
- BIBLE.md + prompts/SYSTEM.md release invariant now distinguishes
author-facing spelling (VERSION / README / tag / ARCHITECTURE) from
the PEP 440 canonical form required by pyproject.toml.
- prompts/SYSTEM.md Immutable Safety Files list matches
SAFETY_CRITICAL_PATHS exactly (BIBLE.md included).
- docs/CHECKLISTS.md row 1 explains the author-facing vs PEP 440 split.
Runtime SSOT:
- server.py::_run_supervisor feeds workers.init the manifest-driven
branch names from _runtime_branch_defaults() instead of hardcoded
ouroboros / ouroboros-stable.
- contracts/skill_manifest.py accepts body-only instruction-skill
markdown that starts with a --- thematic break.
Regression coverage added for every real fix above across
test_extension_loader, test_extensions_api, test_build_scripts,
test_build_repo_bundle, test_launcher_sync, test_onboarding_wizard,
test_packaging_sync, test_release_sync, test_contracts, plus a new
test_release_workflow guard for the CI release path.
VERSION carriers synchronized:
- VERSION = 4.50.0-rc.3 (author-facing)
- pyproject.toml = 4.50.0rc3 (PEP 440 canonical)
- README.md badge + Version History row = 4.50.0-rc.3
- docs/ARCHITECTURE.md header = 4.50.0-rc.3
Wires Phase 4's extension registration surface into the actual runtime
dispatch paths and adds a Skills lifecycle UI.
ouroboros/extensions_api.py (new):
- GET /api/extensions → catalogue of every discovered skill
(bundled + OUROBOROS_SKILLS_REPO_PATH) + extension_loader.snapshot()
- GET /api/extensions/<skill>/manifest → parsed metadata
- ALL /api/extensions/<skill>/<rest:path> → catch-all dispatcher
honouring the methods tuple registered via PluginAPI.register_route
- POST /api/skills/<skill>/toggle → UI-direct enable/disable (wraps
extension_loader.load_extension / unload_extension)
- POST /api/skills/<skill>/review → tri-model review via
asyncio.to_thread so the event loop stays responsive during the
30-60s blocking LLM call
ouroboros/tools/registry.py ToolRegistry:
- execute() falls back to extension_loader.get_tool for ext.* names
(uses self._ctx, matching built-in tool dispatch)
- schemas() merges extension tool schemas so the agent's system
prompt includes ext.<skill>.<name> tools
- get_timeout() honours the extension's declared timeout_sec
server.py:
- ws_endpoint dispatches type: "ext.*" WS messages via
extension_loader.list_ws_handlers, sending one-shot <type>.reply
responses
- /review command now only matches exactly /review or /review <args>
— /review-skill and other /review-* commands route through the
normal chat pipeline so they can hit their own tools
ouroboros/skill_loader.py discover_skills:
- new include_bundled=True default merges the shipped
repo/skills/ reference directory with the user's configured
external path, so the bundled weather skill is discoverable out
of the box. Tests opt out via conftest.py's autouse
_hide_bundled_skills fixture
skills/weather/ (new bundled reference):
- type: script, runtime: python3, permissions: [net], one
fetch.py using stdlib urllib with a host-allowlist guard
(wttr.in only) — the minimum viable reviewed + enabled skill
web/modules/skills.js + index.html + style.css (new Skills page):
- nav-rail button + #page-skills with per-skill status cards
- Review / Enable+Disable / Refresh actions all HTML-escape
manifest-derived values (XSS guard against malicious external
skills)
- Banner feedback for success / failure with tone classes
docs/ARCHITECTURE.md:
- Section 1 module tree lists extensions_api.py
- Section 3 module list lists skills.js; sidebar updated to 8 pages
- Section 4 API table lists all four new endpoints
tests/test_extensions_api.py (9 Starlette TestClient tests) pins:
- catalogue + manifest endpoint shapes
- toggle endpoint's load+unload integration (enable triggers
extension_loader.load_extension, disable triggers unload)
- catch-all dispatcher actually forwards to registered handlers
- 404 for unknown routes
- review endpoint routes through the tri-model pipeline
(mocked to stay hermetic)
- AST guards: ws_endpoint dispatches ext.* and ToolRegistry.execute
falls back to extension_loader
tests/test_skill_loader.py adds test_discover_skills_includes_bundled_by_default
that explicitly overrides the autouse suppression fixture.
Rolloff: v4.44.0 was removed in this release to respect the P7 5-minor
cap. Full body at git tag v4.44.0.
Tri-model review: triad + scope over 4 rounds. Per-round verification
against the real diff; all substantive findings fixed (XSS escape,
self._ctx private attribute access, api_skill_review asyncio.to_thread
offload + event_queue attribute, /review-skill command shadowing,
bundled-skills ALWAYS surfaced, CSS class correctness).