- domains.toml: D14 ten leaves (extension_loader six + skill_review four) and
D08 sixteen leaves mapped; extension_loader/skill_review split rows fully
retired; control/events/queue/workers leaves-rows shrunk to their
hot-deferred remainders (cancel/custody family, D07 half of control).
- D08-pick conflict resolved: the reload_all suite lives in the D14 sibling
test_extension_reload_all.py on this tree, so its worker_main clause got
the D08 adaptation there (worker_process.py owner); the giant kept HEAD.
- test_gaia_events_serializer_carries_web_search_sources re-homed from the
test_devtools_benchmarks giant into tests/test_events_llm_usage.py (its
post-split thematic owner): the byte-debt ratchet correctly refused the
in-place retarget (+79 bytes on a shrink-only file) and the re-home is the
designed pressure valve - the giant shrinks, the pin gains its family.
- LEDGER_CORRECTIONS: coordinator section with the five D13 dispositions
(live protected safety.py deltas incl. the UNROWED _safety_drive_root fix
that must gain a carried row at F5; shell_guards rebind pending with D05
wave; runtime_mode_policy remainder returns with its leaves).
- Quotient report regenerated: 1316 strict module edges, 163 domain edges.
Module side. 22 D08 owners classified: 4 byte-identical across tip/reference
(schedule_contract, supervisor/__init__, active_activity, schedule_time);
6 pure upstream drift - tip bytes stand (promotion_source, tools/followup,
message_bus, queue_transitions, state, task_admission); 4 NEW post-cutoff
upstream modules with no ledger rows - untouched (cognitive_operations,
log_addressing, subagent_task_truth, task_dispatch); 4 monoliths split.
The split: 16 leaves, 142 moved spans, every span transplant-tool proof-green
against git show HEAD:<monolith> (ast=tokens=byte-roundtrip on every symbol,
leaf_invariants=[], unread_declared=[]). supervisor/events.py 4476->1947:
eight handler families (chat_delivery, subagent_admission, schedule_task,
project_routing, coop_checkpoint, budget, worker_reports, runtime_controls).
supervisor/queue.py 1587->1265: queue_schedules on the _queue handle.
supervisor/workers.py 4396->2808: worker_promotion/chat_lane/pool_lifecycle
on the _pool handle + worker_process (the child, no pool state).
tools/control.py 3225->2110 (SHARED D07/D08): only the D08 leaves
control_events/control_routing/control_runtime; D07 rows untouched.
Facades = tip parent - moved spans + re-export block, facade audit green
(every kept span byte-identical to tip, every moved name re-exported).
Drift-probe first per leaf: 4 reference leaves fully byte-true, 12 with
41 upstream-drifted spans re-emitted from tip bytes; no oracle semantics
replayed over drift.
HOT-DEFERRED with evidence (cancel/custody D09-class; upstream 65b5d19f
re-decomposed this ownership): events_task_done family, _handle_cancel_task,
_close_campaign_after_owner_stop, events_evolution_done, queue_snapshot,
queue_timeouts, queue_evolution (upstream's own evolution_lifecycle.py
supersedes), worker_assignment, worker_health. Deferred semantic-delta rows:
D06 events taxonomy (dispatch/EVENT_HANDLERS keep tip bytes), D04 retired
timeout knobs, six retired-name rows. Row 2016 (_handle_schedule_task)
deferred on a mechanism finding: its >300-line FUNCTION_DEBT entry is
(path, qualname)-keyed and this tree's transition validator has no D11
same-qualname relocation rule - the handler stays with its debt key.
Test side: identity suites test_events_extraction (5, pins the deferred
inventory as the F2 work order), test_control_extraction (5),
test_worker_process_extraction (5, reference-verbatim); LEAVES table in
test_module_handle_extraction gains six tool-derived declared-set rows.
Dead-patch class re-pointed to owner leaves mirroring reference adaptations
across 13 files (coop quiescence, schedules, worker_process trio, promote
admission seam -> control_events, pool-disabled probe -> control_routing,
evolution restart claims -> control_runtime alias, budget append_jsonl ->
events_budget, ephemeral-turn source scan -> worker_chat_lane, routing
producers AST scan -> control_routing, worker_main scan -> worker_process,
resource-leak source contracts reference-verbatim); all 15 touched test
files lossless (test-name multisets equal), no new ast-identical dup bodies.
HOT_CODE_PATHS mirrors the 12 hot leaves (D04-block precedent).
size-ratchet manifest regenerated with the official tool (queue.py enters
the 1001-1500 band with rationale; giants shrink in place); ratchet lane
5 passed. ruff check . --select F clean. CI-shape battery on this tree:
parallel 11965 passed rc=0, serial 609 passed rc=0. Import smoke fwd+rev
green; worker_main stays picklable from supervisor.worker_process.
docs/v7next/LEDGER_CORRECTIONS.md: D08 lane section (11 entries).
(cherry picked from commit 1426ea3ae2413580992575f5ab890497469a6657)
Module side (41 D14 owners at tip): 15 byte-identical across tip/reference/
base (event_bus, extension_isolated_deps, extension_reconcile_queue,
marketplace/{__init__,adapter,clawhub,fetcher,install,install_specs,
isolated_deps}, skill_owner_attestation, skill_readiness,
skill_repair_admission, skill_review_status, skill_token); 16 pure upstream
drift - upstream bytes stand (extension_companion, extension_health,
extension_process_runner, extension_ui_validation,
marketplace/{ouroboroshub,provenance}, skill_dependencies,
skill_lifecycle_queue, skill_loader, skill_publish_eligibility,
skill_review_history, skill_review_passes, skill_review_runner,
tools/skill_{exec,preflight,publish}); 8 NEW post-cutoff upstream modules
with no ledger rows - untouched (betterleaks_runtime, skill_payload_binding,
skill_publish_{github,result,scanner,snapshot} of 8cc2ac69,
skill_review_cycles of 386e9417, skill_review_usage of f18da8c3). The
reference's unrowed failure_kind delta on extension_process_runner is NOT
replayed (typed-dispatch family, F3 territory); the supervised-future leak
in PluginAPIImpl is preserved per plan (F3-acceptance carries the direct
regression test). PluginAPI semantics untouched - mechanical byte-preserving
splits only.
The splits: extension_loader.py 2194 -> 960 into six leaves per ledger rows
2467-2519 (registry_state 87, surface_names 132, child_catalog 109,
import_staging 173, liveness 242, plugin_api 744); 49/53 spans
byte-identical to the oracle leaves, 4 byte-falsified by pure upstream drift
(widget-geometry promotion, durable companion-health overlay) and re-emitted
from tip bytes; two unrowed tip riders travel with their readers
(_widget_geometry_from_render -> surface_names beside its rowed sibling,
_apply_durable_extension_health -> liveness). skill_review.py 1600 -> 841
into four leaves (packs 166, rebuttals 127, prompt 376, output 245); 25/31
ledger rows executed from tip bytes, 6 SUPERSEDED by upstream's own re-home
into skill_review_cycles (386e9417) - upstream ownership stands, the facade
keeps the historical aliases and the identity suite pins them.
_ws_broadcaster deliberately not aliased on the facade (rebindable global
keeps one binding at its owner - the reference contract, pinned). Proof
green: transplant-tool --check per leaf - ast=tokens=bytes=True on all 80
spans, undeclared_top_level=[], leaf_invariants=[], exit 0; facade audit:
17+14 retained top-level statements byte-identical to tip, every tip name
accounted, every moved name re-imported (minus the pinned _ws_broadcaster);
import smoke + shared-registry object-identity checks green.
Test side: identity suites carried - test_extension_loader_extraction.py
(+2 rider rows) and test_skill_review_extraction.py (tool_module_inventory
clauses reverse-mapped out: v7-only mechanism absent at this tip;
cycles-alias identity clause added; facade size bound 800 -> 900 for the
tip-retained lifecycle members, which join the patchable-seams pin).
test_extension_loader.py 1717 -> 359 split into 4 siblings +
_extension_loader_shared.py per 45 ledger rows (zero upstream drift on the
file; 43/45 moved bodies byte-identical, 1 oracle dual-supervisor-patch
adaptation kept, 1 reverse-mapped to the tip spelling supervisor/workers.py
- the reference's worker_process.py is the still-pending D08 split).
test_skill_review.py 1943 -> 579 into 5 siblings + _skill_review_shared.py
per 65 rows (58 byte-identical, 3 re-emitted from tip test drift, 3
patch-retarget adaptations kept; the upstream-deleted advisory test is not
resurrected - its f8d87c69 successors stay in the remainder on tip bytes,
theme re-home deferred to F5). Lossless both families: 52==52 and 74==74
test names, zero duplicate names, no new ast-identical bodies (10
pre-existing review_cycles duplicates at base noted, untouched). Dead-patch
class closed: the remainder's advisory pre-review patch -> prompt owner,
extension_companion supervisor patch doubled onto the plugin_api owner
(single-module patch proven dead by a red run); every other facade patch
site of moved names verified live (production consumers do call-time facade
imports).
size-ratchet manifest regenerated with the official tool (extension_loader,
test_extension_loader, test_skill_review leave GIANT_PATHS; no new band
entries); --check exact; ratchet 5 passed; ruff F clean tree-wide.
Receipts: 12 identity + 126 split-family + 1751 targeted/-n16-loadscope +
full CI-shape battery 11950 parallel + 618 serial, all rc=0; HEAD held and
worktree status unchanged through every pytest run. Ledger corrections:
docs/v7next/LEDGER_CORRECTIONS.md D14 section, entries 1-10.
(cherry picked from commit c7e3d04df95eb9470f5ca3ad824485083a201fe4)
Wave-3 conformance (sol at 92238298) passed byte parity, lossless names,
facade identity and the seam; two test-quality residuals closed:
- the one-attempt pin now counts at the true physical boundary: the real
_execute_candidate runs and execute_physical_attempt is the stub, so an
internal retry anywhere in the lane shows up as calls > 1. Proven by a
mutation probe (a 2-try loop injected into _execute_candidate turns the
pin red).
- test_llm_extraction owner map gains the two re-exported names it missed
(_RESPONSE_METADATA_LABEL_MAX_CHARS, _bounded_response_metadata_label),
making the every-moved-name identity claim complete.
domains.toml: the ten D02 leaves enter [modules]; llm.py leaves both
split_pending sections (D05 carried its own toml rows in-lane). Quotient
report regenerated. D19 contributed a zero-work classification (7
byte-identical + 4 pure-drift contracts, no commit - sha256 evidence in its
lane report), nothing to merge.
Module side (23 D02 owners): 5 byte-identical across tip/reference/base
(fallback_cooldown, local_model, local_model_autostart, model_concurrency,
pricing); 2 pure upstream drift - upstream bytes stand (llm_observability,
vision_routing); 12 NEW post-cutoff upstream modules with no ledger rows -
upstream bytes stand untouched (anthropic_native_custody, net_transport,
openai_chat_custom/dispatch, openrouter_attribution, request_wire_* x6,
route_spec, transport_custody; transport/timeout contracts not touched).
The split: llm.py 4414 -> 727 composes LLMClient from ten mixins per ledger
rows 1666-1793/4001-4003 (131 rows). 100 spans byte-identical to the oracle
leaves, 28 byte-falsified by pure upstream drift (request-wire custody
041e6e39, attribution 9a20df6a, custody hardening 802f1056/f702439f) and
re-emitted from tip bytes. Proof green: transplant-tool verify per leaf
(ast=tokens=bytes on every module-level span, undeclared_top_level=[],
leaf_invariants=[]) + member-level byte proof (117 mixin members == tip
LLMClient members) + facade audit (14 kept members and 4 kept top-level
defs byte-identical to tip; every tip top-level name and member accounted).
Two non-tip-byte spans, both ledger-sanctioned: row 1674 (the documented
one-identifier requalification) and row 1784 - the LIVE D09 delta: the local
lane's reintroduced `for attempt in range(3)` retry loop is deleted (one
physical attempt per call; transient failures surface to call_llm_with_retry),
RE-DERIVED on tip bytes to preserve upstream 802f1056's exception-owned
capture custody clause that a verbatim oracle replay would have reverted.
The D09 typed-policy-refusal subfamily (rows 1706/1749/1751/1759/1760 +
ProviderPolicyRefusal machinery) is HOT-DEFERRED with evidence: zero
raisers/classifiers at tip, all consuming rungs upstream-reworked.
provider_models: ledger note-contract completed (lazy config imports ->
top-level model_slots/settings_defaults leaf imports, cycle-free); the
reference pin restored under its ledger name. llm_probe: reference delta
adopted (executor import named at its owner leaf, tip==base). Unrowed tip
symbols _RESPONSE_METADATA_LABEL_MAX_CHARS/_bounded_response_metadata_label
ride with their only reader into llm_openai_compatible; facade re-exports.
Facade keeps the full tip import surface (noqa discipline).
Test side: identity suite tests/test_llm_extraction.py (7) and provider-route
goldens tests/test_llm_provider_golden.py + 9 fixtures carried; goldens
re-baselined from tip behaviour via the suite's own --write (drift classes:
attribution headers, request_wire disclosure, response metadata labels,
effort-ladder; the volatile request_wire.attempt_id projected to a presence
flag; the 2 typed-refusal cases removed - they pin the deferred organ, as does
test_llm_typed_policy_refusal.py, not carried). D09 pin
test_local_transport_makes_exactly_one_physical_attempt added; the three
local-lane retry tests re-pinned 3 -> 1 attempts. Dead-patch class closed:
execute_physical_attempt patches -> llm_attempt (6 files),
chat-path _execute_candidate/last_physical_attempt_capture -> llm_fallback
(2 post-cutoff files, disclosed in-file), local executor -> llm_local;
TTL-consumer path pin -> llm_attempt. D01-owned oracle adaptations
(loop_messages/loop_round_limits) reverse-mapped to tip spellings, not
carried. Test-name accounting: 11 touched files lossless, +1 pin test,
1 pin renamed to its ledger name, 2 new files.
size-ratchet manifest regenerated with the official tool (llm.py leaves
GIANT_PATHS; no new band entries); ratchet 5 passed; ruff F clean tree-wide;
receipts: 324 (targeted battery) + 1788/-n16-loadscope + 1 serial across all
65 llm-importing test files, all rc=0; HEAD held through every pytest run.
Ledger corrections: docs/v7next/LEDGER_CORRECTIONS.md D02 section, entries
1-10.
(cherry picked from commit 1c25ee661df2745c6d5d1346d996d036cb031f40)
Shell: shell_process (11 spans) / shell_effects (12) / shell_outputs (16) by
verbatim extraction with transplant proof (ast=tokens=bytes, exit 0); the
facade re-exports every moved identity. Ten shell_outputs rows were already
extracted upstream into tools/shell_audit.py (c7315c57) — superseded, ledger
correction recorded. Core: core_file_tools (30 spans) / core_artifacts (9)
emitted from tip bytes; the reference's typed-result cutover deltas on ten
producers are deliberately NOT ported (F2 organ). tools/core.py keeps a
re-export facade (§5.3-Δ2 item 12) — divergence from the reference's
no-facade cutover disclosed in docs/v7next/LEDGER_CORRECTIONS.md; rowed test
bindings 361-371 landed. Path-keyed mirrors: _POPEN_ALLOWLIST row moved to
shell_process; load_settings/module-handle monkeypatches retargeted to the
leaf owners per the reference adaptations. Size ratchet regenerated with a
band rationale for the core.py band re-entry (2283 -> 1373).
(cherry picked from commit 999184e6ee4e6973883720961b4c728811b14163)
Wave-2 conformance review (sol at 0859b681) passed byte parity, lossless
splits, protection closure and ledger cardinality; this closes its two
NEEDS-FIXES clusters:
- domains.toml seam repair: the first seam commit wrote the registry leaf
list into [split_pending] (whose values are domain IDs) and left the full
7-leaf list in [split_pending_leaves]; now ['D04'] in the former, the two
hot-deferred leaves in the latter; report regenerated (unchanged counts).
- verifier target unfold is now depth-recursive and shared (_unfold_target)
by extraction and the leaf gate: with only A requested
fails CLOSED at extraction naming X/Y; a non-Name leaf (obj.attr) marks
the statement complex and is always an extra. 3 new probes (49 passed).
- LEDGER_CORRECTIONS: append-only coordinator section supersedes the D04
'four leaves' count typo and D12's stale 'tuple gate exits 2' claim.
- domains.toml: 15 landed leaves mapped (D18 launcher_windows_runtime; D12
model_slots/review_model_routes/runtime_limits/settings_defaults/
settings_scales; D04 four tool_access_* + five registry leaves); completed
split_pending rows retired; registry.py row shrunk to its two hot-deferred
leaves (registry_core, tool_result). Quotient report regenerated (1123
strict module edges, 163 domain edges).
- transplant verifier: unfold tuple/list assignment targets in the
undeclared-top-level gate (A, B = 500, 10 with both names requested is a
moved span, not an extra - the D12 lane's proven false positive); foreign
name in a tuple still fails. 2 new tests (44 passed).
- Disclosure: the D04 pick extends SAFETY_CRITICAL_PATHS and HOT_CODE_PATHS
(two protected files) to the five registry leaves - strictly additive,
mirrors the owner-approved reference delta (label parity: guard bodies
moved out of the protected parent keep its protection), pinned by
test_registry_split_leaves_keep_protected_label_parity.
Module side: config.py (1598) hands its vocabulary to the five oracle leaves,
re-emitted from TIP bytes with the transplant proof green on every span:
settings_defaults (361; SHARED leaf emitted FULL from BOTH parents per the
ledger - config.py rows 840-854 + provider_models.py rows 3238-3241, final
--check against the concatenated parents), settings_scales (111), model_slots
(113), review_model_routes (131), runtime_limits (192); the facade shrinks
1598 -> 883 with the oracle's re-export import block. The drift-probe
falsified the reference as a copy source on 7 spans (SETTINGS_DEFAULTS,
RETIRED_SETTING_KEYS, ENDPOINT_AUTHORED_SETTINGS, OPENROUTER_REVIEW_DEFAULTS,
get_websearch_timeout_sec, get_search_code_wall_sec, get_max_subagent_depth)
plus one structural reshape: upstream's tuple statement also binds the
unrowed MAX_SUBAGENT_DEPTH_HARD_CAP, which rides into runtime_limits and is
re-exported (F5 row needed; the tool's undeclared-top-level gate has a
Tuple-target blind spot - span proofs green, documented). provider_models.py
is touched by span removal + the re-export import ONLY (D02 module). The
reference's D04 retirement of the SOFT/HARD timeout knobs is
upstream-diverged and NOT replayed.
launcher_onboarding.py lands the launcher HALF of the approved D03 settings
seam (rows 918-920): zero upstream drift since merge-base, reference bytes
verbatim - startup stops persisting the pre-server provider normalization;
the server-side mirror stays for the D11 lane. The config.py in-place seam
rows 914/916 (normalize_settings_raw / serialize_settings) are HOT-DEFERRED:
upstream rewrote the read path through the new settings_integrity module, so
a verbatim replay would revert it; their pin test_settings_read_seam.py
defers with the machinery.
Test side: test_config_extraction.py + test_settings_env_on_disk.py
transplanted (adaptations: the tuple-bound cap in the owner inventory, one
literal re-pinned to tip's ENDPOINT_AUTHORED_SETTINGS, the provider_models
clause narrowed to this tree - the no-config-import half returns with D02);
the onboarding pins follow the launcher delta (row 920 rename, wizard and
server_runtime launcher clauses; every server clause keeps tip bytes).
Lossless: 19==19 / 63==63 / 28==28 test functions on the touched suites, one
ledger-mandated rename. Two path-keyed mirrors updated (prompt-cache TTL
definition-sites, hotreload key-paths set).
CI-shape battery on the final tree: parallel 11807 tests 0 failed (3
skipped) rc=0; serial 618 tests 0 failed (9 skipped) rc=0; -m size_ratchet 5
passed (official manifest regeneration = byte no-op); ruff --select F clean;
HEAD held through every pytest run. Ledger corrections: entries 13-19 in
docs/v7next/LEDGER_CORRECTIONS.md.
(cherry picked from commit 3a6a0344eabb9b933352c5fc32cabb0a79b186e5)
Module side: 4 of 8 D04 owners are byte-identical to the reference and merge
base (protected_artifacts, tool_policy, tools/__init__, tool_discovery), one
is pure upstream drift (tool_capabilities - upstream bytes stand), one keeps
upstream bytes because its reference delta is the typed-dispatch cutover
(extension_dispatch, rows 187/188 deferred). The two splits land from TIP
bytes with the transplant tool's triple proof green on every symbol
(ast=tokens=bytes, leaf_invariants=[], exit 0):
- tools/registry.py 3960 -> 2686 (PROTECTED file: pure byte-preserving span
relocation + one re-export block + noqa on historical imports - proven by
line-accounting audit): tool_context (2 symbols), tool_catalog (1),
tool_resolution (28), registry_guards (16), registry_guard_process (27).
13 spans were byte-falsified as oracle copy sources by pure upstream drift
and re-emitted from tip bytes; 5 rows whose reference destination carries
typed-result semantics moved their TIP bodies verbatim (typed deltas ride
with F2). HOT-DEFERRED with evidence: registry_core.py (ToolRegistry is a
2252-line class; the reference slimmed it via 17 method->function
extractions, which are not byte-preserving), tool_result.py (32/33 symbols
are the D02 typed organ, absent at tip).
- tool_access.py 1591 -> 782: tool_access_types (14), tool_access_paths (10),
tool_access_roots (9), tool_access_user_files (8); 39/41 spans byte-equal
to the reference, _skill_payload_base and ResolvedResourceBinding
re-emitted from tip (upstream refactor/field would have been reverted by an
oracle copy). The D1 mirror-path defect travels UNFIXED per lane orders.
Protection closure (protective-only, mirrors the reference and the tree's own
LC2 parity rule): the five landed registry leaves join SAFETY_CRITICAL_PATHS
and HOT_CODE_PATHS - code that moved out of the protected, hot registry keeps
both labels; pinned by the new parity test.
Test side: tests/test_tool_capabilities.py 1991 -> 681 splits into 4 siblings
per ledger rows 784-825 from tip bytes (34/42 spans oracle-equal, 8 re-emitted
from tip); lossless 61==61 test functions, tree-wide AST dup scan clean.
Pins carried with disclosed adaptations: test_tool_owner_facades.py (+alias_for
row), test_tool_access_extraction.py (4 adaptations in docstring),
tool_resolution identity test appended to test_workspace_authority_binding.py
(typed companion deliberately not carried).
size-ratchet manifest regenerated with the official tool (test_tool_capabilities
leaves GIANT_PATHS; no new band entries); ratchet lane 5 passed; ruff F clean;
90+518+586+257 tests green in isolation at the lane base; HEAD held through
every pytest run. Ledger corrections: docs/v7next/LEDGER_CORRECTIONS.md D04
section, entries 1-11.
(cherry picked from commit 2321514369b6f003e92bcd84e6a9a7efda6857df)
The one ledger-derived D18 split: launcher.py's Windows pythonnet/pywebview
preparation (_prepare_windows_webview_runtime, _show_windows_message,
_windows_dll_dir_handles) moves whole to ouroboros/launcher_windows_runtime.py
(MIGRATION_v7 rows 3998-4000). Drift-probe of the reference leaf against tip
monolith bytes was green on every span (ast=tokens=bytes=True, exit 0), so the
leaf is tip bytes; the facade re-exports the same objects and differs from the
reference facade by exactly upstream dc4c0204's delegated-restart hunk.
launcher.py 1582 -> 1484 lines; band re-entry carries an official rationale.
Reference pin test_launcher_reexports_the_windows_runtime_leaf appended to
tests/test_launcher_sync.py (byte-identical to the reference file afterwards).
Everything else the domain owns is classification, recorded in ledger
corrections 13-18: cli.py, ouroboros/__init__.py, launcher_server_reaper.py,
packaged_cli_install.py byte-identical across tip/reference/base;
launcher_bootstrap.py, platform_layer.py pure upstream drift (zero v7 delta);
packaged_cli.py::_save_settings HOT-DEFERRED with the half-absorbed settings
seam; the utils.py O_BINARY and reaper-test cross-OS deltas superseded by
upstream's own class fixes; packaged_runtime/packaging_sync reference deltas
deferred to their owning lanes (D33/F2) or F5 (unrowed).
(cherry picked from commit 6aad318d5a4100662ed6b31877f94d342c409b6f)
- scripts/x.py-not-real no longer matches by its existing .py prefix: the
token regex gained a trailing lookahead - a partial token is prose, not a
reference (a sentence period after a path still parses).
- Containment uses pathlib parents membership instead of '/'-suffixed string
prefixing, portable across Windows separators.
- tests/../README.md traversal: paths are resolved and must stay inside
their top directory; any '..' segment is rejected outright.
- not-scripts/x.py misread: the token regex now requires a non-path
character (or start) before tests//scripts//docs/ via lookbehind.
- smuggled bogus reference: existence is enforced for ANY extension - a
missing tests/missing.js next to a valid reference is an error, not
ignored.
Probes: all three bypasses red, legit single/multi-reference hooks and the
all-done synthetic manifest stay green.
Round 3 closed F1/F3/F4/F5 and both lane/dedup sweeps; the one residual was
F2's hook clause: verification hooks stayed free prose and the release gate
never checked them. Now at --release every non-post-release row's hook must
carry at least one repo-path reference (tests/, scripts/ or docs/ file) and
every referenced path must exist - prose-only hooks cannot ship. Outside
--release hooks remain prose (they legitimately name future suites while
work is pending). Probes: all-resolvable synthetic manifest passes; a
prose-only hook and a missing-path hook each turn release red.
The D09 lane added tests/test_preflight_process_containment.py (the reference
thematic sibling) without shrinking the upstream origin - the same class the
F0 review flagged for the D15 pilot (F5). A whole-tree duplicate-name sweep
(AST-identical bodies) found the 9 copies; they leave
tests/test_preflight_runner.py (the sibling owns them), 155 passed across both
suites. The same sweep found 10 upstream-native duplicates between
test_review_cycles_dispatch.py and test_review_cycles_skill_dispatch.py -
those are upstream's own bytes (both files untouched by this campaign), left
as an upstream issue candidate, not deduped here (no ledger row would cover
the delta).
context_runtime_facts.py -> D03, headless_status.py + workspace_patch_capture.py
-> D17 in [modules]; their completed split_pending rows retired. Report
regenerated on this tree (1097 strict module edges after the three new leaves,
163 domain edges unchanged, fingerprint header).
F0 review round 2, F5: the D15 pilot copied post-task reflection/synthesis
tests into tests/test_post_task_reflection.py and
tests/test_root_post_task_synthesis.py without shrinking the origin. The
copies are the owners; this removes the 16 AST-identical duplicates and the
_capture_summary_and_reflection_prompts helper (dead in atp once its only
three callers moved) from tests/test_agent_task_pipeline.py (1515->837
lines). Zero duplicate test names remain between the origin and the copied
suites (the context-side 15 were removed by the D03 lane commit). Affected
suites: 37 passed. Ratchet manifest regenerated (origin leaves the giant
list); -m size_ratchet re-run in the battery that follows.
headless.py (1573 at tip) gives up its two ledger-assigned leaves again:
ouroboros/headless_status.py (50, 11 symbols) and
ouroboros/workspace_patch_capture.py (668, 19 symbols). All 30 spans are
byte-identical between the reference leaves and git show HEAD bytes — the
hardened transplant --check (mandatory byte gate, undeclared-top-level check,
def681bd) is green on every symbol of both leaves. The facade (947) replays
the oracle's exact edit script over tip bytes: its only divergence from the
reference facade is genuine upstream residue drift (child_ref promotion
machinery, import changes), verified hunk by hunk.
Test splits per the ledger, upstream bytes as truth:
- test_headless_cli.py 2824 -> 462 + five themed siblings + shared fixtures;
93 test functions preserved exactly (lossless set equality), one adapted
span kept (the _PATCH_MAX_UNTRACKED_FILE_BYTES monkeypatch retargeted to
the new leaf, ledger row 739's own adaptation); nine oracle spans carrying
OTHER domains' v7 spellings reverse-mapped to upstream signatures keyed to
git show HEAD (registry._run_shell_safety_check string form, tools.core
_repo_read, queue.init 3-arg, queue.QUEUE_SNAPSHOT_PATH).
- test_workspace_executor.py 1995 -> 541 + three siblings + shared builder;
two reference spans byte-falsified by upstream drift (06339bb7 readiness
truth, a849c9a6 probe uncertainty) re-emitted from tip bytes and recorded
in docs/v7next/LEDGER_CORRECTIONS.md.
- test_agent_task_pipeline.py 1658 -> 1515: only the five ledger-assigned
_store_task_result rows carved into tests/test_store_task_result.py; the
other siblings belong to their own lanes.
Thirteen post-cutoff upstream tests have no ledger rows; placed by the
split's theme rule (task_api x4, task_artifacts x1, docker x6, services x2),
disclosed in LEDGER_CORRECTIONS for F5 row-minting.
Pins and mirrors: test_headless_extraction.py transplanted with the
tool_module_inventory clause reduced under an oracle-SHA note (that leaf
belongs to the tools lane); conftest serial table gains the executor family;
the process-custody Popen allowlist row moves headless.py ->
workspace_patch_capture.py with the oracle's justification. All 14 non-split
D17 runtime modules re-proven zero-v7-delta (task_results.py included:
upstream-hot drift stands, no ledger split assigned).
size-ratchet manifest regenerated with the official tool (three test giants
leave GIANT_PATHS, no new band entries); --check green. ruff F clean.
135/105/244/113 tests green in 4-var isolation; HEAD held after every run.
(cherry picked from commit 8dac8303006085bfdb636bacbfc402d6905fae7c)
Cancellation/custody sits on the F2 border and the drift probe drew the
line sharply. Quiet side, landed: the D07 Emergency Stop delta on
ouroboros/server_control.py - execute_panic_stop takes bound_port as a
keyword instead of lazy-importing the composition root, server.py hands
down _actual_bound_port() - applied in place onto tip bytes with an exact
two-way byte proof (vs tip: only the D07 hunks; vs the reference leaf:
only upstream's dc4c0204 reconcile_delegate_custody hunk). Four reference
pins ride along: test_panic_stop_port_sweep.py and
test_owner_stop_fences_s6.py re-pinned to tip bytes where upstream
genuinely moved (kill_workers kwargs, 5 host leaves not 11, the bounded
re-sweep's two token-less calls - each noted in-line),
test_cascade_chatless_residual_s6.py and
test_preflight_process_containment.py verbatim-green.
Hot side, HOT-DEFERRED with evidence in docs/v7next/LEDGER_CORRECTIONS.md
(entries 7-12): the 16-symbol cancel_custody extraction (upstream 65b5d19f
re-decomposed the same ownership differently), the cancel_intents D08
corruption rule (upstream absorbed it for request_cancel/claim_intent,
four mutators still fail open), the S7b test split bound to the deferred
extraction, and two D09-listed pins that belong to D07's module delta and
the delegation organ (delegate_start now refuses selectorless starts).
The E-suite stays with F2/F4 per the lane charter; cross-domain test
adaptations (events_task_done, loop_delivery, shell_process/git_ops_reset)
are left for their owning lanes. scripts/v7next_transplant.py +
tests/test_v7next_transplant.py byte-synced from the hardened def681bd (no
tool-emitted leaves in this lane; the in-place proof above is the byte
gate). 204 tests green in isolation at the lane base; size-ratchet
manifest regeneration is a no-op and lane 5 passed; ruff F clean.
(cherry picked from commit f1a37d3b413b66a9167fcf7e1a6fee5e229e60be)
Module side: 4 of 10 D03 owners are byte-identical to the reference and 4 more
are pure upstream drift (fit/budget/compaction/capability_evidence - nothing to
do, upstream bytes stand); the one split lands again from tip bytes:
ouroboros/context_runtime_facts.py (241 lines, 4 fact builders, 0 handle
rewrites - the leaf is projection-only by design) with the triple proof green
on every symbol (ast/tokens/bytes), and the facade shrinks 1590 -> 1380 with
the oracle's in-place re-export block and noqa discipline. The drift-probe
caught the reference leaf byte-falsified as a copy source: upstream b14ba397
rewrote _delegation_capability_fact (configured_route dropped, profile
evidence added) - copying the oracle leaf verbatim would have reverted that
feature; recorded in docs/v7next/LEDGER_CORRECTIONS.md (entries 7-11).
Test side: the 2080-line tests/test_context.py giant completes its S7a split
from TIP bytes (the giant drifted upstream, so oracle files could not be
carried blind): runtime_section (502) / memory (the D15-carried file
re-derived from tip and proven identical - not stale) / drive_state (233,
==oracle) / advisory_review (300, ==oracle) / _context_shared (49, ==oracle),
remainder 612. Lossless: 73==73 test functions, zero lost, zero added, and
the 15 memory tests that ran TWICE since the D15 carry are deduplicated. Two
upstream test rewrites (actor-scoped backlog digest, profile-evidence
delegation fact) travel to the oracle's destination as identity
continuations; 8 unrowed upstream recent-chat tests stay in the remainder
pending F5 rows. Oracle adaptations of test_context_fit_integration /
test_loop_compaction (loop -> loop_model_call) are foreign-domain spellings
whose leaf does not exist on this tree - reported, not applied.
size-ratchet manifest regenerated with the official tool: test_context.py
leaves GIANT_PATHS, context.py enters the band from the 1501-1600 zone with a
FLAGGED rationale (not self-approved). 74+110+228 tests green in isolation at
the lane base; ratchet lane 5 passed; ruff F clean; HEAD held through every
pytest run.
(cherry picked from commit 2021008bdb35436818b9ef3146ca69cb3a317b23)
Round 2 (gpt-5.6-sol at a0d8f2f7) verified F6 CLOSED and narrowed the rest:
- F1: handle shape validation now rejects async defs, positional-only
parameters, constant returns, and handles whose only returns live in
nested functions; every own-body return must be a Name/Attribute module
reference (generate_handle_def shape). 5 new shape tests (42 passed).
- F2: REQUIRED_PHASE pins the phase of every required row (rescheduling
needs a new owner decision); post-release is allowed only for
OWNER_DEFERRED={ABI-8} (Q5=A/Q16=A). Mutation probes: ABI-1 phase flip
and D02 post-release flip both turn rc=1.
- F3: the quotient report header no longer claims a commit SHA it cannot
know pre-commit; it states the base HEAD explicitly and carries a
content fingerprint (sha256 over sorted tracked runtime files,
path+bytes) so any tree can verify freshness by regenerate-and-compare.
- F4: try_swallows_import_failure() skips handlers that re-raise
(except ImportError: raise stays a strict witness); conditional
re-raise errs toward STRICT by design.
- F5 (residual duplicate tests vs test_agent_task_pipeline.py) is being
closed in the integration sequence: the queued D03 lane removes the
test_context duplicates and a dedicated dedup commit follows the D17
pick; round 3 reviews the combined tree.
F0 phase review (gpt-5.6-sol, adversarial, at def681bd) returned NEEDS FIXES;
this lands the dispositions:
- F1 CRITICAL: verify_transplant() proved span fidelity but not leaf
runnability. New _validate_leaf_invariants(): the handle must exist exactly
once with a canonical parent-returning body whenever the leaf reads through
it (projection-only leaves may omit it), declared and preamble-bound names
must be disjoint, declared names must actually be read, and the leaf-owned
allowlist may not absorb declared names. 6 new mutation/positive tests
(37 passed). Already-landed leaves are covered by the green full battery
(runtime import subsumes the static handle check).
- F2 HIGH: adoption validator now requires the full ABI-1..10 + CPL-1..7
inventory (mutation-probed: deleting ABI-3 turns rc=1), couples the
post-release row state (disposition/status/phase together), and lets
owner-deferred post-release rows through the release bar; docstring 17->18.
- F3 HIGH: domains.toml maps ouroboros/usage_legacy_import.py -> D16 and
retires the completed usage_accounting split-pending rows; quotient report
regenerated at this tree (163 domain edges).
- F4 MEDIUM: import collector now classifies __main__-guarded and
failure-swallowing-try imports as 'guarded' (import-time but tolerant),
removing false strict cycle witnesses (extension_process_runner:996,
launcher:133).
- F6 HIGH: ADOPTION rows synced to owner decisions - ABI-8 handler-ABI is
post-release backlog (Q5=A, Q16=A), ABI-6 drops the rejected
_updater_imports change (baseline pin is by design).
- F5 (duplicated D15 test copies) is closed by the queued D03 lane
integration which shrinks test_context.py and dedups; not double-fixed here.
Size ratchet: transplant tool entered the 1001-1500 band with rationale via
the official regenerator; -m size_ratchet lane 5 passed, ruff F clean.
The 15:03 audit proved a false-green: byte_identical was computed but never
gated ok (token equality is blind to inter-token whitespace - 'x + y' and
'x + y' tokenize identically), and the verifier checked only the requested
symbols, so an extra top-level definition or an import-time side effect in a
leaf rode an ok=true proof. Now: (1) the byte round trip is MANDATORY per
symbol, failing with the first divergent offset; (2) every top-level node of
the leaf must be attributable - requested symbol, docstring, imports, the
handle def, or a deliberate leaf-owned assignment allowlist - anything else
fails named with its line. Four mutation tests pin exactly the classes the
audit named (whitespace edit, smuggled def, import-time side effect) plus the
allowlisted-preamble pass. 32 tests, exit 0 preserved and shown.
The D15 pilot commit (5d3398c1) shrank tests/test_evolution_state_integrity_v3
without regenerating the size-ratchet manifest, so its manifest does not match
its own tree and the pairwise history transition condemns it permanently (a
linear descendant cannot heal history). Per the documented recovery, this
merge re-ties the mainline: first parent = the last clean commit (23cacfa2),
second parent = the condemned-but-now-healed chain (1a17218d, whose tree
carries the regenerated manifest from the D16 lane), tree = that healed tree.
Nothing is rewritten; the branch fast-forwards onto this merge.
Process fix already landed in the campaign plan: every integration commit
runs the size_ratchet lane BEFORE push.
usage_accounting.py (sitting exactly at the 1600 hard cap upstream) gives up
its L-C2 extraction again: ouroboros/usage_legacy_import.py (270 lines, 5
symbols, 3 handle rewrites matching the oracle's declared set exactly) emitted
by scripts/v7next_transplant.py with the triple proof green on every symbol
(ast/tokens/bytes), and the facade shrinks 1600 -> 1384 with the oracle's EOF
re-export block and noqa discipline. The one residual diff vs the reference
leaf is a genuine upstream comment rewrite (e9bf6f14) - upstream bytes win;
the reference ledger row is byte-falsified as a copy source and recorded in
docs/v7next/LEDGER_CORRECTIONS.md (transform stays valid).
Pins transplanted with the campaign's reverse-mapping discipline: LEAVES /
LC2 owner tables reduced to the rows whose leaves exist on this tree (32
v7-only modules absent, noted with oracle-SHA keys); the settings-writer
exempt table is re-keyed to the leaf exactly as the oracle's own adaptation
did. size-ratchet manifest regenerated with the official tool: the new 1384
band entry carries a FLAGGED rationale (not self-approved), and the
regeneration also heals the stale giant entry the D15 pilot left behind
(the base was red on the size_ratchet lane - caught by this lane).
512/110/206/33 tests green in isolation at the lane base; ratchet lane 5
passed; independently re-verified.
tests/system_e2e/: SystemHarness over IsolatedServer with (1) a genuinely
KEYLESS lane - every provider credential family and proxy var is stripped
from the child env AND the settings builder pins every model slot to the
loopback stub, with a regression test planting live-shaped keys and proving
none reach the child, plus a companion test that PINS the upstream
pass-through hole itself so its eventual fix collapses our override loudly;
(2) ScriptedStubModel - ordered per-scenario tool scripts and a review-organ
branch placed BEFORE the finalization check, classification pinned against
the tree's own parsers and prompt-marker literals (drift = red);
(3) ArtifactOracle durable readers incl. per-task forked drive roots.
Smokes on this exact tree: S1 boot/identity/contract 15s; S2 review-organ
62s - a doc-only commit_reviewed LANDS in the isolated clone under BLOCKING
enforcement with stub triad ([]+NO_FINDINGS) and scope (8-item matrix
validated by normalize_scope_items), advisory-bypass ledger and events
asserted. 11 passed re-verified independently; ruff F clean.
Operational facts recorded for F4: the stub must advertise a large context
window (capability-evidence gates commit_reviewed pre-dispatch), and an
observed publish-before-persist race (/api/tasks says completed before the
durable result lands) is a candidate upstream finding, not worked around
silently - wait_durable_result() names it.
All 15 D15 runtime modules are unsplit by design - upstream bytes stand as-is:
7 byte-identical to the reference, 6 pure upstream drift (nothing to do), and
2 (consciousness, reflection) carry v7 deltas of the D02 family that must NOT
be replayed verbatim - reflection's status-set delta would invert over
upstream's own handling (the re-prove trap is documented in
docs/v7next/LEDGER_CORRECTIONS.md together with a MIGRATION row already
superseded by upstream).
The test side transplants the oracle's D15 split: the 2386-line evolution
integrity giant becomes six themed suites (lossless: 65==65 test functions,
zero upstream drift of the giant since merge-base) plus three siblings carried
verbatim; adaptations are exclusively reverse-mappings of OTHER domains' v7
spellings back to upstream signatures, each keyed to the original monolith.
98 passed in isolation (re-verified independently); every file <=617 lines;
HEAD held through six pytest runs.
- scripts/v7next_domains.toml: 368 modules -> 20 domains 1:1; derivation
hard-validated by reproducing every DOMAIN_MAP index count exactly; the 80
new upstream modules carry classification=proposed (5 disputed rows named);
D20 Presence proposed as a new domain (9 modules, ~3.3k lines) - owner call.
- scripts/v7next_domain_report.py + docs/v7next/DOMAIN_QUOTIENT_REPORT.md:
report-only quotient (never a gate before the owner rules on edges): 164
domain edges, one 20-domain SCC pre-transplant, 357/724 witnesses marked as
split-pending monolith glue that the transplant dissolves, 367 genuine
residue enumerated with witnesses; lazy/TYPE_CHECKING/dynamic classified
separately (one unresolved dynamic site named).
- ADOPTION_v7next.md + scripts/v7next_adoption.py: the parseable reverse-
direction manifest - 36 rows (F1=13/F2=6/F3=10/F5=7) over the exact 18
approved-delta families; D09 evidence refreshed at this tip (the range(3)
retry loop is alive); one delta superseded by upstream #384; four honestly
pending-decision for the F2 three-column matrices; validator green,
--release correctly red while rows stay unresolved.
Known collision recorded for F1: upstream's new ouroboros/delegate_terminal.py
vs the reference leaf tools/delegate_terminal.py.
Owner decision (batch: Q4=A): upstream bytes are the truth, and every leaf
transplant must PROVE it changed nothing but the declared parent-owned reads.
scripts/v7next_transplant.py (903 lines, stdlib-only): extracts symbol spans
byte-exactly, rewrites ONLY module-scope Load references of the declared set
into call-time handle reads (symtable scope walk in compiler order, asserted),
and proves every emission three ways - inverse-normalized AST equality with
upstream, lockstep token comparison outside rewritten references, and an
informational byte roundtrip. Fail-closed on import-time reads (decorators,
defaults, class bodies, module assigns), conditional/duplicate definitions,
wildcard imports, f-string reads, and any name that is neither local, leaf
imports, declared, nor builtin - the refusal names the exact dependencies,
which IS the per-leaf declared-set recalculation workflow.
Proven on real bodies at b9f7597f: queue_snapshot pair byte-identical to the
v7 reference leaf; drifted persist/restore converge after ONE mechanical
recalculation round (43 rewrites, proof green); loop_messages and
git_ops_remotes byte-identical including declared-and-moved names; --check
honestly exit-2s on upstream drift naming the first diverged token.
28 tests green under isolation; ruff F clean.
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · ubuntu-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · windows-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · codex · API key only, subscription NOT covered (push) Waiting to run
The round-five triad caught the reference-minted owner card presenting
itself as live foreground work: unlike attachReviewGroup, the mint
never marked the card as a review anchor, so a failed hydration left a
Working/typing card that kept chat activity awake. The reference path
now applies the same anchor-eligibility predicate (extracted as the
shared reviewAnchorEligible) before minting, and the card stays an
inert Reviews anchor until real task truth arrives. The repeated
String(x || '').trim() normalization collapses into one taskKey helper
(24 sites), which also keeps chat.js under its byte-ratchet baseline.
Incident class: a configured-session nanny's metered round dies
provider_outcome_unknown (dispatched request, no terminal provider fact —
never resent, per custody doctrine) while its one physical delegated leaf
is alive; terminalizing the nanny let the cause-blind terminal cleanup
cancel the healthy leaf.
Main fix (D1-min): the round gate latches a durable hold and the next
round top parks the task in the same $0 supervised_wait the nanny would
have chosen. Eligibility is narrow and fail-closed: exact-route configured
sessions, exactly one open run, no pending invocations, no open
containment fault, and a READ-ONLY engine poll proving a live
non-terminal state. A meaningful leaf wake resumes with a NEW round whose
transcript carries the wake receipt (bound to the unknown attempt id);
owner dialogue drained at the round top resumes the same way. One wake =
one dispatch: an unacknowledgeable wake fails closed to the no-resend
terminal with its receipt removed. Control wakes (Stop, deadline,
finalize_now — re-checked at the source), daemon refusals, round-limit
boundaries, and budget exits all close the hold into no-call terminals,
never a paid dial; in-process loop exits clear the latch (a worker crash
preserves it for recovery). Repeated unknown cycles re-latch behind a
bounded backoff floor kept under the idle-rail minimum.
Custody companions: every provider-death arm of the rail stamps
terminal_origin=host_salvage through one wrapper (deadline grace finals
and scheduled swarm handoffs keep their legacy shape; budget/round-limit
rails stay untouched); durable llm_api_error events bind the physical
attempt (capture resolved through the explicit cause chain) and the
bounded transport cause type, which also leads the unresolved-attempt
reason; and the periodic sweep, after settling runs, re-runs the
read-only terminal custody audit so a stale delegated_runs_unreconciled
disclosure heals instead of lying forever (retry-lineage projections read
the original row live). Doctrine docs updated to state the resend
boundary precisely.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
The final reviewers caught two real residuals. A plan review_reference
whose owner card does not exist yet had nowhere to land a hydration
error: the reference is consumed, no card is minted on failure, and the
promised error/Retry shell never appears — handleReviewReference now
mints the owner card before hydrating (the reference itself proves a
review exists). And a non-string value under a KNOWN finding key (e.g.
evidence carrying a nested object) was str()-flattened before
redaction, bypassing structural key-based secret masking the same way
the unknown-shape fallback used to; non-string values now go through
object-first redaction, pinned with a nested-secret regression. Plus
mechanical residue: text() on the wave reason, the EOF blank line that
failed git diff --check, and comment tightening that keeps chat.js
under its byte-ratchet baseline.
The scope reviewer caught the fallback for unknown finding shapes
serializing the object before redaction: string-level redaction masks
secret-looking VALUES but not values held under secret KEY NAMES, so a
{password: ...} row would ship verbatim. The object is now redacted
first (structural key masking), then serialized and bounded; regression
pinned with a secret-bearing unknown shape.
Round two of the review panel: the zero-group error shell rendered its
status inside the hidden groups container, so a collapsed section still
hid the failure, and the retry's own loading pass unmounted the shell
mid-flight. The status node now sits outside the collapsible container
(readable on a collapsed section; the keyed node survives loading↔error
in place) and the controller remembers a failure across the retry's
loading pass (hadHydrateError) so the shell stays mounted until the
retry settles. Finding lines stop treating item and summary as
alternatives — a row carrying both renders both — and the finding id is
emitted beside the severity label.
The triad+scope panel converged on three real gaps. A failed FIRST
hydration had nothing to hang its error on — renderReviewsSection
returned empty before the status node, so the section silently did not
exist; the shell now renders for the error state (quiet zero-group
loading passes stay invisible because every card expand hydrates).
Duplicate dispositions for one finding rendered first-wins, presenting
a decision the backend refuses as contradictory intent as if it were
operative; every disposition row now renders. The formatter dropped the
reason/summary/verdict fields the projection deliberately carries for
array-contract reviewers, and the durable observability pointer fired
only on omitted rows; finding lines now render the full projected key
set and the pointer is unconditional. The ABI row and api_types JSDoc
document the complete row shape.
Adversarial review found the flagship error state never displayed its
message: the keyed status node was patched in place across
loading→error, and the DOM patcher syncs text only through
childless-element innerHTML, so the stale 'Loading…' text node survived
next to a cloned Retry button. The status message now rides inside a
<span> (patched as a childless element), the error variant carries
role=alert, and Retry moves keyboard focus onto the refreshed status
node. A reconcile-level test pins the transition in both directions.
Absent is no longer conflated with failed: the hydrator consumes the
new fetchTaskDetailStrict (404 → null stays a quiet no-op; a failed
read rejects and renders the error state) while the lenient
fetchTaskDetail keeps its every-miss→null contract for the cancel/stop
reconcile flows.
The projected finding row moves to review_execution_projection.py (the
read-side presentation leaf) with reviewer-lexicon keys extended by
verdict/summary/reason, so an array-contract triad finding no longer
loses its substantive reason text; review_substrate returns under the
1600-line gate and chat.js ends 50 bytes below its byte-ratchet
baseline. The new checkpoint tests live in review_findings_detail
.test.js, the skill_review typedef documents history_omitted, and the
size-ratchet manifest records the tightened chat.js byte debt.
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · claude · API key only, subscription NOT covered (push) Has been cancelled
Claudexor platform gate (API keys — subscription auth NOT covered) / live · ubuntu-latest · claude · API key only, subscription NOT covered (push) Has been cancelled
Claudexor platform gate (API keys — subscription auth NOT covered) / live · windows-latest · claude · API key only, subscription NOT covered (push) Has been cancelled
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · codex · API key only, subscription NOT covered (push) Has been cancelled
CI / build (zip, windows-latest, windows-x64, syft_1.50.0_windows_amd64.zip, syft.exe, 815ee6973ec5dff6a671d7f41b0e78835a8c45b91d5a39f4743ea1cee833d3be) (push) Has been cancelled
CI / build (dmg, macos-latest, macos-arm64, syft_1.50.0_darwin_arm64.tar.gz, syft, e32fdb9d47823fa633748a1efca2528fd77c37469ea93c9e40ab835da44e4cce) (push) Has been cancelled
CI / build (tar.gz, ubuntu-latest, linux-x86_64, syft_1.50.0_linux_amd64.tar.gz, syft, bf7b29ff57f06da30918266a0e1c2885a8f99784798d1bdb1628886aa015d788) (push) Has been cancelled
Keep the full browser lifecycle race assertions on the active per-family host after account-list rebuilds, and ship the green fix-forward as 6.113.4.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
The checkpoint smoke asserted only formatter line prefixes, so it stayed
green while the cards carried no review content. The fixtures now
include acceptance actor findings (with an omitted count and a
response_ref) and a full plan wave with a blocking finding plus its
agent disposition; the smoke asserts the finding bodies, the omitted
disclosure and the observability pointer are visible in the expanded
chat card and in Dashboard Logs, and walks both review groups on the
two-group card.
The task-card Reviews checkpoint answered only "a review happened,
verdict X": plan attempts never read wave.findings/dispositions and the
shared projection formatter never read actor findings, so the owner had
to dig through task results to learn what reviewers actually said.
Plan attempt detail now renders each finding as text lines (class,
summary, broken spec id, locator, recommendation, slot/model) with the
agent's disposition mapped under its finding, unmatched dispositions in
a General block, one line per reviewer that produced no parseable
verdict, the counts row, and the capture honesty stamps (per-slot page
cap, text/spec truncation). Compacted waves name their recorded counts
and the immutable wave-artifact remainder (sha256/bytes, no filesystem
path). The shared formatReviewProjection renders the new bounded actor
findings rows with their exact omitted count and, where rows are absent
or bounded, the observability call id as the durable pointer — fixing
the checkpoint, Dashboard Logs, and the task_done log body through the
one seam.
Review hydration stops swallowing failures: the hydrator reports typed
first-load/error/recovery states (background refreshes stay silent),
the section renders a keyed loading/error node with Retry, and Retry
invalidates the applied-revision receipt so it cannot dedup itself away.
Disclosure gaps of the same class: the skill card trusts the runner's
disclosed ten-row history window instead of re-slicing it and labels
"(N of M)" when rows were omitted; the marketplace compact hint says
"(+N more)" instead of presenting the first finding as the only one.
Task-acceptance actor projections kept only a findings COUNT while the
structured rows (severity/item/evidence/recommendation) were dropped at
projection time, so no owner surface could show what reviewers actually
found. The actor projection now ships a disclosed bounded findings list
(at most MAX_PROJECTED_ACTOR_FINDINGS=8 rows, strings redacted and
bounded at 2000 chars with explicit omission markers) plus the exact
findings_omitted count, emitted only when a parsed response exists —
absence stays a transport/parse hole, never "zero findings". The skill
review UI projection likewise discloses its ten-row history window with
an exact group-scoped history_omitted count instead of a silent [-10:].
Contract mirrors: api_types.js JSDoc and the ARCHITECTURE frozen-ABI row
for review_projection document the additive keys.
Bundle the exact Claudexor 3.9.2 runtime so Claude subscription quota reads refresh expired access tokens automatically through Claude Code before polling usage.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Sol delta-review finding: the merged picker's advisory handler cleared
state.advisory.effort on ANY reference pick, so switching a saved
{subagent_id, effort} row to another roster row silently dropped the
owner's explicit override on the next save. The transition is now a pure
node-tested helper: crossing from an inline branch still sheds the inline
effort ('' = the roster row's own effort decides), a reference-to-
reference switch keeps the override — matching the triad/scope handlers.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Owner decision 1=B: each triad/scope/advisory row now picks its reviewer
from ONE select — the Available-subagents roster rows lead as references
(facts-first labels), then the inline channels (API model, one entry per
harness). The two-level source+route pair, its Direct-model label that
falsely implied API-only delivery, and the separate roster select are
gone; the stored forms, the survive-the-save rules for missing references
and undiscovered harnesses, and the advisory effort defaults are
unchanged. Picker captions also strip directional marks (LRM/RLM/ALM) and
trim by code points so a cut can never split a surrogate pair (fable
review finding 7).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Finding 1: a preset receipt written before the name retirement hashed a
serialization that carried the retired key, so its fingerprint can never
match again on an untouched install and legacy-migrated provenance (and
with it the model_lane/executor compatibility window) lapsed on mere
upgrade. _materialized_source now falls back to the receipt's own embedded
rows re-read through the CURRENT parser — provenance survives the
migration, an owner edit still downgrades.
Finding 2: the family-mounted login card lives inside the groups container,
whose innerHTML rebuild destroyed a focused paste-code/name input before
preserveCardFocus could see it; the capture now wraps the whole rebuild.
Findings 4/5: the roster-cap guard reads MAX_CONFIGURED_SUBAGENTS instead
of a literal, the overflow inline fallback gains a real test, and the
extended roster re-validates through make_configured_subagents so a
minting bug reports as a roster-precise error.
Finding 6: drop the stale name mention from the child-snapshot doc row.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
The rate-limit terminal of the safety supervisor now BLOCKS the one
unchecked call with the typed non-verdict `SAFETY_UNAVAILABLE` outcome
instead of executing it unchecked: `full` keeps its owner contract (an
unchecked guarded call never runs; the documented fail-open cases stay
exactly the three no-backend ones in prompts/SYSTEM.md), while the false
SAFETY_VIOLATION accusation this PR removes stays removed - the
`_UNAVAILABLE` first line classifies downstream as a plain tool error,
and the message carries the retry contract so the agent retries the same
call instead of "fixing" a command that was never judged.
Around that terminal:
- the local-FALLBACK lane keeps its documented fail-open contract
(SYSTEM.md case (c)): a 429 there is not stricter than the
RuntimeError beside it (`_rate_limited_outcome` splits the lanes);
- a real double-429 arms a short process-local storm latch: further
checks in the window answer immediately with the same typed outcome
and zero provider calls - the safety lane is the highest-frequency
LIGHT consumer and must not amplify the storm it reports; the latched
short-circuit itself never re-arms the window, or the advised retry
would keep the latch alive forever;
- the task deadline bounds the RETRY, not only the backoff sleep: an
expired deadline permits no second paid attempt;
- `_body_is_quota_refusal` reads only structured code/type fields and
skips numeric markers: `402`/`billing` occur inside request ids and
hostnames of free-form messages, and a false quota match re-labelled
a plain throttle as a verdict - the exact bug this lane removes;
- the durable audit append is checked (`append_jsonl` answers whether
the row landed) and falls back to the event queue on a failed write;
- the shared rate-limit vocabulary is imported under public aliases
(`is_rate_limit_text`, `NON_RETRYABLE_PROVIDER_MARKERS`) instead of
private names;
- the OB-02 test section moved whole to tests/test_safety_rate_limit.py
at test_safety_policy.py's 1600-line module gate, and the five
rationale docstrings the previous head had compressed to their first
lines are restored verbatim.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>