- The alarm no longer re-arms on the owner's own direct turn ending (B13) nor on a
project digest of a tree consciousness started; notify keeps the boot floor after
a restart (max(last wake, boot) + floor).
- A wake carries what is left of its allowance as the tree's GRACEFUL ceiling
(`root_cost_ceiling_usd`, now honored for the root itself by the in-task stop)
instead of narrowing the ledger fence: one Main attempt reserves ~$8 up front, and
the narrowed fence refused every wake of a nearly spent day before its first call,
posting a budget error into Main at each heartbeat (stand: 15 such wakes). A
remainder at or below the planning margin is skipped as allowance_exhausted.
- A wake the lane could not admit backs off like a failed wake; a closed budget door
is retried quietly at the interval, a transient door at the floor.
- The wake message lists unanswered cards of any task, the previous wake's own
included (an expired_terminal card still takes a late answer).
- Only an owner's stop is sticky against toggle_evolution: a stop the agent placed
remembers its source (`evolution_stop_source`) and stays undoable by the agent.
- The status snapshot carries `unknown_unmetered`/`integrity_degraded`; the allowance
line says "at least $X" and "ledger integrity degraded" when they apply.
- Docs: a Presence cycle a wake starts is outside the allowance; benchmark profiles
drop the retired OUROBOROS_BG_MAX_ROUNDS key.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Add caller-owned Claudexor model calls, unified subscription/API setup, explicit role accounts and live quota/auth continuation. Preserve prompt/tool ownership, physical-attempt custody and manual assignments. Keep the existing task process through waits and post-work.
The reviewed source checkpoint retains the current dependency pin. Published signed Claudexor bytes, cross-platform CI and the agreed merge ordering remain delivery prerequisites. No Ouroboros version, tag or release is created.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Second adversarial round on the absorb-wait rewrite.
Reservation: after the previous commit only SM1 can promote (SW1/SK1 pin
OUROBOROS_POST_TASK_EVOLUTION=false), yet RunBudget.reservation still
added the evolution root for every attempt under --self-mod, so SW1/SK1
reserved a ceiling they could never spend and run-cap admission refused
or serialized paid attempts on it (the SK1-only mini run at cap 130 was
refused by exactly this over-reservation: 50 x (2 + 1) = 150). The rule is
now per_task x (root_tasks + int(self_mod and absorbs)); admit,
budget_preflight, dispatch_order and run_lane pass the scenario's
expects_absorb. Owner configuration (cap 300, per-task 50, 3 attempts,
self-mod): SM1 100, SK1 100, SW1 50 — realistic spends admit 9/9 for
$159, pessimistic 8/9 for $219 (SK1_a3 refused), every scenario keeping
two; the CI e2e-live arithmetic comment and summary header are rewritten
(full set $225, not $315) inside the D-12 job, and the CI-lane pins
re-derive the numbers with the scenario flag.
Idle reasons: absorb_idle_reason types relative to the wait's start
(history length snapshot) so a resumed campaign's older cycles never
speak for this boundary; a paused/stopped/completed status wins;
no_promotion means an every_n post-task tick was recorded (llm cadences
write none: no_decision). Tests cover the resumed-campaign boundary and
the reason table with the cycle written during the wait.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Adversarial review of the corrected --self-mod seeding (2026-09-06) traced
the post-task path end to end and found that IsolatedServer.wait_for_absorb
ended the wait as no_promotion on ONE idle /api/state sample once the grace
had passed: no request file, pending_count 0, running_count 0. Those are
exactly the readings of the absorb path itself. A cycle that committed keeps
its campaign transaction as waiting_for_restart while the supervisor
restarts synchronously (RUNNING already popped, the counter unchanged), and
the re-exec'd server answers /api/state with zero counts before its
supervisor is up; absorbed_cycles_done moves only when the worker boot
verifies the restart. The pre-seeded benchmark campaign used to mask this
(its t=0 cycle kept running_count above zero and the kept request file
blocked the exit); with the owner-id-only seeding every SM1 lane would have
failed a healthy absorb as no_promotion.
wait_for_absorb now needs PROOF that no cycle is pending, held on six
consecutive polls after the grace: the queue idle AND supervisor_ready AND
no post_task_evolution_request.json AND no campaign active_transaction. The
reason is typed from the durable campaign state (campaign_summary /
absorb_idle_reason): no_promotion (no campaign although the post-task
decision ran), no_decision, cycle_no_op, cycle_not_absorbed,
campaign_<status>, cycle_not_enqueued; the summary travels in the wait dict
the stand records. Benchmark callers (evolve_smoke, the CLB adapter) keep
their early exit on a genuinely idle campaign.
Two more findings from the same review: scenarios that commit nothing
(SW1, SK1) now pin OUROBOROS_POST_TASK_EVOLUTION=false in their lane
settings — a one-shot cycle promoted from their own roots could commit and
re-exec the server in the middle of the lifecycle under test; and the SK1
gate also requires the review call's own executable_review (a pending
duplicate job never passes) while _skill_entry reads the listing through
the non-raising helper so a failed listing is a failed check with facts, not
an infra_error. Tests: a new module pins the wait on a fake clock
(waiting_for_restart and a booting server do not end it, one idle sample
does not, a pending request blocks it, the absorb confirms, every idle
reason), the SK1 module pins the product gate rule and the per-scenario
promotion pin, the runner module's settings pin flips for SW1/SK1.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
The commit gate's hermetic pytest pass runs inside each lane server and
resolves `-n auto` to os.cpu_count() (128 on this host) with no ceiling; the
2026-09-04 paid run started >= 104 xdist workers per self-mod lane with three
lanes overlapping. The runtime already has the lever
(OUROBOROS_PREFLIGHT_TEST_WORKERS, floor 2, read by
preflight_runner._preflight_worker_count and scrubbed from the candidate
suite), but IsolatedServer's settings-authoritative sweep dropped it before
the lane server started, so the stand could not use it.
Devtools-only wiring (evidence action A):
- server_runner: keep OUROBOROS_PREFLIGHT_TEST_WORKERS through the
authoritative sweep as an operational host-load lever (comment reworded;
it is not a model, credential or settings key).
- run_live_lanes: derive max(2, 16 // lanes) at argument time (shared-host
rule: at most 16 pytest workers across the stand), set it in the launcher
process before the first lane starts so an ambient shell value never wins,
and record it as extra.preflight_test_workers in run_manifest.json and as
preflight_test_workers in every lane row.
- docs/DEVELOPMENT.md: one sentence in the live E2E stand section.
Pins: the lane sees the computed value (ambient 128 overridden), the runtime's
_preflight_worker_count reads it, the manifest and lane row record it, the
floor and key names match preflight_runner's constants, and IsolatedServer
forwards the key while still stripping an ambient OUROBOROS_MODEL. No
runtime code changed.
Size ratchet: tests/test_e2e_live_runner.py enters the 1001-1500 band with this
commit (996 -> 1042 lines); the band rationale is recorded in the manifest.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 1824f8910509c65b9fc94041c7a76f98a9aa617d)
run_live_lanes.py admits through the benchmark family's seams (dirty seed refused with the
refusal persisted; the key by NAME from the environment, never a pool file; the credit
preflight takes min(key limit remaining, account credits) via the new
manifests.openrouter_account_credits, the second bound only), writes the effective settings
from the tree's own defaults with TOTAL_BUDGET/OUROBOROS_PER_TASK_COST_USD as settings keys,
names the model in the run manifest from the applied file, and fans out --lanes (default 4,
max 6) isolated servers 2-3 s apart. scenarios.py is the table: SM1 lands a web/style.css
token through commit_reviewed under advanced+blocking (S2 set + computed style read from the
committed CSS after a restart), SW1 arms Swarm in the browser (force_plan, roster, >=2
children with causal lineage, fanout receipt, cost rollup, /proc no-orphans), SK1 has the
model author a skill then reviews, grants, enables, dispatches and deletes it; acceptance is
callable over durable artifacts only. ui_probe.py resolves the suite's PlaywrightUIClient when
it carries the surface, else headless Chromium, else a typed ui_unavailable. --stub rehearses
every scenario for $0 on the loopback stub of tests/system_e2e/harness.py (stub_lane.py routes
the swarm wire by role). Per-lane result.json carries checks, settings sha256 plus a
secret-free config digest, seed describe, pre/post HEAD and the diff digest, grants by
fingerprint and the runtime terminal disclosure; a watcher prints lane states, free disk on /
and /mnt/data and the key headroom. Tests pin the launcher gate by source, the table shape,
lane/stagger bounds, the TMPDIR guard, the dirty-seed and credential refusals, the two-plane
credit floor, manifest-model-equals-applied-file, secret-free artifacts, and a gated stub
rehearsal of SM1 on a real server. DEVELOPMENT and ARCHITECTURE describe the stand.
(cherry picked from commit ab36206a2a7ccbad01bf6b25a69181a6a69d9aa6)
Second absorption of the frozen upstream line (23ab428f..db6d7cf8: 89
commits, 47 files) on top of the rc.10 hotfix tip, by the F2 rules (S1
upstream body in the owning leaf, S2 hand-merge, S3 only with proof;
retired 7.0 surfaces never return):
- PR #609 net-resilience: interactive transport-wait episodes bounded by
the task idle timeout with the typed task_incident/toast_once pair, a
bounded paid repeat after a typed post-dispatch transport death with a
round-keyed record that fences every other send, the shutdown-aware
supervisor crash counter and bounded lifespan join, Darwin keepalive
tuning. loop_transport/transport_custody/net_transport/loop_llm_call
land verbatim (same shapes on both sides); the loop.py deltas are
relocated into the v7 leaves (loop_round_limits, loop_model_call,
loop_delivery, loop_forced_finalization, loop_nudges, loop_messages,
loop_budget) with bodies AST-equal to upstream modulo the call-time
handles; _emit_overflow_retry_skipped stays a public helper (the v7
facade contract) and upstream's nested _skipped delegates to it.
- PR #614 update letter: ouroboros/update_letter.py and its web module
land verbatim; the new OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC key and
get_update_letter_timeout_sec live in their v7 owners
(settings_defaults.py, runtime_limits.py, re-exported by config.py);
_supervisor_stop lives in server_process.py beside the restart events;
docs/PERSISTENCE.md gains the state/update_letter.json row and the
inventory pin moves to 286.
- Tests: the relocated run_llm_loop tests take the emit_progress
incident keyword (every one-argument progress fake in tests/ swept, a
gap upstream itself left in test_tree_cost_ceiling); the official-update
runtime-section test lands in tests/test_context.py; _MOVED_OWNERS
registers the relocated getter.
- Docs: ARCHITECTURE/DEVELOPMENT hunks land on the upstream text; the two
legacy timeout rows upstream's context still carries stay retired (7.0).
- Size ratchet regenerated; the band rationale for tests/test_update_letter.py
is carried verbatim from upstream.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
devtools/benchmarks/terminal_bench/test_run_tb_methodology.py (upstream body,
absorbed by F2) looked the shipped triad up through
SETTINGS_DEFAULTS["OUROBOROS_REVIEW_MODELS"]; ABI 7.0 retired that settings
key, so the rc.9 tag CI job benchmark-methodology failed with KeyError in two
tests. The launcher itself derives the same list from
ouroboros.settings_defaults.OPENROUTER_REVIEW_DEFAULTS["triad"] (the SSOT the
retired key was joined from), so the tests now read that list. Test-only;
no runtime change.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Second parent is the frozen upstream `ouroboros` head (23ab428f, 407 commits
since the merge base a76961de); first parent is v7.0.0-rc.8 (18b9832e).
Every upstream change lands in v7's owning leaf: S1 transplants keep upstream's
bodies (comments verbatim) under the call-time handle idiom, S2 hand-merges keep
both intents, S3 keeps v7 only with proof (retired 7.0 ABI surfaces, superseded
mechanisms). Per-symbol relocation ledger: docs/archive/v7next/LEDGER_CORRECTIONS.md
(F2 absorption section). Provisional decisions awaiting owner ratification:
D-18 (two-destination symbols), D-19 (acceptance rows follow upstream R2),
D-20 (acceptance_dialogue stays deleted), D-21 (tools/registry.py: facade
import block only).
Docs: upstream ARCHITECTURE/DEVELOPMENT as the base with compact v7 deltas;
bookkeeping moved to docs/archive/v7next. Size-ratchet manifest, domain
manifest and generated inventories regenerated; new leaves: tools/write_shape
walker, gateway/cost_breakdown, tools/core_secret_paths; provider_catalogs.py
and acceptance_dialogue.py removed (v7 owners).
The continual-learning runbook and the SWE-bench Pro methodology said
the transport-death repeats sit on top of the transient burst. The
primary dispatch grants a repeat only while the outer attempt budget
has room, so both now say: one outer attempt budget per call, within
which up to three rows can be typed transport-death failures (the
first death plus at most two repeats).
Five conflicts, resolved by reading both sides in full:
- loop.py: the trace-touched skill-name scan moved into skill_readiness.py
upstream; its import replaces this branch's inline copy, whose removeprefix
form was a byte fold of the same behaviour.
- loop.py: the no-resend terminal keeps upstream's block (its short comment and
its live_trace re-read); this branch's own two-stamp fold is re-applied on top,
and the record-fenced source stays.
- loop_llm_call.py: _send_main_candidate binds whenever a physical context OR a
candidate predicate is present, upstream's semantics, expressed through the
binding this file already uses. The context manager is only constructed there,
never entered, so the effect order is upstream's.
- loop_llm_call.py: call_llm_with_retry keeps both new parameters,
transport_death_retries and initial_messages, on the two lines the file's
line cap already pays for.
- ARCHITECTURE.md: upstream's producer-word sentence, with this branch's
no-resend source clause re-applied, so the paragraph states both.
loop.py carries shrink-only byte debt and upstream absorbed both places where
this branch had paid for its own additions, so the payment is made again inside
the functions this branch owns: the stamp fold above, the fifth site of the fit
key tuple now calls _fit_key, and _emit_overflow_retry_skipped is folded into
_skipped, its only remaining caller since this branch collapsed the other three.
No comment, docstring, diagnostic or test was shortened. 271928 -> 271855 bytes.
The manifest is regenerated on the merged tree; it is upstream's manifest with
loop.py's exact byte count. Upstream's three new band rationales are carried
across verbatim because the generator reads the pre-merge committed manifest.
Target drift since the sprint base (PR #557-#591: the agentic-review synthesis
moved the acceptance machinery whole into acceptance_dialogue.py, three-delivery
rows, Claudexor 3.9.7, ibl fixes). Resolutions: the acceptance-packet changes
(children debt, dialogue history and the packet budget passed INTO the bounded
builder; per-slot input caps on the panel request; packet sizing from the same
triad delivery rows the panel dispatches) are carried into the relocated module;
a partial tool-result projection withholds packet rows only for a genuinely
unavailable source (the legacy truthy sentinel still refuses) and never
retrieving rows; a report-shaped native episode keeps its draft on a budget
refusal after the failed send is observed; ARCHITECTURE/DEVELOPMENT merge both
sides' rows and paragraphs; the upstream acceptance-delivery test now asserts
the documented `not_dispatched` refusal shape (a refusal is a transport state,
never a verdict).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
The PR target moved from 85c1e386 to f437408b (PRs #565, #590: the task
acceptance surface moved into acceptance_dialogue.py, deep self-review runs
on the configured reviewer row, the finalization nudges fold into one
_inject helper, and the project last-task-result reader grew in server.py).
One file conflicted: ouroboros/size_ratchet_manifest.py. Upstream shrank
ouroboros/loop.py to 272,905 bytes while this branch held 284,335, so the
BYTE_DEBT row was the only unmerged hunk. Resolved by regenerating the
manifest on the merged tree: BYTE_DEBT["ouroboros/loop.py"] = 272,805, which
is the merged file's exact size (upstream's 272,905 minus this branch's own
100-byte shrink) and below both parents' debts. Upstream's six new 1001-1500
band rationales and its shrunken tests/test_devtools_benchmarks.py debt ride
along unchanged.
Every other shared file auto-merged on disjoint regions: the transport-death
and wait-episode seams in loop.py, the incident= keyword on
OuroborosAgent._emit_progress, the supervisor stop event and lifespan
teardown in server.py, and the transport paragraphs in ARCHITECTURE and
DEVELOPMENT are all preserved beside the upstream hunks.
"Up to three llm_api_error rows for one logical call" read as an absolute
bound; the transient burst shares the same attempt loop, so the numeral
now names the transport-death rail alone, on top of that burst.
A dispatched request whose socket died with a typed transport death (httpx
ReadError/WriteError/RemoteProtocolError or the requests ProtocolError/
RemoteDisconnected shape, via the explicit __cause__ chain; never a timeout,
status/body error, pre-dispatch failure, local provider or loopback route)
is provider_outcome_unknown and stays billed at its upper bound. The PRIMARY
main-loop round dispatch alone may now send it again at most twice per round,
each repeat a NEW execute_physical_attempt with its own ledger row, with a
4 s then 8 s deadline-aware backoff and a counter keyed by the round so the
wait episode's free redial cannot re-arm it. The retry decision is made
before the durable llm_api_error row is written, so retry_same_request is
truthful on every attempt and llm_non_retryable_same_request marks only the
exhaustion. Forced-final, fallback candidates, review, safety, probes, web
search, consolidation and Background Consciousness keep the default 0; the
global classifier is unchanged; BudgetExceeded propagates untouched.
Ratchet paydown in the same files: loop_llm_call.py folds the four copies of
the error-event identity block into one, drops the redundant kwargs that
_handle_main_llm_call_exception/_stop_after_llm_error received beside the
context that already carries them, and folds five consecutive pops into a
loop; loop.py hoists two function-local imports of an already-imported
module and reuses the direct_chat value it had already computed.
Integrates the moved target (193 commits over the sprint base b9bcc2da,
release 6.114.0 — the DeepSeek landing PR #563 and the docs-consolidation
PR #556 included) into the synthesis branch without rewriting history.
Four files conflicted; each was resolved by substance so BOTH sides' facts
and behaviours survive:
- ouroboros/size_ratchet_manifest.py (generated): UNION of both sides'
BAND_PATHS rationales (the target's ouroboros/gateway/extensions.py,
tests/test_provider_contract_ci.py, tests/test_ui_smoke_project_continuity.py
and web/modules/settings_ui.js beside our acceptance_dialogue.py,
deep_self_review.py and test-suite rows), then regenerated for the merged
tree: ouroboros/gateway/history.py left the band on the target line,
BYTE_DEBT is the live merged size everywhere (loop.py 272905,
tests/test_devtools_benchmarks.py 328068 — below both parents — and
web/modules/chat.js 206949). `scripts/regenerate_size_ratchet.py --check`
is green; no new module, function or byte debt.
- docs/ARCHITECTURE.md and docs/DEVELOPMENT.md: the target's consolidated
structure is the frame; every agentic-review fact of this branch is placed
in the target's section or row — the three deliveries of task acceptance
and deep self-review, the reviewer-row schema (`deep_review` singleton,
roster references, `profile_id`), pacing simplified to the one admission
floor (R52/R55; the EWMA sentences are gone), poll purity, the native
read receipts and BIBLE.md coverage, the retrieving work order and the
CI methodology job. The module-tree rows keep the target's condensed form
extended with our contracts; the full contracts live once in §6 (Task
lifecycle, Review delivery, Deep self-review). Stale target sentences
that our side retired (the `api_chat` acceptance pin, the round cap, the
legacy/default API panel residual) are replaced, never duplicated.
- ouroboros/review_native_episode.py: our side already measures the send
bound as the wire size of the serialized message list, which carries the
WHOLE assistant dict — `reasoning_content` included — so the target's
fix (count replayed reasoning in the fail-closed bound, 295c9062) is
subsumed; the target's regression test passes unchanged. The comment
above the append records the invariant.
Auto-merged both-changed files were checked for silent overlap:
tests/conftest.py gained the same autouse os.environ snapshot/restore
fixture on both sides — the target's tested `_os_environ_isolation` is
kept and our redundant `_restore_process_environment_between_tests` is
dropped (its rationale folded into the surviving docstring); our
gateway-settings binding restore fixture stays. .github/workflows/ci.yml,
ouroboros/config.py, ouroboros/tools/control.py, web/modules/settings_ui.js
and web/modules/reviewer_slots.js carry both sides' changes exactly once
(our deep self-review block sits in the target's new `.reviewer-slots-group`
container like its advisory sibling).
The session rows a Terminal-Bench container cannot run were disclosed on the
run manifest and as a metadata comment — artifacts an operator reads after
the run, when the money is spent and the acceptance seat has already
degraded on every task. The owner decided (R40, 2026-09-02) that admission
says it once, loudly: when the configured triad carries agent-session rows,
the launcher prints one stderr warning naming each row (`codex=…`,
`cursor=…`), that a container has no harness CLI/daemon or credentials, that
the rows are not declared as used models and their acceptance seat degrades
typed, and how to configure a submittable run. The run continues
(`command_generated`); the manifest field and the metadata comment stay. The
methodology test asserts the warning text appears exactly once beside the
manifest and metadata facts, and an api-only panel admits silently with an
empty disclosure list.
The environment fixture's docstring no longer claims to mirror conftest's
`_scrub_inherited_subagent_selection`: the key sets differ on purpose —
conftest drops the subagent roster, account pin and structured panel; this
suite reads both panel forms and drops the structured panel and the legacy
comma-list keys — and it now says so.
The TB methodology suite lives outside tests/ and so outside the conftest
scrub that keeps every other test independent of the operator's shell. Its
environment fixture restored what a test wrote but still let an inherited
`OUROBOROS_REVIEWER_SLOTS` or legacy comma-list key reach a panel-reading
test (`test_metadata_omits_web_search_when_web_disabled` read the shell's
panel). The fixture now also drops those keys before each test, mirroring
`_scrub_inherited_subagent_selection`; the gate run for this commit exported
a poisoned shell panel and comma lists on purpose and stayed green.
METHODOLOGY states precisely what the manifest carries per delivery: on a
fixed-model run `harness.fixed_model_actor.reviewer_slots` holds `slot_id`,
`route{kind, target_id}` and `effort` per row (a fixed-model panel is always
direct api rows; no subagent binding exists there); a plain `--model` run
records only the session rows, and its api-vs-native split survives only
when the panel is persisted in the forwarded host settings — the container
adapter resolves the environment first, so a panel supplied only through the
operator's environment leaves no durable per-delivery record.
Two small hardening points on the retrieving work order. The slot-label
trim compared the renderer's raw tail against `Slot: <id>`; a future
trailing newline in the renderer would have left the label in place and the
executor — which labels the slot itself — would have emitted two `Slot:`
lines. The tail is stripped before the comparison, and the work-order
contract test pins that a work order carries no `Slot:` line of its own.
The native packet projection normalized a malformed `omissions_manifest`
only when it also had something to omit; a packet with nothing to omit
travelled with the malformed value as-is. The manifest is normalized
whenever the source value is present and not a list, and an absent
manifest is still never invented.
Terminal-Bench METHODOLOGY states what `metadata.yaml` cannot say: an api
packet row and a configured-subagent native inspection row are both
`commit_review_triad` by model id and dedupe onto the measured model under
the container's one-model roster. The per-delivery record is the run
manifest — `harness.fixed_model_actor.reviewer_slots` on a fixed-model run,
`extra.triad_rows_not_executable_in_container` for the session rows — and
for a plain `--model` run the api-vs-native split is read from the forwarded
host settings the container adapter used.
`run_tb.apply_all_model` — and `main --all-model` through it — writes the
fixed-model contract into `os.environ` directly; that is the launcher's real
behaviour and stays. But a test exercising it handed the written
`OUROBOROS_REVIEWER_SLOTS` and forwarded slot keys to every later test of
the same xdist worker (`monkeypatch.delenv(raising=False)` records nothing
for a key that did not exist, so nothing removed it afterwards), which is
the `benchmark-scope-1` contamination that made scope-review identity tests
and the legacy metadata test fail depending on worker order — on the base
commit as much as here. The class is closed where the other between-tests
hygiene lives, `tests/conftest.py`: every test runs on a snapshot of the
process environment restored afterwards whatever it wrote, and the inherited
reviewer panel is scrubbed with the actor list and account pin, so a test
that pins the legacy comma-list branch (the metadata test in
test_devtools_benchmarks.py, a shrink-only byte-debt module that therefore
stays untouched) never reads the operator's shell. The TB methodology suite
lives outside tests/ and carries the same snapshot fixture itself.
The container-provenance disclosure is now pinned end to end through
`main` (command generation, no harbor): `run_manifest.json` carries
`extra.triad_rows_not_executable_in_container` with the session rows'
targets verbatim and in row order — including a target with its own `/`
(`cursor=openai/gpt-5`) — and `metadata.yaml` carries the same list as a
comment while declaring no session row as a model. The gate run for this
commit exported a poisoned shell `OUROBOROS_REVIEWER_SLOTS` on purpose and
stayed green.
The previous commit declared every configured triad row in metadata.yaml,
including agent-session rows under a `commit_review_triad_agent_session`
role. A Terminal-Bench task container cannot run such a row: the image has
no harness CLI or daemon, the forwarded-env allowlist carries no harness
credentials, and the container secret policy forbids them — so declaring the
row as a model the run used reinstated the exact "declared but never run"
class the base docstring forbade. The stronger provenance clause wins:
metadata names only what the container executes.
`_effective_helper_models` now declares api packet rows and
configured-subagent native inspection rows only (decided by the typed
`row.is_session`, not a role-string substring), and restores the base
docstring's "never a declared-but-never-run model" clause. Session rows are
carried by the new typed disclosure `triad_rows_not_executable_in_container`
(their `harness[=model]` targets in row order) on `run_manifest.json`, and as
a comment line in `metadata.yaml` — a comment, not a key, because the
leaderboard schema owns that file's keys. Both read the panel through one
`_container_triad` seam (operator env, else host settings, parsed under the
container's one-model roster). The `"agent_session" in role` provider branch
goes with the class, which also retires the `cursor=openai/gpt-5` split
hazard: no session target is ever rendered as a model row. The methodology
test asserts the disclosure, not a declaration, for env, all-session and
settings-file panels; METHODOLOGY states the rule and that an all-session
triad's acceptance seat degrades typed inside the container.
With the triad rows reaching the acceptance panel as configured, a retrieving
row (a configured-subagent native inspection episode, or an agent session)
still had nothing to run: the acceptance request carried no `session_task`,
so both retrieving executors refused typed, and the gates around the panel
assumed every row was a packet row — the wave budget gate priced a
subscription session as API money, the partial-projection refusal turned a
row that reads the exact source away, and the packet rows' format-repair
resend would have bought a second episode.
Owner decisions R1/R4/R5/R15/R23 (2026-09-01) settle the work order.
`acceptance_dialogue.acceptance_retrieving_work_order` writes one per
retrieving row onto the new `ReviewRequest.slot_session_tasks` (per-slot,
falling back to the shared `session_task`, consumed by both retrieving
executors): the same task-stable contract the packet rows render, the same
output contract — `review_execution.review_output_contract` is now the ONE
governance text, rendered into the api pack's byte-stable segment (bytes
unchanged, pinned by the golden digest) and handed to retrieving rows as
`policy["output_contract"]`, so they never fall back to the generic object
form — absolute retrieval pointers over the task's ACTIVE workspace
(`review_repo_dirs_for`'s subject root, never the governance repo), and the
packet in the form the delivery can use. A session row gets the FULL packet
(its run is unobserved by the host, so the packet is its only attested view)
plus the disclosure that access outside the workspace is not guaranteed and a
refused read is absence of evidence, not of the artifact. A native row gets
the packet WITHOUT its freely degradable tail — the trajectory rows and
artifact previews the api ladder spends first, manifested as
`retrieving_delivery` omissions — plus the real data root
(`policy["native_data_root"]`, R5), because its episode reads those sources
itself. The FULL packet stays the `evidence_refs` authority on every
delivery; route-owned policy keys are filtered from the rendered Policy JSON
so the api pack states the contract once. The owner deadline rides the
request (R23).
The gates are route-aware: the wave budget gate prices API money only (a
packet row by its real message pair, a native row as one episode send of its
work order, a session row not at all); the partial-projection refusal spares
retrieving rows while the immutable-core overflow refuses every delivery; the
format-repair resend is packet-row only — a retrieving row's executor
canonicalizes its own answer. The wallet stamp is unchanged and fires on
every delivery through the same captured stamp (R11), pinned by an
api/native/mixed/session matrix that also proves a spent wallet refuses a new
paid identity before any send. The trap test lands first: a retrieving row
citing a real `verification_receipts[0]` resolves clean against the full
packet, with the fabricated sibling ref disclosed.
Docs: ARCHITECTURE (acceptance section, module rows, review delivery, the
governance matrix row for the three deliveries), DEVELOPMENT (wallet stamp
paragraph route-agnostic; acceptance checklist item), GAIA/OSWorld
METHODOLOGY acceptance-axis comparability notes.
Task acceptance was the one review surface still reading a PROJECTION of the
reviewer panel: `reviewer_slots()` with no route list rebuilt api-pinned rows
from the legacy comma key, which `project_reviewer_slots_into_env` filled with
the non-retrieving api rows only and floored to the shipped defaults when
none remained (the D15 pin). An owner whose triad ran on a subscription paid
API money for every substantive task's acceptance panel, per-row effort,
credential pins and stable slot ids never reached it, and a malformed
structured value left acceptance running a silently projected default panel
while every other surface refused typed.
Owner decisions R0/R2/R3 (2026-09-01) end that. `reviewer_slot_config.
triad_delivery_slots` is THE triad-row builder: plan review calls it with its
own slot properties, skill/commit review consume it as the aligned vectors of
`commit_triad_delivery`, and the acceptance panel calls it directly — so no
surface can read the panel through a projection again, and the per-row
construction that `structured_scope_review_slots` duplicated is one function.
A malformed configuration refuses acceptance with a typed DEGRADED panel
(`reviewer_slot_config_invalid`), the same posture as plan and skill review.
Child-task and `off`-mode acceptance stay advisory and run the packet rows
only, refusing typed when none remain. The API-pin apparatus is deleted:
`api_fallback_disclosure`, `reviewer_slot_api_fallback_warning`,
`_fallback_warning_text`, `_record_api_fallback_substitution` and the durable
`state/reviewer_slot_api_fallback.json` are gone; the legacy comma keys stay
a runtime projection of api MODEL IDS for legacy readers (external review
tooling, benchmark manifests), no longer filtering configured-subagent rows
out because a surface that no longer exists needed it.
Terminal-Bench metadata follows the run it describes: every triad row is
declared — api rows by model id, a session row by its `harness[=model]`
target under `commit_review_triad_agent_session` with the harness as its
provider — never a shipped default the container does not execute. UI copy
(Review lanes note, onboarding ladder and footnote) states the rule instead
of the retired pin. Retrieving rows now reach the acceptance panel as rows;
their route-owned work order lands in the next commit, so until then they
refuse typed (`session_task_missing`) rather than being silently dropped.
Tests: the D15 pins are rewritten; a new suite pins the shared builder,
per-row identity/effort/pin on the acceptance panel, the typed malformed
refusal, the legacy comma-key parity (GAIA/CLB/SWE-Pro class) and the
child/off packet-only rule; stub sites move from `reviewer_slots` to
`triad_delivery_slots`. Docs: ARCHITECTURE rows and settings table, TB
METHODOLOGY comparability note.
Two halves of the same money boundary.
The ledger fence priced every anthropic-family round as if the whole prompt
were a fresh cache write. On a large transcript that reserved several dollars
per round where the real send read almost all of it from cache, so a task with
a healthy cap could be refused at the fence and end with no answer at all --
the inverse of the intended order, where the graceful in-task stop lands
first. The reservation now prices the task's own last observed split (same
task, same model, inside that split's cache TTL horizon); anything missing,
stale, or from another route keeps today's full-write reservation, which is
the conservative direction. One guarded line in the scope merge copies the
bound task id onto the request, without which the split could never be found
on the live path. The split memory itself is a small process-local seam module
beside the ledger rows memo.
Second half: the in-task ceiling could still be crossed by a margin that is a
constant, so the task could reach the fence with less money than one wrap-up
call needs. After the ceiling comparison the loop now asks the pacing SSOT
whether one more wrap-up call would still be admitted, using the fence's own
reservation function -- same split, same arithmetic -- and soft-lands with a
typed `wrapup_reservation_unaffordable` stamp when it would not. Unknown
price, no root cap, no bound scope, an unknown prompt size and a disabled
ceiling all fail open and keep the axis silent; the planning margin constant
is untouched, and benchmark profiles that disable the hard stop disable this
rail with it.
Disclosed residuals: the cost axis is still checked only after tool-call
rounds, and after a cache expiry between rounds the reservation may
under-reserve by one write -- the fence at the full cap still binds.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
The launcher structural gate (test_every_migrated_launcher_passes_the_structural_gate)
refused the previous commit: reading the host settings file before
admit_benchmark_run() is a refusal with no durable manifest. The metadata is
now rendered FIRST inside the finalize seam — a malformed reviewer panel is a
typed refusal recorded on the durable manifest (stage leaderboard_metadata),
before any submission tree exists — and the two side-effect-ordering comments
describe the real seam (assert_outside_repo is pure; run_root materializes
with the manifest).
Round-2 finding: scope-review models were still declared although scope review
is a commit-time gate that never fires inside a Terminal-Bench task — the same
overstatement as the advisory. Scope rows and defaults are no longer declared
on either the structured or the legacy path; the older assertion that the scope
default "assists the run" pinned that overstatement and is rewritten. The
shrink-only test file's byte debt ratchets down with it.
Review findings (Ф1 wave) on the structured-panel derivation: metadata read
OUROBOROS_REVIEWER_SLOTS from the operator env only while the container
adapter forwards env, else the host settings file; a subagent-bound row was
resolved against the OPERATOR roster although the container gets a one-model
roster (the row does not resolve there, so the declared model never ran); and
inside a Terminal-Bench task nothing commits — the panel reaches the run
through task acceptance, which today executes the panel's non-retrieving api
rows (else the shipped defaults), so declaring every row named session models
the container never executed and omitted the defaults it did.
Read the panel in the adapter's order, parse it under the container roster
(reviewer_slot_config.roster_env_override, the save handler's existing
context-local overlay made reusable), declare exactly the acceptance
projection, and render metadata BEFORE any run directory exists so a malformed
panel is a typed refusal rather than a traceback in the manifest finalizer.
An enabled advisory is deliberately not declared: the advisory pre-review is
invoked only by the commit gate and skill review, never inside a task.
Fixed-model provenance refused hosted-session reviewer rows but accepted a
configured-subagent api row on the measured model, although that row RETRIEVES
the subject (native tool rounds) — a different delivery class from the packet
panel every published number was produced with. The validator now tests the
row's delivery class (retrieves) and names native_tool_rounds in its message.
Terminal-Bench metadata derived the assisting reviewer models from the legacy
comma keys while the container runs the structured reviewer panel the adapter
forwards; with both present, metadata could name models the runtime never
executed. Declare the structured panel when the operator env carries one
(session rows name their harness target with the route in the role), falling
back to the legacy keys exactly as before.
OUROBOROS_SOFT_TIMEOUT_SEC and OUROBOROS_HARD_TIMEOUT_SEC stopped terminating
anything when the activity model (idle window + subtree liveness + absolute
ceiling) replaced them. What survived was five surfaces discussing a value none
of them obeyed: SETTINGS_DEFAULTS offered it, the Settings UI accepted a number,
the save response apologised for it, queue.init compared the caller's value
against the constant it then wrote anyway and logged a deprecation row, and
/status printed "legacy_timeouts_ignored: soft=600s, hard=1800s" on every
request. A knob discussed everywhere and obeyed nowhere reads as a live tunable.
Retired through the existing idiom - RETIRED_SETTING_KEYS, stripped on load. No
successor knob (the activity model already governs), so nothing to seed. Gone
with them: both globals and init parameters in queue/workers, the
_emit_timeout_deprecation_once emitter and its latch, the gateway's
_RETIRED_NO_EFFECT_KEYS bucket (a retired key cannot reach an effect bucket at
all, so _effect_buckets no longer needs the warnings parameter), the status_text
parameters and legacy line, the server reads and ctx fields, the bench settings
carriers and the TB forwarded-env allowlist, and the two ARCHITECTURE rows.
rc_audit's `since` stopped being a one-key special case: RETIRED_IN_THIS_ABI
names the distinction, so an upgrading install still learns the difference
between "stopped working in THIS upgrade" and "was already inert".
Pins: tests/test_legacy_timeout_retirement.py (10 cases, incl. a grep-class
sweep and the auditor's since/behavior). The N-1 fixture carries the pair at its
DEFAULT values - a default-valued ghost is the one nobody looks for - so the
rc_audit fixture suite now pins that both produce a retired-setting finding.
Two tests that asserted the old no-op semantics are reshaped, not deleted.
Disclosed: saving the key through POST /api/settings no longer returns an
explicit "Retired setting(s) saved" warning; it is merged away silently like
every other retired key. Restoring it would mean reading the raw body for keys
the merge deliberately never looks at.
write_json now persists through the byte-exact atomic text writer (LF on every
platform, v7next audit #14-5), so the producer-write verification must not
translate its expected bytes to os.linesep any more: on Windows that raised
CyberGymIntegrationUnavailable("applied settings changed during producer
write") on every run and in test_cybergym_protocol.
The CPL4-C1/C2 train started rotating events.jsonl and tools.jsonl on the
supervisor tick, but the trajectory builder still read those two as single
from-birth files while chat/progress already went through the chain reader.
A rotated trial therefore published a trajectory missing its early tool
calls and its usage/startup events - a FALSE trajectory, not a short one.
Upstream = semantic truth, campaign = structural truth. Every upstream
semantic delta lands in its campaign owner leaf; upstream duplicate
extractions do not survive as twins:
- acceptance_dialogue.py -> folded into loop_acceptance{,_review}.py
(A-material paid identity, free replay, identical-refusal terminal,
dialogue history, inconclusive-dialogue reducer semantics)
- delivery_protocol.py -> folded into loop_delivery.py (hold-control
literals, RecursionError-degraded and trailing-object protocol parsers
over the shared strip_protocol_fence normalization)
- chat_delivery_events.py -> folded into events_chat_delivery.py
(unified _delivery_chat_id incl. chat-0 media, send_links/send_quiz,
registry merged via **_CDE)
- events review-wave handler retired for telemetry_events.py registry
- python_interpreter.py -> process_interpreters.py (upstream as-is);
registry/tool_resolution/shell retargeted; interpreter_attestation
scope in registry_core; node post-gates predispatch
- R5 typed process facts: process_facts.py channel + loop-side merge
into the typed result_meta; describe_returncode SSOT retargeted
- deadline_utils.deadline_expired public name adopted (rename-class);
test pins moved off the private spelling
- protected surfaces (BIBLE.md, safety.py, CHECKLISTS.md, registry.py,
gateway/contracts.py incl. endpoint_index extraction) landed as-is
- web wave, VERSION 6.113.5, package data landed as-is
- size_ratchet_manifest regenerated via scripts/regenerate_size_ratchet.py
(band rationales recorded for the F6-grown leaves); tools/core.py
link/quiz/escalate spans moved to core_artifacts.py to stay under the
giant gate; _run_shell and registry dispatch shaved under the
function gate via shell_process/registry_core helpers
Review fix batch 5 (wave 2 on 75c78ca2..295c9062; triad-codex
run-18ba6271145f, scope-sol run-7a1f9dd087ee, audit run-1c5c25e16be7).
Critical fixes:
- llm.py: the OpenRouter-branch reasoning_content scrub keyed on value
truthiness, so a legal empty-string echo (or a legacy null) rode the OR
wire; the guard is now key-presence (+falsy-residue regression).
- devtools/benchmarks/common/server_runner.py: DEEPSEEK_ joins
_AUTHORITATIVE_ENV_PREFIXES — an ambient DEEPSEEK_API_KEY survived the
settings-authoritative sweep that promises to strip provider families.
- devtools/benchmarks/programbench: the _active_direct_provider mirror of
config._exclusive_direct_remote_provider_env gains the minimax and
deepseek rows it silently omitted.
- devtools/benchmarks/terminal_bench: _network_preflight gains deepseek
(fixed DEEPSEEK_BASE_URL) and minimax (resolve_minimax_base_url over
MINIMAX_REGION, now forwarded by _container_env) branches — the agent
injected both keys but probed neither endpoint.
- tests/test_deepseek_provider.py: the density-witness regression clears
the process-global _DENSITY_MEMO in a finally (the memo key omits
drive_root, so leftovers poison co-located tests).
Registry-first regressions in the new
tests/test_benchmark_provider_touchpoints.py pin all three benchmark
surfaces against PROVIDER_CREDENTIAL_GROUPS, and both test-side provider
scrub tuples now derive from that registry instead of hardcoding keys
(test_devtools_benchmarks.py stays under its byte-debt ceiling: 328100).
Advisory closures: deepseek rows in the model-catalog and onboarding test
pins (deepseek-only setup, suggestions, profile derivation); honest doc
comments (ci.yml optional-provider rows, pyproject integration marker,
colab collection docstring, server_runtime provider lists, the stale
_EFFORT_CARRYING_PROVIDERS reference in llm.py); the
estimate_message_chars docstring and ARCHITECTURE row now name its real
consumer (local-context compaction proxy) instead of claiming the shared
fit/density basis; size manifest regenerated for the shrunken debt file.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The structured OUROBOROS_REVIEWER_SLOTS is the ONE reviewer configuration
surface. The legacy migration-read block (reviewer_slot_config.py:
_shared_session_route_spec / _legacy_rows / _legacy_config — comma-lists,
phase-5 per-row/advisory route envs, global-effort copying) is DELETED; with
no structured value the loader serves the SHIPPED DEFAULT PANEL: api_chat
triad/scope rows over get_review_models()/get_scope_review_models() (the
derived env plane — identical models to the old legacy defaults on every
config class, incl. single-direct-provider adaptation and a bench launcher's
env override), default advisory row, source="default", historical positional
slot ids (receipts keep lining up) and the unchanged legacy fingerprint
identity for the unconfigured panel (skill-review replay authority survives
the upgrade).
Settings vocabulary: OUROBOROS_REVIEW_MODELS / OUROBOROS_SCOPE_REVIEW_MODELS /
OUROBOROS_SCOPE_REVIEW_MODEL plus the phase-5 route envs
(OUROBOROS_REVIEW_ROUTES / OUROBOROS_SCOPE_REVIEW_ROUTES /
OUROBOROS_ADVISORY_REVIEW_ROUTE) leave SETTINGS_DEFAULTS and join
RETIRED_SETTING_KEYS (ghost purge on load; an install that configured
reviewers only through them gets the default panel — the RC auditor names
this migration, per plan). The derived runtime projection
(project_reviewer_slots_into_env) STAYS, floored from
OPENROUTER_REVIEW_DEFAULTS; get_review_models/get_scope_review_models keep
serving the API-pinned surfaces from that env plane. server_runtime's
provider-defaults migration now normalizes a ghost comma value it is FED but
never INTRODUCES a retired key (the direct-provider review adaptation lives
on the read side); the singular→plural promotion in model_slots is gone
(purged before it could run). gateway/settings candidate probes read the
structured candidate else the live derived config.
Bench templates migrated to structured slots with the SAME models
(continual_learning, gaia, swe_bench_pro base/example/probe/profile; comma
keys dropped everywhere incl. osworld/cybergym/programbench which already
carried structured values); run_tb metadata defaults now come from
OPENROUTER_REVIEW_DEFAULTS; manifests record OUROBOROS_REVIEWER_SLOTS.
tests/test_comma_list_sweep.py is the phase CI gate (ABI-10 hook): no
migration-read branches, retired vocabulary pinned, bench templates clean,
prose rewritten, derived projection alive. ARCHITECTURE deltas same commit.
(cherry picked from commit e7578133758f01eac584e8e6522b3f03c4d09818)
BREAKING (ABI 7.0 window, owner Q10=A verbatim: 'FLOOR/fail_tasks/
deadline-алиасы — удалить'). Both knobs were deprecated one-minor aliases
whose own event announced 'removal: next_major' - this is that major.
- contracts/task_contract.py: VALID_IMPROVEMENT_POLICIES = (fixed, adaptive);
an unknown policy (incl. the retired spelling) normalizes to 'fixed';
stall_rounds_threshold leaves the normalized profile shape (it was
normalized but consumed by nothing).
- task_pacing.py: the deprecated_task_pacing_alias event machinery in
resolve_budget_profile is gone; the until_deadline count-axis lift in
effective_max_improvement_passes is gone, and with it the has_deadline
parameter that existed solely for that branch (callers in task_results
and the rails line updated; the deadline/reserve TIME rail is untouched).
- bench adapters (sanctioned explicitly): programbench schemas.py + README
and swe_bench_pro entrypoint_pro.sh + METHODOLOGY switch to
improvement_policy=fixed - behavior-identical for those runs because their
explicit max_improvement_passes=6 was ALWAYS the binding count axis under
every policy; stall_rounds_threshold=12 dropped (never consumed).
- docs: ARCHITECTURE task_pacing module line and the acceptance passage now
state the removal.
- tests: the aliases' own tests removed (review_cycles alias test + rails A3
clause, two v6544 until_deadline tests); vehicle tests adapted to the
surviving semantics (wallet-authority test now derives its uncapped lane
from the unlimited shared cap; v664 deprecation-noise test now pins that
NO deprecation events are emitted at all; headless CLI forwards 'adaptive';
contract-shape and PB/SWE-Pro expectation pins updated). NEW removal pins
in tests/test_abi5_q10_removals.py: policy tuple, normalize-away shape,
signature no longer takes has_deadline, functional-remnant sweep.
Disclosed consequence: a PRE-7.0 stored root contract whose normalized
profile says until_deadline is judged malformed by the acceptance-wallet
authority (pre-existing unknown-policy behavior); pre-7.0 task-result
history is quarantined wholesale by ABI-2 (Q8=B) in this same release, and
the ABI-7 RC auditor names the migration.
Gates: ruff F clean; size_ratchet 5 passed; affected suites 478+90 passed.
Seven reference-named families leave the D08 monoliths as tip-byte leaves
(events_task_done, events_evolution_done, queue_snapshot, queue_timeouts,
worker_assignment, worker_health; the cancel ingress joins
events_runtime_controls and the owner-stop campaign closure moves beside
stop_evolution_tasks in queue_transitions), each proven by the transplant
tool (ast=tokens=True per symbol, whole-leaf invariants clean) with MAXIMAL
declared sets - every facade name tests monkeypatch routes through the
call-time handle, swept by an ast.walk over all test setattr forms.
Delta-D08 re-derived on tip bytes (owner Q-d=A): the four remaining
fail-open cancel_intents mutators (mark_finalize_control_drained,
mark_intent_scope, release_claim, settle_intent) now refuse a corrupt
projection with the typed CancelIntentProjectionCorrupt instead of
answering "no intent"; every tip caller audited for the raise path in the
lane ledger. Owner decisions Q-a=A/Q-b=A/Q-c=A applied: the settle owner
stays in task_lifecycle (no cancel_custody.py), the evolution remainder
stays on the queue facade, the closure home is queue_transitions.
Durable pins return re-keyed to the tip form: C1/C2 corruption suite,
C7-C10 protocol inventory (manifests re-keyed by symbol, upstream
retry/depth lanes added), the E1-E12 owner-control scenario suite with its
driver extensions (typed cancel/hurry + _api_status in the bench server
runner; E8 retired, E13 carries the live pause semantics), and the C9
owed-before-persist docstring correction. The phase-a test giant is split
by the S7b themes from tip bytes, lossless (107 functions / 112 items).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 95bf2fa60b1133aa058bdd11d42bee50219a1cac)
The full LinksOutbound-pattern stack for a new typed outbound: QuizOutbound
contract + WS discriminator, chat.quiz host topic, shared validate_quiz_payload
(assumption REQUIRED per owner decision 27=A; options capped by the pinned
Python<->JS MAX_QUIZ_OPTIONS), bridge.send_quiz with a base64-free persisted
row, supervisor delivery handler with project binding, typed history replay
(a quiz row can never read as a bare final), api_types.js mirrors, the
chat_decision.js card family (question / at-stake / option buttons / the
assumption line that survives settlement as the record of the default path;
answers post to the unified decision contract), token-born CSS, a DESIGN.md
'Quiz card' section, a Telegram display subscription, and an ARCHITECTURE ABI
row. Delivery wiring for media/links/quiz is extracted from chat.js into
chat_media.wireDeliveries, shrinking chat.js 219170 -> 217710 bytes with the
quiz wiring included (BYTE_DEBT ratcheted down). The surface stays dormant
until the escalate tool lands: nothing emits send_quiz worker events yet.
Resolution: the Claude-runtime retirement (c510de72) deleted
ouroboros/gateways/claude_code.py and tests/test_claude_code_gateway.py on
the mainline; the PR's gateway clamp commits target that retired runtime, so
both files stay deleted and the clamp dies with the runtime it guarded.
A second human-facing label beside recommended_use rotted against route
edits (the shipped 'Fast scout' name survived a re-route to a cursor
session and misdescribed the row in the reviewer picker and the delegation
prompt). The parser accepts legacy 'name' values and DROPS them — the next
serialize omits the key, which is the whole migration; no fabricated
title-case fallback either. The delegation catalog and dispatch snapshot
carry facts first (subagent_id, route_class, requested target/model,
effort) with recommended_use LAST as bounded owner intent; presets and the
legacy-derived rows mint no names; committed bench profiles re-serialized
through the SSOT. prompts/SYSTEM.md states the identity rule (facts true
when description disagrees) and the LLM-first hygiene: an agent editing a
row's route rewrites its recommended_use in the same change.
Native episode: transcript bound enforced BEFORE every send (an oversized
initial prompt or final round can no longer slip past the rail); per-tool-
result 120K bound with disclosed truncation and an offset/limit continuation
handle; episode rollup carries ledger_attempt_ids/resolved_model/provider
(add_usage drops them — the substrate's actor records were echoing the
requested model dressed as resolved); tool receipts get an outcome field and
native_tool_calls counts ALL calls, not the capped receipt store; episode
instructions and the advisory prompt teach the REAL inspection tool names
and chunked reads (the retired SDK's Read/Grep/Glob vocabulary burned
capped rounds on refusals). Default transcript cap 400K->900K (real-repo
E2E evidence).
Scope authority: a native-retrieving actor row now rides the SAME retrieving
authority decision as sessions (BIBLE P3 alternate delivery, sourced >=200K
floor) instead of falling into the api 1M-pack branch it never uses.
Contract identity: commit/skill/plan reviewer fingerprints bind the actor
reference (delivery class) — added only when a row carries one, so unchanged
legacy rosters keep their exact bytes.
Settings save atomicity: reviewer references validate against the roster the
SAME save produces (add row+reference in one save works; a roster-only save
removing a still-referenced actor is refused).
Migration honesty: advisory_gate_unavailability_reason surfaces the typed
disabled_reason (migration force-disable never reads as a standing owner
choice); the three legacy-target migration branches are now pinned by tests.
Also: retrieves predicate consolidated at object-level call sites; cybergym
README retired-transport wording; managed-update affordability floor's
one-pack pricing convention documented.
- A7: finalize_now (deadline/ceiling/owner stop) arriving during an active
transport-wait episode no longer dispatches a forced provider call over the
proven-dead egress: _maybe_early_finalize routes the exit through the
transport no-resend terminal (finalize_now_transport_terminal helper closes
the episode's durable evidence with detail=finalize_now); mailbox-drain
control bookkeeping is unchanged and the task terminalizes promptly.
- A8: is_pre_dispatch_transport_failure types the standard unreachable-proxy
chain (requests ProxyError -> MaxRetryError -> urllib3 ProxyError ->
NewConnectionError/ConnectTimeoutError); proxy HTTP responses and
post-dispatch read failures stay untyped.
- A9: additive route_is_loopback fact on AttemptRequest/PhysicalAttemptCapture
(set from the target base_url host in _attempt_request via
is_loopback_base_url); classify_llm_exception excludes loopback routes from
transport_unavailable alongside provider=='local', so a stopped Ollama/LM
Studio/vLLM server fails fast instead of waiting out a phantom outage.
- A10: provider_terminal_fallback_text keys fail-fast wording on NOT
wait_eligible, adds an eligible-but-zero-wait wording, and replaces
INTERRUPTED with outage phrasing that cannot collide with the supervisor's
STATUS_INTERRUPTED lifecycle term.
- A11: the Q14 final free redial reserves a named 3s margin
(_FINAL_REDIAL_MARGIN_SEC) so round-top overhead cannot get the granted
redial refused by the admission gate.
- A12: contract tests — error_kind_changed episode exit resumes ordinary
fallback; unknown-outcome redial takes provider_outcome_unknown_no_resend
with zero further dials; episode window carries zero dispatched paid
attempts; nonstandard OUROBOROS_TRANSIENT_RETRY_MAX keeps the one-attempt
transport contract; review actors' physical_attempt_limit(2) rail pinned.
- A13: _WAIT_BACKOFF_START_SEC promoted to config.py as
NETWORK_WAIT_BACKOFF_START_SEC (config stays 1600 lines); generalized
_self_check_round comment; dropped redundant capture-None guard in the
classify branch; CLB RUNBOOK notes headless tasks keep the idle rail as an
additional bound. Size-ratchet manifest regenerated (loop.py byte debt
shrinks 312847 -> 312831).
A REMOTE pre-dispatch transport failure (released custody + typed predicate +
non-local provider) now classifies as its own kind, takes exactly one physical
attempt per call, and is paced by a round-level wait episode in the new
ouroboros/loop_transport.py: durable network_wait events, owner progress notes
(first immediately, then min(NETWORK_WAIT_NOTE_INTERVAL_SEC, idle/2) cadence),
an owner-signal-interruptible backoff sleep capped at the existing 60s bound,
and free redials of the same round until the task's own rails (deadline minus
the dispatch-admission reserve with one last free redial, budget, Stop, the
supervisor ceiling). The fallback chain runs at most once per episode and only
when USE_LOCAL_FALLBACK makes it local; a chain candidate dying pre-dispatch
stops the walk. Direct-chat/ephemeral turns fail fast instead of waiting.
Exhaustion takes a deterministic no-resend terminal keyed on the episode's
latched cause (transport_unavailable_no_resend, reason provider_unavailable,
infra_failed) with no forced-final provider call; the in-episode deadline
sliver suppresses the paid deadline-local finalize so the terminal story
cannot fork. Provider-failure hint/salvage helpers move to loop_transport
(loop.py sits at its byte-ratchet cap). ARCHITECTURE and the CL-Bench runbook
disclose the changed outage-time semantics of OUROBOROS_TRANSIENT_RETRY_MAX.
The whole 3-OS matrix predates this branch in red; every leg is now
root-caused (four investigation lanes, all dispositions verified on this
host where reproducible).
Production fixes:
- web/modules/chat.js: merge #358 over-narrowed the durable-replay
conclusion gate to typed terminal facts only, but the wire never stamps
one on an ordinary bare final row - on replay a subagent final leaked
into main chat, orphaned cards never converged, and a bare-final card
sat on Working forever (three ui-smoke contracts). A plain untyped
final row concludes again; every marked row class the checkpoint
defended (system_type/msg_type) still never concludes. Verified by the
full ui_browser sweep (50 passed) and the checkpoint's own pins.
- cybergym: two real Windows portability bugs - virtual symlink targets
are container-side POSIX paths and are now classified with PurePosixPath
semantics instead of host Path; the applied-settings integrity hash now
binds the platform-newline serialization write_text_atomic actually
persists (CRLF on Windows always tripped "changed during producer
write").
Platform truth (with the files' own precedents): fifteen cybergym tests
assert POSIX-exclusive capabilities of a Linux-container docker benchmark
(descriptor-safe extract primitives, POSIX-absolute mount paths, POSIX
mode bits) and now carry the established capability guards/skipif
patterns already used in the same files - they reactivate automatically
wherever the capability exists.
Deterministic test seams (no assertion weakened, no production change):
process-global time.monotonic patching replaced with bounded injection;
float-precision and microsecond-truncation clock comparisons made
representable; hidden-deadline waits and fixed mock-LLM sleeps replaced
with event-gated synchronization and explicit in-flight drains; the
process-lifetime supervisor event-bus shutdown latch is isolated per the
in-repo fixture precedent (the xdist pollution behind the "local-green,
CI-red" attachment staging pair and the inflight-seam trio); the
ui-smoke status test now gates on history hydration before emitting
(per the project-continuity precedent).
chat.js and the manifest stay inside their byte ratchets (comment
compaction on this branch's own additions); the size-ratchet lane and
manifest --check are green against the drifted base.