Commit graph

252 commits

Author SHA1 Message Date
Anton Razzhigaev
f6711d3369 feat: integrate subscription accounts into model setup and execution
Add caller-owned Claudexor model calls, unified subscription/API setup, explicit role accounts and live quota/auth continuation. Preserve prompt/tool ownership, physical-attempt custody and manual assignments. Keep the existing task process through waits and post-work.

The reviewed source checkpoint retains the current dependency pin. Published signed Claudexor bytes, cross-platform CI and the agreed merge ordering remain delivery prerequisites. No Ouroboros version, tag or release is created.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-07 10:54:51 +00:00
Ouroboros
53815d6305 feat(e2e_live): the stand runs a cheap three-family review panel at effort low
The owner's choice of 2026-09-06 for the live stand (questions 1-4 = A):
keep the product's defaults untouched for installs and give the STAND
lanes a cheaper panel that still covers three model families. run3 on the
default panel cost 141.63 USD, 75% of it review, claude-opus-5 alone 39%.

Paid lanes now write OUROBOROS_REVIEWER_SLOTS (the one structured
reviewer surface the product reads): triad google/gemini-3.8-flash,
openai/gpt-5.6-luna, deepseek/deepseek-v4-pro; scope
deepseek/deepseek-v4-pro (the scope pack is input-heavy, so the cheapest
1M-context input wins); advisory anthropic/claude-sonnet-5 (the Anthropic
family stays in the loop where it is cheapest, an inspection rather than
a packet); every reviewer at effort low, OUROBOROS_EFFORT_TASK and
OUROBOROS_EFFORT_EVOLUTION at medium. scenarios.STAND_REVIEW_PANEL /
STAND_PANEL_SETTINGS carry the document; --production-panel runs the
tree's own defaults instead; the stub lane keeps its loopback rows.
The manifest's model snapshot already records OUROBOROS_REVIEWER_SLOTS,
so the panel is part of the run's provenance.

tests/test_e2e_live_panel.py parses the document with the product's own
parser (a document it rejects would fall back to the default panel
silently), pins the three families and the efforts, and pins the
production-panel and stub exceptions. DEVELOPMENT.md states the panel.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 13:54:24 +00:00
Ouroboros
8e847e7982 fix(e2e_live): spend from the product ledger; skill fixture declares its runtime; landing skip flags recorded and checked
Three stand corrections from the deep read of rc.15 run3.

Money: lane_spend summed the llm_usage telemetry rows, which the product
itself calls compatibility telemetry (pricing.py); the money authority is
the settled physical-attempt ledger state/usage_attempts.jsonl. run3 read
114.81 by telemetry and 141.63 by the ledger — the skill-review triad, the
advisory pre-review through the agent SDK and the post-task synthesis
write no telemetry row — so the run cap admitted on a short count. The
reader now sums settled, cost-final ledger rows (cost_usd), counting an
unpriced one as unknown; the docstring's claim that the telemetry carried
every row is gone.

Skill fixture: the stand's reference SKILL.md declares type extension with
no runtime, which CREATING_SKILLS.md says is required for extensions; the
model copied the fixture byte for byte in all three lanes and the terra
reviewer failed manifest_schema (a hard-critical item) in two of them. The
fixture now declares runtime: python3.

Landing path: DEVELOPMENT.md describes the SM1 landing as commit_reviewed
with no skip flags and a real advisory row, but the stand only checked
that some advisory run was real. run2 SM1_a1 and run3 SM1_a1 landed
through review_rebuttal + skip_advisory_review=True and still passed.
commit_refusal_facts now records the landing call's truthy skip_* args as
landing_skip_flags and run_sm1 checks landed_without_skip_flags, so a
bypass landing is a typed failed check with the facts beside it.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 11:08:02 +00:00
Ouroboros
19128ad7f6 fix(e2e_live): SM1 clean-worktree check tolerates the runtime's transient scratch and records the porcelain
rc.15 run3 on b6ac8c13, lane SM1_a1: the reviewed commit landed, and the
post-task promotion enabled a one-shot evolution cycle that started
seconds later in the same clone. Its run_script wrote
<clone>/.ouroboros/tmp_scripts/script_<uuid>.py (the active-workspace
scratch root, tools/shell.py, unlinked in a finally), nothing ignores that
directory, and the stand's worktree_clean_after_commit check — a bare
`git status --porcelain == ""` right after the commit — saw
"?? .ouroboros/" and failed a lane whose commit had left the tree clean.
run2's lanes used run_script 33 and 13 times without tripping it: a race
with the neighbouring task, not a constant. The product side is issue
#701 (the runtime's scratch inside the self-modified worktree).

worktree_after_commit() records the porcelain and the transient entries
as lane facts and counts only the runtime's own `.ouroboros/` scratch as
transient; any other untracked or modified path still fails the check.
A focused test module pins both directions on a throwaway repository.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 03:34:03 +00:00
Ouroboros
fa2977423b fix(e2e_live): reserve the evolution root only for the scenario that absorbs; typed idle reasons relative to the wait
Second adversarial round on the absorb-wait rewrite.

Reservation: after the previous commit only SM1 can promote (SW1/SK1 pin
OUROBOROS_POST_TASK_EVOLUTION=false), yet RunBudget.reservation still
added the evolution root for every attempt under --self-mod, so SW1/SK1
reserved a ceiling they could never spend and run-cap admission refused
or serialized paid attempts on it (the SK1-only mini run at cap 130 was
refused by exactly this over-reservation: 50 x (2 + 1) = 150). The rule is
now per_task x (root_tasks + int(self_mod and absorbs)); admit,
budget_preflight, dispatch_order and run_lane pass the scenario's
expects_absorb. Owner configuration (cap 300, per-task 50, 3 attempts,
self-mod): SM1 100, SK1 100, SW1 50 — realistic spends admit 9/9 for
$159, pessimistic 8/9 for $219 (SK1_a3 refused), every scenario keeping
two; the CI e2e-live arithmetic comment and summary header are rewritten
(full set $225, not $315) inside the D-12 job, and the CI-lane pins
re-derive the numbers with the scenario flag.

Idle reasons: absorb_idle_reason types relative to the wait's start
(history length snapshot) so a resumed campaign's older cycles never
speak for this boundary; a paused/stopped/completed status wins;
no_promotion means an every_n post-task tick was recorded (llm cadences
write none: no_decision). Tests cover the resumed-campaign boundary and
the reason table with the cycle written during the wait.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 01:31:11 +00:00
Ouroboros
4f3cda85d5 fix(e2e_live): absorb wait proves no cycle is pending; non-absorbing lanes do not promote
Adversarial review of the corrected --self-mod seeding (2026-09-06) traced
the post-task path end to end and found that IsolatedServer.wait_for_absorb
ended the wait as no_promotion on ONE idle /api/state sample once the grace
had passed: no request file, pending_count 0, running_count 0. Those are
exactly the readings of the absorb path itself. A cycle that committed keeps
its campaign transaction as waiting_for_restart while the supervisor
restarts synchronously (RUNNING already popped, the counter unchanged), and
the re-exec'd server answers /api/state with zero counts before its
supervisor is up; absorbed_cycles_done moves only when the worker boot
verifies the restart. The pre-seeded benchmark campaign used to mask this
(its t=0 cycle kept running_count above zero and the kept request file
blocked the exit); with the owner-id-only seeding every SM1 lane would have
failed a healthy absorb as no_promotion.

wait_for_absorb now needs PROOF that no cycle is pending, held on six
consecutive polls after the grace: the queue idle AND supervisor_ready AND
no post_task_evolution_request.json AND no campaign active_transaction. The
reason is typed from the durable campaign state (campaign_summary /
absorb_idle_reason): no_promotion (no campaign although the post-task
decision ran), no_decision, cycle_no_op, cycle_not_absorbed,
campaign_<status>, cycle_not_enqueued; the summary travels in the wait dict
the stand records. Benchmark callers (evolve_smoke, the CLB adapter) keep
their early exit on a genuinely idle campaign.

Two more findings from the same review: scenarios that commit nothing
(SW1, SK1) now pin OUROBOROS_POST_TASK_EVOLUTION=false in their lane
settings — a one-shot cycle promoted from their own roots could commit and
re-exec the server in the middle of the lifecycle under test; and the SK1
gate also requires the review call's own executable_review (a pending
duplicate job never passes) while _skill_entry reads the listing through
the non-raising helper so a failed listing is a failed check with facts, not
an infra_error. Tests: a new module pins the wait on a fake clock
(waiting_for_restart and a booting server do not end it, one idle sample
does not, a pending request blocks it, the absorb confirms, every idle
reason), the SK1 module pins the product gate rule and the per-scenario
promotion pin, the runner module's settings pin flips for SW1/SK1.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 01:04:49 +00:00
Ouroboros
d70f93c3c1 fix(e2e_live): SK1 verdict is the product review gate, clean state a fact
The rc.15 SK1-only rerun on 560f7d71 (three attempts, 2026-09-05) passed
every lifecycle check three times — preflight, review job, auto-grant equal
to the request, enable to a live-loaded dispatchable skill, a physical
extension call echoed into the owner chat, cleanup — while the stand's own
"all-PASS review" criterion passed once: the model-authored payload came
back clean, then with warnings (a SKILL.md/plugin.py contradiction), then
with blockers (manifest without runtime). The criterion measured the
author model, not the product.

Owner decision (2026-09-06, question 3 = A): SK1 counts when the review is
executable by the PRODUCT gate — the /api/extensions row's
executable_review (clean, warnings, or blockers under advisory enforcement
by operator choice, skill_review_gate) — plus the full lifecycle. The
check is now review_executable via sk1_review_gate(); the review status,
enforcement, blocking reason, non-PASS items and the clean/all-PASS state
travel as recorded facts. DEVELOPMENT.md says so; a parametrized test pins
the rule on the three rerun outcomes and the failure shapes (blocking
enforcement, failed review call, no findings, no gate fact).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 00:42:09 +00:00
Ouroboros
c6f002c122 fix(e2e_live): --self-mod seeds the owner id only, never a campaign
The rc.15 paid run2 (2026-09-05, SM1_a1/a2) showed the stand's --self-mod
lanes pre-seeding the benchmark helper's ACTIVE campaign ("benchmark
evidence" objective, evolution_mode_enabled) instead of relying on the
post-task promotion path the flag is meant to exercise:

- generic cycles ran from t=0 next to the scenario task (three no_op
  cycles, $16 on one lane), each counted as a consecutive failure;
- the promotion the scenario task wrote (post_task_evolution_request.json)
  was refused by apply_pending_request because evolution was already
  enabled, and the kept request file also blocked wait_for_absorb's
  no_promotion exit — every SM1 lane waited the full task timeout with
  absorbed_cycles_done at 0 although its reviewed commit had landed.

seed_owner_state(data_root, evolution_enabled=False) under every profile:
--self-mod stays a settings fact (OUROBOROS_POST_TASK_EVOLUTION + cadence
every_n:1) and the scenario task's own promotion enables the one-shot
campaign whose cycle lands, restarts and absorbs. The reservation rule
keeps +1 root (the one post-task cycle); its wording, the RunBudget
docstring and the DEVELOPMENT.md stand section say so, and a runner test
pins the seeded state (owner_chat_id only, no evolution_campaign.json,
no evolution_mode_enabled) with and without --self-mod.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 23:25:35 +00:00
Ouroboros
df9c140ade Live E2E stand: the SK1 probe plugin presents the host token to loopback only
The skill review blocked the stand's own probe plugin twice on the rc.15
paid stand (host_token_handling, CHECKLISTS skill item 12): it built the
Host Service base from HOST_SERVICE_URL unvalidated, so the token could
be sent to any host. The plugin now refuses a base that is not a loopback
http URL before any request is built; the echo behaviour is unchanged. A
small test module drives the plugin text the model is told to write.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 21:16:12 +00:00
Ouroboros
ed8ecd216f Scope the live stand's absorb wait and check to scenarios that land a commit
Under --self-mod the runner applied the post-task absorb wait and the
self_mod_absorb_confirmed check to every lane, but only SM1 commits
anything: SW1 and SK1 failed by construction and waited for the evolution
campaign to promote nothing (rc.15 paid stand, 2026-09-05: SK1_a1 passed
twelve of thirteen lifecycle checks, the thirteenth being a real reviewer
finding on the model-authored plugin, then waited about 27 minutes until
the evolution cycle ended and the wait returned no_promotion).

Scenario gains expects_absorb (True for SM1 only). The absorb snapshot,
the wait in restart() and after the scenario, and the lane check follow
that flag; the run-level gate lists only absorbing lanes and the manifest
names how many lanes were expected to absorb. A lane that expects no
absorb stops its server right after the scenario and records
self_mod_absorb: {"expected": false}. OUROBOROS_POST_TASK_EVOLUTION and
the per_task x (roots + self_mod) reservation stay as they were: the
campaign may still run and spend during the scenario; the stand just
does not wait for it.

Pins: the scenario table flag, SM1 waiting and carrying the check while
SW1/SK1 finish with no confirm_absorb call and evolution still on in
their applied settings, the run-level gate and manifest counters. The
suite stays at its 1500-line band ceiling by folding its two fake servers
into one and the argument-refusal cases into a loop; the runner stays at
the 1000-line limit. Handbook: the scoping and the incident.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 21:04:08 +00:00
Ouroboros
e1036afe08 Live E2E stand: FIFO admission by dispatch index, exact feasibility pins, no largest-first driver
Re-sorting the pool ordered submission only: RunBudget.admit was not FIFO, and
on the real runner the freed lane's next job takes the lock before the woken
waiter every time (300/300: settle in run_attempt's finally, return, the
executor's next job, admit on the same thread), so a round that overflows the
cap (owner configuration 150 + 100 + 100 = 350 > 300) let later attempts
leapfrog the waiter and could leave SW1 at 0/3 under pessimistic spends.

admit now takes the job's dispatch index and keeps the attempts still asking
in a line under the existing Condition: an attempt reserves only when no
earlier-dispatched attempt is still asking and the cap has room beside the
reservations in flight; one that can never fit is refused at once and leaves
the line; reserve, refusal and settle wake every waiter and the line predicate
re-parks the later ones. No deadlock: a waiter waits only while something is in
flight or an earlier attempt is asking, the earliest never waits on the line,
and with nothing in flight it either fits or is refused. The wait log line
names its reason (an earlier attempt vs reservations in flight). main() passes
the index from the new dispatch_order (round-robin by attempt, largest
reservation first within a round), which the test driver now shares.

Tests: the driver parks the ledger's wait and releases threads in the runner's
observed order (newcomer asks first, then the line in dispatch order), so the
owner configuration (cap 300 / per-task 50 / attempts 3 / pass-of 2 / 3 lanes /
--self-mod) is pinned exactly: realistic spends admit all nine in dispatch
order for $159; pessimistic spends admit SK1_a1, SM1_a1, SW1_a1, SK1_a2,
SM1_a2, SW1_a2, SM1_a3 and refuse SK1_a3 then SW1_a3 for $211 (every scenario
keeps two). The largest-first driver option and its comparison pin (red 4/15)
are gone; a RunBudget unit pin covers the line (a later attempt that would fit
waits behind an earlier one; the refused head frees it). Handbook and
docstrings name the waiter-vs-newcomer race, the head-of-line cost, the exact
traces (largest-first under the same admission: SW1 0/3 at $225) and that
round_worst_case_usd sums one attempt per scenario.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 17:22:44 +00:00
Ouroboros
80d49bd7ca Live E2E stand: round-robin dispatch by attempt, no bench budget profile, per-round preflight worst case
The verdict is pass-of PER scenario, so the dispatch order now maximises the
minimum admitted attempts per scenario: a1 of every scenario, then a2, then a3,
largest reservation first within a round (requested_task_ids keep the argument
order). At the owner's configuration (cap 300, per-task 50, attempts 3, pass-of 2,
3 lanes, --self-mod) with pessimistic spends, largest-first left SW1 with at most
one admitted attempt; round-robin keeps every scenario at two or more.

The cost_hard_stop_pct=100 projection is dropped (scenarios.STAND_BUDGET_PROFILE,
LaneContext.submit metadata, its pins and handbook text): with a reservation of
at least 2 x per-task the product's per-task axis already binds, and a bench
profile would change the pacing path under test. Every stand root runs under the
product's default ceiling, including the evolution root and SW1's UI-path root.

budget_preflight takes --lanes and records the per-round worst case (the lanes
largest reservations of one attempt per scenario at once); the handbook states
that outcomes among equal waiters depend on wake order, so a feasibility number
is a range at the margin. Docstrings now say rc.14 showed up to two evolution
cycles per lane (t=0 and post-task) and that the lane TOTAL_BUDGET is the fence.

Tests: the deterministic driver over the real RunBudget takes the dispatch order
and pins the owner configuration under realistic and pessimistic spends for
round-robin (admitted set, per-scenario counts, spend range) and shows the
rejected largest-first order leaving SW1 below pass-of; the run-wide cap manifest
pin is re-parametrised for round-robin (cap 16, pass-of 2); configuration A/B pins
and the projection pin are removed within the band ceiling.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 17:22:43 +00:00
Ouroboros
ebc713c379 Live E2E stand: reserve per root (evolution counted), project cost_hard_stop_pct=100, budget preflight, largest-first dispatch
The rc.14 artifacts show the 2x reservation of 90f02b02 was covering the SECOND
ROOT, not the 50% ceiling: with --self-mod the evolution cycle starts at t=0
next to the scenario task and shares the lane fence (SM1_a1: task $3.84 +
evolution $12.41 + $2.84 = $19.08 of $20). The audit driver over the real
RunBudget then showed that under the planned rc.15 flags (cap 200, per-task 50,
SK1 = 2 x 50 x 2 = the whole cap) SK1 could be admitted at most once in any
order, and never in runner order, after SM1/SW1 had already spent.

- RunBudget.reservation = max(floor, per_task x (root_tasks + int(self_mod))):
  the evolution root is counted as a root when --self-mod; HARD_STOP_INVERSE
  and the ouroboros.task_pacing import are gone; RESERVATION_RULE, the
  --per-task-usd help, the docstrings and the handbook state the root-count
  rule.
- The in-task halving that factor stood in for is switched off on the roots
  the stand submits itself: LaneContext.submit projects
  metadata.budget_profile.cost_hard_stop_pct=100 (scenarios.STAND_BUDGET_PROFILE,
  the seam ProgramBench uses, in the shape normalize_budget_profile accepts).
  The evolution root keeps the product default (its task metadata carries only
  the transaction; no seam) and so does SW1's root on the UI path
  (ui_probe.send_chat is a WS chat frame without metadata; acceptable: SW1
  spends ~$8 against a ceiling of >= $25 at per-task $50) - both stated in the
  handbook.
- budget_preflight (the credit preflight's typed shape, recorded as
  extra.budget_preflight): a reservation row per scenario, the worst case
  sum(reservation x attempts) against the cap, and a fail-closed refusal
  {stage: budget_preflight, reason: reservation_unreachable} (exit 3, before
  the key, the seed or any lane) when a reservation exceeds the cap or equals
  it with attempts >= 2. No override flag: the operator changes the flags.
- Jobs are dispatched largest reservation first (stable among equals); the
  manifest's requested_task_ids keep the argument order.

Tests: the ledger pins keep every admission number (per-task $8 is the $8 unit
now); the reservation pin asserts per-task x roots, +1 with self_mod, and that
the exact number reaches the lane's settings.json; a submit pin captures the
POST body and resolves the product ceiling with and without the projection
($47 vs $25 for per-task $50 in a $50 lane); preflight pins cover the direct
function, the manifest refusal and the dispatch order; a deterministic driver
over the real RunBudget (3 lanes, SCENARIOS.root_tasks, stub spends visible at
settle, waiting detected by instrumenting the ledger's wait) pins
configuration A (cap 200 / per-task 50 / attempts 2: realistic spends all six
admitted, $106; pessimistic spends five admitted, exactly one SW1 attempt
refused after waiting, $158) and configuration B (cap 260 / attempts 3:
realistic all nine admitted, $159; pessimistic six admitted, every SW1 refused,
$225). The run-wide cap pin is re-parametrised for the root-count rule and the
largest-first order (SK1_a2 refused, SM1 still runs after that refusal).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 17:22:43 +00:00
Ouroboros
e5f61f5852 Live E2E stand: the orphan cmdline pin is procfs-gated; the lazy-open note names SW1
The orphan-detail pin asserted the survivor's cmdline text, which is read
from /proc and is the typed empty string on macOS and Windows (the module
runs on the 3-OS matrix). The text assertion is now gated on
PROCFS_AVAILABLE; the shape and cap pins stay unconditional. The handbook
says the use-time open applies to SM1 (SW1 drives the task through the UI
and holds its browser), and UIProbe.close carries the thread-affinity
invariant a cross-thread close would silently violate.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:57:14 +00:00
Ouroboros
800159039d e2e_live: open the UI probe at use time, type a dead browser target, name orphan survivors
The v7.0.0-rc.14 paid run opened each lane's headless Chromium at lane start
and first used it after the task and the self-mod absorb wait; lanes that
waited 10-23 minutes met TargetClosedError on the first goto, the exception
escaped run_sm1 and the whole lane became infra_error although its task-side
checks were complete. Both lanes also recorded no_orphans_after_stop=false
without naming the survivors.

- ui_probe.GuardedUI wraps the client: any browser-call failure becomes the
  typed reason ui_unavailable:<ExceptionType> (recorded as ui_reason and in the
  lane facts), later calls are no-ops, the exception never reaches the scenario.
- LaneContext.ui is a lazy property: the client opens on first use against the
  current server; restart() closes any open probe and re-resolves on the next
  use (a non-self-mod restart may change the port). The scenario surface
  (ctx.ui truthiness, goto/computed_property/send_chat/screenshot) is unchanged.
- run_live_lanes: lane start only probes availability (open + close); the lane
  row's ui verdict is taken from the context at the end; the orphan scan hands
  the surviving pids to _apply_orphan_scan, which names them in row["orphans"]
  (pid + cmdline head, capped at 20 with orphans_omitted).
- tests pin the lazy open / restart reopen, the typed open failure, the closed
  target degrading the SM1 UI tail to checks_failed at both context and lane
  level, and the orphans row shape and cap.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:57:14 +00:00
Ouroboros
db805c5b73 Live E2E stand: SM1 pins that both sheets are in the commit and skips taken versions
Two reviewer findings on the realistic SM1: (1) the diff-scope check could
fail a commit the scope reviewer legitimately widened (another accent
touchpoint such as the unlock page's inline colour or the site stylesheet),
so the stand now pins only that both stylesheets are IN the landed commit
and records the companions as a fact; the reviewers own the scope judgment.
(2) A seed cloned from an older ref carries the newer tags and the review
binding refuses a staged version whose tag exists, so the stub's next
version skips every version whose tag the clone already has.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:52:21 +00:00
Ouroboros
e0d81f4ec1 Live E2E stand: SM1 asks for a design-system-consistent accent change landed as a reviewed release
The first paid stand run on v7.0.0-rc.14 failed all three SM1 attempts on the
commit gate, and the reviewers were right: the prompt changed only the
--accent value (a mixed blue/red palette against docs/DESIGN.md "Brand
accent"), forbade the version bump (BIBLE P9, CHECKLISTS item 8) and routed
around the advisory preflight, and claimed a visible UI change with no
vision-inspected evidence (development_compliance 2(i)). The stand must
describe the realistic task and let the product's policy shape the work, not
fight it.

SM1 now asks for the brand accent changed consistently with the design system
(the one --accent value, its named roles and the alpha ladder) in both
stylesheets, verified on a real consumer flow in the browser, and landed as a
reviewed release through preflight_review and commit_reviewed with the
synchronized carriers bumped in the same diff - no skip flags, no "do not
bump". The acceptance pins only the observable contract: the commit landed,
both committed sheets carry the new accent and keep their whole shared :root
token set in parity, VERSION is a strictly-greater release version whose
carriers pass the product's release admission gate read at the landed commit,
the diff is confined to the stylesheets, the carriers the release_sync SSOT
names, docs/DESIGN.md and comment-only web CSS companions, the advisory
ledger row is a real reviewer run rather than a skip/bypass row, and the
browser/vision tool rows are recorded as a vision_evidence_present fact for
the reviewers to judge. The stub rehearsal projects the bump offline through
sync_release_metadata plus the README history row (trimmed to the P9 limits
the gate enforces) and takes the same full path without skip flags. Docs and
the hermetic pins follow the new contract.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:52:21 +00:00
Ouroboros
90f02b0228 e2e_live: reserve 2x per-task per root task (inverse of the product's hard stop)
The rc.14 paid run (per-task $20) failed lane SM1_a3 with budget_exhausted at
$10.21 accounted while its cap was $20: the product's in-task ceiling is
min(cost_hard_stop_pct of the GLOBAL remaining at task start, per-task cap -
planning margin), and in a lane the global remaining IS the lane budget = the
1x reservation, so every root task's ceiling collapsed to 50% of its cap. On a
real install the global budget is far larger and the root-cap axis binds.
Second, --self-mod runs the post-task evolution cycle as another root task
under the same per-task fence (SM1_a3: task $10.2 + cycle ~$8 of $20), so a 1x
reservation under-reserves the lane's legitimate spend.

The reservation is now max(LANE_BUDGET_FLOOR_USD, HARD_STOP_INVERSE x
per_task_usd x root_tasks), where HARD_STOP_INVERSE = 100 / the product's
_DEFAULT_COST_HARD_STOP_PCT imported from ouroboros.task_pacing (2 today; it
follows the product's default rather than copying the number). The lane's
TOTAL_BUDGET stays equal to its reservation, so the ceilings in flight remain
disjoint and settled spend + in-flight ceilings <= cap. RESERVATION_RULE (the
manifest's reservation_rule), the docstrings, the --per-task-usd help and the
handbook paragraph state the rule and both reasons.

Tests: the ledger pins use per-task $4 so the $8 unit and every admission
number are unchanged; a new EQUALITY pin drives run_lane up to the written
settings file and asserts that per-task $20 with one root reserves $40 and that
exactly $40 reaches the lane's settings.json as TOTAL_BUDGET (never the
template's $100 run cap), and that the factor equals 100 / the product's
default cost_hard_stop_pct.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:39:34 +00:00
Ouroboros
2f6d6e59c2 Live E2E stand: the handbook states the wait-then-refuse admission rule
docs/DEVELOPMENT.md still described the retired contract ("halts scheduling
at the first refusal"); it now states the per-attempt rule the runner
implements (wait while blocked only by reservations in flight, refuse only
what can never fit, manifest keys refusals/first_refused). The RunBudget
docstring names the non-FIFO fairness policy and the on_wait-under-lock
contract the reviewers asked to pin.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:34:02 +00:00
Ouroboros
88509e2cea Live E2E stand: budget admission waits on in-flight reservations, refuses per attempt
The first paid run (v7.0.0-rc.14, cap $100, per-task $20, 3 lanes, SM1/SW1/SK1 x3)
wrote SW1_a1..a3 and SK1_a1..a3 off as not_run within one second at t=+21 min: the
first refusal (spent 40.79 + reserved 40 + needed 20 > 100) set a run-wide halt while the
blocker was the two SM1 reservations still in flight, not the cap — once they settled
SW1 x3 and at least one SK1 would have fit.

RunBudget.admit now asks two questions per attempt. "Can never fit" (spent + reservation
> cap) refuses THIS attempt only and records it not_run; a later attempt with a smaller
reservation is asked on its own. "Cannot fit yet" (fits the cap but not the reservations
in flight) waits on the ledger (threading.Condition on the existing lock; settle()
notifies all) and re-asks after every settle, re-reading the durable spend each time, so
a waiter can end refused. A waiter only waits while something is in flight, so the lane
that holds the reservation always wakes it. The lane pool is unchanged: a waiting attempt
idles its lane thread. The wait is logged once per wait with the numbers and shown in the
watcher state; the refusal line keeps its format and names the wait when there was one.

Durable facts: cap_usd, spent_usd, reserved_usd, reservation_usd, unknown_cost_rows and
attempts_not_run stay; every admission fact now carries waited_sec (the not_run row's
budget facts included). The snapshot's global halt is replaced by per-refusal facts:
`refusals` (attempt, reason, at, the facts at refusal) and `first_refused` (the first
attempt's name, for the report). There is no `halted` flag any more — the run never halts,
so keeping the key would be a false fact; stop_reason=budget_cap is written when at least
one attempt was refused. The module and class docstrings state the new rule.

Tests: the two tests that pinned the old semantics
(test_run_budget_reservation_rule_halts_new_attempts_at_the_cap and
test_run_wide_cap_stops_new_attempts_and_records_not_run_rows) are rewritten to the new
contract — the unit test drives a waiter thread that is re-asked on the first settle, admitted
on the second, and a second waiter that is refused after its wait with waited_sec recorded;
the micro/fractional pins now show the cap-filling attempt waiting (then admitted when the
blockers settle) instead of being refused; the manifest-level test runs SM1,SK1,SW1 so that
SW1_a1 runs AFTER the SK1 refusals, pinning the not_run rows, refusals, first_refused, the
absence of `halted` and stop_reason.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:34:02 +00:00
Ouroboros
6dddf31a09 Live E2E stand: the SK1 echo line is bounded on the final text and is one line
The plugin capped the message and then added the "echo: " prefix (up to
206 characters) and did not normalize line breaks, so "one bounded line,
at most 200 characters" was false and the test pinned the 206 (codex M3
finding). Line breaks now collapse to spaces and the cap applies to the
final prefixed text; the pins assert the cap on the final text and the
single-line shape.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 10:06:36 +00:00
Ouroboros
594f0a66e8 devtools: measure the FULL scope input with the real assembler (P4)
The measurer's scope arm called only the private touched-section renderer
and reported it as `scope_touched`, so its scope figure was one fragment of
the reviewer's input: it excluded the byte-stable prefix (scope checklist +
the five canonical docs in full), the intent scaffolding, the staged diff
and the generated repo atlas with the owed-in-full class. It also resolved
paths with `git diff --cached --name-only` and passed `deleted_paths=[]`,
so a staged deletion was packed as a current path that resolves to nothing
and the deleted-file HEAD content the real pack inlines was never counted.

The scope arm now runs `scope_review_pack._build_scope_prompt` with a
`_ScopePromptContext` (governance from the measured checkout,
`drive_root=None`: write-free, no obligations read) and reports
`scope_full` — the prompt split at the assembler's own stable-prefix
boundary, the ladder facts from the scope context manifest (atlas status,
selected/tracked counts, ladder steps, unassembled REQUIRED artifacts), the
scope input cap and the headroom under it; the touched section stays as the
labelled sub-number `scope_full.scope_touched`. Staged entries come from
`--name-status` through the assembler's own parser and the `D` entries ride
the deleted-paths channel. The tool stays offline: the assembler resolves
its cap through `scope_window` (a lazy provider-metadata fetch), so for the
duration of the build its call-time seams `_effective_scope_input_limit` /
`_scope_window` are bound to a cap computed by that helper's own formula
(`_scope_input_limit`: window-scaled reserves, density-calibrated cap read
from the evidence store, REVIEW_PROMPT_TOKEN_BUDGET ceiling) on the
cache-only window or on `--scope-window N`, disclosed as such in the report.

Measured on this worktree with the change staged (o200k from the local
cache): stable prefix 863,088 chars / 188,680 o200k; at the 920k cap
(`--scope-window 2000000`) the full input is 3,379,355 chars / 844,839
chars-4 / 774,334 o200k, of which the touched section is 32,980 chars. At
the cold-density default cap (545,454 chars/4, no evidence) the pack does
NOT assemble on this checkout (`fixed_overflow`, 42 required artifacts
unassembled) — the measurer now prints that refusal instead of a number.

Tests: the synthetic checkout gains an "Intent / Scope Review Checklist"
section (the real assembler fails closed without it); new pins for the
full-input split, the ladder facts, the offline seams (every fetch seam
raises), the write-free data root, the explicit-window cap formula, the
staged-deletion split and the printed labels. No existing assertion changed.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 9a24f5ddff51aad9c55bf2ecd5e36381863de8d5)
2026-09-05 09:52:35 +00:00
Ouroboros
5ad8019dd4 e2e_live: cap the nested preflight xdist fan-out per lane (W4-preflight-workers)
The commit gate's hermetic pytest pass runs inside each lane server and
resolves `-n auto` to os.cpu_count() (128 on this host) with no ceiling; the
2026-09-04 paid run started >= 104 xdist workers per self-mod lane with three
lanes overlapping. The runtime already has the lever
(OUROBOROS_PREFLIGHT_TEST_WORKERS, floor 2, read by
preflight_runner._preflight_worker_count and scrubbed from the candidate
suite), but IsolatedServer's settings-authoritative sweep dropped it before
the lane server started, so the stand could not use it.

Devtools-only wiring (evidence action A):
- server_runner: keep OUROBOROS_PREFLIGHT_TEST_WORKERS through the
  authoritative sweep as an operational host-load lever (comment reworded;
  it is not a model, credential or settings key).
- run_live_lanes: derive max(2, 16 // lanes) at argument time (shared-host
  rule: at most 16 pytest workers across the stand), set it in the launcher
  process before the first lane starts so an ambient shell value never wins,
  and record it as extra.preflight_test_workers in run_manifest.json and as
  preflight_test_workers in every lane row.
- docs/DEVELOPMENT.md: one sentence in the live E2E stand section.

Pins: the lane sees the computed value (ambient 128 overridden), the runtime's
_preflight_worker_count reads it, the manifest and lane row record it, the
floor and key names match preflight_runner's constants, and IsolatedServer
forwards the key while still stripping an ambient OUROBOROS_MODEL. No
runtime code changed.

Size ratchet: tests/test_e2e_live_runner.py enters the 1001-1500 band with this
commit (996 -> 1042 lines); the band rationale is recorded in the manifest.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 1824f8910509c65b9fc94041c7a76f98a9aa617d)
2026-09-05 09:52:35 +00:00
Ouroboros
b7672d2823 Live E2E stand: make the SK1 probe fixture honest about inject_chat
The SK1 scenario declared permissions ["tool", "inject_chat"] over a plugin
that only registered an echo tool, with prose denying any host access. The
skill review refused it 3/3 in the first paid run on exactly the two checklist
items that exist to catch that (permissions_honesty, inject_chat_minimization);
the reviewers were right and the scenario could never go green as written.

The fixture now performs what it declares while keeping the scenario's
purpose (the owner grants exactly ONE privileged permission, enables, and
dispatches):

- the echo tool relays its reply as one bounded line (<= 200 chars) into the
  owner's own Main chat through the loopback Host Service /chat/inject route,
  token in the request header only (SkillToken.use_in_request()), destination
  a module constant (never a tool argument), sender left unidentified, env
  proxies ignored, no retry, no listener; a Host Service refusal surfaces as a
  tool error instead of a silent success;
- SKILL.md states that narrow user-facing purpose and declares "net" for the
  single loopback request per call (checklist item 2 names urllib as a net
  user); "tool"/"net" are loader-enforced declarations, the only owner GRANT
  stays inject_chat, so grants_exactly_requested keeps its meaning;
- acceptance adds dispatch_relayed_one_line_per_call_into_owner_chat: the
  durable chat.jsonl row the host stamped source=skill:e2e_live_probe under
  chat_id 1 with the exact echo text, one per successful call;
- the stub script gains a second closing final because the relayed line opens
  one owner-chat turn on the same wire.

Pins: the manifest's permissions map one-to-one to plugin source that
exercises them, requested_skill_permissions == SK1_GRANTS, the token appears
at exactly one request site, the tool schema exposes no chat_id, the prose
names the route/chat and no longer denies host access; and the plugin text is
executed against a loopback sink (202 path: body/header/truncation, exactly
one POST per call; 403 path: raises). The stub-script shape pin extends from
index 4 to the full role queue because the queue grew by one final.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 0eff1c70f8665b962cb28a5ac34dc262b0b4f0d0)
2026-09-05 09:43:36 +00:00
Ouroboros
69554dc1c1 Scrub every projected settings key from the hermetic tests preflight
`_preflight_env` dropped only OUROBOROS_*, GH_*, PYTEST_* and secret-suffixed
keys, while `config.apply_settings_to_env` projects two dozen further settings
keys into the server process by bare name: provider base URLs
(OPENAI_COMPATIBLE_BASE_URL, OPENAI_BASE_URL, GIGACHAT_*, ...), USE_LOCAL_*,
LOCAL_MODEL_*, MCP_ENABLED/MCP_TOOL_TIMEOUT_SEC, GITHUB_REPO, TOTAL_BUDGET.
The nested candidate suite inherited them, so its verdict depended on the
operator's install profile: with OPENAI_COMPATIBLE_BASE_URL set,
tests/test_settings_effort.py lost four test_get_review_models_* tests; with
USE_LOCAL_MAIN a different four. The scrub now also drops every key in
`settings_env_keys()` — derived, so a new settings key is covered without a
hand-kept list — before the disposable data/settings/repo triple is
re-injected. The existing scrubs are unchanged.

Pin: tests/test_preflight_runner.py parametrizes over `settings_env_keys()`
and asserts each key is absent from `_preflight_env(...)`, next to the
PYTEST_* wholesale-drop pin.

The live E2E stub rehearsal no longer passes `skip_tests` to
`commit_reviewed`: that flag was the documented residual of exactly this leak
(the loopback stub lane projects OPENAI_COMPATIBLE_BASE_URL), and the $0
rehearsal now runs the hermetic suite as its tests preflight like the paid
prompt. tests/test_e2e_live_runner.py now pins the stub commit shape; its old
docstring said the stub's skip_tests was "deliberately NOT pinned" only
because the residual was open, which no longer holds. DEVELOPMENT.md's Live
E2E stand paragraph and the preflight contributor rule describe the closed
residual instead of documenting it.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 95c77a1fa0a0838329ed9b6cd91e3547b7c590ef)
2026-09-05 09:43:36 +00:00
Ouroboros
7726fca3ff Fix the rc.11 tag CI: Windows and macOS full-test
Four classes, none Linux-visible (the exact-SHA battery on this host was
green):

- Windows has no os.fchmod: the live stand's write_settings raised on
  five tests. The 0600-before-content write keeps fchmod where it exists
  and falls back to chmod after the write elsewhere (the shape the
  system_e2e harness already uses).
- The orphan scan reads /proc environ (Linux only): run_lane's finally
  raised FileNotFoundError on macOS and Windows. Without procfs the lane
  records a typed fact (orphan_scan=unavailable:no_procfs) and no check —
  never a passed check that did not run.
- supervisor/evolution_lifecycle._write_evolution_campaign treated a
  campaign file that exists but cannot be read as "no campaign" and let
  a stale write win (windows-latest: test_stale_campaign_cannot_overwrite_
  a_new_campaign — a transient read failure is enough). Present-but-
  unreadable now refuses the write with a warning; an absent file stays
  writable.
- macos-latest: the password resolution pin answered '' with settings
  patched to a password — the module wrapper had been replaced on that
  xdist worker by a started-and-never-stopped patch from an earlier
  module. The resolution order is now a pure function
  (resolve_network_password(env_value, settings_loader)) that the
  wrapper calls with os.environ and load_settings; the pin exercises the
  pure function and cannot be reached by either polluter class (a leaked
  environment writer or a leaked patch), and a second test pins the
  wrapper through the resolver seam.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 02:54:15 +00:00
Ouroboros
091825f7d6 Live E2E stand: one effective ceiling per attempt, orphan scan pinned
The delta-4 review found the floor and the rounding applied only in
`ceiling()`: admission reserved the raw `per_task_usd x root_tasks`
while the lane received `max(0.01, round(x, 4))`, so micro or fractional
inputs (cap $0.005 with $0.001 per task; $0.01006) handed out more
ceiling than was admitted. `reservation()` is now the ONE effective
number — floored at LANE_BUDGET_FLOOR_USD, never rounded upward — and
admission, the stored reservation, the lane's TOTAL_BUDGET and the
reports all carry it; `ceiling()` returns exactly what admission
reserved. The invariant is worded precisely everywhere: settled spend +
in-flight ceilings <= cap (a live lane's spend is already inside its
ceiling).

The orphan-after-stop flip lives in `_apply_orphan_scan` and is pinned
in result.json and result_index.jsonl (fail + reason_code=checks_failed,
never an empty reason). Tests add the micro (5 x 0.01 fill a 0.05 cap,
the 6th refused; a floored reservation above a sub-cent cap is refused)
and fractional (two exact quarters fill a half-dollar cap) boundaries.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 02:20:28 +00:00
Ouroboros
932ad7c118 Live E2E stand: disjoint lane ceilings, orphan reason_code, honest stub pins
The delta-3 review of the runner found the run-wide cap still not
enforced under concurrency: a lane's TOTAL_BUDGET was the cap minus the
OTHER lanes' reservations, so two lanes admitted against a $20 cap with
$8 reservations each received a $12 ceiling — $24 of authority for a $20
run. A lane's ceiling is now its OWN reservation (per_task_usd x root
tasks): immutable, disjoint from every other lane's, and by the admission
rule the ceilings in flight plus the spend never exceed the cap. The
reservation rule string, the runner docstrings and DEVELOPMENT say so.

- An orphan left after stop flips a passing lane to fail; the row now
  carries reason_code=checks_failed instead of an empty reason in the
  index.
- tests: the SK1 echo pin was `A and B or True` (always true, and B was
  false: the message travels as a tool argument, not in the plugin
  source) — replaced by direct equalities; the stub end-to-end docstring
  no longer claims the tests preflight runs (it is the disclosed
  residual); the budget tests pin the disjoint ceilings, their sum under
  the cap, the not-admitted floor and the micro-reservation floor.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 02:09:34 +00:00
Ouroboros
010afd969b Live E2E stand: run-wide budget, detached seed, confirmed absorb, per-task checks
The first paid run of the stand exposed the runner, not the product: the
run cap was copied into every lane, the seed was the operator's live
worktree (staggered lanes cloned different HEADs, one of them dirty),
SK1 overwrote its author checks with the dispatch ones and accepted a
stamped generation digest as a successful call, SM1 could not land
because web/onboarding.css mirrors --accent by value, --self-mod could
pass with no restart, the key sat in the run-root template, and the
watcher's key probe turned two transport hiccups into alarm lines.

- --total-budget is a RUN-WIDE cap (RunBudget): spend is re-read from
  the lanes' durable llm_usage rows through the harness oracle, each
  attempt reserves --per-task-usd x root tasks (SK1 = 2), scheduling
  halts at the first refusal with not_run rows (reason_code budget_cap),
  each lane's TOTAL_BUDGET is the headroom left at its start, the
  watcher prints the running total and the manifest records cap, spend,
  reservation rule and stop reason; money and interval arguments must
  be finite and positive (--watch-interval >= 5 s).
- The seed is a clean DETACHED clone of --seed (a ref of --source-repo)
  materialized once under the run root; every lane asserts its clone is
  at the admitted sha and clean; the source's dirtiness is disclosed.
- --self-mod requires a CONFIRMED absorb per lane (pre-task snapshot of
  clone HEAD/served sha/uptime/absorbed cycles; afterwards the counter
  advanced, the sha moved, uptime reset, server ready) and a run-level
  gate over every self-mod lane.
- SK1: wait_task namespaces checks per task (author_/dispatch_),
  LaneContext.check refuses a duplicate key, the dispatch counts only on
  a tools.jsonl row with status ok and the extension's exact echo.
- SM1: the accent change lands in both web/style.css and
  web/onboarding.css (prompt, stub, acceptance parity check); the stub
  rehearsal still skips the tests preflight (documented residual: the
  loopback base URL leaks into the hermetic suite), and the typed refusal trail
  (ledger block_reasons, PREFLIGHT_BLOCKED / TESTS_PREFLIGHT_BLOCKED /
  SCOPE_REVIEW_BLOCKED tool codes, terminal reason_code) is a fact.
- The run-root effective_settings.json is redacted; the key reaches
  disk only in each lane's 0600 settings file and is disclosed by
  fingerprint as the runtime grant.
- Lane infra failures carry a typed refusal {type, code, message} and a
  reason_code in result.json and result_index.jsonl.
- The key probe runs on its own thread with an 8 s HTTP bound, at most
  once a minute, backing off on failure; a failed probe is informational
  and never delays a tick.

DEVELOPMENT "Live E2E stand" describes the new contracts and the
--per-task-usd sizing guidance from the first paid run.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit d333d573b01de8f7ba2ecc01f4eded634c7332db)
2026-09-05 01:58:58 +00:00
Ouroboros
8b2d1e6b6d rc.11 delta review: api-pack rows sized as production filters them, the advisory preview slot on its resolved route, a checked one-change invariant in the measurer, W4-F4 notes mirror the register
- devtools/measure_review_pack.py: the quorum limit and headroom are sized over
  the rows that RECEIVE the api pack — api_chat rows without a configured-
  subagent binding, decided by review_execution.delivery_retrieves exactly as
  review._prepare_unified_review decides before fit_triad_prompt; an all-
  retrieving panel reports 'no API pack is assembled for this panel' instead
  of a number (codex MAJOR). The 'index IS the working tree' comment became a
  checked invariant: an unstaged edit or an untracked file is the typed
  MeasuredCheckoutDirty refusal (exit 2), and the advisory arms are re-checked
  to resolve the same path set (fable MINOR).
- ouroboros/tools/claude_advisory_review.py: the native advisory slot is built
  by the dispatch builder (reviewer_slot_config.reviewer_slots), so use_local
  comes off the resolved route and a LOCAL advisory model previews its
  MANDATORY READ bound on its own window with the typed shortfall disclosure,
  instead of the remote/unknown route's (codex MAJOR); file stays at 1492 lines.
- ADOPTION_v7next.md / scripts/v7next_adoption.py: the Notes no longer call
  W4-F4 operator-disclosed; a 'Deferral authorities:' declaration mirrors
  DEFERRED_OUT_OF_V70 and the validator refuses a declaration that disagrees
  with the register (codex MINOR).
- docs/archive/v7next/LEDGER_CORRECTIONS.md: one-row note that the immutable
  band rationale for claude_advisory_review.py cites 1434 lines while the file
  stands at 1492 (fable minor).

Tests: tests/test_measure_review_pack.py (3 new, red-first), tests/test_advisory_route_pack.py (1 new, red-first), tests/test_v7next_adoption.py (2 new).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit eca294f4520802f2f3b64eaaa7e236d19899b4a5)
2026-09-04 23:55:32 +00:00
Ouroboros
92a8b3f1b1 Live E2E stand: devtools/e2e_live runs K staggered isolated servers over the SM1/SW1/SK1 scenario table (F4-A)
run_live_lanes.py admits through the benchmark family's seams (dirty seed refused with the
refusal persisted; the key by NAME from the environment, never a pool file; the credit
preflight takes min(key limit remaining, account credits) via the new
manifests.openrouter_account_credits, the second bound only), writes the effective settings
from the tree's own defaults with TOTAL_BUDGET/OUROBOROS_PER_TASK_COST_USD as settings keys,
names the model in the run manifest from the applied file, and fans out --lanes (default 4,
max 6) isolated servers 2-3 s apart. scenarios.py is the table: SM1 lands a web/style.css
token through commit_reviewed under advanced+blocking (S2 set + computed style read from the
committed CSS after a restart), SW1 arms Swarm in the browser (force_plan, roster, >=2
children with causal lineage, fanout receipt, cost rollup, /proc no-orphans), SK1 has the
model author a skill then reviews, grants, enables, dispatches and deletes it; acceptance is
callable over durable artifacts only. ui_probe.py resolves the suite's PlaywrightUIClient when
it carries the surface, else headless Chromium, else a typed ui_unavailable. --stub rehearses
every scenario for $0 on the loopback stub of tests/system_e2e/harness.py (stub_lane.py routes
the swarm wire by role). Per-lane result.json carries checks, settings sha256 plus a
secret-free config digest, seed describe, pre/post HEAD and the diff digest, grants by
fingerprint and the runtime terminal disclosure; a watcher prints lane states, free disk on /
and /mnt/data and the key headroom. Tests pin the launcher gate by source, the table shape,
lane/stagger bounds, the TMPDIR guard, the dirty-seed and credential refusals, the two-plane
credit floor, manifest-model-equals-applied-file, secret-free artifacts, and a gated stub
rehearsal of SM1 on a real server. DEVELOPMENT and ARCHITECTURE describe the stand.

(cherry picked from commit ab36206a2a7ccbad01bf6b25a69181a6a69d9aa6)
2026-09-04 23:37:49 +00:00
Ouroboros
eb7941d914 review economics: the span-only release-carrier cut with the same disclosure in the scope and advisory packs (F3 Q4 = A)
Owner decision (question 4 = A): the version-carrier cut F3-A applied to the
commit triad's touched pack applies to the scope review pack and to the
advisory pre-review's changed-context pack too, each over the pair it
reviews and through its own seam — never a foreign predicate bolted on.

review_file_pack: span_only_release_carriers is the ONE predicate the three
packs share (carrier_only_change over the release_sync carrier SSOT; VERSION
in the change; HEAD->index by default, HEAD->working tree with worktree=True)
and pack_exclusion_note / CARRIER_CUT_REASON its one disclosure;
triad_pack_exclusions now composes them (same two classes, same note).
build_advisory_changed_context applies the cut over the pair the advisory
actually reviews — HEAD vs the working-tree text the pack reads, staged or
not — withholding a span-only carrier once through the builder's
exclude_paths marker + omitted list (the prompt's existing omission note) and
appending the PACK EXCLUSION NOTE inside the touched-files section; the
governance pointers are untouched and the native episode keeps read_file.

scope_review_pack: _carrier_span_only_paths cuts the same carriers over the
scope pack's own HEAD->index pair — no snapshot, named in the CURRENT FILE
CONTEXT DEDUPLICATION NOTE with the reason, declared diff-only to the atlas
with a by-design reason (ReviewContextAtlasRequest.diff_only_reasons: the
same already_included disposition, a truthful row; the required-artifact
escalation ignores the override) and recorded as the ladder's first typed
entry (carrier_span_only_omitted). Never for a managed subject (its delta is
M0->staged) and never for an artifact the atlas owes in full — a structural
guard, so the cut can never turn into a refusal. Canonical docs keep riding
the existing dedup class; the governance prefix is not trimmed anywhere.

Measured on the synthetic button-colour release commit (o200k): triad touched
pack 536,010 -> 56,244 (F3-A: 535,746 -> 56,237), scope touched section
429,848 -> 56,003, advisory changed-context pack 429,754 -> 56,144;
devtools/measure_review_pack.py now reports all three packs.

Docs: ARCHITECTURE rows for preflight_review_prompt/run, review_file_pack,
scope_review_pack and the guaranteed-fit paragraph. Tests: test_scope_review
(scope cut, manifest row, ladder entry, managed / owed-in-full refusals),
test_advisory_route_pack (live-tree pair, outside-span keep, no-VERSION
no-cut, prompt placement), test_review_context_atlas (reason override), the
facade extraction maps.

(cherry picked from commit dd68a89be7cc2194aead8828b96c04170dd7e9fd)
2026-09-04 23:26:37 +00:00
Ouroboros
4e82bc57b1 measure_review_pack: offline by construction, exact zero-diff headroom, one checkout
Closes the second-absorption review findings on the F3-A pack meter:
- reviewer windows are read cache-only through capability_evidence.probe(
  allow_fetch=False); an unknown window is disclosed and sized at the fit
  ladder's default instead of probed, so the meter never fetches provider
  metadata nor writes capability evidence (proved byte-identical output
  with and without network, no state/capability_evidence.json created);
- the o200k tokenizer never downloads: a cache miss is a typed
  TokenizerUnavailable disclosed in the report;
- headroom is derived from the exact zero-diff serialized message
  (constitutional preamble + BIBLE, stable prefix, dynamic scaffolding with
  an empty pack, the fixed user turn) in the fit ladder's units, with the
  parts the ladder itself does not count named separately;
- --repo controls every corpus read (BIBLE, checklist + archive, docs).
tests/test_measure_review_pack.py pins the cache-only seams, the tokenizer
refusal, the headroom formula and the --repo selection.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 23:24:29 +00:00
Ouroboros
95dea461d3 review economics: cut span-only release carriers and prefix-duplicated governance docs from the triad touched pack; type the advisory native transcript-bound end (D-06 a/d)
Triad touched pack (D-06a): build_touched_file_pack takes the advisory
seam's exclude_paths (set[str]) and marks a withheld path once through its
own omission marker + the caller's OMISSION NOTE; triad_pack_exclusions
names the two disclosed classes — release carriers on a VERSION-staged
commit whose HEAD->staged change sits inside their declared version spans
(release_sync carrier SSOT: the span primitive substitute_carrier_spans is
lifted out of supervisor/update_carriers.py and carrier_only_change is built
on it, no second table), and governance docs byte-identical to the copy the
prompt's governance prefix already inlines — and returns the PACK EXCLUSION
NOTE the call site appends. A managed subject keeps every full text.
Measured on a synthetic button-colour release commit: touched pack 535,746
-> 56,237 o200k tokens; diff headroom on the default panel 27,807 -> 363,070
(chars/4). devtools/measure_review_pack.py is the offline, $0 recipe.

Advisory (D-06d): a native episode exception keeps failure_custody() as the
result usage and the exception's structured code; the preflight caller
classifies native_transcript_cap_exceeded on that code (never message text,
not the provider window vocabulary) into the typed non-blocking
ADVISORY_SKIPPED: native_transcript_bound_exceeded carrying the bound, the
refused chars and the paid rounds, instead of a generic ADVISORY_ERROR with
an empty {}. Pinned: the window-derived bound is applied on the advisory's
native episode with the advisory's own model.

Docs: ARCHITECTURE rows for release_sync, claude_advisory_review and
review_file_pack. Tests extended in test_update_carriers, test_scope_review,
test_git_review_enforcement and test_advisory_route_pack.

(cherry picked from commit 742b0ebab2045e1be1bc7214de81d9a94f07d435)
2026-09-04 22:53:33 +00:00
Ouroboros
62f87cc94c Merge upstream ouroboros db6d7cf8 into the v7 line: absorb PR #609 net-resilience and PR #614 update letter
Second absorption of the frozen upstream line (23ab428f..db6d7cf8: 89
commits, 47 files) on top of the rc.10 hotfix tip, by the F2 rules (S1
upstream body in the owning leaf, S2 hand-merge, S3 only with proof;
retired 7.0 surfaces never return):

- PR #609 net-resilience: interactive transport-wait episodes bounded by
  the task idle timeout with the typed task_incident/toast_once pair, a
  bounded paid repeat after a typed post-dispatch transport death with a
  round-keyed record that fences every other send, the shutdown-aware
  supervisor crash counter and bounded lifespan join, Darwin keepalive
  tuning. loop_transport/transport_custody/net_transport/loop_llm_call
  land verbatim (same shapes on both sides); the loop.py deltas are
  relocated into the v7 leaves (loop_round_limits, loop_model_call,
  loop_delivery, loop_forced_finalization, loop_nudges, loop_messages,
  loop_budget) with bodies AST-equal to upstream modulo the call-time
  handles; _emit_overflow_retry_skipped stays a public helper (the v7
  facade contract) and upstream's nested _skipped delegates to it.
- PR #614 update letter: ouroboros/update_letter.py and its web module
  land verbatim; the new OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC key and
  get_update_letter_timeout_sec live in their v7 owners
  (settings_defaults.py, runtime_limits.py, re-exported by config.py);
  _supervisor_stop lives in server_process.py beside the restart events;
  docs/PERSISTENCE.md gains the state/update_letter.json row and the
  inventory pin moves to 286.
- Tests: the relocated run_llm_loop tests take the emit_progress
  incident keyword (every one-argument progress fake in tests/ swept, a
  gap upstream itself left in test_tree_cost_ceiling); the official-update
  runtime-section test lands in tests/test_context.py; _MOVED_OWNERS
  registers the relocated getter.
- Docs: ARCHITECTURE/DEVELOPMENT hunks land on the upstream text; the two
  legacy timeout rows upstream's context still carries stay retired (7.0).
- Size ratchet regenerated; the band rationale for tests/test_update_letter.py
  is carried verbatim from upstream.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 22:53:15 +00:00
Ouroboros
cfccf12c06 tests: the TB methodology tests read the shipped triad defaults from their SSOT
devtools/benchmarks/terminal_bench/test_run_tb_methodology.py (upstream body,
absorbed by F2) looked the shipped triad up through
SETTINGS_DEFAULTS["OUROBOROS_REVIEW_MODELS"]; ABI 7.0 retired that settings
key, so the rc.9 tag CI job benchmark-methodology failed with KeyError in two
tests. The launcher itself derives the same list from
ouroboros.settings_defaults.OPENROUTER_REVIEW_DEFAULTS["triad"] (the SSOT the
retired key was joined from), so the tests now read that list. Test-only;
no runtime change.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 21:49:28 +00:00
Ouroboros
9698e2e077 Merge upstream ouroboros 23ab428f into the v7 line: absorb 407 commits into the module split
Second parent is the frozen upstream `ouroboros` head (23ab428f, 407 commits
since the merge base a76961de); first parent is v7.0.0-rc.8 (18b9832e).

Every upstream change lands in v7's owning leaf: S1 transplants keep upstream's
bodies (comments verbatim) under the call-time handle idiom, S2 hand-merges keep
both intents, S3 keeps v7 only with proof (retired 7.0 ABI surfaces, superseded
mechanisms). Per-symbol relocation ledger: docs/archive/v7next/LEDGER_CORRECTIONS.md
(F2 absorption section). Provisional decisions awaiting owner ratification:
D-18 (two-destination symbols), D-19 (acceptance rows follow upstream R2),
D-20 (acceptance_dialogue stays deleted), D-21 (tools/registry.py: facade
import block only).

Docs: upstream ARCHITECTURE/DEVELOPMENT as the base with compact v7 deltas;
bookkeeping moved to docs/archive/v7next. Size-ratchet manifest, domain
manifest and generated inventories regenerated; new leaves: tools/write_shape
walker, gateway/cost_breakdown, tools/core_secret_paths; provider_catalogs.py
and acceptance_dialogue.py removed (v7 owners).
2026-09-04 19:32:55 +00:00
Ouroboros
952c39db12 Benchmark disclosures: repeats live inside the call's attempt budget
The continual-learning runbook and the SWE-bench Pro methodology said
the transport-death repeats sit on top of the transient burst. The
primary dispatch grants a repeat only while the outer attempt budget
has room, so both now say: one outer attempt budget per call, within
which up to three rows can be typed transport-death failures (the
first death plus at most two repeats).
2026-09-04 14:17:57 +00:00
Ouroboros
6802a822a3 Merge managed/ouroboros (5b4546ea) into the network-resilience branch
Five conflicts, resolved by reading both sides in full:

- loop.py: the trace-touched skill-name scan moved into skill_readiness.py
  upstream; its import replaces this branch's inline copy, whose removeprefix
  form was a byte fold of the same behaviour.
- loop.py: the no-resend terminal keeps upstream's block (its short comment and
  its live_trace re-read); this branch's own two-stamp fold is re-applied on top,
  and the record-fenced source stays.
- loop_llm_call.py: _send_main_candidate binds whenever a physical context OR a
  candidate predicate is present, upstream's semantics, expressed through the
  binding this file already uses. The context manager is only constructed there,
  never entered, so the effect order is upstream's.
- loop_llm_call.py: call_llm_with_retry keeps both new parameters,
  transport_death_retries and initial_messages, on the two lines the file's
  line cap already pays for.
- ARCHITECTURE.md: upstream's producer-word sentence, with this branch's
  no-resend source clause re-applied, so the paragraph states both.

loop.py carries shrink-only byte debt and upstream absorbed both places where
this branch had paid for its own additions, so the payment is made again inside
the functions this branch owns: the stamp fold above, the fifth site of the fit
key tuple now calls _fit_key, and _emit_overflow_retry_skipped is folded into
_skipped, its only remaining caller since this branch collapsed the other three.
No comment, docstring, diagnostic or test was shortened. 271928 -> 271855 bytes.

The manifest is regenerated on the merged tree; it is upstream's manifest with
loop.py's exact byte count. Upstream's three new band rationales are carried
across verbatim because the generator reads the pre-merge committed manifest.
2026-09-04 08:41:52 +00:00
Ouroboros
77b9df3ed5 Merge managed/ouroboros (b9d39ac7) into the catlife route fixes
Target drift since the sprint base (PR #557-#591: the agentic-review synthesis
moved the acceptance machinery whole into acceptance_dialogue.py, three-delivery
rows, Claudexor 3.9.7, ibl fixes). Resolutions: the acceptance-packet changes
(children debt, dialogue history and the packet budget passed INTO the bounded
builder; per-slot input caps on the panel request; packet sizing from the same
triad delivery rows the panel dispatches) are carried into the relocated module;
a partial tool-result projection withholds packet rows only for a genuinely
unavailable source (the legacy truthy sentinel still refuses) and never
retrieving rows; a report-shaped native episode keeps its draft on a budget
refusal after the failed send is observed; ARCHITECTURE/DEVELOPMENT merge both
sides' rows and paragraphs; the upstream acceptance-delivery test now asserts
the documented `not_dispatched` refusal shape (a refusal is a transport state,
never a verdict).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 08:58:47 +03:00
Ouroboros
5da2a81acd Merge managed/ouroboros (f437408b) into the network-resilience branch
The PR target moved from 85c1e386 to f437408b (PRs #565, #590: the task
acceptance surface moved into acceptance_dialogue.py, deep self-review runs
on the configured reviewer row, the finalization nudges fold into one
_inject helper, and the project last-task-result reader grew in server.py).

One file conflicted: ouroboros/size_ratchet_manifest.py. Upstream shrank
ouroboros/loop.py to 272,905 bytes while this branch held 284,335, so the
BYTE_DEBT row was the only unmerged hunk. Resolved by regenerating the
manifest on the merged tree: BYTE_DEBT["ouroboros/loop.py"] = 272,805, which
is the merged file's exact size (upstream's 272,905 minus this branch's own
100-byte shrink) and below both parents' debts. Upstream's six new 1001-1500
band rationales and its shrunken tests/test_devtools_benchmarks.py debt ride
along unchanged.

Every other shared file auto-merged on disjoint regions: the transport-death
and wait-episode seams in loop.py, the incident= keyword on
OuroborosAgent._emit_progress, the supervisor stop event and lifespan
teardown in server.py, and the transport paragraphs in ARCHITECTURE and
DEVELOPMENT are all preserved beside the upstream hunks.
2026-09-04 04:22:14 +00:00
Ouroboros
eeddd95e48 benchmarks: scope the transport-death error-count numeral to its own rail
"Up to three llm_api_error rows for one logical call" read as an absolute
bound; the transient burst shares the same attempt loop, so the numeral
now names the transport-death rail alone, on top of that burst.
2026-09-04 02:01:49 +00:00
Ouroboros
60a17a0a53 loop: bounded paid repeat after a typed post-dispatch transport death
A dispatched request whose socket died with a typed transport death (httpx
ReadError/WriteError/RemoteProtocolError or the requests ProtocolError/
RemoteDisconnected shape, via the explicit __cause__ chain; never a timeout,
status/body error, pre-dispatch failure, local provider or loopback route)
is provider_outcome_unknown and stays billed at its upper bound. The PRIMARY
main-loop round dispatch alone may now send it again at most twice per round,
each repeat a NEW execute_physical_attempt with its own ledger row, with a
4 s then 8 s deadline-aware backoff and a counter keyed by the round so the
wait episode's free redial cannot re-arm it. The retry decision is made
before the durable llm_api_error row is written, so retry_same_request is
truthful on every attempt and llm_non_retryable_same_request marks only the
exhaustion. Forced-final, fallback candidates, review, safety, probes, web
search, consolidation and Background Consciousness keep the default 0; the
global classifier is unchanged; BudgetExceeded propagates untouched.

Ratchet paydown in the same files: loop_llm_call.py folds the four copies of
the error-event identity block into one, drops the redundant kwargs that
_handle_main_llm_call_exception/_stop_after_llm_error received beside the
context that already carries them, and folds five consecutive pops into a
loop; loop.py hoists two function-local imports of an already-imported
module and reuses the direct_chat value it had already computed.
2026-09-04 02:00:40 +00:00
Ouroboros
c50c2c944c Merge managed/ouroboros (85c1e386) into the agentic-review synthesis branch
Integrates the moved target (193 commits over the sprint base b9bcc2da,
release 6.114.0 — the DeepSeek landing PR #563 and the docs-consolidation
PR #556 included) into the synthesis branch without rewriting history.
Four files conflicted; each was resolved by substance so BOTH sides' facts
and behaviours survive:

- ouroboros/size_ratchet_manifest.py (generated): UNION of both sides'
  BAND_PATHS rationales (the target's ouroboros/gateway/extensions.py,
  tests/test_provider_contract_ci.py, tests/test_ui_smoke_project_continuity.py
  and web/modules/settings_ui.js beside our acceptance_dialogue.py,
  deep_self_review.py and test-suite rows), then regenerated for the merged
  tree: ouroboros/gateway/history.py left the band on the target line,
  BYTE_DEBT is the live merged size everywhere (loop.py 272905,
  tests/test_devtools_benchmarks.py 328068 — below both parents — and
  web/modules/chat.js 206949). `scripts/regenerate_size_ratchet.py --check`
  is green; no new module, function or byte debt.

- docs/ARCHITECTURE.md and docs/DEVELOPMENT.md: the target's consolidated
  structure is the frame; every agentic-review fact of this branch is placed
  in the target's section or row — the three deliveries of task acceptance
  and deep self-review, the reviewer-row schema (`deep_review` singleton,
  roster references, `profile_id`), pacing simplified to the one admission
  floor (R52/R55; the EWMA sentences are gone), poll purity, the native
  read receipts and BIBLE.md coverage, the retrieving work order and the
  CI methodology job. The module-tree rows keep the target's condensed form
  extended with our contracts; the full contracts live once in §6 (Task
  lifecycle, Review delivery, Deep self-review). Stale target sentences
  that our side retired (the `api_chat` acceptance pin, the round cap, the
  legacy/default API panel residual) are replaced, never duplicated.

- ouroboros/review_native_episode.py: our side already measures the send
  bound as the wire size of the serialized message list, which carries the
  WHOLE assistant dict — `reasoning_content` included — so the target's
  fix (count replayed reasoning in the fail-closed bound, 295c9062) is
  subsumed; the target's regression test passes unchanged. The comment
  above the append records the invariant.

Auto-merged both-changed files were checked for silent overlap:
tests/conftest.py gained the same autouse os.environ snapshot/restore
fixture on both sides — the target's tested `_os_environ_isolation` is
kept and our redundant `_restore_process_environment_between_tests` is
dropped (its rationale folded into the surviving docstring); our
gateway-settings binding restore fixture stays. .github/workflows/ci.yml,
ouroboros/config.py, ouroboros/tools/control.py, web/modules/settings_ui.js
and web/modules/reviewer_slots.js carry both sides' changes exactly once
(our deep self-review block sits in the target's new `.reviewer-slots-group`
container like its advisory sibling).
2026-09-03 22:07:20 +00:00
Ouroboros
ee49f9ff4a terminal-bench: warn loudly at admission when the triad carries rows the container cannot run
The session rows a Terminal-Bench container cannot run were disclosed on the
run manifest and as a metadata comment — artifacts an operator reads after
the run, when the money is spent and the acceptance seat has already
degraded on every task. The owner decided (R40, 2026-09-02) that admission
says it once, loudly: when the configured triad carries agent-session rows,
the launcher prints one stderr warning naming each row (`codex=…`,
`cursor=…`), that a container has no harness CLI/daemon or credentials, that
the rows are not declared as used models and their acceptance seat degrades
typed, and how to configure a submittable run. The run continues
(`command_generated`); the manifest field and the metadata comment stay. The
methodology test asserts the warning text appears exactly once beside the
manifest and metadata facts, and an api-only panel admits silently with an
empty disclosure list.

The environment fixture's docstring no longer claims to mirror conftest's
`_scrub_inherited_subagent_selection`: the key sets differ on purpose —
conftest drops the subagent roster, account pin and structured panel; this
suite reads both panel forms and drops the structured panel and the legacy
comma-list keys — and it now says so.
2026-09-03 19:34:20 +00:00
Ouroboros
279a21f3e3 terminal-bench: scrub the shell's panel keys before each methodology test; name what the manifest carries
The TB methodology suite lives outside tests/ and so outside the conftest
scrub that keeps every other test independent of the operator's shell. Its
environment fixture restored what a test wrote but still let an inherited
`OUROBOROS_REVIEWER_SLOTS` or legacy comma-list key reach a panel-reading
test (`test_metadata_omits_web_search_when_web_disabled` read the shell's
panel). The fixture now also drops those keys before each test, mirroring
`_scrub_inherited_subagent_selection`; the gate run for this commit exported
a poisoned shell panel and comma lists on purpose and stayed green.

METHODOLOGY states precisely what the manifest carries per delivery: on a
fixed-model run `harness.fixed_model_actor.reviewer_slots` holds `slot_id`,
`route{kind, target_id}` and `effort` per row (a fixed-model panel is always
direct api rows; no subagent binding exists there); a plain `--model` run
records only the session rows, and its api-vs-native split survives only
when the panel is persisted in the forwarded host settings — the container
adapter resolves the environment first, so a panel supplied only through the
operator's environment leaves no durable per-delivery record.
2026-09-03 19:34:07 +00:00
Ouroboros
5a6c772464 acceptance: trim the slot label robustly, never carry a malformed manifest; TB methodology names the per-delivery record
Two small hardening points on the retrieving work order. The slot-label
trim compared the renderer's raw tail against `Slot: <id>`; a future
trailing newline in the renderer would have left the label in place and the
executor — which labels the slot itself — would have emitted two `Slot:`
lines. The tail is stripped before the comparison, and the work-order
contract test pins that a work order carries no `Slot:` line of its own.
The native packet projection normalized a malformed `omissions_manifest`
only when it also had something to omit; a packet with nothing to omit
travelled with the malformed value as-is. The manifest is normalized
whenever the source value is present and not a list, and an absent
manifest is still never invented.

Terminal-Bench METHODOLOGY states what `metadata.yaml` cannot say: an api
packet row and a configured-subagent native inspection row are both
`commit_review_triad` by model id and dedupe onto the measured model under
the container's one-model roster. The per-delivery record is the run
manifest — `harness.fixed_model_actor.reviewer_slots` on a fixed-model run,
`extra.triad_rows_not_executable_in_container` for the session rows — and
for a plain `--model` run the api-vs-native split is read from the forwarded
host settings the container adapter used.
2026-09-03 19:34:01 +00:00
Ouroboros
61612a18dc tests: restore the process environment between tests; pin the container disclosure end to end
`run_tb.apply_all_model` — and `main --all-model` through it — writes the
fixed-model contract into `os.environ` directly; that is the launcher's real
behaviour and stays. But a test exercising it handed the written
`OUROBOROS_REVIEWER_SLOTS` and forwarded slot keys to every later test of
the same xdist worker (`monkeypatch.delenv(raising=False)` records nothing
for a key that did not exist, so nothing removed it afterwards), which is
the `benchmark-scope-1` contamination that made scope-review identity tests
and the legacy metadata test fail depending on worker order — on the base
commit as much as here. The class is closed where the other between-tests
hygiene lives, `tests/conftest.py`: every test runs on a snapshot of the
process environment restored afterwards whatever it wrote, and the inherited
reviewer panel is scrubbed with the actor list and account pin, so a test
that pins the legacy comma-list branch (the metadata test in
test_devtools_benchmarks.py, a shrink-only byte-debt module that therefore
stays untouched) never reads the operator's shell. The TB methodology suite
lives outside tests/ and carries the same snapshot fixture itself.

The container-provenance disclosure is now pinned end to end through
`main` (command generation, no harbor): `run_manifest.json` carries
`extra.triad_rows_not_executable_in_container` with the session rows'
targets verbatim and in row order — including a target with its own `/`
(`cursor=openai/gpt-5`) — and `metadata.yaml` carries the same list as a
comment while declaring no session row as a model. The gate run for this
commit exported a poisoned shell `OUROBOROS_REVIEWER_SLOTS` on purpose and
stayed green.
2026-09-03 19:33:57 +00:00
Ouroboros
5a4f3521d2 terminal-bench: disclose session triad rows the container cannot run, never declare them
The previous commit declared every configured triad row in metadata.yaml,
including agent-session rows under a `commit_review_triad_agent_session`
role. A Terminal-Bench task container cannot run such a row: the image has
no harness CLI or daemon, the forwarded-env allowlist carries no harness
credentials, and the container secret policy forbids them — so declaring the
row as a model the run used reinstated the exact "declared but never run"
class the base docstring forbade. The stronger provenance clause wins:
metadata names only what the container executes.

`_effective_helper_models` now declares api packet rows and
configured-subagent native inspection rows only (decided by the typed
`row.is_session`, not a role-string substring), and restores the base
docstring's "never a declared-but-never-run model" clause. Session rows are
carried by the new typed disclosure `triad_rows_not_executable_in_container`
(their `harness[=model]` targets in row order) on `run_manifest.json`, and as
a comment line in `metadata.yaml` — a comment, not a key, because the
leaderboard schema owns that file's keys. Both read the panel through one
`_container_triad` seam (operator env, else host settings, parsed under the
container's one-model roster). The `"agent_session" in role` provider branch
goes with the class, which also retires the `cursor=openai/gpt-5` split
hazard: no session target is ever rendered as a model row. The methodology
test asserts the disclosure, not a declaration, for env, all-session and
settings-file panels; METHODOLOGY states the rule and that an all-session
triad's acceptance seat degrades typed inside the container.
2026-09-03 19:33:44 +00:00
Ouroboros
6a5c9b8270 acceptance: route-owned work order for retrieving rows and route-aware gates
With the triad rows reaching the acceptance panel as configured, a retrieving
row (a configured-subagent native inspection episode, or an agent session)
still had nothing to run: the acceptance request carried no `session_task`,
so both retrieving executors refused typed, and the gates around the panel
assumed every row was a packet row — the wave budget gate priced a
subscription session as API money, the partial-projection refusal turned a
row that reads the exact source away, and the packet rows' format-repair
resend would have bought a second episode.

Owner decisions R1/R4/R5/R15/R23 (2026-09-01) settle the work order.
`acceptance_dialogue.acceptance_retrieving_work_order` writes one per
retrieving row onto the new `ReviewRequest.slot_session_tasks` (per-slot,
falling back to the shared `session_task`, consumed by both retrieving
executors): the same task-stable contract the packet rows render, the same
output contract — `review_execution.review_output_contract` is now the ONE
governance text, rendered into the api pack's byte-stable segment (bytes
unchanged, pinned by the golden digest) and handed to retrieving rows as
`policy["output_contract"]`, so they never fall back to the generic object
form — absolute retrieval pointers over the task's ACTIVE workspace
(`review_repo_dirs_for`'s subject root, never the governance repo), and the
packet in the form the delivery can use. A session row gets the FULL packet
(its run is unobserved by the host, so the packet is its only attested view)
plus the disclosure that access outside the workspace is not guaranteed and a
refused read is absence of evidence, not of the artifact. A native row gets
the packet WITHOUT its freely degradable tail — the trajectory rows and
artifact previews the api ladder spends first, manifested as
`retrieving_delivery` omissions — plus the real data root
(`policy["native_data_root"]`, R5), because its episode reads those sources
itself. The FULL packet stays the `evidence_refs` authority on every
delivery; route-owned policy keys are filtered from the rendered Policy JSON
so the api pack states the contract once. The owner deadline rides the
request (R23).

The gates are route-aware: the wave budget gate prices API money only (a
packet row by its real message pair, a native row as one episode send of its
work order, a session row not at all); the partial-projection refusal spares
retrieving rows while the immutable-core overflow refuses every delivery; the
format-repair resend is packet-row only — a retrieving row's executor
canonicalizes its own answer. The wallet stamp is unchanged and fires on
every delivery through the same captured stamp (R11), pinned by an
api/native/mixed/session matrix that also proves a spent wallet refuses a new
paid identity before any send. The trap test lands first: a retrieving row
citing a real `verification_receipts[0]` resolves clean against the full
packet, with the fabricated sibling ref disclosed.

Docs: ARCHITECTURE (acceptance section, module rows, review delivery, the
governance matrix row for the three deliveries), DEVELOPMENT (wallet stamp
paragraph route-agnostic; acceptance checklist item), GAIA/OSWorld
METHODOLOGY acceptance-axis comparability notes.
2026-09-03 19:32:19 +00:00