Commit graph

23 commits

Author SHA1 Message Date
Ouroboros
d70f93c3c1 fix(e2e_live): SK1 verdict is the product review gate, clean state a fact
The rc.15 SK1-only rerun on 560f7d71 (three attempts, 2026-09-05) passed
every lifecycle check three times — preflight, review job, auto-grant equal
to the request, enable to a live-loaded dispatchable skill, a physical
extension call echoed into the owner chat, cleanup — while the stand's own
"all-PASS review" criterion passed once: the model-authored payload came
back clean, then with warnings (a SKILL.md/plugin.py contradiction), then
with blockers (manifest without runtime). The criterion measured the
author model, not the product.

Owner decision (2026-09-06, question 3 = A): SK1 counts when the review is
executable by the PRODUCT gate — the /api/extensions row's
executable_review (clean, warnings, or blockers under advisory enforcement
by operator choice, skill_review_gate) — plus the full lifecycle. The
check is now review_executable via sk1_review_gate(); the review status,
enforcement, blocking reason, non-PASS items and the clean/all-PASS state
travel as recorded facts. DEVELOPMENT.md says so; a parametrized test pins
the rule on the three rerun outcomes and the failure shapes (blocking
enforcement, failed review call, no findings, no gate fact).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 00:42:09 +00:00
Ouroboros
c6f002c122 fix(e2e_live): --self-mod seeds the owner id only, never a campaign
The rc.15 paid run2 (2026-09-05, SM1_a1/a2) showed the stand's --self-mod
lanes pre-seeding the benchmark helper's ACTIVE campaign ("benchmark
evidence" objective, evolution_mode_enabled) instead of relying on the
post-task promotion path the flag is meant to exercise:

- generic cycles ran from t=0 next to the scenario task (three no_op
  cycles, $16 on one lane), each counted as a consecutive failure;
- the promotion the scenario task wrote (post_task_evolution_request.json)
  was refused by apply_pending_request because evolution was already
  enabled, and the kept request file also blocked wait_for_absorb's
  no_promotion exit — every SM1 lane waited the full task timeout with
  absorbed_cycles_done at 0 although its reviewed commit had landed.

seed_owner_state(data_root, evolution_enabled=False) under every profile:
--self-mod stays a settings fact (OUROBOROS_POST_TASK_EVOLUTION + cadence
every_n:1) and the scenario task's own promotion enables the one-shot
campaign whose cycle lands, restarts and absorbs. The reservation rule
keeps +1 root (the one post-task cycle); its wording, the RunBudget
docstring and the DEVELOPMENT.md stand section say so, and a runner test
pins the seeded state (owner_chat_id only, no evolution_campaign.json,
no evolution_mode_enabled) with and without --self-mod.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 23:25:35 +00:00
Ouroboros
df9c140ade Live E2E stand: the SK1 probe plugin presents the host token to loopback only
The skill review blocked the stand's own probe plugin twice on the rc.15
paid stand (host_token_handling, CHECKLISTS skill item 12): it built the
Host Service base from HOST_SERVICE_URL unvalidated, so the token could
be sent to any host. The plugin now refuses a base that is not a loopback
http URL before any request is built; the echo behaviour is unchanged. A
small test module drives the plugin text the model is told to write.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 21:16:12 +00:00
Ouroboros
ed8ecd216f Scope the live stand's absorb wait and check to scenarios that land a commit
Under --self-mod the runner applied the post-task absorb wait and the
self_mod_absorb_confirmed check to every lane, but only SM1 commits
anything: SW1 and SK1 failed by construction and waited for the evolution
campaign to promote nothing (rc.15 paid stand, 2026-09-05: SK1_a1 passed
twelve of thirteen lifecycle checks, the thirteenth being a real reviewer
finding on the model-authored plugin, then waited about 27 minutes until
the evolution cycle ended and the wait returned no_promotion).

Scenario gains expects_absorb (True for SM1 only). The absorb snapshot,
the wait in restart() and after the scenario, and the lane check follow
that flag; the run-level gate lists only absorbing lanes and the manifest
names how many lanes were expected to absorb. A lane that expects no
absorb stops its server right after the scenario and records
self_mod_absorb: {"expected": false}. OUROBOROS_POST_TASK_EVOLUTION and
the per_task x (roots + self_mod) reservation stay as they were: the
campaign may still run and spend during the scenario; the stand just
does not wait for it.

Pins: the scenario table flag, SM1 waiting and carrying the check while
SW1/SK1 finish with no confirm_absorb call and evolution still on in
their applied settings, the run-level gate and manifest counters. The
suite stays at its 1500-line band ceiling by folding its two fake servers
into one and the argument-refusal cases into a loop; the runner stays at
the 1000-line limit. Handbook: the scoping and the incident.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 21:04:08 +00:00
Ouroboros
e1036afe08 Live E2E stand: FIFO admission by dispatch index, exact feasibility pins, no largest-first driver
Re-sorting the pool ordered submission only: RunBudget.admit was not FIFO, and
on the real runner the freed lane's next job takes the lock before the woken
waiter every time (300/300: settle in run_attempt's finally, return, the
executor's next job, admit on the same thread), so a round that overflows the
cap (owner configuration 150 + 100 + 100 = 350 > 300) let later attempts
leapfrog the waiter and could leave SW1 at 0/3 under pessimistic spends.

admit now takes the job's dispatch index and keeps the attempts still asking
in a line under the existing Condition: an attempt reserves only when no
earlier-dispatched attempt is still asking and the cap has room beside the
reservations in flight; one that can never fit is refused at once and leaves
the line; reserve, refusal and settle wake every waiter and the line predicate
re-parks the later ones. No deadlock: a waiter waits only while something is in
flight or an earlier attempt is asking, the earliest never waits on the line,
and with nothing in flight it either fits or is refused. The wait log line
names its reason (an earlier attempt vs reservations in flight). main() passes
the index from the new dispatch_order (round-robin by attempt, largest
reservation first within a round), which the test driver now shares.

Tests: the driver parks the ledger's wait and releases threads in the runner's
observed order (newcomer asks first, then the line in dispatch order), so the
owner configuration (cap 300 / per-task 50 / attempts 3 / pass-of 2 / 3 lanes /
--self-mod) is pinned exactly: realistic spends admit all nine in dispatch
order for $159; pessimistic spends admit SK1_a1, SM1_a1, SW1_a1, SK1_a2,
SM1_a2, SW1_a2, SM1_a3 and refuse SK1_a3 then SW1_a3 for $211 (every scenario
keeps two). The largest-first driver option and its comparison pin (red 4/15)
are gone; a RunBudget unit pin covers the line (a later attempt that would fit
waits behind an earlier one; the refused head frees it). Handbook and
docstrings name the waiter-vs-newcomer race, the head-of-line cost, the exact
traces (largest-first under the same admission: SW1 0/3 at $225) and that
round_worst_case_usd sums one attempt per scenario.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 17:22:44 +00:00
Ouroboros
80d49bd7ca Live E2E stand: round-robin dispatch by attempt, no bench budget profile, per-round preflight worst case
The verdict is pass-of PER scenario, so the dispatch order now maximises the
minimum admitted attempts per scenario: a1 of every scenario, then a2, then a3,
largest reservation first within a round (requested_task_ids keep the argument
order). At the owner's configuration (cap 300, per-task 50, attempts 3, pass-of 2,
3 lanes, --self-mod) with pessimistic spends, largest-first left SW1 with at most
one admitted attempt; round-robin keeps every scenario at two or more.

The cost_hard_stop_pct=100 projection is dropped (scenarios.STAND_BUDGET_PROFILE,
LaneContext.submit metadata, its pins and handbook text): with a reservation of
at least 2 x per-task the product's per-task axis already binds, and a bench
profile would change the pacing path under test. Every stand root runs under the
product's default ceiling, including the evolution root and SW1's UI-path root.

budget_preflight takes --lanes and records the per-round worst case (the lanes
largest reservations of one attempt per scenario at once); the handbook states
that outcomes among equal waiters depend on wake order, so a feasibility number
is a range at the margin. Docstrings now say rc.14 showed up to two evolution
cycles per lane (t=0 and post-task) and that the lane TOTAL_BUDGET is the fence.

Tests: the deterministic driver over the real RunBudget takes the dispatch order
and pins the owner configuration under realistic and pessimistic spends for
round-robin (admitted set, per-scenario counts, spend range) and shows the
rejected largest-first order leaving SW1 below pass-of; the run-wide cap manifest
pin is re-parametrised for round-robin (cap 16, pass-of 2); configuration A/B pins
and the projection pin are removed within the band ceiling.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 17:22:43 +00:00
Ouroboros
ebc713c379 Live E2E stand: reserve per root (evolution counted), project cost_hard_stop_pct=100, budget preflight, largest-first dispatch
The rc.14 artifacts show the 2x reservation of 90f02b02 was covering the SECOND
ROOT, not the 50% ceiling: with --self-mod the evolution cycle starts at t=0
next to the scenario task and shares the lane fence (SM1_a1: task $3.84 +
evolution $12.41 + $2.84 = $19.08 of $20). The audit driver over the real
RunBudget then showed that under the planned rc.15 flags (cap 200, per-task 50,
SK1 = 2 x 50 x 2 = the whole cap) SK1 could be admitted at most once in any
order, and never in runner order, after SM1/SW1 had already spent.

- RunBudget.reservation = max(floor, per_task x (root_tasks + int(self_mod))):
  the evolution root is counted as a root when --self-mod; HARD_STOP_INVERSE
  and the ouroboros.task_pacing import are gone; RESERVATION_RULE, the
  --per-task-usd help, the docstrings and the handbook state the root-count
  rule.
- The in-task halving that factor stood in for is switched off on the roots
  the stand submits itself: LaneContext.submit projects
  metadata.budget_profile.cost_hard_stop_pct=100 (scenarios.STAND_BUDGET_PROFILE,
  the seam ProgramBench uses, in the shape normalize_budget_profile accepts).
  The evolution root keeps the product default (its task metadata carries only
  the transaction; no seam) and so does SW1's root on the UI path
  (ui_probe.send_chat is a WS chat frame without metadata; acceptable: SW1
  spends ~$8 against a ceiling of >= $25 at per-task $50) - both stated in the
  handbook.
- budget_preflight (the credit preflight's typed shape, recorded as
  extra.budget_preflight): a reservation row per scenario, the worst case
  sum(reservation x attempts) against the cap, and a fail-closed refusal
  {stage: budget_preflight, reason: reservation_unreachable} (exit 3, before
  the key, the seed or any lane) when a reservation exceeds the cap or equals
  it with attempts >= 2. No override flag: the operator changes the flags.
- Jobs are dispatched largest reservation first (stable among equals); the
  manifest's requested_task_ids keep the argument order.

Tests: the ledger pins keep every admission number (per-task $8 is the $8 unit
now); the reservation pin asserts per-task x roots, +1 with self_mod, and that
the exact number reaches the lane's settings.json; a submit pin captures the
POST body and resolves the product ceiling with and without the projection
($47 vs $25 for per-task $50 in a $50 lane); preflight pins cover the direct
function, the manifest refusal and the dispatch order; a deterministic driver
over the real RunBudget (3 lanes, SCENARIOS.root_tasks, stub spends visible at
settle, waiting detected by instrumenting the ledger's wait) pins
configuration A (cap 200 / per-task 50 / attempts 2: realistic spends all six
admitted, $106; pessimistic spends five admitted, exactly one SW1 attempt
refused after waiting, $158) and configuration B (cap 260 / attempts 3:
realistic all nine admitted, $159; pessimistic six admitted, every SW1 refused,
$225). The run-wide cap pin is re-parametrised for the root-count rule and the
largest-first order (SK1_a2 refused, SM1 still runs after that refusal).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 17:22:43 +00:00
Ouroboros
e5f61f5852 Live E2E stand: the orphan cmdline pin is procfs-gated; the lazy-open note names SW1
The orphan-detail pin asserted the survivor's cmdline text, which is read
from /proc and is the typed empty string on macOS and Windows (the module
runs on the 3-OS matrix). The text assertion is now gated on
PROCFS_AVAILABLE; the shape and cap pins stay unconditional. The handbook
says the use-time open applies to SM1 (SW1 drives the task through the UI
and holds its browser), and UIProbe.close carries the thread-affinity
invariant a cross-thread close would silently violate.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:57:14 +00:00
Ouroboros
800159039d e2e_live: open the UI probe at use time, type a dead browser target, name orphan survivors
The v7.0.0-rc.14 paid run opened each lane's headless Chromium at lane start
and first used it after the task and the self-mod absorb wait; lanes that
waited 10-23 minutes met TargetClosedError on the first goto, the exception
escaped run_sm1 and the whole lane became infra_error although its task-side
checks were complete. Both lanes also recorded no_orphans_after_stop=false
without naming the survivors.

- ui_probe.GuardedUI wraps the client: any browser-call failure becomes the
  typed reason ui_unavailable:<ExceptionType> (recorded as ui_reason and in the
  lane facts), later calls are no-ops, the exception never reaches the scenario.
- LaneContext.ui is a lazy property: the client opens on first use against the
  current server; restart() closes any open probe and re-resolves on the next
  use (a non-self-mod restart may change the port). The scenario surface
  (ctx.ui truthiness, goto/computed_property/send_chat/screenshot) is unchanged.
- run_live_lanes: lane start only probes availability (open + close); the lane
  row's ui verdict is taken from the context at the end; the orphan scan hands
  the surviving pids to _apply_orphan_scan, which names them in row["orphans"]
  (pid + cmdline head, capped at 20 with orphans_omitted).
- tests pin the lazy open / restart reopen, the typed open failure, the closed
  target degrading the SM1 UI tail to checks_failed at both context and lane
  level, and the orphans row shape and cap.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:57:14 +00:00
Ouroboros
db805c5b73 Live E2E stand: SM1 pins that both sheets are in the commit and skips taken versions
Two reviewer findings on the realistic SM1: (1) the diff-scope check could
fail a commit the scope reviewer legitimately widened (another accent
touchpoint such as the unlock page's inline colour or the site stylesheet),
so the stand now pins only that both stylesheets are IN the landed commit
and records the companions as a fact; the reviewers own the scope judgment.
(2) A seed cloned from an older ref carries the newer tags and the review
binding refuses a staged version whose tag exists, so the stub's next
version skips every version whose tag the clone already has.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:52:21 +00:00
Ouroboros
e0d81f4ec1 Live E2E stand: SM1 asks for a design-system-consistent accent change landed as a reviewed release
The first paid stand run on v7.0.0-rc.14 failed all three SM1 attempts on the
commit gate, and the reviewers were right: the prompt changed only the
--accent value (a mixed blue/red palette against docs/DESIGN.md "Brand
accent"), forbade the version bump (BIBLE P9, CHECKLISTS item 8) and routed
around the advisory preflight, and claimed a visible UI change with no
vision-inspected evidence (development_compliance 2(i)). The stand must
describe the realistic task and let the product's policy shape the work, not
fight it.

SM1 now asks for the brand accent changed consistently with the design system
(the one --accent value, its named roles and the alpha ladder) in both
stylesheets, verified on a real consumer flow in the browser, and landed as a
reviewed release through preflight_review and commit_reviewed with the
synchronized carriers bumped in the same diff - no skip flags, no "do not
bump". The acceptance pins only the observable contract: the commit landed,
both committed sheets carry the new accent and keep their whole shared :root
token set in parity, VERSION is a strictly-greater release version whose
carriers pass the product's release admission gate read at the landed commit,
the diff is confined to the stylesheets, the carriers the release_sync SSOT
names, docs/DESIGN.md and comment-only web CSS companions, the advisory
ledger row is a real reviewer run rather than a skip/bypass row, and the
browser/vision tool rows are recorded as a vision_evidence_present fact for
the reviewers to judge. The stub rehearsal projects the bump offline through
sync_release_metadata plus the README history row (trimmed to the P9 limits
the gate enforces) and takes the same full path without skip flags. Docs and
the hermetic pins follow the new contract.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:52:21 +00:00
Ouroboros
90f02b0228 e2e_live: reserve 2x per-task per root task (inverse of the product's hard stop)
The rc.14 paid run (per-task $20) failed lane SM1_a3 with budget_exhausted at
$10.21 accounted while its cap was $20: the product's in-task ceiling is
min(cost_hard_stop_pct of the GLOBAL remaining at task start, per-task cap -
planning margin), and in a lane the global remaining IS the lane budget = the
1x reservation, so every root task's ceiling collapsed to 50% of its cap. On a
real install the global budget is far larger and the root-cap axis binds.
Second, --self-mod runs the post-task evolution cycle as another root task
under the same per-task fence (SM1_a3: task $10.2 + cycle ~$8 of $20), so a 1x
reservation under-reserves the lane's legitimate spend.

The reservation is now max(LANE_BUDGET_FLOOR_USD, HARD_STOP_INVERSE x
per_task_usd x root_tasks), where HARD_STOP_INVERSE = 100 / the product's
_DEFAULT_COST_HARD_STOP_PCT imported from ouroboros.task_pacing (2 today; it
follows the product's default rather than copying the number). The lane's
TOTAL_BUDGET stays equal to its reservation, so the ceilings in flight remain
disjoint and settled spend + in-flight ceilings <= cap. RESERVATION_RULE (the
manifest's reservation_rule), the docstrings, the --per-task-usd help and the
handbook paragraph state the rule and both reasons.

Tests: the ledger pins use per-task $4 so the $8 unit and every admission
number are unchanged; a new EQUALITY pin drives run_lane up to the written
settings file and asserts that per-task $20 with one root reserves $40 and that
exactly $40 reaches the lane's settings.json as TOTAL_BUDGET (never the
template's $100 run cap), and that the factor equals 100 / the product's
default cost_hard_stop_pct.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:39:34 +00:00
Ouroboros
2f6d6e59c2 Live E2E stand: the handbook states the wait-then-refuse admission rule
docs/DEVELOPMENT.md still described the retired contract ("halts scheduling
at the first refusal"); it now states the per-attempt rule the runner
implements (wait while blocked only by reservations in flight, refuse only
what can never fit, manifest keys refusals/first_refused). The RunBudget
docstring names the non-FIFO fairness policy and the on_wait-under-lock
contract the reviewers asked to pin.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:34:02 +00:00
Ouroboros
88509e2cea Live E2E stand: budget admission waits on in-flight reservations, refuses per attempt
The first paid run (v7.0.0-rc.14, cap $100, per-task $20, 3 lanes, SM1/SW1/SK1 x3)
wrote SW1_a1..a3 and SK1_a1..a3 off as not_run within one second at t=+21 min: the
first refusal (spent 40.79 + reserved 40 + needed 20 > 100) set a run-wide halt while the
blocker was the two SM1 reservations still in flight, not the cap — once they settled
SW1 x3 and at least one SK1 would have fit.

RunBudget.admit now asks two questions per attempt. "Can never fit" (spent + reservation
> cap) refuses THIS attempt only and records it not_run; a later attempt with a smaller
reservation is asked on its own. "Cannot fit yet" (fits the cap but not the reservations
in flight) waits on the ledger (threading.Condition on the existing lock; settle()
notifies all) and re-asks after every settle, re-reading the durable spend each time, so
a waiter can end refused. A waiter only waits while something is in flight, so the lane
that holds the reservation always wakes it. The lane pool is unchanged: a waiting attempt
idles its lane thread. The wait is logged once per wait with the numbers and shown in the
watcher state; the refusal line keeps its format and names the wait when there was one.

Durable facts: cap_usd, spent_usd, reserved_usd, reservation_usd, unknown_cost_rows and
attempts_not_run stay; every admission fact now carries waited_sec (the not_run row's
budget facts included). The snapshot's global halt is replaced by per-refusal facts:
`refusals` (attempt, reason, at, the facts at refusal) and `first_refused` (the first
attempt's name, for the report). There is no `halted` flag any more — the run never halts,
so keeping the key would be a false fact; stop_reason=budget_cap is written when at least
one attempt was refused. The module and class docstrings state the new rule.

Tests: the two tests that pinned the old semantics
(test_run_budget_reservation_rule_halts_new_attempts_at_the_cap and
test_run_wide_cap_stops_new_attempts_and_records_not_run_rows) are rewritten to the new
contract — the unit test drives a waiter thread that is re-asked on the first settle, admitted
on the second, and a second waiter that is refused after its wait with waited_sec recorded;
the micro/fractional pins now show the cap-filling attempt waiting (then admitted when the
blockers settle) instead of being refused; the manifest-level test runs SM1,SK1,SW1 so that
SW1_a1 runs AFTER the SK1 refusals, pinning the not_run rows, refusals, first_refused, the
absence of `halted` and stop_reason.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 14:34:02 +00:00
Ouroboros
6dddf31a09 Live E2E stand: the SK1 echo line is bounded on the final text and is one line
The plugin capped the message and then added the "echo: " prefix (up to
206 characters) and did not normalize line breaks, so "one bounded line,
at most 200 characters" was false and the test pinned the 206 (codex M3
finding). Line breaks now collapse to spaces and the cap applies to the
final prefixed text; the pins assert the cap on the final text and the
single-line shape.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 10:06:36 +00:00
Ouroboros
5ad8019dd4 e2e_live: cap the nested preflight xdist fan-out per lane (W4-preflight-workers)
The commit gate's hermetic pytest pass runs inside each lane server and
resolves `-n auto` to os.cpu_count() (128 on this host) with no ceiling; the
2026-09-04 paid run started >= 104 xdist workers per self-mod lane with three
lanes overlapping. The runtime already has the lever
(OUROBOROS_PREFLIGHT_TEST_WORKERS, floor 2, read by
preflight_runner._preflight_worker_count and scrubbed from the candidate
suite), but IsolatedServer's settings-authoritative sweep dropped it before
the lane server started, so the stand could not use it.

Devtools-only wiring (evidence action A):
- server_runner: keep OUROBOROS_PREFLIGHT_TEST_WORKERS through the
  authoritative sweep as an operational host-load lever (comment reworded;
  it is not a model, credential or settings key).
- run_live_lanes: derive max(2, 16 // lanes) at argument time (shared-host
  rule: at most 16 pytest workers across the stand), set it in the launcher
  process before the first lane starts so an ambient shell value never wins,
  and record it as extra.preflight_test_workers in run_manifest.json and as
  preflight_test_workers in every lane row.
- docs/DEVELOPMENT.md: one sentence in the live E2E stand section.

Pins: the lane sees the computed value (ambient 128 overridden), the runtime's
_preflight_worker_count reads it, the manifest and lane row record it, the
floor and key names match preflight_runner's constants, and IsolatedServer
forwards the key while still stripping an ambient OUROBOROS_MODEL. No
runtime code changed.

Size ratchet: tests/test_e2e_live_runner.py enters the 1001-1500 band with this
commit (996 -> 1042 lines); the band rationale is recorded in the manifest.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 1824f8910509c65b9fc94041c7a76f98a9aa617d)
2026-09-05 09:52:35 +00:00
Ouroboros
b7672d2823 Live E2E stand: make the SK1 probe fixture honest about inject_chat
The SK1 scenario declared permissions ["tool", "inject_chat"] over a plugin
that only registered an echo tool, with prose denying any host access. The
skill review refused it 3/3 in the first paid run on exactly the two checklist
items that exist to catch that (permissions_honesty, inject_chat_minimization);
the reviewers were right and the scenario could never go green as written.

The fixture now performs what it declares while keeping the scenario's
purpose (the owner grants exactly ONE privileged permission, enables, and
dispatches):

- the echo tool relays its reply as one bounded line (<= 200 chars) into the
  owner's own Main chat through the loopback Host Service /chat/inject route,
  token in the request header only (SkillToken.use_in_request()), destination
  a module constant (never a tool argument), sender left unidentified, env
  proxies ignored, no retry, no listener; a Host Service refusal surfaces as a
  tool error instead of a silent success;
- SKILL.md states that narrow user-facing purpose and declares "net" for the
  single loopback request per call (checklist item 2 names urllib as a net
  user); "tool"/"net" are loader-enforced declarations, the only owner GRANT
  stays inject_chat, so grants_exactly_requested keeps its meaning;
- acceptance adds dispatch_relayed_one_line_per_call_into_owner_chat: the
  durable chat.jsonl row the host stamped source=skill:e2e_live_probe under
  chat_id 1 with the exact echo text, one per successful call;
- the stub script gains a second closing final because the relayed line opens
  one owner-chat turn on the same wire.

Pins: the manifest's permissions map one-to-one to plugin source that
exercises them, requested_skill_permissions == SK1_GRANTS, the token appears
at exactly one request site, the tool schema exposes no chat_id, the prose
names the route/chat and no longer denies host access; and the plugin text is
executed against a loopback sink (202 path: body/header/truncation, exactly
one POST per call; 403 path: raises). The stub-script shape pin extends from
index 4 to the full role queue because the queue grew by one final.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 0eff1c70f8665b962cb28a5ac34dc262b0b4f0d0)
2026-09-05 09:43:36 +00:00
Ouroboros
69554dc1c1 Scrub every projected settings key from the hermetic tests preflight
`_preflight_env` dropped only OUROBOROS_*, GH_*, PYTEST_* and secret-suffixed
keys, while `config.apply_settings_to_env` projects two dozen further settings
keys into the server process by bare name: provider base URLs
(OPENAI_COMPATIBLE_BASE_URL, OPENAI_BASE_URL, GIGACHAT_*, ...), USE_LOCAL_*,
LOCAL_MODEL_*, MCP_ENABLED/MCP_TOOL_TIMEOUT_SEC, GITHUB_REPO, TOTAL_BUDGET.
The nested candidate suite inherited them, so its verdict depended on the
operator's install profile: with OPENAI_COMPATIBLE_BASE_URL set,
tests/test_settings_effort.py lost four test_get_review_models_* tests; with
USE_LOCAL_MAIN a different four. The scrub now also drops every key in
`settings_env_keys()` — derived, so a new settings key is covered without a
hand-kept list — before the disposable data/settings/repo triple is
re-injected. The existing scrubs are unchanged.

Pin: tests/test_preflight_runner.py parametrizes over `settings_env_keys()`
and asserts each key is absent from `_preflight_env(...)`, next to the
PYTEST_* wholesale-drop pin.

The live E2E stub rehearsal no longer passes `skip_tests` to
`commit_reviewed`: that flag was the documented residual of exactly this leak
(the loopback stub lane projects OPENAI_COMPATIBLE_BASE_URL), and the $0
rehearsal now runs the hermetic suite as its tests preflight like the paid
prompt. tests/test_e2e_live_runner.py now pins the stub commit shape; its old
docstring said the stub's skip_tests was "deliberately NOT pinned" only
because the residual was open, which no longer holds. DEVELOPMENT.md's Live
E2E stand paragraph and the preflight contributor rule describe the closed
residual instead of documenting it.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 95c77a1fa0a0838329ed9b6cd91e3547b7c590ef)
2026-09-05 09:43:36 +00:00
Ouroboros
7726fca3ff Fix the rc.11 tag CI: Windows and macOS full-test
Four classes, none Linux-visible (the exact-SHA battery on this host was
green):

- Windows has no os.fchmod: the live stand's write_settings raised on
  five tests. The 0600-before-content write keeps fchmod where it exists
  and falls back to chmod after the write elsewhere (the shape the
  system_e2e harness already uses).
- The orphan scan reads /proc environ (Linux only): run_lane's finally
  raised FileNotFoundError on macOS and Windows. Without procfs the lane
  records a typed fact (orphan_scan=unavailable:no_procfs) and no check —
  never a passed check that did not run.
- supervisor/evolution_lifecycle._write_evolution_campaign treated a
  campaign file that exists but cannot be read as "no campaign" and let
  a stale write win (windows-latest: test_stale_campaign_cannot_overwrite_
  a_new_campaign — a transient read failure is enough). Present-but-
  unreadable now refuses the write with a warning; an absent file stays
  writable.
- macos-latest: the password resolution pin answered '' with settings
  patched to a password — the module wrapper had been replaced on that
  xdist worker by a started-and-never-stopped patch from an earlier
  module. The resolution order is now a pure function
  (resolve_network_password(env_value, settings_loader)) that the
  wrapper calls with os.environ and load_settings; the pin exercises the
  pure function and cannot be reached by either polluter class (a leaked
  environment writer or a leaked patch), and a second test pins the
  wrapper through the resolver seam.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 02:54:15 +00:00
Ouroboros
091825f7d6 Live E2E stand: one effective ceiling per attempt, orphan scan pinned
The delta-4 review found the floor and the rounding applied only in
`ceiling()`: admission reserved the raw `per_task_usd x root_tasks`
while the lane received `max(0.01, round(x, 4))`, so micro or fractional
inputs (cap $0.005 with $0.001 per task; $0.01006) handed out more
ceiling than was admitted. `reservation()` is now the ONE effective
number — floored at LANE_BUDGET_FLOOR_USD, never rounded upward — and
admission, the stored reservation, the lane's TOTAL_BUDGET and the
reports all carry it; `ceiling()` returns exactly what admission
reserved. The invariant is worded precisely everywhere: settled spend +
in-flight ceilings <= cap (a live lane's spend is already inside its
ceiling).

The orphan-after-stop flip lives in `_apply_orphan_scan` and is pinned
in result.json and result_index.jsonl (fail + reason_code=checks_failed,
never an empty reason). Tests add the micro (5 x 0.01 fill a 0.05 cap,
the 6th refused; a floored reservation above a sub-cent cap is refused)
and fractional (two exact quarters fill a half-dollar cap) boundaries.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 02:20:28 +00:00
Ouroboros
932ad7c118 Live E2E stand: disjoint lane ceilings, orphan reason_code, honest stub pins
The delta-3 review of the runner found the run-wide cap still not
enforced under concurrency: a lane's TOTAL_BUDGET was the cap minus the
OTHER lanes' reservations, so two lanes admitted against a $20 cap with
$8 reservations each received a $12 ceiling — $24 of authority for a $20
run. A lane's ceiling is now its OWN reservation (per_task_usd x root
tasks): immutable, disjoint from every other lane's, and by the admission
rule the ceilings in flight plus the spend never exceed the cap. The
reservation rule string, the runner docstrings and DEVELOPMENT say so.

- An orphan left after stop flips a passing lane to fail; the row now
  carries reason_code=checks_failed instead of an empty reason in the
  index.
- tests: the SK1 echo pin was `A and B or True` (always true, and B was
  false: the message travels as a tool argument, not in the plugin
  source) — replaced by direct equalities; the stub end-to-end docstring
  no longer claims the tests preflight runs (it is the disclosed
  residual); the budget tests pin the disjoint ceilings, their sum under
  the cap, the not-admitted floor and the micro-reservation floor.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-05 02:09:34 +00:00
Ouroboros
010afd969b Live E2E stand: run-wide budget, detached seed, confirmed absorb, per-task checks
The first paid run of the stand exposed the runner, not the product: the
run cap was copied into every lane, the seed was the operator's live
worktree (staggered lanes cloned different HEADs, one of them dirty),
SK1 overwrote its author checks with the dispatch ones and accepted a
stamped generation digest as a successful call, SM1 could not land
because web/onboarding.css mirrors --accent by value, --self-mod could
pass with no restart, the key sat in the run-root template, and the
watcher's key probe turned two transport hiccups into alarm lines.

- --total-budget is a RUN-WIDE cap (RunBudget): spend is re-read from
  the lanes' durable llm_usage rows through the harness oracle, each
  attempt reserves --per-task-usd x root tasks (SK1 = 2), scheduling
  halts at the first refusal with not_run rows (reason_code budget_cap),
  each lane's TOTAL_BUDGET is the headroom left at its start, the
  watcher prints the running total and the manifest records cap, spend,
  reservation rule and stop reason; money and interval arguments must
  be finite and positive (--watch-interval >= 5 s).
- The seed is a clean DETACHED clone of --seed (a ref of --source-repo)
  materialized once under the run root; every lane asserts its clone is
  at the admitted sha and clean; the source's dirtiness is disclosed.
- --self-mod requires a CONFIRMED absorb per lane (pre-task snapshot of
  clone HEAD/served sha/uptime/absorbed cycles; afterwards the counter
  advanced, the sha moved, uptime reset, server ready) and a run-level
  gate over every self-mod lane.
- SK1: wait_task namespaces checks per task (author_/dispatch_),
  LaneContext.check refuses a duplicate key, the dispatch counts only on
  a tools.jsonl row with status ok and the extension's exact echo.
- SM1: the accent change lands in both web/style.css and
  web/onboarding.css (prompt, stub, acceptance parity check); the stub
  rehearsal still skips the tests preflight (documented residual: the
  loopback base URL leaks into the hermetic suite), and the typed refusal trail
  (ledger block_reasons, PREFLIGHT_BLOCKED / TESTS_PREFLIGHT_BLOCKED /
  SCOPE_REVIEW_BLOCKED tool codes, terminal reason_code) is a fact.
- The run-root effective_settings.json is redacted; the key reaches
  disk only in each lane's 0600 settings file and is disclosed by
  fingerprint as the runtime grant.
- Lane infra failures carry a typed refusal {type, code, message} and a
  reason_code in result.json and result_index.jsonl.
- The key probe runs on its own thread with an 8 s HTTP bound, at most
  once a minute, backing off on failure; a failed probe is informational
  and never delays a tick.

DEVELOPMENT "Live E2E stand" describes the new contracts and the
--per-task-usd sizing guidance from the first paid run.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit d333d573b01de8f7ba2ecc01f4eded634c7332db)
2026-09-05 01:58:58 +00:00
Ouroboros
92a8b3f1b1 Live E2E stand: devtools/e2e_live runs K staggered isolated servers over the SM1/SW1/SK1 scenario table (F4-A)
run_live_lanes.py admits through the benchmark family's seams (dirty seed refused with the
refusal persisted; the key by NAME from the environment, never a pool file; the credit
preflight takes min(key limit remaining, account credits) via the new
manifests.openrouter_account_credits, the second bound only), writes the effective settings
from the tree's own defaults with TOTAL_BUDGET/OUROBOROS_PER_TASK_COST_USD as settings keys,
names the model in the run manifest from the applied file, and fans out --lanes (default 4,
max 6) isolated servers 2-3 s apart. scenarios.py is the table: SM1 lands a web/style.css
token through commit_reviewed under advanced+blocking (S2 set + computed style read from the
committed CSS after a restart), SW1 arms Swarm in the browser (force_plan, roster, >=2
children with causal lineage, fanout receipt, cost rollup, /proc no-orphans), SK1 has the
model author a skill then reviews, grants, enables, dispatches and deletes it; acceptance is
callable over durable artifacts only. ui_probe.py resolves the suite's PlaywrightUIClient when
it carries the surface, else headless Chromium, else a typed ui_unavailable. --stub rehearses
every scenario for $0 on the loopback stub of tests/system_e2e/harness.py (stub_lane.py routes
the swarm wire by role). Per-lane result.json carries checks, settings sha256 plus a
secret-free config digest, seed describe, pre/post HEAD and the diff digest, grants by
fingerprint and the runtime terminal disclosure; a watcher prints lane states, free disk on /
and /mnt/data and the key headroom. Tests pin the launcher gate by source, the table shape,
lane/stagger bounds, the TMPDIR guard, the dirty-seed and credential refusals, the two-plane
credit floor, manifest-model-equals-applied-file, secret-free artifacts, and a gated stub
rehearsal of SM1 on a real server. DEVELOPMENT and ARCHITECTURE describe the stand.

(cherry picked from commit ab36206a2a7ccbad01bf6b25a69181a6a69d9aa6)
2026-09-04 23:37:49 +00:00