mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
The first paid run of the stand exposed the runner, not the product: the
run cap was copied into every lane, the seed was the operator's live
worktree (staggered lanes cloned different HEADs, one of them dirty),
SK1 overwrote its author checks with the dispatch ones and accepted a
stamped generation digest as a successful call, SM1 could not land
because web/onboarding.css mirrors --accent by value, --self-mod could
pass with no restart, the key sat in the run-root template, and the
watcher's key probe turned two transport hiccups into alarm lines.
- --total-budget is a RUN-WIDE cap (RunBudget): spend is re-read from
the lanes' durable llm_usage rows through the harness oracle, each
attempt reserves --per-task-usd x root tasks (SK1 = 2), scheduling
halts at the first refusal with not_run rows (reason_code budget_cap),
each lane's TOTAL_BUDGET is the headroom left at its start, the
watcher prints the running total and the manifest records cap, spend,
reservation rule and stop reason; money and interval arguments must
be finite and positive (--watch-interval >= 5 s).
- The seed is a clean DETACHED clone of --seed (a ref of --source-repo)
materialized once under the run root; every lane asserts its clone is
at the admitted sha and clean; the source's dirtiness is disclosed.
- --self-mod requires a CONFIRMED absorb per lane (pre-task snapshot of
clone HEAD/served sha/uptime/absorbed cycles; afterwards the counter
advanced, the sha moved, uptime reset, server ready) and a run-level
gate over every self-mod lane.
- SK1: wait_task namespaces checks per task (author_/dispatch_),
LaneContext.check refuses a duplicate key, the dispatch counts only on
a tools.jsonl row with status ok and the extension's exact echo.
- SM1: the accent change lands in both web/style.css and
web/onboarding.css (prompt, stub, acceptance parity check); the stub
rehearsal still skips the tests preflight (documented residual: the
loopback base URL leaks into the hermetic suite), and the typed refusal trail
(ledger block_reasons, PREFLIGHT_BLOCKED / TESTS_PREFLIGHT_BLOCKED /
SCOPE_REVIEW_BLOCKED tool codes, terminal reason_code) is a fact.
- The run-root effective_settings.json is redacted; the key reaches
disk only in each lane's 0600 settings file and is disclosed by
fingerprint as the runtime grant.
- Lane infra failures carry a typed refusal {type, code, message} and a
reason_code in result.json and result_index.jsonl.
- The key probe runs on its own thread with an 8 s HTTP bound, at most
once a minute, backing off on failure; a failed probe is informational
and never delays a tick.
DEVELOPMENT "Live E2E stand" describes the new contracts and the
--per-task-usd sizing guidance from the first paid run.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit d333d573b01de8f7ba2ecc01f4eded634c7332db)
|
||
|---|---|---|
| .. | ||
| __init__.py | ||
| run_live_lanes.py | ||
| scenarios.py | ||
| stub_lane.py | ||
| ui_probe.py | ||