Commit graph

263 commits

Author SHA1 Message Date
Ouroboros
b463bb3d93 Add the Z.ai (GLM) direct provider with effort projection at the send boundary
zai:: joins the direct providers exactly the way deepseek:: did: prefix and
credential registry, ZAI_API_KEY plus a ZAI_PLAN endpoint selector
(provider_models.resolve_zai_base_url: empty/payg = api.z.ai/api/paas/v4,
coding = the Coding Plan endpoint), the routing target, live catalog fetch,
provider Test, settings card, onboarding contract, review-fallback roles,
single-provider startup and review detection, secret masking, benchmark
env hygiene, and docs.

Reasoning effort now reaches Z.ai. The provider serves an ABSENT
reasoning_effort at its maximum tier, so every call on the old generic
compatible route was billed at max regardless of the configured effort.
The canonical scale is projected onto Z.ai's own low/high/max enum
(ZAI_REASONING_EFFORT_ALIASES: none/minimal -> low, medium -> high,
xhigh/ultra -> max), disclosed as reasoning_effort_clamped when the tier
changes; GLM-5.3 rejects every other value and cannot disable thinking
(HTTP 400 code 1210), and forced tool_choice works with thinking on, so
there is no DeepSeek-style suppression arm. The projection is keyed on the
provider id the owner configured, never on a model name: a GLM served
from an owner's own OpenAI-compatible endpoint keeps today's behavior.

The provider port is the contributor's own work from the closed PR #1194,
narrowed to Z.ai (the DashScope and Moonshot lanes were not measured and
stay out). 07-configuration gains two settings rows and one route
paragraph (budget 37300 -> 38400), 02-naming records the dated Z.ai
probe in the external-fact inventory, and the onboarding bootstrap
fixture and data-layout inventory are regenerated.

Co-authored-by: josephsteuerjr <josephsteuerjr@gmail.com>
2026-09-25 17:41:54 +03:00
Ouroboros
26aef771b4 fix: bind Cowork evaluator score to run protocol provenance 2026-09-25 07:09:54 +03:00
Ouroboros
720fd98152 feat: add claimed Cowork evaluation and opt-in residual audit 2026-09-25 06:00:18 +03:00
Ouroboros
7f72ad0e9b Merge ouroboros after Cowork evaluator receipt PR-A
Preserve independent official receipt/status on all execution branches while retaining #1259 meter-bound interruption and paid-activity semantics.
2026-09-25 04:18:32 +03:00
Ouroboros
af1c18f95b fix(cowork): preserve cleanup and interrupted evidence on write failures 2026-09-25 03:31:43 +03:00
Ouroboros
aef127346b fix(benchmarks): retain official evaluator evidence independently of execution 2026-09-25 03:21:08 +03:00
Ouroboros
d9b9157fa4 fix(cowork): bound meter blindness and retain interrupted activity 2026-09-25 01:22:52 +03:00
Anton
3fb5ddd518 Merge origin/ouroboros c9180ef8 (#1252 cards, #1244 safe-check, #1254 shutdown, v7.4.10) into feature/1196-continuity
Conflicts: generated inventories/DOMAIN_MAP taken from upstream and to be regenerated; chapter 06 money paragraph taken from upstream (ours was a compression of the same text). contracts.py/chat.js trimmed back under their ratchet limits. KNOWN RED: size_ratchet_manifest.py stale — runtime function count 10036 > 10000 after the merge (upstream grew ~42 functions since 87fd00f4); to be paid down by simplification, not a cap raise.
2026-09-24 17:21:56 +03:00
Anton
612067385c WIP: preserve #1196 continuity repairs before upstream integration
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-24 05:09:55 +03:00
Ouroboros
f04fb1ebfe Truthful cards and history: presentation batch #1011 #1087 #1110 #1154 #1061 #931 #1073 #516 #869 #498 (candidate) 2026-09-24 00:48:30 +03:00
Ouroboros
0abab76225 ouroboros: checkpoint after task c6207d8d665040b0 — Долгая работа #1196 — продолжение до реализации 2026-09-23 23:52:17 +03:00
Ouroboros
db955324e6 test(review): await settled progress and qualify benchmark semantics 2026-09-18 18:19:19 +03:00
Ouroboros
75425053bf Merge current development while preserving workflow review contracts 2026-09-18 16:39:33 +03:00
Ouroboros
5ae3839434 Integrate workflow review improvements with the current development branch 2026-09-18 14:31:09 +03:00
Anton Razzhigaev
7e2cae933e Merge PR #1044: preserve Cowork budget and interrupted work
Integrate the reviewed Cowork changes with the current development branch.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 11:30:07 +00:00
Ouroboros
2c1a93fe5b Merge frozen release repairs into development
Preserve both the explicit platform-lock calls and development audit binding checks. Adapt the existing negative fixtures to the same calls without changing their assertions.
2026-09-18 13:51:38 +03:00
Ouroboros
91eb4304e8 fix: close remaining frozen release gate regressions 2026-09-18 13:28:04 +03:00
Anton Razzhigaev
8232705a82 Merge landed Windows benchmark fixes into Cowork recovery
Carry the unchanged Cowork patch onto the current development branch, including
the reviewed Windows path and campaign-lock fixes and their structural audit.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 10:09:15 +00:00
Ouroboros
c9543aca5c Recognize the shared platform lock in benchmark admission audit
Keep the existing canonical campaign-lock exception bound to its exact owner, blocking selection and sole hashed host-temp open. Preserve shadowing and additional-I/O rejection without changing launcher or lock behavior.
2026-09-18 12:48:34 +03:00
Ouroboros
6a9e396ae1 fix(bench): recognize shared platform locks in the launcher audit 2026-09-18 12:47:37 +03:00
Ouroboros
befe5424e9 fix(bench): preserve Windows campaign paths and exclusive execution 2026-09-18 12:09:21 +03:00
Ouroboros
27d2429fea Integrate landed review wake repairs and document author finality 2026-09-18 04:40:06 +03:00
Anton Razzhigaev
484d241cfd Merge Ouroboros 7.1 development into Cowork recovery
Keep the reviewed Cowork adapter changes on the latest frozen development base.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 01:32:54 +00:00
Anton Razzhigaev
acf312d441 fix(bench): bound complete Cowork meter reads by the shared deadline
Run the read-only HTTP probe in a short-lived standard-library worker with
stdin-only credentials. Kill and reap it at the deadline, reject late values,
and test delayed response headers/body through a local HTTP endpoint.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 01:08:29 +00:00
Anton Razzhigaev
805d35517b fix(bench): preserve interrupted Cowork work and confirm budget readings
Confirm suspect usage samples within one bounded read window without accepting
lower counters. Verify exact-label resource absence after competing cleanup and
retain removal diagnostics. Checkpoint scrubbed task traces during polling so
abrupt container removal preserves partial evidence.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 00:29:33 +00:00
Ouroboros
d8e1c627cb Merge landed target into CyberGym integration preparation
Preserve the reviewed CyberGym source, Cowork launcher registration, and relocated generated inventories. Regenerate only the data-layout source digest after the independent documentation changes.
2026-09-18 02:44:57 +03:00
Ouroboros
484f6617e7 Fix CyberGym late delivery and terminal accounting
Recover final measured results below historically held bounds without refunding or repeating paid work. Normalize coherent terminal finality for failure paths, retain custody and append-only evidence, and correct the documented benchmark treatment.
2026-09-18 02:35:37 +03:00
Anton Razzhigaev
225dbd4b0f fix(bench): honor the explicit Cowork billing reserve
Keep task lifetime limits separate from the allowance for unsettled charges.
A 32-task campaign must not stop solely because its task budgets sum to $800.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-17 23:12:12 +00:00
Ouroboros
26159dc644 Simplify CyberGym settlement and consolidate gateway custody 2026-09-18 00:49:42 +03:00
Ouroboros
2fd8ed5f78 Merge current ouroboros into CyberGym recovery 2026-09-18 00:41:54 +03:00
Anton Razzhigaev
d6e55f7325 fix(bench): preserve recovery scope and complete Cowork evidence
Bind recovery to its retained selection, ancestry, and immutable image. Preserve setup failure semantics, runtime log history and artifact/cost finality; monitor preparation storage and synchronize integration checks.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-17 21:24:51 +00:00
Anton Razzhigaev
e41b00cdc4 feat(bench): add isolated Cowork Bench adapter and campaign controls
Run Ouroboros through the pinned official task/evaluation runner with persistent MCP services, cumulative spending controls, bounded Docker resources, and auditable results.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-17 20:28:36 +00:00
Andrew
e552635d1e Allow explicit Cyber Pro mode for CyberGym runs
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-17 13:50:33 +03:00
Ouroboros
56ff4eede9 Consciousness: review round 1 fixes (alarm re-arm exclusions, graceful wake ceiling, own cards, honest status)
- The alarm no longer re-arms on the owner's own direct turn ending (B13) nor on a
  project digest of a tree consciousness started; notify keeps the boot floor after
  a restart (max(last wake, boot) + floor).
- A wake carries what is left of its allowance as the tree's GRACEFUL ceiling
  (`root_cost_ceiling_usd`, now honored for the root itself by the in-task stop)
  instead of narrowing the ledger fence: one Main attempt reserves ~$8 up front, and
  the narrowed fence refused every wake of a nearly spent day before its first call,
  posting a budget error into Main at each heartbeat (stand: 15 such wakes). A
  remainder at or below the planning margin is skipped as allowance_exhausted.
- A wake the lane could not admit backs off like a failed wake; a closed budget door
  is retried quietly at the interval, a transient door at the floor.
- The wake message lists unanswered cards of any task, the previous wake's own
  included (an expired_terminal card still takes a late answer).
- Only an owner's stop is sticky against toggle_evolution: a stop the agent placed
  remembers its source (`evolution_stop_source`) and stays undoable by the agent.
- The status snapshot carries `unknown_unmetered`/`integrity_degraded`; the allowance
  line says "at least $X" and "ledger integrity degraded" when they apply.
- Docs: a Presence cycle a wake starts is outside the allowance; benchmark profiles
  drop the retired OUROBOROS_BG_MAX_ROUNDS key.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-16 10:09:18 +03:00
Ouroboros
9a8f63d7ed Align composite test contracts after feature integration
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-13 11:56:17 +03:00
Andrew
eef44c3e86 Preserve terminal root accounting evidence
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-13 08:28:17 +03:00
Anton Razzhigaev
ffa33ccf75 cybergym: heal reconcile delivery for the interrupted full-1507 run
Two defects blocked recovering the 91 server-side-completed tasks of run
20260907T233516Z:

1. reconcile_task treated a cached sparse (terminal but cost-unverifiable)
   frame as a permanent refusal.  The cache is only a refusal snapshot: it
   now never suppresses the live gateway re-poll and stands solely when
   neither the gateway nor the isolate disk can produce a frame.  This lets
   an operator-acknowledged quarantine heal (renaming the isolate's
   usage_attempts.quarantine.jsonl clears the read-time integrity_degraded
   flag) take effect on the next pass instead of being masked by the cache.

2. run_dispatched stranded completed rows behind a budget-refused position:
   the source-order drain keyed on next_record never advances past a
   row-less refused position, so under a lost claim race the campaign
   reported zero landed rows.  Final settlement now returns the disjoint
   union of completed and dispatched maps in source order.  The
   campaign-level cap test is made winner-agnostic (either lane can win the
   first claim) and a deterministic run_dispatched regression test pins the
   earlier-refusal/later-completion ordering.
2026-09-08 11:40:49 +00:00
Anton Razzhigaev
add19cc535 cybergym: redeliver terminal evidence for non-terminal claims in reconcile
Run 20260907T233516Z left 90 gateway-terminal tasks with infra_failed rows
and unresolved claims.  The reconcile arm only redelivered supersedeable
rows whose claim had already settled; unresolved/reserved claims took the
settle-only recorded-recovery path, forfeiting terminal evidence that the
isolate had already produced and settling at the reservation upper bound
instead of the measured cost.  Reconcile now redelivers for any
non-terminal claim state, skips the settled-cost equality check for claims
that never settled, and settles those claims at the delivered measured
cost in the durability order (row, settle, workspace release).
2026-09-08 11:17:17 +00:00
Anton Razzhigaev
1a0e5b9710 cybergym: fix budget phantom liability, custody latch, and claim-refusal flush
Run 20260907T233516Z died on a compound failure: one poisoned workspace
custody entry burned 107 tasks, 127 post-dispatch attempts with unknowable
spend leaked their $20 reservations as eternal unresolved liability and
tripped the $3000 cap on phantom spend, and the dispatcher then flushed
1145 never-dispatched tasks into infra_failed rows on claim refusals.

- settle_finished_attempt: attempts with no terminal gateway frame and no
  measured cost evidence now settle terminally at their held reservation
  (projection-neutral); measured non-final costs stay unresolved for
  reconcile.
- cybergym_docker: _heal_unresolved_workspace_custody re-inspects latched
  entries and releases provably terminal owned containers; running or
  unproven entries stay latched.
- cybergym_dispatch: a claim-time BudgetRefused now pauses admission
  through a budget gate, probes in-flight settlements for freed headroom,
  and ends the campaign with BudgetCapReached only when the drained pool
  still refuses — undispatched tasks stay row-free for resume.
- cybergym_wire: model-hallucinated unknown_tool outcomes classify as
  agent-attributable fair completions, not infrastructure failures.
- run_cybergym: finalize budget_cap_reached as a typed outcome.
2026-09-08 11:10:07 +00:00
Anton Razzhigaev
43afe25a57 cybergym: recovery layer on v7 for benchmark completion
Squashed devtools layer from cursor/cybergym-mega-20260901 (PR #456)
onto managed/ouroboros a5e6b983, with core conflicts resolved in favor
of v7 (stale core patches intentionally not carried over).

Benchmark-layer fixes on top:
- dispatch: settle budget and record results at task completion, not in
  source order (fixes reservation pile-up that produced 642 BudgetRefused
  in 15s on the 2026-09-05 run); result_index.jsonl is completion-ordered,
  catalog/summary keep source order
- wire: strict adapter-level detector for DeepSeek DSML tool-call markup
- executor/sidecar: optional read-only mount of the task's vulnerable
  runtime dir into the workspace; server sidecar accepts an external
  binary dir via read-only bind mount at the identical absolute path
- regrade: inventory builder + append-only CLI driver that re-verifies
  historical final.poc files through the official verifier without model
  or gateway calls, with per-task cleanup and resume
- run config: 3h timeout, 600 rounds, $3000 campaign budget, task
  acceptance review disabled

Local provenance anchor for the 2026-09-08 full run; not for merge.
2026-09-07 23:24:02 +00:00
Anton Razzhigaev
f6711d3369 feat: integrate subscription accounts into model setup and execution
Add caller-owned Claudexor model calls, unified subscription/API setup, explicit role accounts and live quota/auth continuation. Preserve prompt/tool ownership, physical-attempt custody and manual assignments. Keep the existing task process through waits and post-work.

The reviewed source checkpoint retains the current dependency pin. Published signed Claudexor bytes, cross-platform CI and the agreed merge ordering remain delivery prerequisites. No Ouroboros version, tag or release is created.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-07 10:54:51 +00:00
Ouroboros
fa2977423b fix(e2e_live): reserve the evolution root only for the scenario that absorbs; typed idle reasons relative to the wait
Second adversarial round on the absorb-wait rewrite.

Reservation: after the previous commit only SM1 can promote (SW1/SK1 pin
OUROBOROS_POST_TASK_EVOLUTION=false), yet RunBudget.reservation still
added the evolution root for every attempt under --self-mod, so SW1/SK1
reserved a ceiling they could never spend and run-cap admission refused
or serialized paid attempts on it (the SK1-only mini run at cap 130 was
refused by exactly this over-reservation: 50 x (2 + 1) = 150). The rule is
now per_task x (root_tasks + int(self_mod and absorbs)); admit,
budget_preflight, dispatch_order and run_lane pass the scenario's
expects_absorb. Owner configuration (cap 300, per-task 50, 3 attempts,
self-mod): SM1 100, SK1 100, SW1 50 — realistic spends admit 9/9 for
$159, pessimistic 8/9 for $219 (SK1_a3 refused), every scenario keeping
two; the CI e2e-live arithmetic comment and summary header are rewritten
(full set $225, not $315) inside the D-12 job, and the CI-lane pins
re-derive the numbers with the scenario flag.

Idle reasons: absorb_idle_reason types relative to the wait's start
(history length snapshot) so a resumed campaign's older cycles never
speak for this boundary; a paused/stopped/completed status wins;
no_promotion means an every_n post-task tick was recorded (llm cadences
write none: no_decision). Tests cover the resumed-campaign boundary and
the reason table with the cycle written during the wait.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 01:31:11 +00:00
Ouroboros
4f3cda85d5 fix(e2e_live): absorb wait proves no cycle is pending; non-absorbing lanes do not promote
Adversarial review of the corrected --self-mod seeding (2026-09-06) traced
the post-task path end to end and found that IsolatedServer.wait_for_absorb
ended the wait as no_promotion on ONE idle /api/state sample once the grace
had passed: no request file, pending_count 0, running_count 0. Those are
exactly the readings of the absorb path itself. A cycle that committed keeps
its campaign transaction as waiting_for_restart while the supervisor
restarts synchronously (RUNNING already popped, the counter unchanged), and
the re-exec'd server answers /api/state with zero counts before its
supervisor is up; absorbed_cycles_done moves only when the worker boot
verifies the restart. The pre-seeded benchmark campaign used to mask this
(its t=0 cycle kept running_count above zero and the kept request file
blocked the exit); with the owner-id-only seeding every SM1 lane would have
failed a healthy absorb as no_promotion.

wait_for_absorb now needs PROOF that no cycle is pending, held on six
consecutive polls after the grace: the queue idle AND supervisor_ready AND
no post_task_evolution_request.json AND no campaign active_transaction. The
reason is typed from the durable campaign state (campaign_summary /
absorb_idle_reason): no_promotion (no campaign although the post-task
decision ran), no_decision, cycle_no_op, cycle_not_absorbed,
campaign_<status>, cycle_not_enqueued; the summary travels in the wait dict
the stand records. Benchmark callers (evolve_smoke, the CLB adapter) keep
their early exit on a genuinely idle campaign.

Two more findings from the same review: scenarios that commit nothing
(SW1, SK1) now pin OUROBOROS_POST_TASK_EVOLUTION=false in their lane
settings — a one-shot cycle promoted from their own roots could commit and
re-exec the server in the middle of the lifecycle under test; and the SK1
gate also requires the review call's own executable_review (a pending
duplicate job never passes) while _skill_entry reads the listing through
the non-raising helper so a failed listing is a failed check with facts, not
an infra_error. Tests: a new module pins the wait on a fake clock
(waiting_for_restart and a booting server do not end it, one idle sample
does not, a pending request blocks it, the absorb confirms, every idle
reason), the SK1 module pins the product gate rule and the per-scenario
promotion pin, the runner module's settings pin flips for SW1/SK1.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 01:04:49 +00:00
Ouroboros
5ad8019dd4 e2e_live: cap the nested preflight xdist fan-out per lane (W4-preflight-workers)
The commit gate's hermetic pytest pass runs inside each lane server and
resolves `-n auto` to os.cpu_count() (128 on this host) with no ceiling; the
2026-09-04 paid run started >= 104 xdist workers per self-mod lane with three
lanes overlapping. The runtime already has the lever
(OUROBOROS_PREFLIGHT_TEST_WORKERS, floor 2, read by
preflight_runner._preflight_worker_count and scrubbed from the candidate
suite), but IsolatedServer's settings-authoritative sweep dropped it before
the lane server started, so the stand could not use it.

Devtools-only wiring (evidence action A):
- server_runner: keep OUROBOROS_PREFLIGHT_TEST_WORKERS through the
  authoritative sweep as an operational host-load lever (comment reworded;
  it is not a model, credential or settings key).
- run_live_lanes: derive max(2, 16 // lanes) at argument time (shared-host
  rule: at most 16 pytest workers across the stand), set it in the launcher
  process before the first lane starts so an ambient shell value never wins,
  and record it as extra.preflight_test_workers in run_manifest.json and as
  preflight_test_workers in every lane row.
- docs/DEVELOPMENT.md: one sentence in the live E2E stand section.

Pins: the lane sees the computed value (ambient 128 overridden), the runtime's
_preflight_worker_count reads it, the manifest and lane row record it, the
floor and key names match preflight_runner's constants, and IsolatedServer
forwards the key while still stripping an ambient OUROBOROS_MODEL. No
runtime code changed.

Size ratchet: tests/test_e2e_live_runner.py enters the 1001-1500 band with this
commit (996 -> 1042 lines); the band rationale is recorded in the manifest.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit 1824f8910509c65b9fc94041c7a76f98a9aa617d)
2026-09-05 09:52:35 +00:00
Ouroboros
92a8b3f1b1 Live E2E stand: devtools/e2e_live runs K staggered isolated servers over the SM1/SW1/SK1 scenario table (F4-A)
run_live_lanes.py admits through the benchmark family's seams (dirty seed refused with the
refusal persisted; the key by NAME from the environment, never a pool file; the credit
preflight takes min(key limit remaining, account credits) via the new
manifests.openrouter_account_credits, the second bound only), writes the effective settings
from the tree's own defaults with TOTAL_BUDGET/OUROBOROS_PER_TASK_COST_USD as settings keys,
names the model in the run manifest from the applied file, and fans out --lanes (default 4,
max 6) isolated servers 2-3 s apart. scenarios.py is the table: SM1 lands a web/style.css
token through commit_reviewed under advanced+blocking (S2 set + computed style read from the
committed CSS after a restart), SW1 arms Swarm in the browser (force_plan, roster, >=2
children with causal lineage, fanout receipt, cost rollup, /proc no-orphans), SK1 has the
model author a skill then reviews, grants, enables, dispatches and deletes it; acceptance is
callable over durable artifacts only. ui_probe.py resolves the suite's PlaywrightUIClient when
it carries the surface, else headless Chromium, else a typed ui_unavailable. --stub rehearses
every scenario for $0 on the loopback stub of tests/system_e2e/harness.py (stub_lane.py routes
the swarm wire by role). Per-lane result.json carries checks, settings sha256 plus a
secret-free config digest, seed describe, pre/post HEAD and the diff digest, grants by
fingerprint and the runtime terminal disclosure; a watcher prints lane states, free disk on /
and /mnt/data and the key headroom. Tests pin the launcher gate by source, the table shape,
lane/stagger bounds, the TMPDIR guard, the dirty-seed and credential refusals, the two-plane
credit floor, manifest-model-equals-applied-file, secret-free artifacts, and a gated stub
rehearsal of SM1 on a real server. DEVELOPMENT and ARCHITECTURE describe the stand.

(cherry picked from commit ab36206a2a7ccbad01bf6b25a69181a6a69d9aa6)
2026-09-04 23:37:49 +00:00
Ouroboros
62f87cc94c Merge upstream ouroboros db6d7cf8 into the v7 line: absorb PR #609 net-resilience and PR #614 update letter
Second absorption of the frozen upstream line (23ab428f..db6d7cf8: 89
commits, 47 files) on top of the rc.10 hotfix tip, by the F2 rules (S1
upstream body in the owning leaf, S2 hand-merge, S3 only with proof;
retired 7.0 surfaces never return):

- PR #609 net-resilience: interactive transport-wait episodes bounded by
  the task idle timeout with the typed task_incident/toast_once pair, a
  bounded paid repeat after a typed post-dispatch transport death with a
  round-keyed record that fences every other send, the shutdown-aware
  supervisor crash counter and bounded lifespan join, Darwin keepalive
  tuning. loop_transport/transport_custody/net_transport/loop_llm_call
  land verbatim (same shapes on both sides); the loop.py deltas are
  relocated into the v7 leaves (loop_round_limits, loop_model_call,
  loop_delivery, loop_forced_finalization, loop_nudges, loop_messages,
  loop_budget) with bodies AST-equal to upstream modulo the call-time
  handles; _emit_overflow_retry_skipped stays a public helper (the v7
  facade contract) and upstream's nested _skipped delegates to it.
- PR #614 update letter: ouroboros/update_letter.py and its web module
  land verbatim; the new OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC key and
  get_update_letter_timeout_sec live in their v7 owners
  (settings_defaults.py, runtime_limits.py, re-exported by config.py);
  _supervisor_stop lives in server_process.py beside the restart events;
  docs/PERSISTENCE.md gains the state/update_letter.json row and the
  inventory pin moves to 286.
- Tests: the relocated run_llm_loop tests take the emit_progress
  incident keyword (every one-argument progress fake in tests/ swept, a
  gap upstream itself left in test_tree_cost_ceiling); the official-update
  runtime-section test lands in tests/test_context.py; _MOVED_OWNERS
  registers the relocated getter.
- Docs: ARCHITECTURE/DEVELOPMENT hunks land on the upstream text; the two
  legacy timeout rows upstream's context still carries stay retired (7.0).
- Size ratchet regenerated; the band rationale for tests/test_update_letter.py
  is carried verbatim from upstream.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 22:53:15 +00:00
Ouroboros
cfccf12c06 tests: the TB methodology tests read the shipped triad defaults from their SSOT
devtools/benchmarks/terminal_bench/test_run_tb_methodology.py (upstream body,
absorbed by F2) looked the shipped triad up through
SETTINGS_DEFAULTS["OUROBOROS_REVIEW_MODELS"]; ABI 7.0 retired that settings
key, so the rc.9 tag CI job benchmark-methodology failed with KeyError in two
tests. The launcher itself derives the same list from
ouroboros.settings_defaults.OPENROUTER_REVIEW_DEFAULTS["triad"] (the SSOT the
retired key was joined from), so the tests now read that list. Test-only;
no runtime change.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 21:49:28 +00:00
Ouroboros
9698e2e077 Merge upstream ouroboros 23ab428f into the v7 line: absorb 407 commits into the module split
Second parent is the frozen upstream `ouroboros` head (23ab428f, 407 commits
since the merge base a76961de); first parent is v7.0.0-rc.8 (18b9832e).

Every upstream change lands in v7's owning leaf: S1 transplants keep upstream's
bodies (comments verbatim) under the call-time handle idiom, S2 hand-merges keep
both intents, S3 keeps v7 only with proof (retired 7.0 ABI surfaces, superseded
mechanisms). Per-symbol relocation ledger: docs/archive/v7next/LEDGER_CORRECTIONS.md
(F2 absorption section). Provisional decisions awaiting owner ratification:
D-18 (two-destination symbols), D-19 (acceptance rows follow upstream R2),
D-20 (acceptance_dialogue stays deleted), D-21 (tools/registry.py: facade
import block only).

Docs: upstream ARCHITECTURE/DEVELOPMENT as the base with compact v7 deltas;
bookkeeping moved to docs/archive/v7next. Size-ratchet manifest, domain
manifest and generated inventories regenerated; new leaves: tools/write_shape
walker, gateway/cost_breakdown, tools/core_secret_paths; provider_catalogs.py
and acceptance_dialogue.py removed (v7 owners).
2026-09-04 19:32:55 +00:00
Ouroboros
952c39db12 Benchmark disclosures: repeats live inside the call's attempt budget
The continual-learning runbook and the SWE-bench Pro methodology said
the transport-death repeats sit on top of the transient burst. The
primary dispatch grants a repeat only while the outer attempt budget
has room, so both now say: one outer attempt budget per call, within
which up to three rows can be typed transport-death failures (the
first death plus at most two repeats).
2026-09-04 14:17:57 +00:00
Ouroboros
6802a822a3 Merge managed/ouroboros (5b4546ea) into the network-resilience branch
Five conflicts, resolved by reading both sides in full:

- loop.py: the trace-touched skill-name scan moved into skill_readiness.py
  upstream; its import replaces this branch's inline copy, whose removeprefix
  form was a byte fold of the same behaviour.
- loop.py: the no-resend terminal keeps upstream's block (its short comment and
  its live_trace re-read); this branch's own two-stamp fold is re-applied on top,
  and the record-fenced source stays.
- loop_llm_call.py: _send_main_candidate binds whenever a physical context OR a
  candidate predicate is present, upstream's semantics, expressed through the
  binding this file already uses. The context manager is only constructed there,
  never entered, so the effect order is upstream's.
- loop_llm_call.py: call_llm_with_retry keeps both new parameters,
  transport_death_retries and initial_messages, on the two lines the file's
  line cap already pays for.
- ARCHITECTURE.md: upstream's producer-word sentence, with this branch's
  no-resend source clause re-applied, so the paragraph states both.

loop.py carries shrink-only byte debt and upstream absorbed both places where
this branch had paid for its own additions, so the payment is made again inside
the functions this branch owns: the stamp fold above, the fifth site of the fit
key tuple now calls _fit_key, and _emit_overflow_retry_skipped is folded into
_skipped, its only remaining caller since this branch collapsed the other three.
No comment, docstring, diagnostic or test was shortened. 271928 -> 271855 bytes.

The manifest is regenerated on the merged tree; it is upstream's manifest with
loop.py's exact byte count. Upstream's three new band rationales are carried
across verbatim because the generator reads the pre-merge committed manifest.
2026-09-04 08:41:52 +00:00