mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
WIP: preserve merged no-tool pause semantics and structural cleanup
Focused verification passed; global function-count gate and full consumer verification remain unresolved. Not publication-ready. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
parent
8759ca946f
commit
bea4049572
20 changed files with 279 additions and 186 deletions
|
|
@ -8,14 +8,14 @@ The manifest is the SSOT of the module→domain assignment (1:1, complete over t
|
|||
|
||||
| domain | name | modules | proposed |
|
||||
|---|---|---:|---:|
|
||||
| D01 | Agent core & main loop | 36 | 0 |
|
||||
| D01 | Agent core & main loop | 37 | 0 |
|
||||
| D02 | LLM client, routing & providers | 38 | 0 |
|
||||
| D03 | Context assembly, fit & compaction | 11 | 0 |
|
||||
| D04 | Tool execution: registry, access & typed results | 21 | 0 |
|
||||
| D05 | Tool surfaces: files, code, shell, media, external | 29 | 0 |
|
||||
| D06 | Review stack | 67 | 0 |
|
||||
| D07 | Delegation, subagents & Claudexor | 53 | 0 |
|
||||
| D08 | Supervisor: queue, workers, events & runtime control | 47 | 0 |
|
||||
| D07 | Delegation, subagents & Claudexor | 54 | 0 |
|
||||
| D08 | Supervisor: queue, workers, events & runtime control | 48 | 0 |
|
||||
| D09 | Cancellation, owner control & process custody | 13 | 0 |
|
||||
| D10 | Git, update & release machinery | 28 | 0 |
|
||||
| D11 | Gateway, server & Web UI | 56 | 0 |
|
||||
|
|
@ -28,7 +28,7 @@ The manifest is the SSOT of the module→domain assignment (1:1, complete over t
|
|||
| D18 | Launcher, packaging, platform & shared substrate | 15 | 0 |
|
||||
| D19 | Frozen contracts (ABI) | 10 | 0 |
|
||||
| D20 | Presence | 10 | 0 |
|
||||
| **total** | | **569** | **0** |
|
||||
| **total** | | **572** | **0** |
|
||||
|
||||
## Dependency direction matrix (strict, pinned)
|
||||
|
||||
|
|
@ -189,6 +189,7 @@ No function body (≥ 10 normalized lines) is shared verbatim across domains. Ne
|
|||
- `ouroboros/agent_dispatch.py`
|
||||
- `ouroboros/agent_startup_checks.py`
|
||||
- `ouroboros/agent_task_pipeline.py`
|
||||
- `ouroboros/budget_pause.py`
|
||||
- `ouroboros/deadline_utils.py`
|
||||
- `ouroboros/focus.py`
|
||||
- `ouroboros/loop.py`
|
||||
|
|
@ -405,6 +406,7 @@ No function body (≥ 10 normalized lines) is shared verbatim across domains. Ne
|
|||
- `ouroboros/claudexor_startup_failure.py`
|
||||
- `ouroboros/configured_subagents.py`
|
||||
- `ouroboros/delegate_containment.py`
|
||||
- `ouroboros/delegate_continuation.py`
|
||||
- `ouroboros/delegate_custody.py`
|
||||
- `ouroboros/delegate_custody_memo.py`
|
||||
- `ouroboros/delegate_custody_reconcile.py`
|
||||
|
|
@ -465,6 +467,7 @@ No function body (≥ 10 normalized lines) is shared verbatim across domains. Ne
|
|||
- `ouroboros/tools/followup.py`
|
||||
- `supervisor/__init__.py`
|
||||
- `supervisor/active_activity.py`
|
||||
- `supervisor/budget_resume.py`
|
||||
- `supervisor/cognitive_operations.py`
|
||||
- `supervisor/direct_roots.py`
|
||||
- `supervisor/event_taxonomy.py`
|
||||
|
|
|
|||
|
|
@ -47,8 +47,8 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
|
|||
│ ├── task_admission.py ← Token-owned admission reservations fence duplicate user-ingress ids before Project/workspace/attachment side effects; queue.py stays the state authority (§5)
|
||||
│ ├── task_lifecycle.py ← Cancellation custody — the ONE settle owner of durable cancel intents — plus the `sweep_cancel_intents` watchdog and the queue-owned root-budget admission fence (flow: §5; rules: §10 invariant 14)
|
||||
│ ├── cancel_publication.py ← Cancellation settlement publication for `task_lifecycle.py`: typed CANCEL_* outcomes, artifact-honest cancelled result fields, ledger cost reconstruction, salvage, owed-before-settle registration, capture-miss terminalization (§5)
|
||||
│ ├── budget_resume.py ← Exact-continuation budget Resume (#1196): the single-use, pause-id and generation-bound grant, its revocation, hold release/re-binding; re-exported by queue_transitions.py
|
||||
│ ├── queue_transitions.py ← Queue-owned transitions outside cancellation custody: acceptance-fence open/inspect/seal, explicit budget resume, typed `stop_evolution_tasks` (an incomplete stop leaves the campaign OPEN under the durable `evolution_owner_stopped` flag, cleared only by an owner start ingress) and fenced Project deletion (lineage ROOTS only; tombstone after provable quiescence); imports nothing from task_lifecycle (§5)
|
||||
│ ├── budget_resume.py ← Exact pause grants, revocation and hold release; queue_transitions re-exports (§6 Exact budget pause and Resume)
|
||||
│ ├── queue_transitions.py ← Acceptance open/inspect/seal, budget Resume, stop_evolution_tasks and fenced Project deletion; separate from cancellation custody. Incomplete evolution stop stays OPEN under evolution_owner_stopped until owner start; deletion tombstones lineage ROOTS only after quiescence (§5)
|
||||
│ ├── terminal_delivery.py ← Durable terminal-answer delivery seam for final answers, cancel salvage, cascade digests and non-retry reaps: restart-surviving `delivery_id` dedupe (the id digests only the stable part of the answer, so a rebuilt replay dedups) + the bounded PENDING outbox `state/terminal_deliveries.json`; typed `terminal_delivery_exhausted` and `terminal_delivery_handoff`; per-origin projection `host_salvage` / `host_notice` / `custody_notice` / `model_final` (§5; §6 Task lifecycle; §10 invariant 15)
|
||||
│ ├── task_reaper.py ← Single-owner off-loop reaper for timeout teardown and health-prepared terminal-file/crash jobs; an unconfirmed death keeps the slot reaping with `task_reaper_wedged`; mints no cancel intents (§5)
|
||||
│ ├── owner_stop.py ← Owner graceful stop: the `finalize_then_cancel` policy axis on the SAME durable cancel intent (monotonic hardening), one typed `finalize_now` control (`owner_requested_finalization`), grace bounded by `OWNER_STOP_OUTER_CAP_SEC`; `running_owner_stop_tasks` bypasses only the idle/finalization-grace rails (§5)
|
||||
|
|
@ -111,7 +111,7 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
|
|||
├── evolution_checkpoints.py ← Append-only campaign/eval checkpoint ledger for evolution progress
|
||||
├── evolution_fingerprint.py ← Canonical fingerprint for evolution-campaign objectives; SSOT for repeat gating
|
||||
├── improvement_backlog.py ← Durable advisory improvement backlog: recurrence-counted dedup (never drop), priority+recurrence+recency ranking, `close_backlog_items`, size-triggered `groom_backlog`; parser-safe locked writer
|
||||
├── loop.py ← High-level LLM tool loop; finalization nudges and the latest typed `FINAL ANSWER:` candidate (§6 Task lifecycle). A harness child with zero durable start attempts gets one nanny nudge; PENDING ≠ FAILED, so a started-but-unsettled run gets a wait reminder instead of a failure accusation that invites a duplicate run. Disclosures, never gates: `nanny_finalized_after_nudge_without_delegation`, `CONFIGURED_ACTOR_INCOMPLETE`/`CONFIGURED_ACTOR_UNKNOWN`, `NANNY_METERED_OVERRUN` (§6 Delegated subagents)
|
||||
├── loop.py ← LLM tool loop, finalization nudges and FINAL ANSWER candidate (§6 Task lifecycle). Nanny without a durable start gets one nudge; unsettled start gets a wait reminder, preventing duplicate runs. Disclosures, not gates: nanny_finalized_after_nudge_without_delegation, CONFIGURED_ACTOR_INCOMPLETE/UNKNOWN, NANNY_METERED_OVERRUN (§6 Delegated subagents)
|
||||
├── acceptance_settlement.py ← A paid acceptance panel that outlives the answer it reviewed: the quorum/completion mailbox wake, delivery under a running panel (`previous_revision_accepted`), the post-terminal `late_settlement` supplement (§6 Task acceptance)
|
||||
├── loop_acceptance.py, loop_acceptance_review.py, acceptance_preparation.py, acceptance_retrieving.py ← Acceptance machinery (externally used names re-exported from `loop`): the fence and its obligations (`ACCEPTANCE_DECISION_REASONS`, the sole decision writer `_set_acceptance_decision`); the run — host evidence packet, the one substantive panel (`_execute_task_acceptance_panel`), `acceptance_dialogue_history`, paid identity and the free replay `_refuse_identical_acceptance`; the retrieving rows' route-owned work order; and the LOCAL pre-binding preparation incident — material identity, one-use source-bound retry, the informed author's own finish/stop and the card/decision projections that make it visible (§6 Task acceptance)
|
||||
├── loop_llm_call.py ← Single-round LLM call + usage accounting
|
||||
|
|
@ -131,7 +131,8 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
|
|||
├── cancel_intents.py ← Durable cancel-intent projection: locked `state/cancel_intents.json` of ACTIVE intents (claim owner/pid + claim GENERATION fencing every mutation, `scope` single-vs-cascade) + forensic `cancel_intent` ledger rows; the ONE ingress `request_cancel`; strict fail-closed reads (`CancelIntentProjectionCorrupt`; a malformed row is disclosed once per row content, so the ~20 s watchdog cannot repeat it forever); owns `claim_is_abandoned` and `allow_settled_target` (§5; §10 invariants 14–15)
|
||||
├── owner_hurry.py ← Owner "hurry": a typed TASK-LOCAL acceleration latch, never a chat message; its durable `owner_hurry` projection is written by `update_json_locked` on its own keys only — never `write_task_result`, whose status-regression guard could drop concurrent terminal fields; effects `acceptance_skip_applied`, zero improvement passes via `effective_budget_profile`, advisory force-plan; dies with the attempt (`retry_reset`; `not_applied_before_terminal`) (§5)
|
||||
├── owner_quiz.py ← Owner-quiz lifecycle projection: `record_asked`, request-id-idempotent first-answer-wins `record_answered` (index validated against the STORED labels), structural `reconcile_terminal` (open → expired_terminal, closing the PAIRED `owner_wait`), `quiz_states` replay (§11.1)
|
||||
├── owner_wait.py ← Native owner-answer continuation: completed-tool source checkpoint, original-process sleep, and same-task recovery only through an acknowledged planned-restart handoff (§5)
|
||||
├── owner_wait.py ← Cognition serialization, native waits and acknowledged restart handoffs (§5, §6)
|
||||
├── budget_pause.py ← Exact monetary pause/Resume (§6)
|
||||
├── routing_wait.py ← SSOT of the durable routing-receipt waits (`wait_for_promotion_admission`, `wait_for_routing_annotation`), so the gateway picker confirms clicks through the SAME receipts the LLM routing tools poll
|
||||
├── outcomes.py ← Typed task-outcome and acceptance-decision authority keeping the lifecycle/execution/objective/review/artifact/verification/child-absorption axes separate; policy denials never masquerade as tool failures; an oversized verification ledger rides as a stub whose `summary` is re-projected from the artifact file, never a source for entries or axes (§6 Task lifecycle; §10.1)
|
||||
├── outcome_receipt_store.py ← Durable verification-receipt path/append/read authority + exact-row union of forked-child and canonical replicas; owns the zero-run WRITE enum (`incomplete`/`unknown` only — a zero-run "complete" is unverifiable self-report)
|
||||
|
|
@ -230,7 +231,7 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
|
|||
├── delegate_start_instructions.py ← Stable host start instructions + a complete separately-hashed coordination appendix; host pre-start sends no appendix
|
||||
├── delegate_target_drift.py ← Read-only authority-tree drift evidence for delegated capture; records changed paths without attributing them to the child or blocking a normal no-change disposition (§6 Delegated subagents)
|
||||
├── delegate_recovery.py ← Narrow exact-leaf recovery for proven crash + planned self-restart; vetoes every no-resume cause
|
||||
├── delegate_continuation.py ← Explicit continuation of a settled run the engine cancelled at its wall-clock cap (`continue_from`): typed gate over durable custody (own run, confirmed `wall_clock_exceeded`, result read, patch disposed, same executor/authority) + the host block; not recovery
|
||||
├── delegate_continuation.py ← Custody-gated continuation of disposed finite-timeout leaves; no crash recovery (§6 Delegated subagents)
|
||||
├── delegate_registration_policy.py ← `persistent_registration` + the STARTED-row field tables
|
||||
├── delegate_pending.py ← Durable pending-invocation replay preserving the original idempotency key + canonical start body
|
||||
├── delegate_custody_memo.py ← Process-local memo of the custody rows (`custody_rows`): an ordered `(st_dev, st_ino, consumed, st_mtime_ns)` fingerprint of the rotated events chain prefix plus a hash of the live file's consumed bytes, advanced by folding only appended bytes, refolded on any doubt, bypassed (never cached) while the chain is unreadable; inline legacy request bodies replaced by a re-readable locator; a warm cache with an exact fallback, not a durable projection
|
||||
|
|
|
|||
|
|
@ -155,7 +155,7 @@ An addressing call (`promote_chat_to_task`, `route_to_project`, `steer_task`, `e
|
|||
|
||||
#### Liveness census and the chat header
|
||||
|
||||
`GET /api/state` unites direct turns and ROOT managed queue tasks in `active_chat_activities`. Managed phases are `queued`, `budget_paused` (PENDING fenced until explicit resume), `working`, and `finalizing` (RUNNING with an open post-task checkpoint); a direct turn paused on its budget rail (#1196) is parked under its own id and reports those phases as `kind="direct_chat"`. Late-mounted chats hydrate from the queue. The census alone inserts live-set entries; finals and census delete them. Only `active_chat_activities_complete === true` with `supervisor_ready === true` clears every absent id, regardless of `kind`; incompleteness retains positive rows and clears nothing. Completeness requires successful reads of every required live-identity source. The page-wide snapshot sequencer projects reads every 3 seconds on Chat, otherwise every 20 seconds, without per-entry generations.
|
||||
`GET /api/state` combines direct turns and managed ROOT tasks in `active_chat_activities`: `queued`, `budget_pausing` (RUNNING saving continuation), `budget_paused` (PENDING until Resume), `working`, `finalizing` (open post-task checkpoint). Parked direct turns keep their ID and `kind="direct_chat"`. Late-mounted chats hydrate from the queue. Only census inserts live entries; finals/census delete them. Absent IDs clear only with `active_chat_activities_complete === true` AND `supervisor_ready === true`, regardless of kind; incomplete reads retain positives. Completeness requires every live-identity source. One page-wide sequencer reads every 3 seconds on Chat, 20 elsewhere, without per-entry generations.
|
||||
|
||||
Typing is a submission receipt: match `client_message_id`, retire local `Sending...`, request census; never insert, revive or extend liveness. `kind` selects `Thinking` for direct turns or `Working`/`Queued`/`Paused` for managed roots, never deletion immunity. Child typing still lacks kind; Telegram native typing ignores it. The header reducer reads connection, census, unconfirmed owner sends and live managed cards. A direct block is not a managed card: the header retains the census verdict, while a mounted unfinished block hosts its running indicator instead of the typing bubble. Terminal failure remains a task result, never reasonless header `Attention`. Only typing receipt, snapshot turn, durable routing receipt, replayed user row, turn conclusion or offline-queue eviction retires `Sending...`; live user echoes and socket writes cannot. Descendants enter the header through their cards, so typing alone cannot expose a child before progress creates its card. Unenumerated roots (including Presence, absent from the direct registry) likewise appear through cards alone. Only the reducer writes the badge, except the panel-boot `Online` seed.
|
||||
|
||||
|
|
|
|||
|
|
@ -12,7 +12,11 @@ A headless task is ADDRESSED when it is admitted, not when it is displayed (`log
|
|||
|
||||
The run is also NAMED at admission and chat promotion, without a new model call: a caller-supplied `title` (`ouroboros run --title` or the top-level contract field; `metadata.title` is refused with a 400 like `metadata.project_id`) is authorship and fills both `title` and `suggested_name`; otherwise the request's first line fills `suggested_name` ALONE, so a truncated prompt never outranks a real name coined later. A `task_named` frame is broadcast on admission so the live card is never born showing its status phrase as a title.
|
||||
|
||||
`queue_snapshot.json` is an atomic recovery and diagnostic projection, not a second scheduler: pending and running rows, acceptance and root-budget fences, worker counts, assignable capacity, any pool-disabled reason, and the latest bounded root focus. Startup restores a recent snapshot into an empty pending queue and never resurrects ordinary RUNNING work: it FENCES every surviving RUNNING row with a durable cancel intent (`reason='server_shutdown'`, ledgered as `terminalized_running`), which cancellation custody terminalizes a watchdog window later, expiring its open quiz and closing the paired owner wait; a PENDING child below it is marked `pending_parent_interrupted` and settled by the boot's `kill_workers`, so a closed window leaves neither ghost nor orphan. Only an owner-wait handoff with an acknowledged planned-restart transaction outlives snapshot age — and an EXACT budget pause (#1196): a `_budget_pause` marker carrying `exact_continuation` and its `checkpoint` locator is never assignable and is restored across any snapshot age without waking; a marker whose task-result `budget_pause` row or source is unreadable, mismatched or missing, and one whose root holds an acceptance fence or whose snapshot has malformed acceptance/budget fences, is retained under a typed `_budget_pause_hold` (never dropped or cancelled); a RUNNING row whose durable pause was complete at shutdown is parked, not fenced; a direct turn's parked record keeps `_is_direct_chat`; its grant rides `_budget_pause_resume`, and a grant that never reached a worker before a restart, or meets an exhausted wallet before dispatch, returns to the pause (`revoke_exact_budget_resume`), never to a replay or a terminal. Assignment carries the grant's `paused_duration_sec` into the RUNNING row as `budget_paused_sec`, a carrier the timeout rail subtracts from execution time beside the quota clock; `started_at` is the original. Assignment rechecks an exact child's root grant and fence identity: a new root pause revokes pending child grants, and another root Resume only restores eligibility for explicit child selection. A legacy zero-dispatch root Resume selects only that root through the existing hold; the root fence remains over its unselected children. Terminal tasks stay terminal, a task with an active cancel intent is left to custody, descendants of an accepted or sealed root finalize as cancelled, and malformed fence evidence fails closed. Assignment mirrors RUNNING into the durable task result for EVERY assigned task, not only a subagent: both orphan healers read the STORED status, and an unmirrored root is a ghost no snapshot-less boot can settle, so orphan reconciliation is a terminal writer for roots too, closing the same quiz and wait the task-done seam closes. Focus updates merge through the existing worker event path under the queue lock; no awareness timer, ledger, or wake is created. `direct_roots.json` carries the symmetric direct-root projection and one aggregate gap/freshness fact.
|
||||
`queue_snapshot.json` projects pending/running rows atomically, acceptance/root-budget fences, worker counts, assignable capacity, pool-disabled reason and bounded root focus; it is not another scheduler. Startup restores recent snapshots into empty PENDING. Ordinary surviving RUNNING rows receive durable `server_shutdown` cancel intents (`terminalized_running`); cancellation custody later terminalizes them, expires quizzes and closes owner waits. Boot `kill_workers` settles their PENDING children as `pending_parent_interrupted`, preventing ghosts/orphans. Only acknowledged planned-restart owner handoffs and exact budget pauses outlive snapshot age.
|
||||
|
||||
Exact pauses restore `_budget_pause`/checkpoint and `_is_direct_chat` without wake; unreadable sources or invalid/acceptance fences retain typed `_budget_pause_hold`. Complete RUNNING pauses park, not crash-fence. `_budget_pause_resume` grants are revoked if unconsumed at restart or money disappears before dispatch. Assignment carries original `started_at` and cumulative `budget_paused_sec`; child selection needs the current root grant/fence. Details: [§6 Exact budget pause and Resume](06-agent-core.md#exact-budget-pause-and-resume).
|
||||
|
||||
Terminal tasks stay terminal; active cancellation belongs to custody; descendants of accepted/sealed roots finalize cancelled except held exact pauses above. Malformed fences fail closed. Assigned tasks mirror RUNNING into its durable result: both orphan healers read STORED status, so an unmirrored root would escape snapshot-less cleanup. Root orphan settlement closes quiz/wait through the same terminal owner. Focus merges under the queue lock via worker events, with no new timer/ledger/wake. `direct_roots.json` projects direct roots and aggregate gap/freshness.
|
||||
|
||||
`supervisor/queue_schedules.py` owns the existing `state/scheduled_tasks.json` table and is its only writer. Every read-modify-write — the tick, the skill resync, the gateway upsert, `schedule_followup`'s cap-and-write, `manage_schedules` — enters through `schedule_transaction`, which takes BOTH locks itself, queue lock then the table's sidecar file lock. The order lives in the transaction, not in its callers: a caller-composed order is what inverted, when the follow-up tool's one-transaction cap-read-and-write reached for the queue lock inside that hold against the tick's queue-then-table. It is reentrant PER TABLE (keyed by the resolved lock path), so a nested call on another drive root still takes that root's lock rather than riding a depth counter. `load_schedule_store` is the strict read every writer and owner surface uses: only an ABSENT path is an empty table it may create, while a present non-regular file, bytes that do not parse and a row that is not an object each raise `ScheduleStoreUnreadable` naming which it was. A write is refused on it — the next atomic write would replace real rows with an empty document or drop the rows it could not parse — and a READ answers unavailable, because "no schedules" is a claim while an unparseable table means the state is unknown.
|
||||
|
||||
|
|
@ -51,7 +55,9 @@ Heartbeat and progress are different evidence: a heartbeat proves a process or l
|
|||
|
||||
A spawned or respawned slot is not assignable until its child's PID-bound `worker_ready` row arrives (`supervisor/worker_pool_lifecycle.py`). A live child's own `worker_starting` row, emitted before extension loading and agent construction, permits one extension of `WORKER_READY_WINDOW_SEC` to `WORKER_READY_CEILING_SEC` (300 seconds from birth, both in `runtime_limits.py`); foreign or pre-spawn rows cannot extend another slot. `worker_ready_window_extended` records that decision. A silent child keeps the original window, and logging failure cannot block startup. After `WORKER_READY_MAX_ATTEMPTS` failed attempts, `Worker.readiness_exhausted` is final for that exact slot — late events cannot reopen it. Total exhaustion, distinguished from busy/booting/reaping capacity and from a live owner-wait stack, closes pooled ingress (owner `/review` included) without blocking direct chat/control or boot/update recovery; once RUNNING completion custody has settled, `disable_exhausted_worker_pool` fails unstarted PENDING work honestly with a Restart hint, and a new task cannot clear the latch. Readiness stays separate from liveness and task idle time; a watcher error releases only still-booting, non-exhausted slots to the crash detector (`worker_ready_released`). Linux workers use forkserver; macOS and Windows use spawn.
|
||||
|
||||
Unexpected worker death reserves exact custody under the queue lock and enqueues `confirmed_dead_worker` on the reaper (`worker_health.recover_confirmed_dead_worker`). A saved terminal source wins even after signal death; unknown or incomplete file publication keeps the same job (`TerminalFileRecoveryPending`); only confirmed absence of one reaches the crash policy: a signal is an infrastructure failure, an otherwise eligible non-signal crash retries within `QUEUE_MAX_RETRIES`, preserving owner-wait replay restrictions and cost. Budget-continuation evidence, including a consumed grant, forbids ordinary retry; a failed revocation of an unconsumed grant retains the exact source and grant in a nonterminal hold, with snapshot persistence failure disclosed rather than claimed durable. A crash storm suppresses respawn while terminal sources settle, then its fence stops pooled admission; direct chat stays available. Startup runs the same terminal-file recovery in `_run_supervisor` after process custody and before `_startup_prune_sweeps` (the no-provider lifespan runs it too, spawning nothing); unknown or still-live ownership defers it rather than racing a writer, and any unresolved or protected source, or an ownership/read error, sets `preserve_task_sources`, skipping task-drive deletion for that pass. For older canonical scheduled rows, `_recover_terminal_task_files` restores that start binding only from a known non-direct child's positive running/started-at record when the existing fresh-queue and later-worker-boot checks prove it orphaned, with no pending queue owner or active cancel; the normal orphan reconciler and terminal guards retain authority, without resuming work. The recovery report includes `rebound`.
|
||||
`worker_health.recover_confirmed_dead_worker` reserves exact custody under the queue lock, then enqueues reaper `confirmed_dead_worker`. Saved terminal bytes win even after signal death; unknown/incomplete publication retains `TerminalFileRecoveryPending`. Only confirmed absence reaches crash policy: signal means infrastructure failure; eligible non-signal crashes retry within `QUEUE_MAX_RETRIES`, preserving cost and owner-wait restrictions. Budget continuation forbids ordinary retry; unconsumed grants re-park, consumed grants remain terminal, and failed revocation retains source/grant with honest persistence status (§6).
|
||||
|
||||
Crash storms suppress respawn while terminal sources settle, then fence pooled admission; direct chat remains. Startup runs the same recovery after process custody, before `_startup_prune_sweeps`, also in no-provider lifespan without spawning. Unknown/live ownership defers to avoid racing writers. Unresolved/protected sources or ownership/read errors set `preserve_task_sources`, skipping drive deletion. `_recover_terminal_task_files` may restore an older canonical scheduled row's start binding only from a known non-direct child's positive running/started-at record, with fresh-queue/later-boot orphan proof, no pending owner or active cancel. Existing orphan/terminal owners remain authoritative; nothing resumes. The report records `rebound`.
|
||||
|
||||
Startup and throttled maintenance reconcile three residue classes, the ~600 s pass off the loop thread (§10). Process custody checks strict PID, start-time, command, owner-task, session and generation evidence before it reaps. Delegated-run reconciliation applies the same owner-gone reasoning to harness rows (§6 Delegated subagents). Task, review and project reconciliation repair records whose producer no longer exists. None of these are command-line-class kill sweeps, and one instance never reaps another. The dedicated watchdog separately observes phase-stamped loop liveness and every native actor; a wedged chat turn alerts with a `/restart` hint, a loop stall only journals, and neither kills a thread. Other owner conversations run on independent native actors, without a second scheduler.
|
||||
|
||||
|
|
|
|||
File diff suppressed because one or more lines are too long
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
Machine extraction of the `docs/ARCHITECTURE.md` "Data layout (`~/Ouroboros/`)" tree — the durable-file orientation carrier (this tree's counterpart of the reference PERSISTENCE_OWNERS derivation checklist) — regenerated by `python scripts/regenerate_inventories.py`. Do not edit. Every entry is probed against reality: repo entries must exist as tracked paths; data-plane entries must appear as a literal in the runtime sources that construct them. A durable file renamed or removed in code while its tree row survives = red (`tests/test_generated_inventories.py`).
|
||||
|
||||
Source: `docs/architecture/01-high-level-architecture.md`, physical LF lines 597-686; UTF-8 SHA-256 `5da0c28ecc68b10a3d7ccea4b035fe4312ddaec1ab4265cd0c1c0c2627281032`.
|
||||
Source: `docs/architecture/01-high-level-architecture.md`, physical LF lines 600-689; UTF-8 SHA-256 `d7355454cd97f535e1519c82606ac2e69ba5c71bf8d8f0edec4c3124c94ed4e3`.
|
||||
|
||||
- entries: **79** (code-ref: 72, repo-dir: 6, repo-path: 1)
|
||||
|
||||
|
|
|
|||
|
|
@ -2,14 +2,14 @@
|
|||
|
||||
AST-derived inventory of compatibility facades, regenerated by `python scripts/regenerate_inventories.py`. Do not edit. A facade row is any runtime module whose top-level `from <population module> import ...` statements carry the `noqa: F401` re-export marker — the codebase's declared "this binding exists for its binding, not for this module's own use" convention (reference FACADE_CONSUMERS method). Leaf domains come from `ouroboros/domains.toml`; a leaf outside the facade's domain is marked ✗ (that edge also appears in the manifest's pinned direction matrix). `tests/test_generated_inventories.py` pins byte-identity, so any re-export surface change must regenerate this file.
|
||||
|
||||
- facade modules: **61**; marked re-export bindings: **2319**; cross-domain facade→leaf pairs: **132**
|
||||
- facade modules: **62**; marked re-export bindings: **2330**; cross-domain facade→leaf pairs: **132**
|
||||
|
||||
| facade | domain | bindings | leaves |
|
||||
|---|---|---:|---|
|
||||
| `launcher.py` | D18 | 2 | `ouroboros/launcher_windows_runtime.py` (2) |
|
||||
| `ouroboros/agent.py` | D01 | 32 | `ouroboros/agent_dispatch.py` (15)<br>`ouroboros/agent_startup_checks.py` (4)<br>`ouroboros/config.py` (2 ✗D12)<br>`ouroboros/subagent_dispatch_notes.py` (4 ✗D07)<br>`ouroboros/subagents.py` (7 ✗D07) |
|
||||
| `ouroboros/agent_task_pipeline.py` | D01 | 28 | `ouroboros/dialogue_provenance.py` (2 ✗D15)<br>`ouroboros/post_task_synthesis.py` (11)<br>`ouroboros/synthesis_cost_text.py` (5)<br>`ouroboros/task_finalization.py` (10) |
|
||||
| `ouroboros/config.py` | D12 | 133 | `ouroboros/model_slots.py` (17)<br>`ouroboros/provider_models.py` (6 ✗D02)<br>`ouroboros/review_model_routes.py` (10)<br>`ouroboros/runtime_limits.py` (62)<br>`ouroboros/settings_defaults.py` (19)<br>`ouroboros/settings_integrity.py` (4)<br>`ouroboros/settings_scales.py` (13)<br>`ouroboros/update_channels.py` (2) |
|
||||
| `ouroboros/config.py` | D12 | 140 | `ouroboros/model_slots.py` (17)<br>`ouroboros/provider_models.py` (6 ✗D02)<br>`ouroboros/review_model_routes.py` (10)<br>`ouroboros/runtime_limits.py` (65)<br>`ouroboros/settings_defaults.py` (19)<br>`ouroboros/settings_integrity.py` (4)<br>`ouroboros/settings_scales.py` (17)<br>`ouroboros/update_channels.py` (2) |
|
||||
| `ouroboros/context.py` | D03 | 4 | `ouroboros/context_runtime_facts.py` (4) |
|
||||
| `ouroboros/delegate_custody.py` | D07 | 10 | `ouroboros/delegate_custody_reconcile.py` (9)<br>`ouroboros/delegate_evidence.py` (1) |
|
||||
| `ouroboros/extension_loader.py` | D14 | 101 | `ouroboros/contracts/plugin_api.py` (7 ✗D19)<br>`ouroboros/extension_child_catalog.py` (8)<br>`ouroboros/extension_companion.py` (3)<br>`ouroboros/extension_import_staging.py` (6)<br>`ouroboros/extension_isolated_deps.py` (4)<br>`ouroboros/extension_liveness.py` (8)<br>`ouroboros/extension_plugin_api.py` (6)<br>`ouroboros/extension_registry_state.py` (20)<br>`ouroboros/extension_surface_names.py` (12)<br>`ouroboros/extension_ui_validation.py` (5)<br>`ouroboros/gateway/host_service.py` (1 ✗D11)<br>`ouroboros/provider_models.py` (1 ✗D02)<br>`ouroboros/skill_loader.py` (15)<br>`ouroboros/skill_token.py` (1)<br>`ouroboros/tools/skill_exec.py` (1)<br>`ouroboros/utils.py` (3 ✗D18) |
|
||||
|
|
@ -59,9 +59,10 @@ AST-derived inventory of compatibility facades, regenerated by `python scripts/r
|
|||
| `ouroboros/tools/subagent_integration.py` | D07 | 13 | `ouroboros/headless.py` (2 ✗D17)<br>`ouroboros/tools/subagent_integration_delegated.py` (11) |
|
||||
| `ouroboros/usage_accounting.py` | D16 | 43 | `ouroboros/_usage_cache_splits.py` (4)<br>`ouroboros/_usage_rows.py` (7)<br>`ouroboros/_usage_rows_memo.py` (6)<br>`ouroboros/usage_ledger.py` (18)<br>`ouroboros/usage_legacy_import.py` (5)<br>`ouroboros/utils.py` (3 ✗D18) |
|
||||
| `server.py` | D11 | 51 | `ouroboros/server_liveness.py` (4)<br>`ouroboros/server_maintenance.py` (14)<br>`ouroboros/server_owner_routing.py` (5)<br>`ouroboros/server_process.py` (6)<br>`ouroboros/server_restart.py` (8)<br>`ouroboros/server_routing_context.py` (14) |
|
||||
| `supervisor/events.py` | D08 | 95 | `ouroboros/config.py` (1 ✗D12)<br>`ouroboros/contracts/task_constraint.py` (1 ✗D19)<br>`ouroboros/cost_projection.py` (3 ✗D16)<br>`ouroboros/subagent_messages.py` (1 ✗D07)<br>`ouroboros/task_results.py` (2 ✗D17)<br>`ouroboros/tool_capabilities.py` (2 ✗D04)<br>`ouroboros/utils.py` (4 ✗D18)<br>`supervisor/cognitive_operations.py` (2)<br>`supervisor/events_budget.py` (4)<br>`supervisor/events_chat_delivery.py` (9)<br>`supervisor/events_coop_checkpoint.py` (6)<br>`supervisor/events_evolution_done.py` (1)<br>`supervisor/events_project_routing.py` (10)<br>`supervisor/events_runtime_controls.py` (6)<br>`supervisor/events_schedule_task.py` (4)<br>`supervisor/events_subagent_admission.py` (15)<br>`supervisor/events_task_done.py` (8)<br>`supervisor/events_worker_reports.py` (8)<br>`supervisor/log_addressing.py` (4)<br>`supervisor/queue_transitions.py` (1)<br>`supervisor/steering.py` (2 ✗D09)<br>`supervisor/task_dispatch.py` (1) |
|
||||
| `supervisor/events.py` | D08 | 96 | `ouroboros/config.py` (1 ✗D12)<br>`ouroboros/contracts/task_constraint.py` (1 ✗D19)<br>`ouroboros/cost_projection.py` (3 ✗D16)<br>`ouroboros/subagent_messages.py` (1 ✗D07)<br>`ouroboros/task_results.py` (2 ✗D17)<br>`ouroboros/tool_capabilities.py` (2 ✗D04)<br>`ouroboros/utils.py` (4 ✗D18)<br>`supervisor/cognitive_operations.py` (2)<br>`supervisor/events_budget.py` (5)<br>`supervisor/events_chat_delivery.py` (9)<br>`supervisor/events_coop_checkpoint.py` (6)<br>`supervisor/events_evolution_done.py` (1)<br>`supervisor/events_project_routing.py` (10)<br>`supervisor/events_runtime_controls.py` (6)<br>`supervisor/events_schedule_task.py` (4)<br>`supervisor/events_subagent_admission.py` (15)<br>`supervisor/events_task_done.py` (8)<br>`supervisor/events_worker_reports.py` (8)<br>`supervisor/log_addressing.py` (4)<br>`supervisor/queue_transitions.py` (1)<br>`supervisor/steering.py` (2 ✗D09)<br>`supervisor/task_dispatch.py` (1) |
|
||||
| `supervisor/git_ops.py` | D10 | 37 | `ouroboros/utils.py` (1 ✗D18)<br>`supervisor/git_ops_remotes.py` (4)<br>`supervisor/git_ops_rescue.py` (8)<br>`supervisor/git_ops_reset.py` (10)<br>`supervisor/git_ops_updates.py` (8)<br>`supervisor/state.py` (4 ✗D08)<br>`supervisor/update_recovery.py` (2) |
|
||||
| `supervisor/queue.py` | D08 | 112 | `ouroboros/config.py` (6 ✗D12)<br>`ouroboros/contracts/task_contract.py` (3 ✗D19)<br>`ouroboros/schedule_contract.py` (2)<br>`ouroboros/skill_loader.py` (1 ✗D14)<br>`ouroboros/utils.py` (3 ✗D18)<br>`supervisor/evolution_lifecycle.py` (12 ✗D15)<br>`supervisor/message_bus.py` (3)<br>`supervisor/queue_schedules.py` (19)<br>`supervisor/queue_snapshot.py` (5)<br>`supervisor/queue_timeouts.py` (8)<br>`supervisor/queue_transitions.py` (3)<br>`supervisor/schedule_lifecycle.py` (3)<br>`supervisor/schedule_time.py` (7)<br>`supervisor/state.py` (6)<br>`supervisor/task_admission.py` (9)<br>`supervisor/task_lifecycle.py` (17 ✗D09)<br>`supervisor/task_reaper.py` (5 ✗D09) |
|
||||
| `supervisor/queue_transitions.py` | D08 | 3 | `supervisor/budget_resume.py` (3) |
|
||||
| `supervisor/state.py` | D08 | 4 | `ouroboros/utils.py` (4 ✗D18) |
|
||||
| `supervisor/task_lifecycle.py` | D09 | 32 | `supervisor/cancel_publication.py` (20)<br>`supervisor/queue_transitions.py` (11 ✗D08)<br>`supervisor/task_admission.py` (1 ✗D08) |
|
||||
| `supervisor/task_model_wait.py` | D08 | 1 | `ouroboros/model_wait.py` (1 ✗D09) |
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
Machine extraction of `docs/ARCHITECTURE.md` §11.1 (the frozen-ABI SSOT), regenerated by `python scripts/regenerate_inventories.py`. Do not edit — edit the owning chapter named in the Source line and regenerate; `tests/test_generated_inventories.py` pins byte-identity and the resolution invariants (a §11.1 row whose owner or anchor file disappeared from the tree = red).
|
||||
|
||||
Source: `docs/architecture/11-frozen-contracts-v1.md`, physical LF lines 7-39; UTF-8 SHA-256 `47aac4a4fb325ab5fe057456a4637cf910c4964a014be62dd0332bbfec368810`.
|
||||
Source: `docs/architecture/11-frozen-contracts-v1.md`, physical LF lines 7-39; UTF-8 SHA-256 `89dc44662785d2089e88e365c2ebf85e260227cc8828973573bd81b6555f52b6`.
|
||||
|
||||
- table rows: **28**
|
||||
- browser-envelope prose owners:
|
||||
|
|
|
|||
|
|
@ -538,25 +538,11 @@ class OuroborosAgent:
|
|||
|
||||
task_metadata = dict(task.get("metadata") or {}) if isinstance(task.get("metadata"), dict) else {}
|
||||
for key in (
|
||||
"parent_task_id",
|
||||
"root_task_id",
|
||||
"session_id",
|
||||
"actor_id",
|
||||
"delegation_role",
|
||||
"role",
|
||||
"workspace_root",
|
||||
"workspace_mode",
|
||||
"memory_mode",
|
||||
"drive_root",
|
||||
"child_drive_root",
|
||||
"budget_drive_root",
|
||||
"root_cost_ceiling_usd",
|
||||
"model_lane",
|
||||
"requested_model_lane",
|
||||
"effective_model_lane",
|
||||
"model",
|
||||
"use_local_model",
|
||||
"requested_executor",
|
||||
"parent_task_id", "root_task_id", "session_id", "actor_id", "delegation_role", "role",
|
||||
"workspace_root", "workspace_mode", "memory_mode",
|
||||
"drive_root", "child_drive_root", "budget_drive_root", "root_cost_ceiling_usd",
|
||||
"model_lane", "requested_model_lane", "effective_model_lane",
|
||||
"model", "use_local_model", "requested_executor",
|
||||
# `effective_executor`/`capability_delta` are deliberately NOT here: this
|
||||
# projection is only READ for `effective_model_lane` (grandchild
|
||||
# inheritance), the child learns its own reduction from the prompt and the
|
||||
|
|
|
|||
|
|
@ -718,6 +718,7 @@ def _exact_continuation_row(limit_ctx: Any, ctx: Any, *, pause_id: str, rail: st
|
|||
trace = limit_ctx.llm_trace if isinstance(limit_ctx.llm_trace, dict) else {}
|
||||
seen = set(limit_ctx.owner_msg_seen or ())
|
||||
point = resume_point(messages, limit_ctx.round_idx)
|
||||
point["budget_tail"] = getattr(limit_ctx, "budget_tail", "tool")
|
||||
# The rail already stamped its terminal projection on the live usage, and a
|
||||
# hold leaves its own transient row there; neither may travel into the
|
||||
# resumed loop's eventual honest terminal.
|
||||
|
|
|
|||
|
|
@ -500,17 +500,16 @@ def run_llm_loop(
|
|||
# Both continuing tool tails and unfinished no-tool rounds owe budget checks.
|
||||
pending_tool_budget, pending_tool_calls, pending_no_tool_budget = bool(saved), None, False
|
||||
if saved_pause:
|
||||
# Same-ID exact continuation after an owner Resume (#1196): the
|
||||
# saved cognition comes back (nothing is re-executed by the host;
|
||||
# an interrupted batch's unanswered calls are closed as execution-
|
||||
# unknown), then the ordinary budget tail decides with the
|
||||
# refreshed threshold.
|
||||
# Restore cognition under the same ID, closing unanswered calls as
|
||||
# execution-unknown without replay. The shared budget tail checks
|
||||
# the owner-refreshed threshold before any new model call.
|
||||
(active_model, active_effort, active_use_local, active_context_mode,
|
||||
round_idx, context_fit_plan) = resume_paused_loop(
|
||||
tools, saved_pause, messages, llm_trace, accumulated_usage, _owner_msg_seen,
|
||||
budget_remaining_usd=budget_remaining_usd)
|
||||
cost_ceiling = _resolve_task_cost_ceiling(tools._ctx, budget_remaining_usd)
|
||||
pending_tool_budget = True
|
||||
pending_no_tool_budget = saved_pause.get("resume_point", {}).get("budget_tail") == "no_tool"
|
||||
pending_tool_budget, free_redial = not pending_no_tool_budget, pending_no_tool_budget
|
||||
while True:
|
||||
if free_redial or pending_tool_budget:
|
||||
free_redial = False # Tool tails and transport waits retain their logical round.
|
||||
|
|
@ -540,14 +539,14 @@ def run_llm_loop(
|
|||
ctx.active_effort = active_effort
|
||||
ctx.active_use_local = active_use_local
|
||||
|
||||
# One forced-wrap-up context per round: consumed by the round-limit
|
||||
# path and supervisor finalize_now control path below.
|
||||
# One context per round, shared by budget, round-limit and finalize_now rails.
|
||||
limit_ctx = _RoundLimitContext(
|
||||
messages, llm, active_model, active_effort, max_retries, drive_logs,
|
||||
task_id, round_idx, event_queue, accumulated_usage, task_type,
|
||||
active_use_local, MAX_ROUNDS, drive_root=drive_root, llm_trace=llm_trace,
|
||||
task_id, round_idx, event_queue, accumulated_usage, task_type, active_use_local, MAX_ROUNDS,
|
||||
drive_root=drive_root, llm_trace=llm_trace,
|
||||
incoming_messages=incoming_messages, owner_msg_seen=_owner_msg_seen, tool_schemas=tool_schemas)
|
||||
_finalize_limit_ctx(limit_ctx, tools, llm_trace)
|
||||
limit_ctx.budget_tail = "tool" if pending_tool_budget else "no_tool"
|
||||
if MAX_ROUNDS is not None and round_idx > MAX_ROUNDS:
|
||||
# Live hold: a paid [ROUND_LIMIT] dial would be a resend (no wake receipt) — no-call unknown terminal.
|
||||
if _delegate_hold_close(tools, drive_logs=drive_logs, task_id=task_id,
|
||||
|
|
@ -625,17 +624,13 @@ def run_llm_loop(
|
|||
seal_task_transcript(messages)
|
||||
|
||||
model_call = _RoundModelCallContext(
|
||||
llm=llm, messages=messages, tools=tools, context_fit_plan=context_fit_plan,
|
||||
active_model=active_model, tool_schemas=tool_schemas,
|
||||
active_effort=active_effort, max_retries=max_retries,
|
||||
drive_logs=drive_logs, task_id=task_id, round_idx=round_idx,
|
||||
event_queue=event_queue,
|
||||
accumulated_usage=accumulated_usage,
|
||||
task_type=task_type,
|
||||
active_use_local=active_use_local,
|
||||
active_context_mode=active_context_mode,
|
||||
drive_root=drive_root, emit_progress=emit_progress,
|
||||
)
|
||||
llm=llm, messages=messages, tools=tools, context_fit_plan=context_fit_plan,
|
||||
active_model=active_model, tool_schemas=tool_schemas,
|
||||
active_effort=active_effort, max_retries=max_retries,
|
||||
drive_logs=drive_logs, task_id=task_id, round_idx=round_idx,
|
||||
event_queue=event_queue, accumulated_usage=accumulated_usage, task_type=task_type,
|
||||
active_use_local=active_use_local, active_context_mode=active_context_mode,
|
||||
drive_root=drive_root, emit_progress=emit_progress)
|
||||
try:
|
||||
msg, cost, active_context_mode = _call_round_model(model_call)
|
||||
except ModelWaitInterrupted as error:
|
||||
|
|
@ -660,13 +655,8 @@ def run_llm_loop(
|
|||
drive_logs=drive_logs, task_id=task_id, model=active_model, emit_progress=emit_progress)
|
||||
if msg is None and _fallback_chain_allowed(ctx, last_error_kind, transport_wait, accumulated_usage):
|
||||
_episode_before_chain = transport_wait is not None
|
||||
(
|
||||
msg,
|
||||
active_model,
|
||||
active_use_local,
|
||||
context_fit_plan,
|
||||
active_context_mode,
|
||||
) = _run_cross_model_fallback_chain(
|
||||
(msg, active_model, active_use_local,
|
||||
context_fit_plan, active_context_mode) = _run_cross_model_fallback_chain(
|
||||
llm=llm, ctx=ctx, tools=tools, messages=messages, active_model=active_model,
|
||||
active_use_local=active_use_local, tool_schemas=tool_schemas, active_effort=active_effort,
|
||||
max_retries=max_retries, drive_logs=drive_logs, task_id=task_id, round_idx=round_idx,
|
||||
|
|
@ -723,10 +713,10 @@ def run_llm_loop(
|
|||
|
||||
if getattr(tools._ctx, "_skill_finalization_injected", False):
|
||||
tools._ctx._skill_finalization_injected = False
|
||||
assistant_msg = dict(msg)
|
||||
assistant_msg.setdefault("role", "assistant")
|
||||
assistant_msg = dict(msg, role=msg.get("role", "assistant"))
|
||||
messages.append(assistant_msg)
|
||||
_emit_round_progress(content, msg, emit_progress, llm_trace)
|
||||
limit_ctx.budget_tail = "tool"
|
||||
handle_tool_calls(
|
||||
tool_calls, tools, drive_logs, task_id, stateful_executor,
|
||||
messages, llm_trace, emit_progress
|
||||
|
|
|
|||
|
|
@ -59,7 +59,7 @@ def _check_budget_limits(
|
|||
finish_reason = "🚫 Task rejected. Total budget exhausted. Please increase TOTAL_BUDGET in settings."
|
||||
accumulated_usage["execution_status"] = "failed"
|
||||
accumulated_usage["reason_code"] = "budget_exhausted"
|
||||
if ctx.round_idx <= 1:
|
||||
if ctx.round_idx <= 1 and not accumulated_usage.get("rounds"):
|
||||
trace = ctx.llm_trace if isinstance(ctx.llm_trace, dict) else {}
|
||||
tool_ctx = getattr(getattr(ctx, "tools", None), "_ctx", None)
|
||||
suffix = (
|
||||
|
|
@ -297,10 +297,9 @@ def _soft_land_exhausted_ceiling(
|
|||
limit_ctx: "_RoundLimitContext",
|
||||
cost_ceiling: "task_pacing.CostCeiling",
|
||||
) -> Optional[Tuple[str, Dict[str, Any], Dict[str, Any]]]:
|
||||
"""Typed soft landing (v6.91): a root cap at or below the planning margin
|
||||
wraps up BEFORE a work round through the same priced candidate as the
|
||||
last-fit rail; an unaffordable wrap-up ends as budget_wrapup_unaffordable
|
||||
instead of a fence pause. None when the ceiling is not exhausted."""
|
||||
"""Pause a root cap at/below the planning margin, even before its first
|
||||
call: owner Resume may use already-authorized headroom. Only an actor
|
||||
without exact continuation retains the priced legacy soft landing."""
|
||||
if cost_ceiling.state != task_pacing.COST_CEILING_EXHAUSTED_SOFT_LAND:
|
||||
return None
|
||||
cap_text = (
|
||||
|
|
@ -315,10 +314,8 @@ def _soft_land_exhausted_ceiling(
|
|||
f"Per-task tree cap {cap_text} leaves no working room above the "
|
||||
f"wrap-up planning margin ({margin_text}). Budget exhausted."
|
||||
)
|
||||
if limit_ctx.round_idx > 1:
|
||||
# Work exists: pause exactly instead of pricing a wrap-up call.
|
||||
budget_pause.request_pause(limit_ctx, rail=budget_pause.RAIL_SOFT_LAND, scope="root",
|
||||
reason_text=soft_land_reason)
|
||||
budget_pause.request_pause(limit_ctx, rail=budget_pause.RAIL_SOFT_LAND, scope="root",
|
||||
reason_text=soft_land_reason)
|
||||
trace = limit_ctx.llm_trace if isinstance(limit_ctx.llm_trace, dict) else {}
|
||||
priced_prompt = _loop()._prepare_forced_prompt(
|
||||
limit_ctx, f"[BUDGET LIMIT] {soft_land_reason} {_loop()._FORCED_BEST_EFFORT_TAIL}", trace,
|
||||
|
|
@ -771,8 +768,10 @@ def _finish_no_tool_round_budget(
|
|||
nothing else: no metered-baseline bookkeeping (no tools ran) and no
|
||||
`_prepare_post_tool_budget_context`, whose delivery-control arming belongs
|
||||
to a tool batch's effects. The current delivery candidate is untouched, so a
|
||||
budget exit still wraps up the answer the round produced.
|
||||
eligible budget exit checkpoints the answer the round produced. Its cold
|
||||
continuation must return to this tail without tool-only control arming.
|
||||
"""
|
||||
ctx.budget_tail = "no_tool"
|
||||
result = _loop()._check_budget_limits(ctx, budget_remaining_usd, cost_ceiling=cost_ceiling)
|
||||
if result is None:
|
||||
return None
|
||||
|
|
|
|||
|
|
@ -744,14 +744,11 @@ def run_delegated_review_session(
|
|||
custody_drive: Any,
|
||||
invocation: SessionInvocation,
|
||||
) -> Dict[str, Any]:
|
||||
"""Start, watch, settle and collect one delegated read-only review.
|
||||
This is every review surface's single session transport. It pins one
|
||||
subscription harness, asks for schema only when the effective adapter can
|
||||
carry it, stores the canonical start request before POST, and replays only
|
||||
an explicit pending invocation token. A bound token joins its existing run;
|
||||
reconcile-only mode never mints a replacement. The nanny owns verified
|
||||
cancellation at ``timeout_sec`` and reads the full primary output before
|
||||
settling through ``delegate_custody``.
|
||||
"""All reviews share this transport: pin one subscription harness, request
|
||||
schema only if its effective adapter supports it, store the canonical body
|
||||
before POST, and replay only an explicit pending token. Bound tokens join
|
||||
existing runs; reconciliation never replaces them. The nanny verifies
|
||||
cancellation at ``timeout_sec`` and reads full output before ``delegate_custody`` settlement.
|
||||
"""
|
||||
from ouroboros import delegate_custody as custody
|
||||
from ouroboros.claudexor_daemon import ensure_owned_gateway
|
||||
|
|
|
|||
|
|
@ -566,8 +566,7 @@ def _delegate_start(ctx: ToolContext, prompt: str, max_seconds: Optional[int] =
|
|||
facts = {key: value for key, value in claim_refusal.items()
|
||||
if key not in {"reason", "detail"}}
|
||||
return _fail(
|
||||
"delegate_start", reason, detail,
|
||||
**facts,
|
||||
"delegate_start", reason, detail, **facts,
|
||||
**_retire_orphaned_registration(ctx, gateway, owned_project_id, project_persistent=project_persistent, history_facts=history_facts,
|
||||
definite_refusal=True,
|
||||
reason=reason, invocation_id=invocation_id, snapshot_id=snapshot_id,
|
||||
|
|
@ -580,21 +579,16 @@ def _delegate_start(ctx: ToolContext, prompt: str, max_seconds: Optional[int] =
|
|||
"NOT started: a run launched without its custody trail would be "
|
||||
"unfindable if this worker died. Fix the drive/event log and retry.",
|
||||
**({"definitely_unrun": True} if not recovering else {}), **_retire_orphaned_registration(ctx, gateway, owned_project_id, project_persistent=project_persistent, history_facts=history_facts,
|
||||
definite_refusal=not recovering,
|
||||
reason="start_request_row_unwritable",
|
||||
invocation_id=invocation_id,
|
||||
snapshot_id=("" if recovering else snapshot_id)))
|
||||
definite_refusal=not recovering, reason="start_request_row_unwritable",
|
||||
invocation_id=invocation_id, snapshot_id=("" if recovering else snapshot_id)))
|
||||
handle = gateway.start_run(request_body, idempotency_key=invocation_id)
|
||||
run_id = str(handle.get("runId") or handle.get("jobId") or "")
|
||||
if not run_id:
|
||||
return _fail("delegate_start", "queued_without_run_id",
|
||||
f"Claudexor returned a queued handle without a run id: {handle!r}",
|
||||
pending_invocation_id=invocation_id,
|
||||
retry_hint=_RETRY_HINT,
|
||||
pending_invocation_id=invocation_id, retry_hint=_RETRY_HINT,
|
||||
**_retire_orphaned_registration(ctx, gateway, owned_project_id, project_persistent=project_persistent, history_facts=history_facts,
|
||||
definite_refusal=False,
|
||||
reason="queued_without_run_id",
|
||||
invocation_id=invocation_id))
|
||||
definite_refusal=False, reason="queued_without_run_id", invocation_id=invocation_id))
|
||||
except ClaudexorUnavailable as exc:
|
||||
# A registration we created BEFORE the start must not outlive a failed start.
|
||||
# It used to be left behind with nothing anywhere naming its id.
|
||||
|
|
@ -611,10 +605,8 @@ def _delegate_start(ctx: ToolContext, prompt: str, max_seconds: Optional[int] =
|
|||
**({"definitely_unrun": True} if not requested else {}),
|
||||
reset_at=getattr(exc, "reset_at", ""), **pending,
|
||||
**_retire_orphaned_registration(ctx, gateway, owned_project_id, project_persistent=project_persistent, history_facts=history_facts,
|
||||
definite_refusal=definite,
|
||||
reason=str(getattr(exc, "code", "")),
|
||||
invocation_id=invocation_id,
|
||||
snapshot_id=("" if recovering else snapshot_id)))
|
||||
definite_refusal=definite, reason=str(getattr(exc, "code", "")),
|
||||
invocation_id=invocation_id, snapshot_id=("" if recovering else snapshot_id)))
|
||||
except BaseException as exc:
|
||||
# EVERY pre-custody exit leaves a durable disposition, including the ones no
|
||||
# typed handler claims (a bug here, a timeout, a signal). NEVER retired: an
|
||||
|
|
@ -651,13 +643,9 @@ def _delegate_start(ctx: ToolContext, prompt: str, max_seconds: Optional[int] =
|
|||
from ouroboros.tools.control import maybe_emit_delegated_run_fanout
|
||||
maybe_emit_delegated_run_fanout(ctx, run_id=run_id, route_id=route.route_id, objective=text, durable=durable)
|
||||
return _started_payload(handle, run_id, route, access, authority, root,
|
||||
durable=durable, recovering=recovering,
|
||||
invocation_id=invocation_id,
|
||||
snapshot_id=snapshot_id, target_root=target_root,
|
||||
baseline_sha=baseline_sha,
|
||||
resource_ref=resource_ref,
|
||||
processing=processing_info,
|
||||
continuation=continuation,
|
||||
durable=durable, recovering=recovering, invocation_id=invocation_id,
|
||||
snapshot_id=snapshot_id, target_root=target_root, baseline_sha=baseline_sha,
|
||||
resource_ref=resource_ref, processing=processing_info, continuation=continuation,
|
||||
max_seconds=seconds, max_seconds_basis=seconds_basis,
|
||||
engine_version=str(getattr(gateway, "engine_version", "") or ""))
|
||||
|
||||
|
|
|
|||
|
|
@ -29,24 +29,12 @@ from ouroboros._usage_response import (
|
|||
from ouroboros.review_dispatch import invoke_bound_api_review_paid_stamp
|
||||
from ouroboros.transport_custody import release_pre_dispatch_attempt
|
||||
from ouroboros.usage_ledger import ( # noqa: F401 — re-exported substrate
|
||||
LEDGER_REL,
|
||||
QUARANTINE_REL,
|
||||
LedgerResumeState,
|
||||
UsageAccountingError,
|
||||
UsageLedgerCorrupt,
|
||||
_append_bytes_fsync,
|
||||
_append_rows_locked,
|
||||
_drive_root,
|
||||
_final_rows,
|
||||
_ledger_resume_state,
|
||||
_locked,
|
||||
_named_lock,
|
||||
_number,
|
||||
_read_new_records_locked,
|
||||
_read_records_locked,
|
||||
_TERMINAL,
|
||||
_validate_records,
|
||||
_write_bytes_atomic_fsync,
|
||||
LEDGER_REL, QUARANTINE_REL, LedgerResumeState,
|
||||
UsageAccountingError, UsageLedgerCorrupt,
|
||||
_append_bytes_fsync, _append_rows_locked,
|
||||
_drive_root, _final_rows, _ledger_resume_state,
|
||||
_locked, _named_lock, _number,
|
||||
_read_new_records_locked, _read_records_locked, _TERMINAL, _validate_records, _write_bytes_atomic_fsync,
|
||||
)
|
||||
from ouroboros.utils import append_jsonl, atomic_write_json, utc_now_iso # noqa: F401 -- the accounting module keeps its historical import surface for the L-C2 leaf
|
||||
from ouroboros._usage_rows import ( # noqa: F401 (re-exported substrate vocabulary)
|
||||
|
|
@ -1497,19 +1485,10 @@ def execute_physical_attempt(
|
|||
except BaseException as exc:
|
||||
if isinstance(exc, PhysicalAttemptLimitExceeded):
|
||||
_record_attempt_capture(
|
||||
reservation,
|
||||
request,
|
||||
"released",
|
||||
candidate_manifest_ref=manifest_ref,
|
||||
exc=exc,
|
||||
)
|
||||
reservation, request, "released", candidate_manifest_ref=manifest_ref, exc=exc)
|
||||
raise
|
||||
failure = _pre_dispatch_failure(
|
||||
reservation,
|
||||
request,
|
||||
exc,
|
||||
candidate_manifest_ref=manifest_ref,
|
||||
)
|
||||
reservation, request, exc, candidate_manifest_ref=manifest_ref)
|
||||
if failure is exc:
|
||||
raise
|
||||
raise failure from exc
|
||||
|
|
@ -1567,19 +1546,10 @@ async def execute_physical_attempt_async(
|
|||
except BaseException as exc:
|
||||
if isinstance(exc, PhysicalAttemptLimitExceeded):
|
||||
_record_attempt_capture(
|
||||
reservation,
|
||||
request,
|
||||
"released",
|
||||
candidate_manifest_ref=manifest_ref,
|
||||
exc=exc,
|
||||
)
|
||||
reservation, request, "released", candidate_manifest_ref=manifest_ref, exc=exc)
|
||||
raise
|
||||
failure = _pre_dispatch_failure(
|
||||
reservation,
|
||||
request,
|
||||
exc,
|
||||
candidate_manifest_ref=manifest_ref,
|
||||
)
|
||||
reservation, request, exc, candidate_manifest_ref=manifest_ref)
|
||||
if failure is exc:
|
||||
raise
|
||||
raise failure from exc
|
||||
|
|
|
|||
|
|
@ -80,7 +80,7 @@ def test_legacy_root_resume_selects_only_root_then_one_child(tmp_path, monkeypat
|
|||
monkeypatch.setattr(state, "budget_remaining", lambda *_a, **_k: 5.0)
|
||||
root = _fenced_member(workers, "root", "root")
|
||||
root.pop("parent_task_id")
|
||||
child = _fenced_member(workers, "child", "root")
|
||||
_fenced_member(workers, "child", "root")
|
||||
sibling = _fenced_member(workers, "sibling", "root")
|
||||
fence = _set_root_budget_pause_locked("root", {})
|
||||
if marker:
|
||||
|
|
|
|||
|
|
@ -10,14 +10,20 @@ ceiling, and a ready answer that still buys no extra wrap-up.
|
|||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
import copy
|
||||
import json
|
||||
import pathlib
|
||||
import queue
|
||||
import time
|
||||
from types import SimpleNamespace
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
import pytest
|
||||
|
||||
from ouroboros import task_pacing
|
||||
from ouroboros.contracts.task_contract import normalize_budget_profile
|
||||
from ouroboros.loop import _RoundLimitContext, _finish_no_tool_round_budget
|
||||
from tests.test_acceptance_async_loop import full_loop as full_loop
|
||||
|
||||
|
||||
def _ctx(**overrides):
|
||||
|
|
@ -131,3 +137,110 @@ def test_only_the_unfinished_no_tool_branch_arms_the_budget_tail():
|
|||
and any(arming in ast.walk(node) for arming in armings)
|
||||
]
|
||||
assert guards, "the arming must be guarded by an unfinished (final_result is None) round"
|
||||
|
||||
|
||||
@pytest.mark.serial
|
||||
def test_preparation_failure_pauses_and_resumes_the_no_tool_tail(full_loop, monkeypatch):
|
||||
"""The merged acceptance preparation owner survives a real loop/queue pause.
|
||||
|
||||
Resume neither buys a final nor mistakes the saved no-tool answer for a
|
||||
tool batch: its preparation incident and opaque identities survive intact.
|
||||
"""
|
||||
from ouroboros import budget_pause, loop, loop_budget, loop_acceptance_review, owner_wait, usage_accounting
|
||||
from ouroboros.artifacts import read_actor_source_bytes
|
||||
from supervisor.events import _handle_budget_pause
|
||||
from tests.test_budget_pause_exact import _install_queue, _supervisor_ctx
|
||||
|
||||
f = full_loop
|
||||
root, task_id = f.ctx.drive_root, f.ctx.task_id
|
||||
q, state, workers = _install_queue(root, monkeypatch)
|
||||
monkeypatch.setattr(state, "budget_remaining", lambda *_a, **_k: 100.0)
|
||||
f.ctx.task_started_at = time.time() - 5
|
||||
f.ctx.context_fit_plan = None
|
||||
# The shared scripted-model fixture has no immutable ContextFit core;
|
||||
# route rebuilding is exercised by the owner-wait/ContextFit suites.
|
||||
monkeypatch.setattr(owner_wait, "rebind_restored_route", lambda *_a, **_k: (None, "max"))
|
||||
f.ctx._cost_ceiling = _ceiling(50.0)
|
||||
tree = {"accounted_usd": 49.0, "root_limit_usd": 50.0, "age_sec": 0.0}
|
||||
monkeypatch.setattr(loop, "_loop_tree_accounting", lambda **_k: dict(tree))
|
||||
monkeypatch.setattr(loop_budget, "_loop_tree_accounting", lambda **_k: dict(tree))
|
||||
monkeypatch.setattr(loop_budget, "_wrapup_global_remaining", lambda: 100.0)
|
||||
monkeypatch.setattr(usage_accounting, "refresh_root_accounting", lambda *_a, **_k: dict(tree))
|
||||
monkeypatch.setattr(loop, "_forced_final_answer", lambda *_a, **_k: pytest.fail("paid budget final"))
|
||||
monkeypatch.setattr(loop, "_prepare_post_tool_budget_context",
|
||||
lambda *_a, **_k: pytest.fail("no-tool continuation armed tool controls"))
|
||||
builders = []
|
||||
|
||||
def broken(*_a, **_k):
|
||||
builders.append(1)
|
||||
raise RuntimeError("local preparation unavailable")
|
||||
|
||||
monkeypatch.setattr(loop_acceptance_review, "_build_host_acceptance_evidence", broken)
|
||||
saved = {}
|
||||
|
||||
def main(*_a, **kw):
|
||||
f.model_step += 1
|
||||
usage = f.ctx._accumulated_usage
|
||||
if f.model_step == 2:
|
||||
for key in ("incident_id", "source_identity", "attempts", "failure_kind"):
|
||||
assert f.ctx._execution_trace["acceptance_preparation"][key] == saved["trace"]["acceptance_preparation"][key]
|
||||
assert usage["cost"] == 1.25 and usage["rounds"] == 1
|
||||
assert f.ctx._delivery_evidence_fingerprint == saved["delivery"]["_delivery_evidence_fingerprint"]
|
||||
assert f.model_step <= 2, "Resume reset the preparation incident or round state"
|
||||
usage.update(cost=1.25 * f.model_step, rounds=f.model_step)
|
||||
return {"content": "The complete report includes the requested budget."}, 1.25
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
try:
|
||||
with pytest.raises(budget_pause.BudgetPauseRequested) as raised:
|
||||
f.run()
|
||||
pause = raised.value.pause
|
||||
saved.update(json.loads(read_actor_source_bytes(root, task_id, pause["source_ref"])))
|
||||
assert pause["rail"] == budget_pause.RAIL_GRACEFUL_CEILING
|
||||
assert saved["resume_point"]["budget_tail"] == "no_tool" and saved["round_idx"] == 2
|
||||
assert saved["trace"]["acceptance_preparation"]["attempts"] == 1
|
||||
assert f.model_step == 1 and builders == [1] and not f.review_sends
|
||||
task = {"id": task_id, "type": "task", "root_task_id": task_id, "chat_id": 1, "_attempt": 1}
|
||||
workers.RUNNING[task_id] = {"task": task, "worker_id": 0, "attempt": 1}
|
||||
workers.WORKERS[0] = SimpleNamespace(busy_task_id=task_id)
|
||||
_handle_budget_pause({**budget_pause.pause_event(task, pause), "worker_id": 0},
|
||||
_supervisor_ctx(root, workers, q, [], []))
|
||||
tree["root_limit_usd"] = 100.0 # owner-authorized headroom, never reset spend
|
||||
assert q.resume_budget_paused_task(task_id)["ok"] is True
|
||||
f.ctx.budget_pause_resume = copy.deepcopy(workers.PENDING[0]["_budget_pause_resume"])
|
||||
f.run_args["messages"] = [{"role": "user", "content": "fresh shell must restore saved cognition"}]
|
||||
result, usage, trace = f.run()
|
||||
assert result == "The complete report includes the requested budget."
|
||||
assert f.model_step == 2 and builders == [1] and not f.review_sends
|
||||
assert usage["rounds"] == 2 and usage["cost"] == 2.5
|
||||
assert trace["acceptance_decision"]["status"] == "finalized_unaccepted"
|
||||
assert trace["acceptance_decision"]["reason"] == "acceptance_preparation_failed"
|
||||
assert budget_pause.budget_pause_row(root, task_id)["state"] == budget_pause.STATE_RESUMED
|
||||
finally:
|
||||
budget_pause.end_dispatch_fence(task_id)
|
||||
|
||||
|
||||
@pytest.mark.serial
|
||||
@pytest.mark.parametrize("rail,completed_rounds", [("global", 1), ("soft_land", 1), ("soft_land", 0)])
|
||||
def test_eligible_monetary_stops_preserve_exact_continuation(tmp_path, monkeypatch, rail, completed_rounds):
|
||||
"""A first batch is work; a soft threshold also pauses before any work."""
|
||||
from ouroboros import budget_pause, loop
|
||||
from tests.test_budget_pause_exact import _loop_ctx, _running_row
|
||||
|
||||
_running_row(tmp_path, "first-work")
|
||||
ctx, limit = _loop_ctx(tmp_path, "first-work")
|
||||
limit.round_idx = 1
|
||||
limit.accumulated_usage["rounds"] = completed_rounds
|
||||
limit.llm_trace = {}
|
||||
monkeypatch.setattr(loop, "_forced_final_answer", lambda *_a, **_k: pytest.fail("paid budget final"))
|
||||
try:
|
||||
with pytest.raises(budget_pause.BudgetPauseRequested):
|
||||
if rail == "global":
|
||||
loop._check_budget_limits(limit, 0.0)
|
||||
else:
|
||||
loop._soft_land_exhausted_ceiling(limit, _ceiling(0.01))
|
||||
saved = budget_pause.budget_pause_row(tmp_path, ctx.task_id)
|
||||
assert saved["exact_continuation"] and saved["resume_point"]["round_idx"] == 1
|
||||
assert limit.accumulated_usage["execution_status"] == "paused"
|
||||
finally:
|
||||
budget_pause.end_dispatch_fence(ctx.task_id)
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ from dataclasses import replace
|
|||
|
||||
import pytest
|
||||
|
||||
from ouroboros import loop, owner_wait, pricing, task_pacing, usage_accounting as accounting
|
||||
from ouroboros import budget_pause, loop, owner_wait, pricing, task_pacing, usage_accounting as accounting
|
||||
from ouroboros.contracts.task_contract import normalize_budget_profile
|
||||
from ouroboros.owner_mailbox import write_owner_message
|
||||
from tests.test_owner_wait_cold_loop import cold_registry
|
||||
|
|
@ -13,7 +13,8 @@ from tests.test_loop_transport_wait import _loop_kwargs
|
|||
|
||||
|
||||
@pytest.mark.parametrize("queued_override", [False, True])
|
||||
def test_cold_grant_checks_saved_budget_before_ordinary_dispatch(tmp_path, monkeypatch, queued_override):
|
||||
@pytest.mark.parametrize("cold_restart", [False, True])
|
||||
def test_cold_grant_checks_saved_budget_before_ordinary_dispatch(tmp_path, monkeypatch, queued_override, cold_restart):
|
||||
(tmp_path / "logs").mkdir()
|
||||
(tmp_path / "fixture.txt").write_text("A real extra tool read")
|
||||
scope = accounting.UsageScope(drive_root=tmp_path, task_id="t-wait", root_task_id="t-wait",
|
||||
|
|
@ -74,27 +75,40 @@ def test_cold_grant_checks_saved_budget_before_ordinary_dispatch(tmp_path, monke
|
|||
def park(ctx, checkpoint):
|
||||
checkpoints.append(checkpoint)
|
||||
owner_wait.set_owner_wait(tmp_path, "t-wait", {**checkpoint, "state": "waiting"})
|
||||
if cold_restart:
|
||||
# End the old actor at its owner wait, before either monetary
|
||||
# tail runs. The cold actor consumes that saved tail once.
|
||||
raise InterruptedError("planned restart")
|
||||
write_owner_message(tmp_path, owner_answer, task_id="t-wait", msg_id="warm-answer")
|
||||
|
||||
warm = registry()
|
||||
warm._ctx.owner_wait_resume, warm._ctx.owner_wait_callback = None, park
|
||||
warm_result, _, warm_trace = loop.run_llm_loop(**{
|
||||
**_loop_kwargs(tmp_path, warm, []), "budget_remaining_usd": 200, "drive_logs": tmp_path / "logs"})
|
||||
assert len(checkpoints) == 1
|
||||
|
||||
phase = "cold"
|
||||
cold = registry()
|
||||
owner_wait.set_owner_wait(tmp_path, "t-wait", {**checkpoints[0], "state": "waiting"})
|
||||
cold._ctx.owner_wait_resume = {**checkpoints[0], "restart_transaction_id": "observed-restart"}
|
||||
cold._ctx.owner_wait_callback = lambda *_: write_owner_message(
|
||||
tmp_path, owner_answer, task_id="t-wait", msg_id="cold-answer")
|
||||
cold_result, _, cold_trace = loop.run_llm_loop(**{
|
||||
**_loop_kwargs(tmp_path, cold, []), "budget_remaining_usd": 152.9, "drive_logs": tmp_path / "logs"})
|
||||
selected, remaining = warm, 200
|
||||
if cold_restart:
|
||||
with pytest.raises(InterruptedError, match="planned restart"):
|
||||
loop.run_llm_loop(**{**_loop_kwargs(tmp_path, warm, []),
|
||||
"budget_remaining_usd": remaining, "drive_logs": tmp_path / "logs"})
|
||||
phase = "cold"
|
||||
selected, remaining = registry(), 152.9
|
||||
owner_wait.set_owner_wait(tmp_path, "t-wait", {**checkpoints[0], "state": "waiting"})
|
||||
selected._ctx.owner_wait_resume = {**checkpoints[0], "restart_transaction_id": "observed-restart"}
|
||||
selected._ctx.owner_wait_callback = lambda *_: write_owner_message(
|
||||
tmp_path, owner_answer, task_id="t-wait", msg_id="cold-answer")
|
||||
try:
|
||||
with pytest.raises(budget_pause.BudgetPauseRequested) as paused:
|
||||
loop.run_llm_loop(**{**_loop_kwargs(tmp_path, selected, []),
|
||||
"budget_remaining_usd": remaining, "drive_logs": tmp_path / "logs"})
|
||||
state = json.loads(owner_wait.read_actor_source_bytes(tmp_path, "t-wait", paused.value.pause["source_ref"]))
|
||||
assert owner_answer in json.dumps(state["messages"])
|
||||
assert state["round_idx"] == 1 and state["usage"]["cost"] == pytest.approx(.2)
|
||||
assert state["route"]["active_model"] == "same-model"
|
||||
assert state["route"]["active_model_override"] == ("after-budget" if queued_override else None)
|
||||
assert [row["tool"] for row in state["trace"]["tool_calls"]] == ["escalate"]
|
||||
finally:
|
||||
budget_pause.end_dispatch_fence("t-wait")
|
||||
held = accounting.reserve_attempt(accounting.AttemptRequest(
|
||||
model="fixture", provider="fixture", reservation_usd=.2))
|
||||
accounting.release_attempt(held, "test did not send") # hard50 still has room
|
||||
assert network == []
|
||||
assert "Current owner choice considered." in warm_result and "Current owner choice considered." in cold_result
|
||||
assert not [row for row in calls if row[0:2] == ("cold", "ordinary")]
|
||||
assert all(row[2:] == (1, "same-model") for row in calls)
|
||||
assert [row["tool"] for row in cold_trace["tool_calls"]] == [row["tool"] for row in warm_trace["tool_calls"]] == ["escalate"]
|
||||
assert len(checkpoints) == 1
|
||||
assert calls == [("warm", "ordinary", 1, "same-model")]
|
||||
|
|
|
|||
|
|
@ -62,9 +62,7 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = {
|
|||
# 106400 -> 106600 (#1195 merge of 32d8dfc6): the base's settings_catalog.js
|
||||
# paragraph (#1214, +319 bytes) landed in the same window; both additions stand,
|
||||
# neither displaces the other's text.
|
||||
# 106600 -> 106800 (#1196): a paused direct turn reports the managed census
|
||||
# phases; the phase sentence is extended, nothing older describes it.
|
||||
"docs/architecture/03-web-ui-pages-and-buttons.md": 106800,
|
||||
"docs/architecture/03-web-ui-pages-and-buttons.md": 106600,
|
||||
"docs/architecture/04-server-api-endpoints.md": 26833,
|
||||
# 27137 -> 30400: the schedule table gains a documented write contract the
|
||||
# chapter had no text for — one transaction owning the lock ORDER, the strict
|
||||
|
|
@ -77,13 +75,7 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = {
|
|||
# displace; two neighbouring sentences were compressed by 168 bytes first.
|
||||
# +300: the queue snapshot and the supervisor focus event carry the root's
|
||||
# bounded authored focus (cross-focus awareness).
|
||||
# 30900 -> 31700 (#1196): the exact-continuation `_budget_pause` marker, its grant
|
||||
# carrier and the separate paused-interval carrier are new snapshot/assignment
|
||||
# facts the chapter lacked; nothing older describes them.
|
||||
# 31700 -> 32200 (#1196): restart parking of a completed pause, the typed
|
||||
# restore/acceptance holds and the parked direct turn replace the restore
|
||||
# sentence they grew from.
|
||||
"docs/architecture/05-supervisor-loop.md": 32200,
|
||||
"docs/architecture/05-supervisor-loop.md": 30900,
|
||||
# 286850 -> 287600: "an answer that has not arrived is a gap" is a new invariant of
|
||||
# plan review and task acceptance (the slot census vocabulary, the `awaiting`
|
||||
# projection, the only-awaited task outcome); the in-flight sentence it grew from is
|
||||
|
|
|
|||
|
|
@ -83,7 +83,7 @@
|
|||
* @property {string} project_id
|
||||
* @property {string} client_message_id // empty for managed queue rows
|
||||
* @property {string} kind // direct_chat | managed_task — presentational label; membership in this census, not kind, decides liveness
|
||||
* @property {string} phase // managed rows: queued | budget_paused | budget_pausing (RUNNING, writing its exact pause record; additive, #1196) | working | finalizing; direct rows: thinking, or unknown when the live wait owner could not be read; a direct turn paused on its budget rail is parked in the queue under the same id and reports the managed phases (budget_paused, then working/finalizing after an explicit Resume)
|
||||
* @property {string} phase // managed rows: queued | budget_pausing | budget_paused | working | finalizing; direct rows: thinking or unknown; budget-paused direct turns retain their ID/kind and use the managed phases after parking
|
||||
* @property {number} started_at
|
||||
*/
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue