mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 12:18:39 +00:00
fix: close nested post-work interruption and stop paths
This commit is contained in:
parent
d50b61313c
commit
2f02539fc0
16 changed files with 348 additions and 47 deletions
|
|
@ -632,7 +632,7 @@ and do not return `PASS` for an item that also has a `FAIL` — the concrete
|
|||
| 7 | extension_namespace_discipline | `type: extension` only: does the extension register its tool/route/ws-handler/ui-tab under the namespace derived from its `name` (e.g. provider-safe tool/ws names like `ext_<len>_<token>_<surface>`, route `/api/extensions/<name>/…`)? Tool and WS short names must be alphanumeric/underscore and at most 24 characters. Namespace collisions with built-in surfaces are a concrete FAIL. If the extension uses `api.send_ws_message`, are emitted event names short/provider-safe and paired with reviewed host-owned widget `subscription` components rather than arbitrary same-origin JavaScript? If the extension declares streaming UI, is it a reviewed extension route consumed by a host-owned `stream` component? A reviewed `module` widget may also consume the skill's own routes (including streaming responses) and the skill's namespaced WebSocket events through the host-mediated bridge (`OuroborosWidget.fetch` / `OuroborosWidget.onEvent`), which is not arbitrary same-origin JavaScript. If the extension owns background resources (threads, sockets, EventSource clients, subprocesses), does it register cleanup with `api.on_unload(callback)`? If the extension declares a widget render block, is it one of the host-owned schemas (`iframe`, `module`, or declarative v1: forms/actions, markdown/code, JSON/kv/table, tabs/chart, stream/subscription, progress/poll, file/gallery/media, map/calendar/kanban, group/metric/callout), with media sourced from extension routes or safe data URLs and no arbitrary same-origin JavaScript? Nested interactive group/tab children must use stable identity and one host-owned lifecycle, while `subscription.render` stays transitively passive. For non-extension skills, verdict PASS with reason "Not applicable — type != extension." | severity-driven for applicable extensions |
|
||||
| 8 | widget_module_safety | **v5.7.0+. ``kind: "module"`` widgets only.** The host fetches reviewed ``widget.js`` through ``GET /api/extensions/<skill>/module/<entry>``, embeds the source into a sandboxed opaque-origin ``<iframe srcdoc sandbox="allow-scripts allow-pointer-lock allow-downloads" allow="autoplay; fullscreen; clipboard-write">`` with no ``allow-same-origin`` — ``document.cookie``, ``localStorage``, and ``sessionStorage`` throw ``SecurityError`` there by construction and need no source review — and injects a parent-mediated ``fetch`` bridge that rejects paths outside the owning skill route prefix. Reviewers confirm at the source level what the sandbox cannot: (a) no ``fetch``/``XMLHttpRequest`` URL outside ``/api/extensions/<skill>/`` and no bespoke ``postMessage`` protocol to ``window.parent`` beyond the host bridge; (b) the declared launch policy ``render.start`` (SSOT ``ouroboros/extension_ui_validation.py::WIDGET_START_MODES``; see CREATING_SKILLS "Launch policy") fits the widget's weight — ``auto`` only for a cheap instrument, ``manual`` for a program that should not run all the time, ``retain`` only for a program that genuinely must keep running while the owner is elsewhere and stays cheap while hidden; (c) a widget with state worth keeping registers ``window.__ouroWidgetOnDispose(fn)`` (never assigns over it) and saves that state through the skill's own routes, because the frame is disposable; (d) a module that declares ``render.appearance: host`` uses the optional ``OuroborosWidget.onTheme(callback)`` bridge; a real consumer should prove both resolved palettes, while ``independent``/``fixed`` modules remain author-owned. The declaration is author/reviewer intent, not a source-level proof or runtime gate for legacy payloads; a source mismatch or missing browser evidence is advisory unless it exposes a concrete runtime failure. Acceptable interactions: ``fetch('/api/extensions/<skill>/...')`` (through the host bridge), ``window.OuroborosWidget.fetch('/api/extensions/<skill>/...')``, ``window.OuroborosWidget.onTheme(callback)``, and host-supplied data attributes. Mark non-module widgets and non-extension skills PASS with reason "Not applicable". | severity-driven when kind=module |
|
||||
| 9 | inject_chat_minimization | Does any use of the `inject_chat` permission have a narrow, user-facing transport purpose? The Host Service enforces token auth, skill-source attribution, rate limits, in-flight limits, fresh executable review, enablement, and explicit content-hash-bound grants. Reviewed chat transports may carry the same raw owner text as direct chat, including slash commands such as `/panic`, `/restart`, `/review`, `/evolve`, `/bg`, and `/status`; reviewers must evaluate whether the transport itself is authorized, attributable, bounded, and user-facing rather than treating slash-shaped text as automatically forbidden. A skill that accepts external inbound traffic must still show local defense-in-depth appropriate to its transport: owner/chat binding or an equivalent access rule, bounded polling/backpressure, and no unaudited broadcast to unrelated parties. Missing local defense-in-depth is a concrete FAIL for network transports. Mark PASS with reason "Not applicable" when `inject_chat` is not declared. | critical |
|
||||
| 10 | event_subscription_minimization | Are `subscribe_event` and `subscribe_events` limited to the minimum host event topics required by the skill? `chat.outbound`, `chat.typing`, `chat.photo`, `chat.video`, `chat.document`, and `chat.links` expose owner/agent conversation data (including delivered file bytes and outbound link actions) and require explicit justification. Wildcards, undeclared topics, or forwarding subscribed chat content to unrelated external services are concrete FAILs. Mark PASS with reason "Not applicable" when `subscribe_event` is not declared. | critical |
|
||||
| 10 | event_subscription_minimization | Are `subscribe_event` and `subscribe_events` limited to the minimum host event topics required by the skill? `chat.outbound`, `chat.typing`, `chat.photo`, `chat.video`, `chat.document`, `chat.links`, `chat.quiz` and `chat.quiz_state` expose owner/agent conversation data (including delivered file bytes, outbound link actions and the owner's verbatim quiz answers) and require explicit justification. Wildcards, undeclared topics, or forwarding subscribed chat content to unrelated external services are concrete FAILs. Mark PASS with reason "Not applicable" when `subscribe_event` is not declared. | critical |
|
||||
| 11 | companion_process_safety | For `companion_process` / `supervised_task` skills: is every command declared as an argument list (not shell string), using an allowlisted runtime, with no writes outside `skill_dir` / `state_dir`, no unbounded restart loop, and cleanup on unload/panic? Does the process avoid inheriting secrets except through reviewed `env_from_settings` grants? Mark PASS with reason "Not applicable" when no long-lived process/task is declared — a transient `subprocess.run`/`subprocess.Popen` invocation of a build tool like `ffmpeg`, `ImageMagick`, or `git` inside a normal request handler is NOT a long-lived companion process and does not trigger this item (its safety belongs under items 4 / 6 / 13). | severity-driven when applicable |
|
||||
| 12 | host_token_handling | If the skill calls the Host Service API, does it use the provided `SkillToken.use_in_request()` only at request construction sites, avoid logging/serializing tokens, and keep all host-service calls on the loopback endpoint? Printing, persisting, exfiltrating, or embedding the token into user-visible output is a concrete FAIL. Mark PASS with reason "Not applicable" when the skill does not access the Host Service API. | critical |
|
||||
| 13 | error_handling | Does the skill surface actionable errors instead of swallowing exceptions, returning success on partial failure, or leaving users to inspect raw logs manually? Are retry/backoff paths bounded and purpose-specific? | advisory |
|
||||
|
|
|
|||
|
|
@ -525,19 +525,19 @@ Availability follows `deep_review_route`, never a window floor. An API row needs
|
|||
|
||||
#### Post-task reflection
|
||||
|
||||
Typed root post-task triggers decide whether a run warrants Experience Review. `reflection.generate_reflection` sends the Light route one open prompt with the EXACT initial text (never a prefix) and its host-recorded `task_inputs.run_origin` beside it (provenance, never by itself the accepted requirement) plus tool-use, error, review and child projections and the same frozen non-final cost snapshot the facts row carries; it runs outside the tool loop, records its own usage, and its failure never erases the delivered result or changes a review verdict. Its execution trace is the ALL-CALLS listing (`build_trace_summary(all_calls=True)`): every call in order with every argument, identical consecutive calls folded into one `×N` row with their rounds, the first line of a failed or repeated call's result, and one header count of rounds whose every call was non-ok — no positional window and no literal cut, because the consolidation seam fits the call to the Light route whenever that route's window is known (an unknown window sends the prompt unchecked — the accepted residual of enlarging it); the STORED `trace_summary` (task card, parents, children) stays the bounded two-argument preview. The trace row carries the `round_id` of the model round that issued the call (absent when unknown), when the listing really cut an argument value or a failed/repeated call's answer, the redacted per-call record is retained through `retain_memory_source` and named in the prompt as OPTIONAL reading (never a required source), and its claim states the STORED bounds on both axes, not completeness: an argument already passed `sanitize_tool_args_for_log` (an oversized value carries a marker with its length and sha) and a result is the stored actor-visible cap — more than the listing, which shows only the first line of a failed or repeated answer — a partial one naming its own `FULL_RESULT_SOURCE_JSON` or `FULL_RESULT_SOURCE_UNAVAILABLE`, with a call's recorded manifest named only when it has one — claiming results "in full" or an unconditional manifest overstated a cognitive artifact. Cut detection reads the ONE shared marker list (`artifacts.SANITIZER_OMISSION_MARKERS`), because a width test over already-sanitized args measured the widest argument in the task as a small one and retained nothing at all, while a hand-rolled subset missed the `_repr` and `_error` shapes whose arguments survive only in the call blob. Unavailable source retention is disclosed, error details group by full redacted content before display clipping, and post-task synthesis — the reflection and its Pattern Register update — thinks at the owner's Task / Chat effort (`settings_scales.resolve_effort("task")`), never a literal. Admission to the Pattern Register is typed, not a word scan: it opens on a call the loop recorded as errored (its stamped `tool_result_code`, or the recorded status for a legacy row), on a producer fact naming a failure the ok status cannot carry (a preserved commit whose post-commit tests failed publishes `post_commit_tests`), on typed codes already stored with an entry, or on a genuinely FAILED child — a cancelled, soft-landed best-effort or degraded child is not a failure. Those failed-child classes reach the root through the child evidence the synthesis walk already collects and make the run error-bearing for both the trigger and the prompt's error details: children do not reflect, so a short clean root that delegated the work is the only place its child's failure can be learned from at all. Deliberately not "a reason code exists", which would open the register on every terminal.
|
||||
Typed root post-task triggers decide whether a run warrants Experience Review. `reflection.generate_reflection` sends the Light route one open prompt with the EXACT initial text (never a prefix) and its host-recorded `task_inputs.run_origin` beside it (provenance, never by itself the accepted requirement) plus tool-use, error, review and child projections and the same frozen non-final cost snapshot the facts row carries; it runs outside the tool loop, records its own usage, and its failure never erases the delivered result or changes a review verdict. Its execution trace is the ALL-CALLS listing (`build_trace_summary(all_calls=True)`): every call in order with every argument, identical consecutive calls folded into one `×N` row with their rounds, the first line of a failed or repeated call's result, and one header count of rounds whose every call was non-ok — no positional window and no literal cut, because the consolidation seam fits the call to the Light route whenever that route's window is known (an unknown window sends the prompt unchecked — the accepted residual of enlarging it); the STORED `trace_summary` (task card, parents, children) stays the bounded two-argument preview. The trace row carries the `round_id` of the model round that issued the call (absent when unknown), when the listing really cut an argument value or a failed/repeated call's answer, the redacted per-call record is retained through `retain_memory_source` and named in the prompt as OPTIONAL reading (never a required source), and its claim states the STORED bounds on both axes, not completeness: an argument already passed `sanitize_tool_args_for_log` (an oversized value carries a marker with its length and sha) and a result is the stored actor-visible cap — more than the listing, which shows only the first line of a failed or repeated answer — a partial one naming its own `FULL_RESULT_SOURCE_JSON` or `FULL_RESULT_SOURCE_UNAVAILABLE`, with a call's recorded manifest named only when it has one. Cut detection reads the ONE shared marker list (`artifacts.SANITIZER_OMISSION_MARKERS`), because a width test over already-sanitized args measured the widest argument in the task as a small one and retained nothing at all, while a hand-rolled subset missed the `_repr` and `_error` shapes whose arguments survive only in the call blob. Unavailable source retention is disclosed, error details group by full redacted content before display clipping, and post-task synthesis — the reflection and its Pattern Register update — thinks at the owner's Task / Chat effort (`settings_scales.resolve_effort("task")`), never a literal. Admission to the Pattern Register is typed, not a word scan: it opens on a call the loop recorded as errored (its stamped `tool_result_code`, or the recorded status for a legacy row), on a producer fact naming a failure the ok status cannot carry (a preserved commit whose post-commit tests failed publishes `post_commit_tests`), on typed codes already stored with an entry, or on a genuinely FAILED child — a cancelled, soft-landed best-effort or degraded child is not a failure. Those failed-child classes reach the root through the child evidence the synthesis walk already collects and make the run error-bearing for both the trigger and the prompt's error details: children do not reflect, so a short clean root that delegated the work is the only place its child's failure can be learned from at all. Deliberately not "a reason code exists", which would open the register on every terminal.
|
||||
|
||||
A reflection lands where it durably belongs: a non-project root appends the full entry to the canonical `logs/task_reflections.jsonl`, a project-scoped root to its project drive with only a bounded pointer row in the canonical log. A project-bound task's context holds a bounded labeled tail of its project's reflections; a split root's headless mirror drive is never the reflection home; the Pattern Register update stays canonical. Entries carry task identity, evidence, lessons, backlog candidates and validated memory actions. `MEMORY_ACTIONS_JSON` permits only `scratchpad_append`, `knowledge_write` and `identity_update_candidate`, bounded in count and size; `apply_memory_actions` applies them via provenance-preserving memory and knowledge APIs; an identity candidate lands in the scratchpad for review, never auto-written to `identity.md`. A project-scoped task applies knowledge actions only (project store by default, explicit global allowed), skipping scratchpad and identity candidates, which the conversation writes from any room with the full view this Light pass lacks; each skipped, empty or topic-less action is a `reflection_memory_action_skipped` event with its reason and `input_ref`: the retained exact task-input prompt, or an action's reflection-entry copy, else the canonical log pointer or `source_unavailable`. Reflection may propose a campaign or backlog item, never enqueue, review, commit or enable one.
|
||||
|
||||
The Pattern Register writer (`reflection._update_patterns`) REPLACES the whole document, so every decision input it reads is complete — the full current register, the exact initial text (`goal_exact`, beside the bounded `goal` display) and its `run_origin` and the whole reflection text; a prefix of a decision input can never authorize the rewrite, because a clipped clause can record the inverse of what the reflection concluded. `patterns` is a reserved global-only topic alongside `improvement-backlog` and `overview`: whichever room writes it, the register has ONE home on the canonical drive, the only one its writer and its readers (context assembly, deep self-review, the headless copy) address. Re-read, exact compare, history append and atomic replacement are one critical section outside the Light call; a register that moved under a losing writer is preserved and that task's learning is DROPPED with a warning naming the task, never retried against a source it did not decide from. A row's count is bumped once per observed episode, so two roots of one owner request bump it twice — the number counts episodes, not distinct requests.
|
||||
|
||||
The advisory rows a reflection or summary reads are ATTRIBUTED. Advisory runs are scoped by repository and several tasks legitimately review one checkout, so `collect_review_evidence` renders as `recent_advisory_runs` only the rows this task owns plus legacy rows carrying no owner (which stay unknown, never re-attributed); another task's rows reach the prompt under their own heading naming the owning task ids, and every row carries its owning task, attempt, phase and complete snapshot hash. Repository readiness (`current_repo`, open obligations, commit-readiness debt, the exact-snapshot match) stays repository-scoped, because that is what "can this checkout be committed" means. The split is keyed on ROW IDENTITY, never on a scope key: an empty repository key widens the candidate list to every advisory run on the drive, and the installation's history is not one task's record.
|
||||
The advisory rows a reflection reads are ATTRIBUTED. Advisory runs are scoped by repository and several tasks legitimately review one checkout, so `collect_review_evidence` renders as `recent_advisory_runs` only the rows this task owns plus legacy rows carrying no owner (which stay unknown, never re-attributed); another task's rows reach the prompt under their own heading naming the owning task ids, and every row carries its owning task, attempt, phase and complete snapshot hash. Repository readiness (`current_repo`, open obligations, commit-readiness debt, the exact-snapshot match) stays repository-scoped, because that is what "can this checkout be committed" means. The split is keyed on ROW IDENTITY, never on a scope key: an empty repository key widens the candidate list to every advisory run on the drive, and the installation's history is not one task's record.
|
||||
|
||||
Only roots synthesize; `root_phase_checkpoint` makes paid synthesis at-most-once across restart, while children contribute evidence. Durable-result persistence owes the answer as `final:<tid>:<digest>` in `supervisor/terminal_delivery.py`'s bounded outbox (§5; normal/cancel/reap). `send_message` delivers immediately; the retained buffered copy shares its ID for durable dedupe. Replays use bounded backoff. Exhaustion/eviction preserves full text on disk, emits `terminal_delivery_exhausted` and a chat notice; external delivery remains at-least-once. Buffered `task_done` stays last to retain the slot/child drive during synthesis; a hung-synthesis reap need not lose the delivered answer. Project roots keep early answers in Project. Their canonical row and deferred Main mirror use `terminal_projection.settle_terminal_projection` via task-done/checkpoint/startup/maintenance; §3 "Main rows and host-stamped card rows" owns readiness, retirement and limits.
|
||||
|
||||
Synthesis receives a sealed final package from the durable result — the submitted final text, its artifact manifest and completion_observations. Full redacted action observations live in the canonical artifact store (`task.budget_drive_root or drive_root`), in the write-once `source_handles/context_checkpoints` store with verified `task_source` refs, before compact publication and outside deliverables and inferred readiness; their native reader `get_task_result(include_completion_source=true)` returns complete length/hash first, then explicit `source_start_char`/`source_end_char` ranges (`artifacts.text_source_range_projection`, the shared work-order range contract), with bytes, kind, path containment and SHA checked before any excerpt. Packet-only reflection receives per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Before context cleanup, `agent_task_pipeline.emit_task_results` also freezes `review_evidence.task_inputs` through `post_task_synthesis.capture_task_inputs`: `run_origin`, the existing task-local owner corpus, intact question/answer provenance and the canonical split-root verification-receipt union. Reflection receives the same complete redacted content through `reflection.task_inputs_prompt_section`, separate from bounded trace/review excerpts. A zero return code is positive evidence; an unrelated later pass cannot resolve another check's failure. Peer proposals stay attributed, and unavailable input is not evidence that approval or verification never existed. Recovery uses these stored observations and inputs, not a later conversation. A free `host_task_facts` row precedes paid stages (or follows result persistence when Stop skips them): no model call or narrative; its metrics, routing and cost serve history. `files_rescued` (TZ-2 C2) is a stat-only file count of the artifact stores: positive, zero or unknown if unreadable, `hash_computed: false`; the stop receipt repeats it so no salvageable text never means no files.
|
||||
|
||||
Pooled workers retain their slot until post-task work settles; early final-answer delivery is independent of that timing. Native work stays on its registered actor; `TaskModelWait` remains reachable through `POST_TASK_SYNTHESIS_INFLIGHT`, and detached work owns a separate live wait. Mailbox cleanup waits for the checkpoint. The solve-phase absolute ceiling does not cut a settled root's running post-work, but Stop, calendar deadline, monetary admission, per-call and idle rails still bind. The stage owner reads returned and raised failures alike: budget, a control or an unresolved attempt on any provider's chain (`transport_custody.outcome_unknown_on_chain`) skips later paid stages, degrading the checkpoint; an ordinary failure stays local (later stages run, completed reflection actions apply), but its stage lost work: `degraded` with no skipped list, never `completed` (TZ-2 C3). A running checkpoint after restart degrades without replaying a paid request.
|
||||
Pooled workers retain their slot until post-task work settles; early final-answer delivery is independent of that timing. Native work stays on its registered actor; `TaskModelWait` remains reachable through `POST_TASK_SYNTHESIS_INFLIGHT`, and detached work owns a separate live wait. Mailbox cleanup waits for the checkpoint. The solve-phase absolute ceiling does not cut a settled root's running post-work, but Stop, calendar deadline, monetary admission, per-call and idle rails still bind. The stage owner reads returned and raised failures alike, the reflection's nested Pattern Register write too: budget, a control or an unresolved attempt on any provider's chain (`transport_custody.outcome_unknown_on_chain`) skips later paid stages, degrading the checkpoint; an ordinary failure stays local (later stages run, completed reflection actions apply) yet reads `degraded` with no skipped list, never `completed`; a split-answered refusal (`resolution`) is history, not failure (TZ-2 C3). A running checkpoint after restart degrades without replaying a paid request.
|
||||
|
||||
#### Project registry and lease
|
||||
|
||||
|
|
|
|||
|
|
@ -28,7 +28,7 @@ This chapter owns the ABI promise: which typed shapes and their parsing, normali
|
|||
| `ChatOutbound.initiator` — additive origin label of a self-initiated turn (`"consciousness"` on every frame and chat/progress/summary row of a consciousness wake-up and of the roots it starts; absent on an owner's turn); stamped by the turn's own event queue and the agent's frame meta, persisted by `log_chat`/the authored summary row/the task result, replayed by history on each row | `ouroboros/gateway/contracts.py`, `supervisor/log_addressing.py`, `ouroboros/subagent_messages.py`, `supervisor/message_bus.py`, `ouroboros/gateway/history.py`, `web/modules/api_types.js` | `tests/test_consciousness_initiator_label.py`, `tests/test_consciousness_wake_lane.py`, `web/tests/consciousness_label.test.js` |
|
||||
| `project_thread` stamp on all seven outbound frame types, stamped at the message-bus broadcast choke; a stamped frame is never adopted by Main (`chat_activity.mainThreadAccepts`) | `supervisor/message_bus.py`, `ouroboros/projects_registry.py` | `tests/test_message_bus.py`, `web/tests/chat_thread_routing.test.js` |
|
||||
| Media/link envelopes — media `task_id`/`size_bytes`/`download_url`; `LinkAction {label,url}` with at most twelve absolute HTTP(S) actions; `links` in `WS_MESSAGE_TYPES`; `chat.links` host topic | `ouroboros/gateway/contracts.py`, `ouroboros/tools/core.py`, `ouroboros/event_bus.py` | `tests/test_contracts.py` |
|
||||
| Owner quiz ABI — `QuizOption {label, detail?, recommended?}` (the asker marks its recommendation on that option; the web card badges it, Telegram stars its button, the durable block keeps `recommended_index`), `QuizOutbound` (quiz_id, question, options (0–6; empty is an open question answered in the owner's words), stake, `assumption` (required for optional clarification), additive `wait_for_answer` for a live pooled or ordinary-conversation root that must wait, lifecycle state open/answered/expired_terminal/superseded, optional host-written `host_facts` sentence, also on history rows and Main pointers), separate `QuizStateOutbound` discriminator, `chat.quiz` host topic (+ optional event-only `project_name`); the producer is the one escalation verb `escalate(question, options, stake, assumption, wait_for_answer=False, max_wait_minutes=None)` (the bound applies to a required wait only, never past the task's own deadline: the wait resumes with a system notice and the card stays open; named on an optional question it takes the omitted path and the asker's receipt says so) — a ROOT asks the owner, a SUBAGENT delivers a typed frame to its nearest LIVE ancestor, which answers via `forward_to_worker` or escalates verbatim, so the owner sees only what no ancestor answered; answers arrive through the ONE ingress `POST /api/decisions` (family ids `quiz:{task_id}:{quiz_id}`, `routing:{client_message_id}:{routing_token}`; `interaction:` reserved), request-id idempotent, first answer wins, validated against the STORED options; `option_index` is optional for the quiz family alone — a comment-only answer writes NO `answered_index`, because a stored 0 would replay as "chose the first option"; injected as the typed `KIND_QUIZ_ANSWER` mailbox control and broadcast as `quiz_state` (carrying the recorded `comment` when the owner answered in their own words, so the live card shows `Owner's answer:` exactly as replay does); expiry is structural only (the task-done seam flips open quizzes to `expired_terminal`, and the SAME reconcile closes the paired `owner_wait` so a terminal task never projects `quiz=expired_terminal` beside `owner_wait=waiting`; `owner_wait.set_owner_wait`'s refusal to continue waiting on a terminal result is preserved, not caught); a LATE answer to such an expired card is ACCEPTED at the same ingress (the projection records `answered_after_terminal`) and, because no mailbox will ever be drained, is delivered as the owner's OWN message into the card's chat through the named ingress `supervisor.message_bus.accept_local_message`, idempotent on `client_message_id = quiz_late_answer:<task>:<quiz>`, its provenance in the message's own `late_answer` metadata rather than a substituted `client_surface`; the 2xx says `forwarded` so no surface claims a delivery that did not happen, and 409 is left for what is settled (an already answered card, a non-root addressee). History replay merges the projection state | `ouroboros/gateway/contracts.py`, `ouroboros/gateway/task_decision.py`, `ouroboros/owner_quiz.py`, `ouroboros/tools/core.py` | `tests/test_gateway_parity.py`, `tests/test_quiz_display.py`, `tests/test_quiz_answer.py`, `web/tests/chat_decision.test.js` |
|
||||
| Owner quiz ABI — `QuizOption {label, detail?, recommended?}` (the asker marks its recommendation on that option; the web card badges it, Telegram stars its button, the durable block keeps `recommended_index`), `QuizOutbound` (quiz_id, question, options (0–6; empty is an open question answered in the owner's words), stake, `assumption` (required for optional clarification), additive `wait_for_answer` for a live pooled or ordinary-conversation root that must wait, lifecycle state open/answered/expired_terminal/superseded, optional host-written `host_facts` sentence, also on history rows and Main pointers), separate `QuizStateOutbound` discriminator, `chat.quiz` host topic (+ optional event-only `project_name`); the producer is the one escalation verb `escalate(question, options, stake, assumption, wait_for_answer=False, max_wait_minutes=None)` (the bound applies to a required wait only, never past the task's own deadline: the wait resumes with a system notice and the card stays open; named on an optional question it takes the omitted path and the asker's receipt says so) — a ROOT asks the owner, a SUBAGENT delivers a typed frame to its nearest LIVE ancestor, which answers via `forward_to_worker` or escalates verbatim, so the owner sees only what no ancestor answered; answers arrive through the ONE ingress `POST /api/decisions` (family ids `quiz:{task_id}:{quiz_id}`, `routing:{client_message_id}:{routing_token}`; `interaction:` reserved), request-id idempotent, first answer wins, validated against the STORED options; `option_index` is optional for the quiz family alone — a comment-only answer writes NO `answered_index`, because a stored 0 would replay as "chose the first option"; injected as the typed `KIND_QUIZ_ANSWER` mailbox control and broadcast as `quiz_state` and on `chat.quiz_state` (carrying the owner's verbatim `comment`, so the live card shows `Owner's answer:` exactly as replay does); expiry is structural only (the task-done seam flips open quizzes to `expired_terminal`, and the SAME reconcile closes the paired `owner_wait` so a terminal task never projects `quiz=expired_terminal` beside `owner_wait=waiting`; `owner_wait.set_owner_wait`'s refusal to continue waiting on a terminal result is preserved, not caught); a LATE answer to such an expired card is ACCEPTED at the same ingress (the projection records `answered_after_terminal`) and, because no mailbox will ever be drained, is delivered as the owner's OWN message into the card's chat through the named ingress `supervisor.message_bus.accept_local_message`, idempotent on `client_message_id = quiz_late_answer:<task>:<quiz>`, its provenance in the message's own `late_answer` metadata rather than a substituted `client_surface`; the 2xx says `forwarded` so no surface claims a delivery that did not happen, and 409 is left for what is settled (an already answered card, a non-root addressee). History replay merges the projection state | `ouroboros/gateway/contracts.py`, `ouroboros/gateway/task_decision.py`, `ouroboros/owner_quiz.py`, `ouroboros/tools/core.py` | `tests/test_gateway_parity.py`, `tests/test_quiz_display.py`, `tests/test_quiz_answer.py`, `web/tests/chat_decision.test.js` |
|
||||
| Managed update ABI — preflight, `UpdateMergePlan`, pinned apply, process-local `update_progress`, `update_progress_changed` invalidation and boot-only `update_status_ready` | `ouroboros/gateway/contracts.py` | `tests/test_update_apply_routing.py` |
|
||||
| `ChatOutbound.review_projection` — bounded actor findings via `utils.truncate_review_artifact`, at most `MAX_PROJECTED_ACTOR_FINDINGS` rows (`review_execution_projection.py`) | `ouroboros/gateway/contracts.py` | `tests/test_review_substrate_v2.py`, `web/tests/review_truth.test.js` |
|
||||
| Skill preflight statuses — `preflight_failed` is fresh-only; a stale failure surfaces as `preflight_failed_stale`; absence means the caller could not know | `ouroboros/skill_review_status.py` | `tests/test_skill_preflight_repair.py`, `web/tests/skill_preflight_repair.test.js` |
|
||||
|
|
|
|||
|
|
@ -982,12 +982,12 @@ and what enforces each.
|
|||
generation. File/diff requests impose no commit-or-revert rule; self-modification
|
||||
keeps reviewed commits (BIBLE P0/P3).
|
||||
- Before cleanup, freeze `review_evidence.task_inputs` and `completion_observations`
|
||||
for summary/reflection (ARCHITECTURE §6 "Post-task reflection"): run origin, whole
|
||||
for reflection (ARCHITECTURE §6 "Post-task reflection"): run origin, whole
|
||||
owner Q/A, peer provenance and canonical split-root verification receipts. Zero exit is positive;
|
||||
absent is unknown; unrelated passes erase no failure. Send content, not pointers;
|
||||
recover the same snapshot. Count delivery via `OWNER_DELIVERY_TOOL_NAMES`, never
|
||||
global skill state. Summary uses `chat_observed` custody and the task-scoped,
|
||||
archive-aware trace reader.
|
||||
global skill state. The free `host_task_facts` row makes no model call; the paid
|
||||
reflection and its Pattern Register write use `chat_observed` custody.
|
||||
- Promoted tasks carry their host-minted root id and role on the queue payload.
|
||||
RUNNING writes preserve the actual `_task_started_ts` as `started_at` and an existing
|
||||
`queued_at`; terminal `ts` stays its own field; missing historical start facts stay
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
Machine extraction of `docs/ARCHITECTURE.md` §11.1 (the frozen-ABI SSOT), regenerated by `python scripts/regenerate_inventories.py`. Do not edit — edit the owning chapter named in the Source line and regenerate; `tests/test_generated_inventories.py` pins byte-identity and the resolution invariants (a §11.1 row whose owner or anchor file disappeared from the tree = red).
|
||||
|
||||
Source: `docs/architecture/11-frozen-contracts-v1.md`, physical LF lines 7-40; UTF-8 SHA-256 `2aa8ff73bf8a7723250c834d62a91bd1f305583b210cea4ab324839d21250042`.
|
||||
Source: `docs/architecture/11-frozen-contracts-v1.md`, physical LF lines 7-40; UTF-8 SHA-256 `1e9783a2f13a1be07e5eb5d6962eabf2df29a27f34ce0d377c694b765d25b018`.
|
||||
|
||||
- table rows: **29**
|
||||
- browser-envelope prose owners:
|
||||
|
|
|
|||
|
|
@ -717,6 +717,11 @@ def finish_exposed_preparation_author(ctx: Any, record: Dict[str, Any]) -> bool:
|
|||
_loop()._supersede_task_acceptance_for_owner_followup(ctx.tools._ctx, ctx.llm_trace)
|
||||
return True
|
||||
ctx.tools._ctx._task_acceptance_reviewed = True
|
||||
if action == "stop":
|
||||
# Like the reviewer-bound stop, this one binds no subject: an earlier panel's
|
||||
# must not reopen review over changed material on a later delivery pass; only
|
||||
# the author's next decision or owner input does (TZ-2 C4).
|
||||
ctx.tools._ctx._task_acceptance_reviewed_subject = ""
|
||||
ctx.tools._ctx._task_acceptance_pending = ""
|
||||
_loop()._mark_root_acceptance_checkpoint(
|
||||
ctx.tools._ctx, ctx.llm_trace, status="preparation_failed", pass_index=ctx.passes_done,
|
||||
|
|
|
|||
|
|
@ -233,7 +233,8 @@ def _run_post_task_processing_async(
|
|||
env, task_memory, llm_client)),
|
||||
("reflection", lambda: result.__setitem__("reflection_entry", _run_reflection(
|
||||
env, llm_client, task_snapshot, usage_snapshot, trace_snapshot,
|
||||
review_evidence_snapshot, sealed_final=sealed_snapshot))),
|
||||
review_evidence_snapshot, sealed_final=sealed_snapshot,
|
||||
publish=lambda entry: result.__setitem__("reflection_entry", entry)))),
|
||||
("promotion", _promotion),
|
||||
]
|
||||
from ouroboros.post_task_synthesis import (
|
||||
|
|
|
|||
|
|
@ -153,9 +153,9 @@ class ChatOutbound(TypedDict):
|
|||
# A cancellation fault names the PHYSICAL task it could not settle when it differs from the logical task id.
|
||||
cancel_physical_task_id: NotRequired[str]
|
||||
toast_once: NotRequired[str]
|
||||
# #628: the incident's valence for the one-shot toast (warn/ok/error),
|
||||
# stamped by the producer that knows whether the boundary is a wait, a
|
||||
# recovery or an exhaustion; absent = the browser keeps its alarm tone.
|
||||
# #628: the one-shot toast's valence (warn/ok/error; the reaper's rail ``warning``
|
||||
# is normalizeTone's existing ``warn`` alias, no new tone), stamped by the producer
|
||||
# that knows wait/recovery/exhaustion; absent = the browser keeps its alarm tone.
|
||||
toast_tone: NotRequired[str]
|
||||
lifecycle: NotRequired[Dict[str, Any]]
|
||||
# C4 multi-chat dedupe: a duplicate lifecycle initiator's typed pointer to
|
||||
|
|
|
|||
|
|
@ -16,7 +16,7 @@ import logging
|
|||
import pathlib
|
||||
|
||||
from dataclasses import replace
|
||||
from typing import Any, Dict
|
||||
from typing import Any, Callable, Dict
|
||||
from ouroboros.dialogue_provenance import presence_provenance_fields
|
||||
from ouroboros.llm_claudexor import propagate_model_error
|
||||
from ouroboros.outcomes import normalize_outcome_axes
|
||||
|
|
@ -593,14 +593,17 @@ def _post_task_paid_interruption(errors: Any) -> str:
|
|||
Memory consolidation returns errors to keep completed chunks. Only this
|
||||
stage adapter interprets those existing facts: a budget or unknown-provider
|
||||
kind wins and stops later paid post-work; any other kind names the last
|
||||
ordinary failure, so a stage that lost a chunk reads ``degraded`` like a
|
||||
stage that raised (TZ-2 C3: unfinished stages are never ``completed``).
|
||||
UNRESOLVED ordinary failure, so a stage that lost a chunk reads ``degraded``
|
||||
like a stage that raised (TZ-2 C3: unfinished stages are never ``completed``).
|
||||
The history keeps every attempt; a refusal its producer answered (a split whose
|
||||
halves carry their own rows, ``resolution``) is not an unfinished stage.
|
||||
"""
|
||||
rows = [row for row in (errors if isinstance(errors, list) else []) if isinstance(row, dict)]
|
||||
for row in rows:
|
||||
if row.get("kind") in POST_TASK_INTERRUPT_KINDS:
|
||||
return str(row["kind"])
|
||||
return str((rows[-1].get("kind") or "stage_error")) if rows else ""
|
||||
unresolved = [row for row in rows if not row.get("resolution")]
|
||||
return str((unresolved[-1].get("kind") or "stage_error")) if unresolved else ""
|
||||
|
||||
|
||||
def _run_chat_consolidation(env, memory, llm, task, drive_logs):
|
||||
|
|
@ -704,12 +707,15 @@ def _run_scratchpad_consolidation(env: Any, memory: Any, llm: Any) -> None:
|
|||
def _run_reflection(env: Any, llm: Any, task: Dict[str, Any],
|
||||
usage: Dict[str, Any], llm_trace: Dict[str, Any],
|
||||
review_evidence: Dict[str, Any],
|
||||
sealed_final: Dict[str, Any] | None = None) -> Dict[str, Any] | None:
|
||||
sealed_final: Dict[str, Any] | None = None,
|
||||
publish: Callable[[Dict[str, Any]], Any] | None = None) -> Dict[str, Any] | None:
|
||||
"""Run execution reflection synchronously (process memory, Bible P1).
|
||||
|
||||
Returns the entry, or None only when there is nothing to reflect on; a
|
||||
failure raises to the post-task stage coordinator, which degrades the
|
||||
checkpoint and still runs the later stages (TZ-2 C3).
|
||||
checkpoint and still runs the later stages (TZ-2 C3). ``publish`` receives
|
||||
the completed entry before its nested paid Pattern Register write, so an
|
||||
interruption there still leaves the coordinator its free memory actions.
|
||||
"""
|
||||
from ouroboros.reflection import (
|
||||
should_generate_reflection, generate_reflection, append_reflection_routed,
|
||||
|
|
@ -750,6 +756,8 @@ def _run_reflection(env: Any, llm: Any, task: Dict[str, Any],
|
|||
knowledge_context=knowledge_context,
|
||||
)
|
||||
entry = {**entry, **presence_provenance_fields(task)}
|
||||
if publish is not None:
|
||||
publish(entry)
|
||||
append_reflection_routed(env, task, entry)
|
||||
return entry
|
||||
return None
|
||||
|
|
|
|||
|
|
@ -814,8 +814,35 @@ def _admits_pattern_register(entry: Dict[str, Any]) -> bool:
|
|||
)
|
||||
|
||||
|
||||
def _update_pattern_register(drive_root: pathlib.Path, entry: Dict[str, Any]) -> None:
|
||||
"""The reflection stage's nested paid write, under the post-task stage protocol.
|
||||
|
||||
It runs after the reflection is persisted and buys nothing when the reflection's
|
||||
own call was interrupted (a budget or unknown-outcome row). A control, the wallet
|
||||
or an unresolved attempt on any provider's chain propagates, so no later paid
|
||||
post-work runs (TZ-2 C3). An ordinary failure is one typed row on the entry's
|
||||
``memory_operation_errors``: learning that silently fails to land is invisible
|
||||
erosion (P1), so the stage reads degraded while permitted later stages still run.
|
||||
"""
|
||||
from ouroboros.post_task_synthesis import POST_TASK_INTERRUPT_KINDS, propagate_paid_interruption
|
||||
|
||||
errors = entry.get("memory_operation_errors") or []
|
||||
if not _admits_pattern_register(entry) or any(
|
||||
isinstance(row, dict) and row.get("kind") in POST_TASK_INTERRUPT_KINDS for row in errors):
|
||||
return
|
||||
try:
|
||||
_update_patterns(drive_root, entry)
|
||||
except Exception as exc:
|
||||
propagate_paid_interruption(exc)
|
||||
from ouroboros.utils import sanitize_tool_result_for_log
|
||||
|
||||
log.warning("Pattern register update failed for task %s: %s", entry.get("task_id", "?"), exc, exc_info=True)
|
||||
entry["memory_operation_errors"] = [*errors, {"kind": "pattern_register_failed", "label": "Pattern Register",
|
||||
"message": sanitize_tool_result_for_log(str(exc)) or type(exc).__name__}]
|
||||
|
||||
|
||||
def append_reflection(drive_root: pathlib.Path, entry: Dict[str, Any]) -> None:
|
||||
"""Persist a reflection entry to the JSONL file."""
|
||||
"""Persist a reflection entry to the JSONL file, then its Pattern Register write."""
|
||||
reflections_path = drive_root / "logs" / REFLECTIONS_FILENAME
|
||||
try:
|
||||
append_jsonl(reflections_path, entry)
|
||||
|
|
@ -823,15 +850,7 @@ def append_reflection(drive_root: pathlib.Path, entry: Dict[str, Any]) -> None:
|
|||
entry.get("task_id", "?"), entry.get("key_markers", []))
|
||||
except Exception:
|
||||
log.warning("Failed to save execution reflection", exc_info=True)
|
||||
|
||||
if _admits_pattern_register(entry):
|
||||
try:
|
||||
_update_patterns(drive_root, entry)
|
||||
except Exception as exc:
|
||||
# Learning that silently fails to land is invisible erosion: the
|
||||
# register simply never hears about this class again (P1).
|
||||
log.warning("Pattern register update failed for task %s: %s",
|
||||
entry.get("task_id", "?"), exc, exc_info=True)
|
||||
_update_pattern_register(drive_root, entry)
|
||||
|
||||
|
||||
def append_reflection_routed(env: Any, task: Dict[str, Any], entry: Dict[str, Any]) -> None:
|
||||
|
|
@ -856,8 +875,10 @@ def append_reflection_routed(env: Any, task: Dict[str, Any], entry: Dict[str, An
|
|||
except Exception:
|
||||
pid = ""
|
||||
if not pid:
|
||||
append_reflection(canonical, entry)
|
||||
_bind_reflection_action_source(canonical, entry)
|
||||
try:
|
||||
append_reflection(canonical, entry)
|
||||
finally: # a paid interruption propagates only after the free source binding
|
||||
_bind_reflection_action_source(canonical, entry)
|
||||
return
|
||||
from ouroboros.project_facts import project_reflections_path
|
||||
|
||||
|
|
@ -871,12 +892,6 @@ def append_reflection_routed(env: Any, task: Dict[str, Any], entry: Dict[str, An
|
|||
except Exception:
|
||||
project_write_failed = True
|
||||
log.warning("Failed to save project execution reflection", exc_info=True)
|
||||
if _admits_pattern_register(entry):
|
||||
try:
|
||||
_update_patterns(canonical, entry)
|
||||
except Exception as exc:
|
||||
log.warning("Pattern register update failed for task %s: %s",
|
||||
entry.get("task_id", "?"), exc, exc_info=True)
|
||||
try:
|
||||
append_jsonl(canonical / "logs" / REFLECTIONS_FILENAME, {
|
||||
"ts": str(entry.get("ts") or utc_now_iso()),
|
||||
|
|
@ -892,6 +907,7 @@ def append_reflection_routed(env: Any, task: Dict[str, Any], entry: Dict[str, An
|
|||
except Exception:
|
||||
log.warning("Failed to write canonical reflection pointer", exc_info=True)
|
||||
_bind_reflection_action_source(canonical, entry)
|
||||
_update_pattern_register(canonical, entry) # paid, last: its interruption loses no free write
|
||||
|
||||
|
||||
def _bind_reflection_action_source(canonical: pathlib.Path, entry: Dict[str, Any]) -> None:
|
||||
|
|
|
|||
|
|
@ -230,6 +230,9 @@ def summarize_source(
|
|||
return False
|
||||
midpoint = start + len(halves[0])
|
||||
pending.extend([(midpoint, end), (start, midpoint)])
|
||||
# The refusal is answered by its halves, each accounted by its own row; the
|
||||
# attempt stays in the usage history, never read as an unresolved failure.
|
||||
failure["resolution"] = "split"
|
||||
return True
|
||||
|
||||
while pending:
|
||||
|
|
|
|||
|
|
@ -76,9 +76,13 @@ def _run_stop_loop(tmp_path, monkeypatch, responses, *, stop_services):
|
|||
answers = iter(responses)
|
||||
model_inputs: list = []
|
||||
|
||||
def fake_call(_llm, request_messages, *_args, **_kwargs):
|
||||
def fake_call(_llm, request_messages, *_args, **kwargs):
|
||||
model_inputs.append([dict(row) for row in request_messages])
|
||||
answer = next(answers)
|
||||
# The scripted Main replaces the transport, including its actual-context
|
||||
# observer: the request whose response returns exposes the feedback it carried.
|
||||
if callable(kwargs.get("model_context_observer")):
|
||||
kwargs["model_context_observer"](request_messages)
|
||||
if isinstance(answer, dict):
|
||||
return {"role": "assistant", **answer}, 0.0
|
||||
return {"role": "assistant", "content": answer}, 0.0
|
||||
|
|
@ -455,3 +459,106 @@ def test_owner_input_consumes_the_stop_so_a_later_exhausted_exit_is_not_read_as_
|
|||
assert later["agent_rationale"] == STOP_RATIONALE # history kept, never the act
|
||||
assert not recorded_author_stop(later) and STOP_RATIONALE.rstrip(".") not in row()
|
||||
assert fx.calls == ["initial answer"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("change", ["material", "evidence"])
|
||||
def test_a_local_preparation_stop_after_a_fail_panel_keeps_its_finality(tmp_path, monkeypatch, change):
|
||||
"""(h) FAIL panel → the next host pass cannot assemble its evidence locally → the
|
||||
informed author stops → material (a new working-tree file) or evidence (a late
|
||||
``services_stopped``) changes and Main restates its stop. The local-preparation stop
|
||||
kept the FAIL panel's reviewed subject, so ``preparation_delivery_choice`` refused
|
||||
the changed material, the stale subject cleared the reviewed latch and the host
|
||||
bought two more panels over a stop. Like the reviewer-bound stop it binds no
|
||||
subject: one panel, one failed assembly, the stop's local cause, act and rationale."""
|
||||
import ouroboros.loop as loop
|
||||
import ouroboros.loop_acceptance_review as review
|
||||
from ouroboros.outcomes import derive_loop_outcome
|
||||
|
||||
real_build, builds = review._build_host_acceptance_evidence, []
|
||||
|
||||
def build(ctx):
|
||||
builds.append(ctx.content)
|
||||
if len(builds) == 2:
|
||||
raise RuntimeError("local evidence assembly failed")
|
||||
return real_build(ctx)
|
||||
|
||||
monkeypatch.setattr(review, "_build_host_acceptance_evidence", build)
|
||||
real_project, late = loop._project_child_result_dispositions, []
|
||||
|
||||
def project(limit_ctx, llm_trace):
|
||||
if (llm_trace.get("acceptance_decision") or {}).get("reason") == "author_stop" and not late:
|
||||
late.append(True)
|
||||
if change == "material":
|
||||
(tmp_path / "repo" / "late.txt").write_text("changed after the stop\n", encoding="utf-8")
|
||||
else:
|
||||
llm_trace.setdefault("verification_events", []).append({"kind": "services_stopped", "services": [
|
||||
{"service_id": "late", "name": "late", "lifecycle": "stopped"}]})
|
||||
return real_project(limit_ctx, llm_trace)
|
||||
|
||||
monkeypatch.setattr(loop, "_project_child_result_dispositions", project)
|
||||
keep = json.dumps({"delivery_control": "keep"})
|
||||
run = _run_stop_loop(tmp_path, monkeypatch, [
|
||||
"The export endpoint ships.", # reviewed: FAIL, capsule fed back
|
||||
"The export endpoint ships, revised.", # the host cannot assemble its evidence
|
||||
_stop_tool_call(), # the informed author stops
|
||||
*[step for _ in range(3) for step in (STOP_TEXT, keep)],
|
||||
], stop_services=False)
|
||||
assert late, "the change must land after the stop was honoured"
|
||||
assert run.panels == ["The export endpoint ships."], "a local-preparation stop must never buy a panel"
|
||||
assert len(builds) == 2, "the stopped material is never assembled again"
|
||||
assert run.result == STOP_TEXT
|
||||
decision = run.trace["acceptance_decision"]
|
||||
assert decision["origin"] == "local_acceptance_preparation"
|
||||
assert decision["reason"] == "author_stop" and decision["author_action"] == "stop"
|
||||
assert decision["author_disposition"]["rationale"] == STOP_RATIONALE
|
||||
axes = derive_loop_outcome(run.result, run.usage, run.trace)["outcome_axes"]
|
||||
assert axes["objective"]["status"] == "fail" and axes["objective"]["reason"] == "author_stop"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("reopen", ["owner_input", "author_finish"])
|
||||
def test_owner_input_or_a_new_author_act_still_reopens_a_local_preparation_stop(tmp_path, monkeypatch, reopen):
|
||||
"""(i) The subject-free local-preparation stop holds over changed text, yet owner
|
||||
input returns the answer to the ordinary host pass and the author's next explicit
|
||||
act is heard: a finish replaces the stop through the same incident, without a panel."""
|
||||
import ouroboros.loop as loop_mod
|
||||
import ouroboros.loop_acceptance_review as review
|
||||
from ouroboros.acceptance_settlement import expose_acceptance_feedback
|
||||
from ouroboros.loop_acceptance import merge_agent_acceptance_stance
|
||||
|
||||
fx = _finality_pass(tmp_path, monkeypatch)
|
||||
assert fx.run("initial answer") is True
|
||||
assert fx.ctx._task_acceptance_reviewed_subject, "the FAIL panel bound its subject"
|
||||
|
||||
def broken(_ctx):
|
||||
raise RuntimeError("local evidence assembly failed")
|
||||
|
||||
monkeypatch.setattr(review, "_build_host_acceptance_evidence", broken)
|
||||
assert fx.run("revised answer") is True
|
||||
assert fx.trace["acceptance_decision"]["reason"] == "acceptance_preparation_failed"
|
||||
expose_acceptance_feedback(fx.trace, fx.messages, "author-root")
|
||||
fx.trace["tool_calls"].append({"tool": "task_acceptance_review", "args": {}})
|
||||
merge_agent_acceptance_stance(fx.trace, {"explicit_finish": True, "author_action": "stop",
|
||||
"rationale": STOP_RATIONALE}, fx.ctx)
|
||||
assert fx.run(STOP_TEXT) is False
|
||||
assert fx.trace["acceptance_decision"]["reason"] == "author_stop"
|
||||
assert fx.ctx._task_acceptance_reviewed_subject == ""
|
||||
assert fx.run("A restated unfinished answer.") is False
|
||||
assert fx.trace["acceptance_decision"]["reason"] == "author_stop"
|
||||
if reopen == "owner_input":
|
||||
loop_mod._supersede_task_acceptance_for_owner_followup(fx.ctx, fx.trace)
|
||||
assert fx.run("The answer to the owner's follow-up.") is False
|
||||
# The ordinary host pass decided again: the same unrepaired material is not
|
||||
# rebuilt, and the host's own honest ending replaced the consumed stop.
|
||||
decision = fx.trace["acceptance_decision"]
|
||||
assert decision["reason"] == "acceptance_preparation_failed" and "author_action" not in decision
|
||||
assert fx.trace["acceptance_preparation"]["attempts"] == 1
|
||||
else:
|
||||
fx.trace["tool_calls"].append({"tool": "task_acceptance_review", "args": {}})
|
||||
merge_agent_acceptance_stance(fx.trace, {"disposition": "partial", "explicit_finish": True,
|
||||
"author_action": "finish",
|
||||
"rationale": "Delivering the available result with its gap."}, fx.ctx)
|
||||
assert fx.ctx._task_acceptance_reviewed is False
|
||||
assert fx.run("A restated unfinished answer.") is False
|
||||
decision = fx.trace["acceptance_decision"]
|
||||
assert decision["reason"] == "author_finish" and decision["author_disposition"]["action"] == "finish"
|
||||
assert fx.calls == ["initial answer"]
|
||||
|
|
|
|||
|
|
@ -192,6 +192,7 @@ def test_oversized_logical_block_splits_complete_source_and_advances_once(tmp_pa
|
|||
assert not c.should_consolidate(meta, chat)
|
||||
refused = usage["_consolidation_errors"][0]
|
||||
assert refused["kind"] == "context_overflow" and not refused["preflight_only"]
|
||||
assert refused["resolution"] == "split" # history kept; the post-task adapter reads it as answered
|
||||
# A refusal that was split and then fully summarized is a recovered attempt, not a
|
||||
# failed run: the block was written, so no stale error may outlive the advance.
|
||||
assert "last_consolidation_error" not in json.loads(meta.read_text())
|
||||
|
|
@ -258,9 +259,10 @@ def test_unknown_capacity_impossible_overhead_does_not_replay_next_cycle(tmp_pat
|
|||
chat, blocks, meta = _paths(tmp_path)
|
||||
_write_chat(chat)
|
||||
llm = _LLM(limit=1)
|
||||
c.consolidate(chat, blocks, meta, llm)
|
||||
usage = c.consolidate(chat, blocks, meta, llm)
|
||||
first_calls = len(llm.calls)
|
||||
assert first_calls > 0
|
||||
assert not usage["_consolidation_errors"][-1].get("resolution") # the unsplittable refusal stays unresolved
|
||||
c.consolidate(chat, blocks, meta, llm)
|
||||
assert len(llm.calls) == first_calls
|
||||
assert not blocks.exists()
|
||||
|
|
|
|||
|
|
@ -855,3 +855,73 @@ def test_split_root_facts_row_counts_the_actor_store_when_synthesis_runs_canonic
|
|||
assert [s["store"] for s in fact["stores"]] == [
|
||||
str(task_artifacts_dir(f.root, f.task["id"], create=False)),
|
||||
str(task_artifacts_dir(child, f.task["id"], create=False))]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("second_chunk", ["recovered", "lost", "budget"])
|
||||
def test_split_recovery_history_is_not_an_unresolved_consolidation_failure(phase, monkeypatch, second_chunk):
|
||||
"""F-R3: the stage adapter read ``_consolidation_errors`` attempt HISTORY as an
|
||||
unresolved failure, so a context refusal that the real consolidator answered by
|
||||
splitting (and then wrote the block and advanced the cursor) turned post-work
|
||||
``degraded``. Through the REAL chat-consolidation adapter and consolidator (only
|
||||
the provider dispatch is substituted): the refusal row stays in the history with
|
||||
its explicit ``resolution``; a later chunk that is lost still reads degraded
|
||||
without a skip (partial success is not success), and the wallet still stops."""
|
||||
import json
|
||||
from ouroboros import consolidator, context_fit, llm_observability, post_task_synthesis
|
||||
from ouroboros.capability_evidence import CapabilityEvidence
|
||||
from ouroboros.usage_accounting import BudgetExceeded
|
||||
|
||||
f = phase
|
||||
monkeypatch.setattr(consolidator, "_consolidation_route", lambda: ("test/model", False))
|
||||
monkeypatch.setattr(context_fit, "resolve_context_fit_route", lambda task, *, allow_fetch: (
|
||||
{"model": task["model"], "provider": "openrouter"},
|
||||
CapabilityEvidence(0, "unknown", "test", "route-test", model=task["model"], provider="openrouter")))
|
||||
monkeypatch.setattr(context_fit, "_route_calibration_ratio", lambda *_: 1.0)
|
||||
monkeypatch.setattr(pipeline, "_run_chat_consolidation", post_task_synthesis._run_chat_consolidation)
|
||||
monkeypatch.setattr(pipeline, "_run_reflection", lambda *a, **k: f.stages.append("reflection") or None)
|
||||
chat = f.root / "logs" / "chat.jsonl"
|
||||
chat.parent.mkdir(parents=True, exist_ok=True)
|
||||
chat.write_text("".join(json.dumps({"ts": f"2026-01-01T{i // 60:02d}:{i % 60:02d}:00Z", "direction": "in",
|
||||
"text": f"entry-{i} " + "x" * 120, "chat_id": 1}) + "\n"
|
||||
for i in range(200)), encoding="utf-8")
|
||||
refused = []
|
||||
|
||||
def dispatch(_client, *, call_type="", messages=(), **_kwargs):
|
||||
prompt = messages[0]["content"]
|
||||
if call_type == "memory_consolidation" and "entry-0 " in prompt and "entry-99 " in prompt and not refused:
|
||||
refused.append(call_type)
|
||||
raise transport.ClaudexorModelError({"code": "invalid_request", "message": "Controlled provider refusal",
|
||||
"context": {"httpStatus": 400, "vendorCode": "context_length_exceeded", "parameter": "input"}})
|
||||
if "entry-150 " in prompt and second_chunk != "recovered":
|
||||
f.stages.append("second-chunk")
|
||||
raise BudgetExceeded("root wallet spent") if second_chunk == "budget" else RuntimeError("provider failed")
|
||||
return {"content": f"summary of {call_type}"}, {"prompt_tokens": 1, "completion_tokens": 1,
|
||||
"total_tokens": 2, "cost": 0.0}
|
||||
|
||||
monkeypatch.setattr(llm_observability, "chat_observed", dispatch)
|
||||
launch(f)
|
||||
assert f.done.wait(10)
|
||||
assert refused, "the first chunk's complete draft was refused for context"
|
||||
checkpoint = load_task_result(f.root, f.task["id"])["root_phase_checkpoint"]
|
||||
meta = json.loads((f.root / "memory" / "dialogue_meta.json").read_text(encoding="utf-8"))
|
||||
blocks = json.loads((f.root / "memory" / "dialogue_blocks.json").read_text(encoding="utf-8"))
|
||||
events = [json.loads(line) for line in (f.root / "logs" / "events.jsonl").read_text(encoding="utf-8").splitlines()]
|
||||
[row] = [event for event in events if event.get("type") == "chat_block_consolidation"]
|
||||
assert len(blocks) == (2 if second_chunk == "recovered" else 1), "the recovered chunk is a published block"
|
||||
assert meta["last_consolidated_offset"] == (200 if second_chunk == "recovered" else 100)
|
||||
if second_chunk == "recovered":
|
||||
assert checkpoint["post_task_synthesis"] == "completed"
|
||||
assert not checkpoint.get("post_task_stop_reason")
|
||||
assert row["last_error_kind"] == "context_overflow", "the attempt history is preserved"
|
||||
assert "last_consolidation_error" not in meta
|
||||
assert f.stages[-3:] == ["scratch", "reflection", "backlog"]
|
||||
elif second_chunk == "lost":
|
||||
assert checkpoint["post_task_synthesis"] == "degraded"
|
||||
assert not checkpoint.get("post_task_stop_reason")
|
||||
assert meta["last_consolidation_error"]["cursor_offset"] == 100
|
||||
assert f.stages[-4:] == ["second-chunk", "scratch", "reflection", "backlog"]
|
||||
else:
|
||||
assert checkpoint["post_task_synthesis"] == "degraded"
|
||||
assert checkpoint["post_task_stop_reason"] == (
|
||||
"budget_exhausted:skipped=scratchpad_consolidation,reflection,promotion")
|
||||
assert f.stages[-1] == "second-chunk"
|
||||
|
|
|
|||
|
|
@ -2,15 +2,21 @@
|
|||
|
||||
Split out of ``tests/test_agent_task_pipeline.py`` when that module was divided
|
||||
by theme; every moved block is verbatim. Covers `_run_reflection` entry
|
||||
generation, `_update_improvement_backlog`, and the project-scoped channel
|
||||
split: project memory stays project-local while backlog promotion goes to the
|
||||
global drive through `_run_global_backlog_promotion_only`.
|
||||
generation, `_update_improvement_backlog`, the project-scoped channel
|
||||
split (project memory stays project-local while backlog promotion goes to the
|
||||
global drive through `_run_global_backlog_promotion_only`) and the reflection
|
||||
stage's nested paid Pattern Register write under the post-task stage protocol.
|
||||
"""
|
||||
|
||||
import json
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
import ouroboros.agent_task_pipeline as pipeline
|
||||
from ouroboros import model_wait
|
||||
from ouroboros.task_results import load_task_result
|
||||
from tests.test_post_task_model_wait import _generic_unknown, phase as phase
|
||||
|
||||
|
||||
def test_project_scoped_post_task_processing_feeds_global_backlog_but_project_memory(tmp_path, monkeypatch):
|
||||
|
|
@ -159,3 +165,82 @@ def test_run_reflection_returns_entry_when_generated(tmp_path, monkeypatch):
|
|||
(tmp_path / "logs" / "task_reflections.jsonl").read_text(encoding="utf-8").splitlines()]
|
||||
assert stored == [entry]
|
||||
assert captured["pattern_root"] == tmp_path and captured["pattern_entry"] == entry
|
||||
|
||||
|
||||
@pytest.mark.parametrize("scope", ["global", "project"])
|
||||
@pytest.mark.parametrize("outcome", ["reflection_unknown", "pattern_unknown", "pattern_budget",
|
||||
"pattern_deadline", "pattern_ordinary", "pattern_ok"])
|
||||
def test_nested_pattern_register_follows_the_post_task_stage_protocol(phase, monkeypatch, tmp_path, scope, outcome):
|
||||
"""F-R1: the reflection stage's nested paid Pattern Register write ran even when the
|
||||
reflection's own call had already recorded an unknown outcome, and both writers
|
||||
(global ``append_reflection``, project branch of ``append_reflection_routed``)
|
||||
swallowed a budget refusal, an unresolved attempt, a control and an ordinary
|
||||
failure alike, so the promotion stage bought its paid calls and the checkpoint
|
||||
read ``completed``. Through the REAL reflection, routing and register adapters
|
||||
(only the provider dispatch is substituted): an interrupted reflection buys no
|
||||
register call; a register interruption stops promotion after the reflection is
|
||||
persisted and its free actions are applied once; an ordinary register failure
|
||||
degrades while promotion still runs; a register success completes."""
|
||||
from ouroboros import consolidator, context_fit, llm_observability, post_task_synthesis, project_facts
|
||||
from ouroboros.capability_evidence import CapabilityEvidence
|
||||
from ouroboros.usage_accounting import BudgetExceeded
|
||||
|
||||
f = phase
|
||||
monkeypatch.setattr(consolidator, "_consolidation_route", lambda: ("test/model", False))
|
||||
monkeypatch.setattr(context_fit, "resolve_context_fit_route", lambda task, *, allow_fetch: (
|
||||
{"model": task["model"], "provider": "openrouter"},
|
||||
CapabilityEvidence(100_000, "confirmed", "test", "route-test", model=task["model"], provider="openrouter")))
|
||||
monkeypatch.setattr(context_fit, "_route_calibration_ratio", lambda *_: 1.0)
|
||||
monkeypatch.setattr(pipeline, "_run_reflection", post_task_synthesis._run_reflection)
|
||||
if scope == "project":
|
||||
monkeypatch.setattr(project_facts, "_project_store_root", lambda pid: tmp_path / "projects" / pid)
|
||||
f.task["project_id"] = "slime"
|
||||
failures = {"reflection_unknown": _generic_unknown("direct"), "pattern_unknown": _generic_unknown("cause"),
|
||||
"pattern_budget": BudgetExceeded("root wallet spent"),
|
||||
"pattern_deadline": model_wait.ModelWaitInterrupted("deadline"),
|
||||
"pattern_ordinary": RuntimeError("pattern provider failed")}
|
||||
|
||||
def dispatch(*_args, call_type="", **_kwargs):
|
||||
f.stages.append(call_type)
|
||||
failing = "task_reflection" if outcome == "reflection_unknown" else "pattern_register_update"
|
||||
if call_type == failing and outcome in failures:
|
||||
raise failures[outcome]
|
||||
content = ("| Error class | Count | Root cause | Structural fix | Status |\n|---|---|---|---|---|\n"
|
||||
"| run_command | 1 | boom | typed | open |") if call_type == "pattern_register_update" else "Lesson: typed."
|
||||
return {"content": content}, {"prompt_tokens": 1, "completion_tokens": 1, "total_tokens": 2, "cost": 0.0}
|
||||
|
||||
monkeypatch.setattr(llm_observability, "chat_observed", dispatch)
|
||||
applied, entries = [], []
|
||||
monkeypatch.setattr(pipeline, "_apply_reflection_memory_actions", lambda *a, **k: applied.append(1))
|
||||
trace = {"tool_calls": [{"tool": "run_command", "status": "error", "is_error": True, "result": "boom"}]}
|
||||
pipeline._run_post_task_processing_async(
|
||||
f.env, f.task, {"rounds": 3}, trace, {}, f.root / "logs", event_queue=f.events,
|
||||
on_reflection=lambda entry, _llm: entries.append(entry))
|
||||
assert f.done.wait(5)
|
||||
checkpoint = load_task_result(f.root, f.task["id"])["root_phase_checkpoint"]
|
||||
reflected = ["facts", "chat", "scratch", "task_reflection"]
|
||||
stop = {"reflection_unknown": "provider_outcome_unknown", "pattern_unknown": "provider_outcome_unknown",
|
||||
"pattern_budget": "budget_exhausted", "pattern_deadline": "deadline"}.get(outcome)
|
||||
if stop:
|
||||
assert checkpoint["post_task_synthesis"] == "degraded"
|
||||
assert checkpoint["post_task_stop_reason"] == f"{stop}:skipped=promotion"
|
||||
assert f.stages == reflected + ([] if outcome == "reflection_unknown" else ["pattern_register_update"])
|
||||
assert entries == [], "no paid promotion after an interruption"
|
||||
else:
|
||||
assert checkpoint["post_task_synthesis"] == ("degraded" if outcome == "pattern_ordinary" else "completed")
|
||||
assert not checkpoint.get("post_task_stop_reason")
|
||||
assert f.stages == reflected + ["pattern_register_update", "backlog"]
|
||||
[entry] = entries
|
||||
kinds = [row["kind"] for row in entry.get("memory_operation_errors") or []]
|
||||
assert kinds == (["pattern_register_failed"] if outcome == "pattern_ordinary" else [])
|
||||
assert applied == [1], "the completed reflection's free actions are applied exactly once"
|
||||
canonical = [json.loads(line) for line in (f.root / "logs" / "task_reflections.jsonl").read_text(
|
||||
encoding="utf-8").splitlines()]
|
||||
if scope == "project":
|
||||
assert [row["type"] for row in canonical] == ["project_reflection_pointer"]
|
||||
persisted = (tmp_path / "projects" / "slime" / "logs" / "task_reflections.jsonl").read_text(encoding="utf-8")
|
||||
assert len(persisted.splitlines()) == 1
|
||||
else:
|
||||
assert len(canonical) == 1 and canonical[0]["task_id"] == f.task["id"]
|
||||
patterns = f.root / "memory" / "knowledge" / "patterns.md"
|
||||
assert patterns.exists() == (outcome == "pattern_ok")
|
||||
|
|
|
|||
|
|
@ -68,9 +68,11 @@ OBSERVE_JS = """tid => { const card = document.querySelector(`#chat-messages .ch
|
|||
|
||||
class _OutageModel(ScriptedStubModel):
|
||||
"""One real (failing) tool round, then a provider refusal (HTTP 401, a permanent
|
||||
class, so no backoff retries) on every later tool round of the marked task. The
|
||||
host's terminal incident preserves the intermediate output if a later call
|
||||
cannot land; the post-task reflection is the one call the gate holds."""
|
||||
class, so no backoff retries) on every later tool round of the marked task AND on
|
||||
the host's forced outage final (``[PROVIDER_UNAVAILABLE]``): the provider is down
|
||||
for that call too, so the host's terminal incident preserves the intermediate
|
||||
output (``host_salvage``). The post-task reflection is never refused — it is the
|
||||
one call the gate holds."""
|
||||
|
||||
def __init__(self, gate):
|
||||
super().__init__([{"tool": "run_command", "arguments": {
|
||||
|
|
@ -98,9 +100,11 @@ class _OutageModel(ScriptedStubModel):
|
|||
|
||||
def _refuse(self, body):
|
||||
text = body_text(body)
|
||||
if not body.get("tools") or MARKER not in text or SALVAGE_MARKER in text:
|
||||
if MARKER not in text or REFLECTION_MARKER in text:
|
||||
return False
|
||||
if not any(isinstance(m, dict) and m.get("role") == "tool" for m in body.get("messages") or []):
|
||||
later_tool_round = bool(body.get("tools")) and any(
|
||||
isinstance(m, dict) and m.get("role") == "tool" for m in body.get("messages") or [])
|
||||
if not (later_tool_round or SALVAGE_MARKER in text):
|
||||
return False
|
||||
with self._lock:
|
||||
self.refused += 1
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue