mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 20:27:56 +00:00
Conflicts: supervisor/log_addressing.py and supervisor/worker_chat_lane.py (the turn event queue carries both the initiator label and the lazy-naming callback), web/modules/chat.js and model_wait.js (the target's chrome-follows-the-work model wins; no lane placeholder survives), tests/test_log_forwarding.py (both blocks), docs/architecture/03 (their chrome paragraph plus the consciousness wording), docs/DOMAIN_MAP.md and the v7next inventories (regenerated). The merged get_tools sits at the 300-line ratchet cap (one entry joined onto one line). Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
commit
505f6f4ae3
65 changed files with 3045 additions and 427 deletions
|
|
@ -219,8 +219,10 @@ That row keeps a neutral owner anchor visible, but hides task status and typing
|
|||
until a real task status or activity arrives; review presence alone never means
|
||||
`Working`, `Done`, or owner attention.
|
||||
|
||||
Local diagnostic failures remain inspectable in details and Logs, but do not
|
||||
relabel the whole still-working task. A failed child keeps a compact factual
|
||||
A reviewer panel that settles after its task already ended adds one System
|
||||
row to the task's room naming the verdict and which revision it covered; it
|
||||
never changes the finished card. Local diagnostic failures remain inspectable
|
||||
in details and Logs, but do not relabel the whole still-working task. A failed child keeps a compact factual
|
||||
`Failed` marker inside its parent while the root continues under its own
|
||||
authoritative status. Internal reason codes belong in details and diagnostics,
|
||||
not compact headlines. Where a card does show a cause, it says it in the owner's
|
||||
|
|
@ -507,20 +509,29 @@ reload and on reconnect (a turn that moved itself into a Project with
|
|||
Project room, and Main replays only the owner message and the Started
|
||||
annotation); only the row content differs by source.
|
||||
|
||||
The block is one component with two chromes, chosen by the host's
|
||||
`_is_direct_chat` fact (census kind, the rebuilt terminal event, history rows),
|
||||
never by the client. A managed or Swarm root keeps the task card: `Task` title,
|
||||
Working/Done chip, `Turn into project` unless its origin is already bound. A
|
||||
direct conversation turn renders a compact activity block: no `Task` title, no
|
||||
Working/Done chip, a small running indicator and Stop while the turn runs, the
|
||||
wait controls while a wait is open, and the tool count, cost and duration in
|
||||
the collapsed header once the turn ends (a replayed header carries the count
|
||||
and cost; duration is a live fact). A direct turn is never offered `Turn
|
||||
into project`; conversion happens only through the model's own scope tools. A
|
||||
block whose only reason to exist was open attention leaves when that attention
|
||||
closes: a wait-only block disappears when the wait resolves, and the resolved
|
||||
episode's no-reopen ledger survives with the record, so a stale revision cannot
|
||||
bring the block back.
|
||||
The block's chrome follows the work it stands on
|
||||
(`web/modules/chat.js::blockHasWork`, the presence facts minus open attention
|
||||
and minus a bare terminal outcome), never the lane that ran the turn (owner
|
||||
decision 16.09: real work is a task card, a greeting is nothing). A block with
|
||||
work — an always-shown kind, a review group, a child card, a tool or narration
|
||||
row, a tool error — is the task card whether a managed root or a direct
|
||||
conversation turn produced it: a title (the coined name, the latest narration
|
||||
headline, or the `Working…`/`Task activity` placeholder), the status chip, Stop
|
||||
while the host attests it, and `Turn into project` in Main unless its origin is
|
||||
already bound (a direct turn's later rows then route to the Project room like a
|
||||
turn that called `ensure_project_scope`). A block that exists only for open
|
||||
attention — a model wait, a pending or host-offered Stop — or only for a
|
||||
non-Done ending of a turn that did no work carries no title placeholder and no
|
||||
conversion; its chip says the state it is in (Waiting…, Cancelling…, Failed),
|
||||
the wait controls stay, and its first row of work gives it the title. The
|
||||
collapsed header carries the tool count, cost and duration once the turn ends
|
||||
(a replayed header carries the count and cost; duration is a live fact). The
|
||||
host's `_is_direct_chat` fact keeps its host jobs (routing, census `kind`, Stop
|
||||
custody, terminal rows) and, on the client, only the header pill (a direct turn
|
||||
keeps the census verdict beside its block). A block whose only reason to exist
|
||||
was open attention leaves when that attention closes: a wait-only block
|
||||
disappears when the wait resolves, and the resolved episode's no-reopen ledger
|
||||
survives with the record, so a stale revision cannot bring the block back.
|
||||
|
||||
An addressing call (`promote_chat_to_task`, `route_to_project`, `steer_task`, `ensure_project_scope` —
|
||||
the routing-verb family `ouroboros/tool_capabilities.py::ROUTING_VERBS` owns) is stamped by
|
||||
|
|
@ -553,7 +564,7 @@ open default behind a closed exception list").
|
|||
|
||||
Quota exhaustion and a confirmed need to sign in again use the same component,
|
||||
`model_wait.js` with `model_wait.css`, inside the turn's existing host: the task
|
||||
card of a managed root, the activity block of a direct turn. Each
|
||||
card of a turn that has done work, the bare block of a turn that has not. Each
|
||||
waiting role has its own row; the model, account and reason are separate facts.
|
||||
The controls stay visible when the task's timeline is collapsed. Waiting carries
|
||||
a quiet warning status and no computation animation, activity counter or invented
|
||||
|
|
|
|||
|
|
@ -8,7 +8,7 @@ The manifest is the SSOT of the module→domain assignment (1:1, complete over t
|
|||
|
||||
| domain | name | modules | proposed |
|
||||
|---|---|---:|---:|
|
||||
| D01 | Agent core & main loop | 32 | 0 |
|
||||
| D01 | Agent core & main loop | 33 | 0 |
|
||||
| D02 | LLM client, routing & providers | 37 | 0 |
|
||||
| D03 | Context assembly, fit & compaction | 11 | 0 |
|
||||
| D04 | Tool execution: registry, access & typed results | 20 | 0 |
|
||||
|
|
@ -28,7 +28,7 @@ The manifest is the SSOT of the module→domain assignment (1:1, complete over t
|
|||
| D18 | Launcher, packaging, platform & shared substrate | 14 | 0 |
|
||||
| D19 | Frozen contracts (ABI) | 10 | 0 |
|
||||
| D20 | Presence | 9 | 0 |
|
||||
| **total** | | **549** | **0** |
|
||||
| **total** | | **550** | **0** |
|
||||
|
||||
## Dependency direction matrix (strict, pinned)
|
||||
|
||||
|
|
@ -181,6 +181,7 @@ No function body (≥ 10 normalized lines) is shared verbatim across domains. Ne
|
|||
|
||||
- `ouroboros/_outcome_receipts.py`
|
||||
- `ouroboros/_outcome_tool_errors.py`
|
||||
- `ouroboros/acceptance_settlement.py`
|
||||
- `ouroboros/agent.py`
|
||||
- `ouroboros/agent_dispatch.py`
|
||||
- `ouroboros/agent_startup_checks.py`
|
||||
|
|
|
|||
|
|
@ -83,7 +83,7 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
|
|||
├── agent_startup_checks.py ← Worker-boot verification: dirty repo, version sync, budget, memory files, health checks
|
||||
├── agent_task_pipeline.py ← Task execution pipeline orchestration; freezes one shared non-final subtree-cost snapshot for summary/reflection before the terminal checkpoint records final spend; hands the summary and reflection prompts the commit/advisory review lens PLUS the task's own acceptance-panel projection, and an absence statement names the lens it describes; calls the swarm-efficiency rollup owned by task_finalization.py at pipeline end
|
||||
├── agent_dispatch.py, post_task_synthesis.py ← The agent's delegated-child dispatch seam, and the post-task synthesis workers the pipeline runs after a terminal result
|
||||
├── task_finalization.py ← Terminal delivery + sealed final ground truth: live final-answer delivery before blocking post-task (final event selected by the finalizing task's id; buffered copy retained under one `delivery_id`), the sealed final package (submitted final text, the durable result's artifact manifest, and task-related completion observations) fed to summary/reflection as a prompt input, never a validator; owns the per-task `swarm_efficiency` rollup (subagent_count / fanout_count / fanout_interval_sec_total / `lanes_requested` — a rollup built from pre-dispatch fanout events cannot truthfully report effective lanes, which are per-child dispatch facts; `planned` stays null, never inferred as 0 from absent events; host-attested Swarm intent is the typed metadata `force_plan_source == "swarm"`, never prompt inspection, and a Swarm task that fanned out nothing records a minimal `no_fanout_observed` block instead of silence); a fanned-out root additionally carries a `depth` block (`requested_depth`/`permitted_depth`/`attempted_depth`/`achieved_depth` plus a typed status, `host_visible_only`) built from the root contract and its subtree's depth provenance, so a root that never carried its own request recovers it from the children it scheduled, and the terminal task_summary row reports it only when a request exists
|
||||
├── task_finalization.py ← Terminal delivery + sealed final ground truth: live final-answer delivery before blocking post-task (final event selected by the finalizing task's id; buffered copy retained under one `delivery_id`), the sealed final package (submitted final text, the durable result's artifact manifest, and task-related completion observations) fed to summary/reflection as a prompt input, never a validator; owns the per-task `swarm_efficiency` rollup (subagent_count / fanout_count / fanout_interval_sec_total / `lanes_requested` — a rollup built from pre-dispatch fanout events cannot truthfully report effective lanes, which are per-child dispatch facts; `planned` stays null, never inferred as 0 from absent events; host-attested Swarm intent is the typed metadata `force_plan_source == "swarm"`, never prompt inspection, and a Swarm task that fanned out nothing records a minimal `no_fanout_observed` block instead of silence); a fanned-out root additionally carries a `depth` block (`requested_depth`/`permitted_depth`/`attempted_depth`/`achieved_depth` plus a typed status, `host_visible_only`) built from the root contract and its subtree's depth provenance, so a root that never carried its own request recovers it from the children it scheduled, and the task result carries it (the terminal task_summary row states no depth sentence)
|
||||
├── mutation_attribution.py ← Root-task baseline capture in the existing task result; clean-at-baseline Git candidate projection; terminal projection includes the committed interval delta
|
||||
├── process_interpreters.py ← Interpreter resolvers for the user process launch surfaces: one-time pre-guard unversioned-Python resolver + post-gates Node ladder (PATH-first health probe, bundled fallback, attested child-env PATH prepend; the probe EXECUTES a candidate, so it runs only after the dispatch gates approve the call)
|
||||
├── post_task_checkpoint.py ← Durable root post-task phase/final-cost checkpoint shared by task finalization and Project naming recovery
|
||||
|
|
@ -103,6 +103,7 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
|
|||
├── evolution_fingerprint.py ← Canonical fingerprint for evolution-campaign objectives; SSOT for repeat gating
|
||||
├── improvement_backlog.py ← Durable advisory improvement backlog: recurrence-counted dedup (bump count/last_seen, never drop), priority+recurrence+recency ranking, close-on-commit `close_backlog_items`, size-triggered `groom_backlog`; parser-safe locked writer; entries carry priority/kind
|
||||
├── loop.py ← High-level LLM tool loop and its one-shot finalization nudges, ordered nanny → red-verification → masked-verification → no-op-attempt; continuous `FINAL ANSWER:` latching captures the latest typed candidate every round (tool-count-stamped, no prose mining) so review/nudge/forced-finalization paths never erase a structured answer; all marker prompting is gated on `task_contract.answer_protocol="final_answer_line"` via the `answer_protocol_active` SSOT (the gate is sufficient — an empty `expected_output` cannot suppress it) while the latch/extractor stay unconditional; `outcomes.extract_final_answer` (re-imported here) structurally rejects the outcome-tier ledger identifiers as answers — internal enum vocabulary is never a deliverable, and `solved` stays extractable as an ordinary English word. The nanny nudge (`loop_nudges._nanny_finalization_message`, re-exported here) fires once for a harness child finalizing with ZERO durable start attempts (blocked and uncustodied attempts count, so an exact-route startup fault is not accused of skipping delegation); it reads custody evidence from the canonical (budget) root via `delegate_custody.custody_root` and branches PENDING ≠ FAILED — a started-but-unsettled run gets a wait reminder, never a failure accusation, which would invite a duplicate concurrent run; an actually-injected nudge is stamped by the WORKER as a durable custody row, and a COMPLETED harness child with zero runs carries the typed `nanny_finalized_after_nudge_without_delegation` disclosure (stamped by `subagents._disclose_native_only_substrate` at the completion seam; visibility, never a gate); a configured session child finalizing without a succeeded leaf run carries the typed `CONFIGURED_ACTOR_INCOMPLETE`/`CONFIGURED_ACTOR_UNKNOWN` fact (`subagent_bootstrap.actor_first_unresolved_fact`; host children ride along as auxiliary `direct_child_statuses`, never a substitute for the leaf), and a successful run's later silence stays proportional to the measured burn (the `NANNY_METERED_OVERRUN` reminder in `loop_nudges`, metered by `nanny_pacing.nanny_metered_since_delegate_activity`/`nanny_burn_phrase`)
|
||||
├── acceptance_settlement.py ← What happens to a paid acceptance panel that outlives the answer it reviewed: the quorum/completion mailbox wake carrying each reviewer's own verdict, the delivery-under-a-running-panel terminal (wait by default, conscious finish, `previous_revision_accepted` on a PASS over the earlier revision), and the post-terminal supplement that republishes the collected panel and announces it once in the task's room
|
||||
├── loop_acceptance.py, loop_acceptance_review.py ← Acceptance machinery (moved whole out of `loop.py`, which keeps the checkpoint, run-record and message rails; every name with an external caller is re-exported from `loop` so external callers, patch sites and the acceptance-writer inventory keep one import surface — a moved name nobody outside references is not re-exported). The fence and its obligations: eligibility, begin/end/supersede, subtree snapshots, the final-answer latch, the closed typed reason set `ACCEPTANCE_DECISION_REASONS`, the sole decision merge point `_set_acceptance_decision`, and obligation collection/reopen/disposition. The run: the host evidence packet (`_build_host_acceptance_evidence` over `review_evidence.build_task_acceptance_evidence`, plus the forced-rail child debt and the unhashed dialogue history), the one substantive panel (`_execute_task_acceptance_panel`: reviewer rows, wave-budget admission, the free zero-physical refusal, the exact-hash wallet stamp, the timing event), the bounded `acceptance_dialogue_history`, and the paid identity + free-replay `_refuse_identical_acceptance`; the `dialogue_status` reducer is `review_verdict.aggregate_dialogue_status` (the vote SSOT, re-exported through `review_substrate`; these are its consumers). `loop_acceptance_review.acceptance_retrieving_work_order` renders the delivery-conditional work order of the retrieving rows (session: FULL packet + absolute pointers + access disclosure; native: packet without its freely degradable tail + real data root) onto `ReviewRequest.slot_session_tasks` — the FULL packet stays the `evidence_refs` authority
|
||||
├── loop_llm_call.py ← Single-round LLM call + usage accounting
|
||||
├── transcript_prefix.py ← Append-only transcript invariant between the sends of one loop execution: the one `sent_in_previous_send` predicate every producer that appends behind a sent row consults (owner follow-ups, task messages, quiz answers, roster notes and the acceptance observation alike — the #929 carve-out generalized); per-message content digests (role, plain text, tool-call identity — cache markers, block shape and private custody keys are not content) and the `prompt_prefix_break` checkpoint fact (`kind` = system_rewritten | tail_replaced | rewritten | shrunk, plus `sanctioned_by`, stamped by the compaction seams through `sanction_rewrite` and read on the transcript each round actually dispatched; a context-fit reprojection after a real overflow is recorded as an ordinary break, a real cache cost); it RECORDS, never blocks — OpenAI-family caches reuse a previous request only when it is a byte-prefix of the next, so a transient tail or an in-place rewrite discards the whole conversation cache (measured 2026-09-14, #906)
|
||||
|
|
@ -113,7 +114,7 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
|
|||
├── vision_routing.py ← Send-time image routing SSOT: inline vision vs generic captions vs placeholders on a per-send message copy (`OUROBOROS_IMAGE_INPUT_MODE`, `OUROBOROS_MODEL_VISION`)
|
||||
├── fallback_cooldown.py ← Per-process 429-aware cooldown for the `OUROBOROS_MODEL_FALLBACKS` chain: a transiently-failed model is parked for a short window so fallback walks and repeated rounds skip it; advisory, default-on, fail-soft, passive heal; per-process only — honestly not a swarm-wide governor
|
||||
├── model_concurrency.py ← Per-(model,use_local) `BoundedSemaphore` capping concurrent provider calls (`OUROBOROS_MODEL_MAX_CONCURRENCY`, default 3) so a task's own loop + subagent threads + status pings cannot self-DoS one model's rate limit; excess threads WAIT deadline-bounded; wraps only the provider call in `loop_llm_call.call_llm_with_retry`; per-process only
|
||||
├── project_naming.py ← SSOT for LLM-first project naming: bounded light-model title with deterministic fallback, shared by admission naming (`admission_names`: headless runs and chat promotion, no model call), turn-into-project conversion (gateway/projects.py), and `ensure_project_scope` — a direct conversation turn is never named; the provider call goes through the model_concurrency slot
|
||||
├── project_naming.py ← SSOT for LLM-first project naming: bounded light-model title with deterministic fallback, shared by admission naming (`admission_names`: headless runs and chat promotion, no model call), turn-into-project conversion (gateway/projects.py), `ensure_project_scope`, and the lazy turn namer (`spawn_turn_namer`: a direct Main turn is named on its first non-addressing tool call, one bounded Light call, never for a greeting); the provider call goes through the model_concurrency slot
|
||||
├── loop_tool_execution.py ← Tool dispatch and tool-result handling
|
||||
├── deadline_utils.py ← Shared deadline parsing/remaining-time helpers + the transport-vs-logical wait seam for loop milestones and process-tool/review timeouts
|
||||
├── observability.py ← Private forensic execution ledger: redaction, gzip CAS blobs, call manifests, trace refs
|
||||
|
|
|
|||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -84,7 +84,7 @@ A registry of `config.SETTINGS_DEFAULTS` (exact defaults stay canonical in `conf
|
|||
| OUROBOROS_MODEL_MAX_CONCURRENCY | 3 | Per-(model,route) concurrent provider-call cap (`model_concurrency.py`) |
|
||||
| OUROBOROS_MODEL_SLOT_MAX_WAIT_SEC | 180 | Concurrency-slot wait bound |
|
||||
| OUROBOROS_PROJECT_NAMING_TIMEOUT_SEC | 60 | Project-naming call ceiling |
|
||||
| OUROBOROS_PROJECT_NAMING_ASYNC_TIMEOUT_SEC | 8 | Bound of the inline naming call when a card is turned into a project (`gateway/projects.py`); direct turns are not named in the background |
|
||||
| OUROBOROS_PROJECT_NAMING_ASYNC_TIMEOUT_SEC | 8 | Bound of the inline naming call when a card is turned into a project (`gateway/projects.py`); a direct Main turn is named in the background once it starts working (`spawn_turn_namer`, bounded by `OUROBOROS_PROJECT_NAMING_TIMEOUT_SEC` + 30 s) |
|
||||
| OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC | 120 | Update-letter LIGHT one-shot ceiling, slot wait and provider call together (`update_letter.py`) |
|
||||
| OUROBOROS_FALLBACK_COOLDOWN_ENABLED | true | 429-aware per-process model cooldown |
|
||||
| OUROBOROS_FALLBACK_COOLDOWN_SEC | 120 | Cooldown window |
|
||||
|
|
|
|||
|
|
@ -1199,15 +1199,23 @@ by "Provider Independence" above. Call-site imperatives:
|
|||
- Nested process wrappers are ordered, never tied: the provider bound settles
|
||||
before its killable child, the child before the generic ToolEntry envelope
|
||||
(fixed structural settlement margin from `config.py`), so a child or
|
||||
provider result cannot arrive after its owner has abandoned custody. The
|
||||
one deliberate early return is plan review's dispatch barrier
|
||||
(`ReviewRequest.drain_deadline`): the wrapper returns while its workers run,
|
||||
but custody is not abandoned — the workers settle into process-local custody
|
||||
and announce the wave through the task mailbox (`plan_review_collect`).
|
||||
Before its effective blocking verdict, `owner_hurry.force_plan_decision`
|
||||
collects once at zero wait and projects the returned state. Context health
|
||||
only reads the canonical wave; its pending count/time describe the recorded
|
||||
snapshot, not live worker progress. Neither path dispatches a second panel.
|
||||
provider result cannot arrive after its owner has abandoned custody. The two
|
||||
deliberate early returns are plan review's and task acceptance's shared
|
||||
dispatch barrier (`ReviewRequest.drain_deadline`): the wrapper returns while
|
||||
its workers run, but custody is not abandoned — the workers settle into
|
||||
process-local custody and announce the wave through the task mailbox
|
||||
(`plan_review_collect`; `acceptance_settlement.announce_acceptance_settlement`,
|
||||
which for task acceptance announces at the wave's own quorum as well as at
|
||||
completion and carries each reviewer's own verdict, not an instruction to
|
||||
collect). Before its
|
||||
effective blocking verdict, `owner_hurry.force_plan_decision` collects once at
|
||||
zero wait and projects the returned state. Task acceptance collects the same
|
||||
way, but through the host's own reconcile at the top of the acceptance seam
|
||||
(`review_dispatch.reconcile_pending_acceptance_runs`) rather than a
|
||||
model-callable verb: it replays the recorded request and roster and sends
|
||||
nothing. Context health only reads the canonical wave; its pending count/time
|
||||
describe the recorded snapshot, not live worker progress. Neither path
|
||||
dispatches a second panel.
|
||||
- Every physical LLM/review/VLM/tool operation that can outlive a logical
|
||||
wait emits typed `cognitive_operation` start and terminal facts; the
|
||||
supervisor uses the active-operation map only to spare the idle rail, and a
|
||||
|
|
@ -1302,7 +1310,26 @@ by "Provider Independence" above. Call-site imperatives:
|
|||
(`outcomes.turn_has_reviewable_effects` plus a typed
|
||||
deliverable/criterion), never keywords (BIBLE P3/P5). The agent-callable nomination is never authoritative (ARCHITECTURE "Task
|
||||
lifecycle"). Freeze its request and resolved roster; use existing review
|
||||
custody and mailbox continuation for pending work and free collection. The
|
||||
custody and mailbox continuation for pending work and free collection. Before
|
||||
it assembles evidence for a NEW panel and before any capacity refusal
|
||||
(`review_cycles_exhausted`), the host reconciles every already-paid panel of
|
||||
the same root still recorded as running, at $0 over the recorded request and
|
||||
roster — a panel whose subject was re-authored mid-flight would otherwise have
|
||||
its bought verdicts discarded. Reviewer verdicts are advice for the author and
|
||||
reach Main whatever its current draft: the settlement wake carries them, and
|
||||
the keep contract is re-offered on every acceptance wake whose contract bytes
|
||||
changed (identical bytes are not repeated, and a spent malformed-control repair
|
||||
stays spent) rather than once per candidate chain. A delivery under a running
|
||||
panel is never a second panel and never a capacity refusal
|
||||
(`acceptance_settlement._deliver_under_running_panel`): waiting is the default
|
||||
and the only option under blocking enforcement (Cyber Pro never waits, as
|
||||
before); under advisory enforcement Main finishes only through an explicit
|
||||
`"pending_review":"finish"` on its delivery control; the trace a pending panel
|
||||
belongs to stays reachable past the loop exit (`remember_settlement_trace`) so
|
||||
the post-terminal supplement can collect it; a PASS that settled on the earlier revision accepts the task
|
||||
as `previous_revision_accepted`; any other settled verdict lets the ordinary
|
||||
path decide. A panel that settles after the task ended is attached to the task
|
||||
result and announced once in the task's room; nothing starts a model turn. The
|
||||
worker never writes Main's live candidate or author decision. Keep
|
||||
subtree/status facts
|
||||
separate from reviewer findings and Cyber's authority under BIBLE P0.
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
Machine extraction of the `docs/ARCHITECTURE.md` "Data layout (`~/Ouroboros/`)" tree — the durable-file orientation carrier (this tree's counterpart of the reference PERSISTENCE_OWNERS derivation checklist) — regenerated by `python scripts/regenerate_inventories.py`. Do not edit. Every entry is probed against reality: repo entries must exist as tracked paths; data-plane entries must appear as a literal in the runtime sources that construct them. A durable file renamed or removed in code while its tree row survives = red (`tests/test_generated_inventories.py`).
|
||||
|
||||
Source: `docs/architecture/01-high-level-architecture.md`, physical LF lines 567-656; UTF-8 SHA-256 `888f5d5115d1f159ecd20688c34226109af5ff87ac91ccbce45b372e6e8ddef6`.
|
||||
Source: `docs/architecture/01-high-level-architecture.md`, physical LF lines 568-657; UTF-8 SHA-256 `505e9fb5ecab5e45d47a74155024030d06d24e41843c93a2b209062990b90c10`.
|
||||
|
||||
- entries: **78** (code-ref: 71, repo-dir: 6, repo-path: 1)
|
||||
|
||||
|
|
|
|||
|
|
@ -103,14 +103,27 @@ def select_current_review_runs(
|
|||
and "FAIL" not in superseded_signals
|
||||
and "DEGRADED" not in superseded_signals
|
||||
)
|
||||
selected = [] if acceptance_gap else (current_runs if current_runs else all_runs)
|
||||
# A host decision that names a panel by id AND binding applied THAT panel's
|
||||
# verdict (an earlier revision accepted on the reviewers' word): when no run
|
||||
# is current, the named run speaks alone and an older superseded FAIL stays
|
||||
# audit evidence instead of outvoting the applied PASS. With no named panel
|
||||
# the conservative rule stands: every stale run is retained.
|
||||
named = [
|
||||
run for run in all_runs
|
||||
if decision.get("panel_id") and decision.get("binding_hash")
|
||||
and str(run.get("panel_id") or "") == str(decision.get("panel_id"))
|
||||
and str(run.get("binding_hash") or "") == str(decision.get("binding_hash"))
|
||||
]
|
||||
selected = [] if acceptance_gap else (current_runs if current_runs else (named or all_runs))
|
||||
return ReviewRunSelection(
|
||||
all_runs=all_runs,
|
||||
current_runs=selected,
|
||||
superseded_only_acceptance_gap=acceptance_gap,
|
||||
superseded_aggregate_signals=superseded_signals,
|
||||
current_candidate_unaccepted=candidate_unaccepted,
|
||||
has_replacement=bool(current_runs),
|
||||
# A decision-named panel replaces the stale runs for the ledger as well:
|
||||
# an older FAIL reads as superseded, not as live failure evidence.
|
||||
has_replacement=bool(current_runs) or bool(named),
|
||||
)
|
||||
|
||||
|
||||
|
|
|
|||
328
ouroboros/acceptance_settlement.py
Normal file
328
ouroboros/acceptance_settlement.py
Normal file
|
|
@ -0,0 +1,328 @@
|
|||
"""What happens to a paid acceptance panel that outlives the answer it reviewed.
|
||||
|
||||
A reviewer panel is advice for its author, never a signature on bytes the
|
||||
reviewers did not read (owner decision D4=A, 2026-09-16). The three moments
|
||||
below share one fact — the panel keeps its custody and paid identity while
|
||||
Main moves on — so they live together:
|
||||
|
||||
* the wave wakes Main at its own quorum and again when the last slot settles,
|
||||
and the wake carries each reviewer's own verdict
|
||||
(``announce_acceptance_settlement``);
|
||||
* a final answer delivered while that panel still runs neither buys a second
|
||||
panel nor is refused one (``_deliver_under_running_panel``): Main waits — the
|
||||
default, and the only option under blocking enforcement — or consciously
|
||||
finishes; a panel that PASSED the earlier revision accepts the task and the
|
||||
owner row says so (fork 1=B), while any other settled verdict hands the
|
||||
delivery to the ordinary acceptance path with the collected verdicts in its
|
||||
dialogue history;
|
||||
* a panel that settles after its task ended is collected at $0, republished on
|
||||
the task's own review projection and announced once in the task's room
|
||||
(``attach_late_acceptance_settlement``); no model turn starts (fork 2=A).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import pathlib
|
||||
from types import SimpleNamespace
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# The host acceptance reason for "the paid panel approved the answer Main has
|
||||
# since rewritten": the task is accepted on the reviewers' word about the
|
||||
# earlier revision and the owner row says so. Not a member of
|
||||
# ``outcomes._ACCEPTANCE_BLOCKED_TERMINAL_REASONS``: nothing was refused and no
|
||||
# review cycle was spent.
|
||||
REASON_PREVIOUS_REVISION_ACCEPTED = "previous_revision_accepted"
|
||||
|
||||
# The chat row a panel writes when it settles after its task already ended.
|
||||
LATE_SETTLEMENT_SYSTEM_TYPE = "acceptance_late_settlement"
|
||||
|
||||
# The mailbox row a settling acceptance wave writes. It carries the reviewers'
|
||||
# OWN verdicts: a wake that only said results existed, and asked Main to reply
|
||||
# ``keep`` to read them, made the verdicts reachable only by not having moved
|
||||
# on. The reducer's quorum verdict still lands through the ordinary collection
|
||||
# at the next acceptance entry; these are the individual reviewers, unreduced.
|
||||
ACCEPTANCE_SETTLEMENT_WAKE = (
|
||||
"Acceptance review {retry_key}: {settled} of {total} reviewer slot(s) have answered on "
|
||||
"the answer that was under review. These are their own verdicts — advice for you, not "
|
||||
"a signature on your current draft, and they do not stop you from finishing. A rewritten "
|
||||
"answer gets a verdict of its own only when you nominate it again."
|
||||
)
|
||||
|
||||
|
||||
def _reviewer_lines(wave: Dict[str, Any]) -> List[str]:
|
||||
"""One line per roster slot: the reviewer's own verdict and note, or pending."""
|
||||
slots = wave.get("slots") or {}
|
||||
verdicts = wave.get("verdicts") or {}
|
||||
lines: List[str] = []
|
||||
for slot_id, status in slots.items():
|
||||
row = verdicts.get(slot_id) if isinstance(verdicts.get(slot_id), dict) else {}
|
||||
verdict = str(row.get("verdict") or "") or str(status or "pending")
|
||||
note = str(row.get("note") or "")
|
||||
lines.append(f"- {slot_id}: {verdict}" + (f" — {note}" if note else ""))
|
||||
return lines
|
||||
|
||||
|
||||
def acceptance_settlement_message(request: Any, wave: Dict[str, Any]) -> str:
|
||||
"""Render the settled wave as the reviewers' own lines, bounded per slot."""
|
||||
slots = wave.get("slots") or {}
|
||||
# Slots that answered before the release were collected by the drain: they
|
||||
# count as answered, and their verdicts sit in the collected panel itself.
|
||||
early = len(wave.get("answered_before_release_ids") or ())
|
||||
head = ACCEPTANCE_SETTLEMENT_WAKE.format(
|
||||
retry_key=str(getattr(request, "retry_key", "") or ""),
|
||||
settled=early + sum(1 for status in slots.values() if status),
|
||||
total=max(int(wave.get("total") or 0), len(slots) + early),
|
||||
)
|
||||
lines = _reviewer_lines(wave)
|
||||
if early:
|
||||
lines.append(f"- {early} slot(s) answered before the release; their verdicts are in the collected panel")
|
||||
return "\n".join([head, *lines])
|
||||
|
||||
|
||||
def _result_root(usage_ctx: Any) -> pathlib.Path:
|
||||
"""The root the task result lives under, resolved the way the acceptance
|
||||
projection writer resolves it (a forked drive keeps its budget root)."""
|
||||
meta = getattr(usage_ctx, "task_metadata", {})
|
||||
meta = meta if isinstance(meta, dict) else {}
|
||||
return pathlib.Path(meta.get("budget_drive_root") or getattr(usage_ctx, "budget_drive_root", None)
|
||||
or usage_ctx.drive_root)
|
||||
|
||||
|
||||
# The loop exit restores the context's ``_execution_trace`` to whatever it was
|
||||
# before the loop ran (``loop_budget._cleanup_loop_resources``), so a wave that
|
||||
# settles after the turn ended would find no trace to attach to. A panel that
|
||||
# went pending registers the trace it belongs to here, keyed by the wave's
|
||||
# retry key; the supplement reads it back and drops it once the wave settled.
|
||||
_SETTLEMENT_TRACE_CAP = 8
|
||||
|
||||
|
||||
def remember_settlement_trace(tools_ctx: Any, llm_trace: Dict[str, Any], run: Dict[str, Any]) -> None:
|
||||
"""Keep the trace a pending panel belongs to reachable past the loop exit."""
|
||||
request = run.get("request") if isinstance(run, dict) else None
|
||||
retry_key = str((request or {}).get("retry_key") or "") if isinstance(request, dict) else ""
|
||||
if not retry_key:
|
||||
return
|
||||
traces = getattr(tools_ctx, "_acceptance_settlement_traces", None)
|
||||
if not isinstance(traces, dict):
|
||||
traces = {}
|
||||
tools_ctx._acceptance_settlement_traces = traces
|
||||
traces.pop(retry_key, None)
|
||||
traces[retry_key] = llm_trace
|
||||
while len(traces) > _SETTLEMENT_TRACE_CAP:
|
||||
traces.pop(next(iter(traces)))
|
||||
|
||||
|
||||
def _settlement_trace(usage_ctx: Any, retry_key: str) -> Optional[Dict[str, Any]]:
|
||||
"""The trace holding this wave's run: the live one while the loop runs, the
|
||||
remembered one after it exited."""
|
||||
def holds(trace: Any) -> bool:
|
||||
return isinstance(trace, dict) and any(
|
||||
isinstance(run, dict) and run.get("authority") == "host_root"
|
||||
and isinstance(run.get("request"), dict)
|
||||
and str(run["request"].get("retry_key") or "") == retry_key
|
||||
for run in (trace.get("review_runs") or []))
|
||||
|
||||
live = getattr(usage_ctx, "_execution_trace", None)
|
||||
if holds(live):
|
||||
return live
|
||||
remembered = (getattr(usage_ctx, "_acceptance_settlement_traces", None) or {}).get(retry_key)
|
||||
return remembered if holds(remembered) else None
|
||||
|
||||
|
||||
def announce_acceptance_settlement(usage_ctx: Any, request: Any, wave: Dict[str, Any]) -> None:
|
||||
"""Deliver a settled acceptance wave to whoever can still act on it.
|
||||
|
||||
While the turn is alive the reviewers' verdicts go to Main's existing
|
||||
mailbox, which is drained before every round and also ends an acceptance
|
||||
park. Once the task is terminal there is nobody to wake: the same verdicts
|
||||
are attached to the task result and announced once in the task's own room,
|
||||
so the owner sees them and the next turn reads them in chat history. Never a
|
||||
model turn, never a second paid dispatch. Runs on the settlement thread,
|
||||
outside custody locks.
|
||||
"""
|
||||
if usage_ctx is None or not getattr(usage_ctx, "drive_root", None):
|
||||
return
|
||||
task_id = str(getattr(request, "task_id", "") or "")
|
||||
try:
|
||||
from ouroboros.task_results import _TRULY_TERMINAL_STATUSES, load_task_result
|
||||
|
||||
row = load_task_result(_result_root(usage_ctx), task_id) or {}
|
||||
if str(row.get("status") or "") in _TRULY_TERMINAL_STATUSES:
|
||||
attach_late_acceptance_settlement(usage_ctx, request, wave, result=row)
|
||||
return
|
||||
from ouroboros.owner_mailbox import write_task_message
|
||||
|
||||
write_task_message(pathlib.Path(usage_ctx.drive_root), acceptance_settlement_message(request, wave),
|
||||
task_id, source_task_id=task_id, provenance="system")
|
||||
except Exception:
|
||||
log.warning("Acceptance settlement delivery failed for %s", task_id, exc_info=True)
|
||||
|
||||
|
||||
def panel_awaiting_this_turn(tools_ctx: Any, llm_trace: Dict[str, Any]) -> Optional[Dict[str, Any]]:
|
||||
"""The panel THIS turn released, found by the binding the host recorded when
|
||||
it went pending — the one identity a re-authored answer cannot move.
|
||||
|
||||
A panel bound under older owner input is not that panel: once Main has
|
||||
acknowledged a newer owner source (the owner's words changed the premises,
|
||||
owner rule 4=A), the latch clears and the ordinary path decides, while the
|
||||
old panel keeps its custody and its verdicts still arrive as advice.
|
||||
"""
|
||||
from ouroboros.loop_messages import owner_source_sha256
|
||||
|
||||
binding = str(getattr(tools_ctx, "_task_acceptance_pending", "") or "")
|
||||
if not binding:
|
||||
return None
|
||||
run = next((run for run in reversed(llm_trace.get("review_runs") or [])
|
||||
if isinstance(run, dict) and run.get("authority") == "host_root"
|
||||
and str(run.get("binding_hash") or "") == binding), None)
|
||||
if run is None:
|
||||
return None
|
||||
reviewed_source = str(run.get("owner_source_sha256") or "")
|
||||
if reviewed_source and reviewed_source != str(owner_source_sha256(tools_ctx) or ""):
|
||||
tools_ctx._task_acceptance_pending = ""
|
||||
return None
|
||||
return run
|
||||
|
||||
|
||||
def acceptance_choice_offered() -> bool:
|
||||
"""Whether the host can honour a wait/finish choice on this install.
|
||||
|
||||
Blocking enforcement always waits; Cyber Pro never does (Main's final
|
||||
response is its decision there, unchanged). Only advisory enforcement on an
|
||||
ordinary runtime mode leaves the choice to Main, so only there is it offered.
|
||||
"""
|
||||
from ouroboros.config import get_review_enforcement, get_runtime_mode
|
||||
from ouroboros.runtime_mode_policy import runtime_mode_at_least
|
||||
|
||||
return (get_review_enforcement() != "blocking"
|
||||
and not runtime_mode_at_least(get_runtime_mode(), "cyber_pro"))
|
||||
|
||||
|
||||
def acceptance_wait_chosen(tools_ctx: Any) -> bool:
|
||||
"""Waiting is the default; only an explicit, current ``finish`` releases an
|
||||
answer over a panel that is still running where the install lets Main choose."""
|
||||
from ouroboros.config import get_review_enforcement, get_runtime_mode
|
||||
from ouroboros.runtime_mode_policy import runtime_mode_at_least
|
||||
|
||||
if runtime_mode_at_least(get_runtime_mode(), "cyber_pro"):
|
||||
return False
|
||||
if get_review_enforcement() == "blocking":
|
||||
return True
|
||||
return str(getattr(tools_ctx, "_acceptance_pending_review_choice", "") or "") != "finish"
|
||||
|
||||
|
||||
def _deliver_under_running_panel(ctx: Any, prior_run: Any) -> Optional[bool]:
|
||||
"""A DELIVERY is not a nomination: it neither buys a panel nor is refused one.
|
||||
|
||||
The panel this turn already paid for keeps its custody and identity. While
|
||||
it runs, Main waits (the default, and the only option under blocking
|
||||
enforcement) or consciously finishes. Once it has settled on the earlier
|
||||
revision: a PASS accepts the task on the reviewers' word and the owner row
|
||||
says so (fork 1=B); any other verdict is not a verdict on this answer, so the
|
||||
ordinary path decides — a new panel while the review cap allows, otherwise
|
||||
its typed capacity refusal — with the collected verdicts already in its
|
||||
dialogue history. ``None`` means "not this case". While the panel is still
|
||||
running, a rewritten answer buys a NEW panel only when Main nominates it
|
||||
again (``_acceptance_review_only``).
|
||||
"""
|
||||
from ouroboros import loop
|
||||
from ouroboros.loop_acceptance_review import (
|
||||
_end_acceptance_terminal, _finish_cyber_acceptance,
|
||||
_set_applied_host_acceptance_impact, acceptance_run_pending,
|
||||
)
|
||||
from ouroboros.outcomes import ACCEPTANCE_ACCEPTED
|
||||
|
||||
tools_ctx = ctx.tools._ctx
|
||||
if prior_run is not None or getattr(tools_ctx, "_acceptance_review_only", False):
|
||||
return None
|
||||
run = panel_awaiting_this_turn(tools_ctx, ctx.llm_trace)
|
||||
if run is None:
|
||||
return None
|
||||
if acceptance_run_pending(run):
|
||||
if acceptance_wait_chosen(tools_ctx):
|
||||
ctx.emit_progress("Task acceptance review is still running; holding the answer for its verdict.")
|
||||
return True
|
||||
return _finish_cyber_acceptance(ctx, SimpleNamespace(**run))
|
||||
tools_ctx._task_acceptance_pending = ""
|
||||
if str(run.get("aggregate_signal") or "").upper() != "PASS":
|
||||
return None
|
||||
result = SimpleNamespace(**run)
|
||||
_set_applied_host_acceptance_impact(run, result, requires_revision=False)
|
||||
ctx.llm_trace.setdefault("review_decision", {}).update({
|
||||
"panel_id": str(run.get("panel_id") or ""),
|
||||
"binding_hash": str(run.get("binding_hash") or ""),
|
||||
})
|
||||
_end_acceptance_terminal(ctx, "pass")
|
||||
loop._set_acceptance_decision(ctx.llm_trace, {
|
||||
"status": ACCEPTANCE_ACCEPTED,
|
||||
"reason": REASON_PREVIOUS_REVISION_ACCEPTED,
|
||||
"source": "task_acceptance_review",
|
||||
"reviewer_signal": "PASS",
|
||||
"reviewed_panel_id": str(run.get("panel_id") or ""),
|
||||
"reviewed_candidate_hash": str(run.get("candidate_hash") or ""),
|
||||
})
|
||||
ctx.emit_progress(
|
||||
"Task acceptance review: PASS on the earlier revision of this answer; accepted on the "
|
||||
"reviewers' word (the rewrite itself was not re-reviewed)."
|
||||
)
|
||||
return False
|
||||
|
||||
|
||||
def _late_settlement_text(run: Dict[str, Any], wave: Dict[str, Any]) -> str:
|
||||
"""The owner row: which verdict, which revision, and the reviewers' own lines."""
|
||||
signal = str(run.get("aggregate_signal") or "").upper()
|
||||
head = {"PASS": "Reviewers later passed this answer.",
|
||||
"FAIL": "Reviewers later rejected this answer."}.get(
|
||||
signal, "Reviewers later returned no settled verdict on this answer.")
|
||||
which = (" They reviewed the earlier version, which was rewritten before delivery."
|
||||
if run.get("superseded_by_revision") else " They reviewed the answer that was delivered.")
|
||||
return "\n".join([head + which, *_reviewer_lines(wave)])
|
||||
|
||||
|
||||
def attach_late_acceptance_settlement(usage_ctx: Any, request: Any, wave: Dict[str, Any],
|
||||
*, result: Dict[str, Any]) -> bool:
|
||||
"""A panel that settled after its task ended is a supplement, never a new turn.
|
||||
|
||||
Its verdicts are collected at $0 over the recorded operation, republished on
|
||||
the task's own review projection through the existing locked writer, and
|
||||
announced ONCE in the task's room through the existing terminal-delivery
|
||||
outbox (durably owed, keyed by ``delivery_id``; a second settlement of the
|
||||
same wave finds nothing left to reconcile and announces nothing). The owner
|
||||
sees it; the next turn reads it in chat history. The acceptance twin of plan
|
||||
review's historical supplement (docs/architecture/06-agent-core.md).
|
||||
"""
|
||||
retry_key = str(getattr(request, "retry_key", "") or "")
|
||||
task_id = str(getattr(request, "task_id", "") or "")
|
||||
trace = _settlement_trace(usage_ctx, retry_key) if retry_key else None
|
||||
if trace is None:
|
||||
log.debug("late acceptance settlement %s: no trace holds this wave (worker rebound?)", retry_key)
|
||||
return False # the worker was rebound to another task; the record stays as published
|
||||
runs = [run for run in (trace.get("review_runs") or [])
|
||||
if isinstance(run, dict) and run.get("authority") == "host_root"
|
||||
and isinstance(run.get("request"), dict)
|
||||
and str(run["request"].get("retry_key") or "") == retry_key]
|
||||
from ouroboros.loop_acceptance_review import acceptance_run_pending
|
||||
from ouroboros.review_dispatch import reconcile_pending_acceptance_runs
|
||||
from ouroboros.review_projection import publish_acceptance_checkpoint
|
||||
from supervisor.terminal_delivery import enqueue_terminal_delivery
|
||||
|
||||
root = pathlib.Path(usage_ctx.drive_root)
|
||||
# Only THIS wave's runs: reconciling every pending panel here would let one
|
||||
# settlement collect a sibling panel's verdicts and leave that panel's own
|
||||
# settlement with nothing to announce (the run objects are shared with the trace).
|
||||
advanced = reconcile_pending_acceptance_runs({"review_runs": runs}, drive_root=root, usage_ctx=usage_ctx)
|
||||
if not any(acceptance_run_pending(run) for run in runs):
|
||||
(getattr(usage_ctx, "_acceptance_settlement_traces", None) or {}).pop(retry_key, None)
|
||||
if not advanced:
|
||||
log.debug("late acceptance settlement %s: nothing reconciled (still pending or already collected)", retry_key)
|
||||
return False
|
||||
publish_acceptance_checkpoint(usage_ctx, trace, task_id=task_id, drive_root=_result_root(usage_ctx),
|
||||
chat_id=result.get("chat_id"))
|
||||
return bool(enqueue_terminal_delivery(root, {
|
||||
"type": "send_message", "chat_id": int(result.get("chat_id") or 0), "task_id": task_id,
|
||||
"text": _late_settlement_text(runs[-1], wave),
|
||||
"role": "system", "system_type": LATE_SETTLEMENT_SYSTEM_TYPE,
|
||||
"delivery_id": f"acceptance-late:{retry_key}",
|
||||
}, event_queue=getattr(usage_ctx, "event_queue", None)))
|
||||
|
|
@ -71,6 +71,7 @@ D20 = "Presence"
|
|||
"ouroboros/_usage_response.py" = "D16"
|
||||
"ouroboros/_usage_rows.py" = "D16"
|
||||
"ouroboros/_usage_rows_memo.py" = "D16"
|
||||
"ouroboros/acceptance_settlement.py" = "D01"
|
||||
"ouroboros/agent.py" = "D01"
|
||||
"ouroboros/agent_dispatch.py" = "D01"
|
||||
"ouroboros/agent_startup_checks.py" = "D01"
|
||||
|
|
|
|||
|
|
@ -322,6 +322,10 @@ def _supersede_task_acceptance_for_owner_followup(
|
|||
)
|
||||
ctx._task_acceptance_reviewed = False
|
||||
ctx._task_acceptance_fence_generation_mismatch = False
|
||||
# The panel in flight reviewed the answer to the OLD requirements: it stays
|
||||
# custodied and its verdicts still arrive as advice, but the answer Main writes
|
||||
# for the follow-up is not a delivery under it — the ordinary path decides.
|
||||
ctx._task_acceptance_pending = ""
|
||||
llm_trace.pop("root_phase_checkpoint", None)
|
||||
llm_trace["review_decision"] = {**dict(llm_trace.get("review_decision") or {}),
|
||||
"eligibility": "pending_owner_followup",
|
||||
|
|
|
|||
|
|
@ -91,29 +91,6 @@ def acceptance_run_pending(run: Any) -> bool:
|
|||
in {"pending_dispatch", "in_flight"} for actor in actors or [])
|
||||
|
||||
|
||||
def announce_acceptance_settlement(usage_ctx: Any, request: Any, wave: dict) -> None:
|
||||
"""Wake the original Main through its existing mailbox, outside custody locks.
|
||||
|
||||
The worker never changes live candidate, transcript, or acceptance decisions.
|
||||
Main collects this exact recorded operation before interpreting the feedback.
|
||||
"""
|
||||
if usage_ctx is None or not getattr(usage_ctx, "drive_root", None):
|
||||
return
|
||||
from ouroboros.owner_mailbox import write_task_message
|
||||
|
||||
try:
|
||||
slots = wave.get("slots") or {}
|
||||
write_task_message(
|
||||
pathlib.Path(usage_ctx.drive_root),
|
||||
f"Task acceptance operation {request.retry_key} settled "
|
||||
f"({len(slots)} released reviewer slots). Its recorded results are ready "
|
||||
"for collection; this notification is not a verdict.",
|
||||
request.task_id, source_task_id=request.task_id, provenance="system",
|
||||
)
|
||||
except Exception:
|
||||
log.warning("Acceptance settlement wake failed for %s", request.task_id, exc_info=True)
|
||||
|
||||
|
||||
def _resolve_ctx_lineage(ctx: Any, task_id: str = "") -> Dict[str, Any]:
|
||||
"""One reader of a live tool context's lineage facts, shared by both
|
||||
eligibility sites so the observation seam and the entrypoint agree."""
|
||||
|
|
@ -182,8 +159,9 @@ def wait_for_acceptance_feedback(tools: Any, limit_ctx: Any, trace: dict,
|
|||
binding = getattr(ctx, "_task_acceptance_pending", "")
|
||||
if not binding:
|
||||
return
|
||||
if not getattr(getattr(ctx, "_delivery_candidate", None), "control_episode_seen", False):
|
||||
_loop()._arm_delivery_control(tools, limit_ctx, trace)
|
||||
# Re-offered on EVERY wake: a replacement candidate inherits
|
||||
# ``control_episode_seen``, which hid the one free route back to the verdicts.
|
||||
_loop()._arm_delivery_control(tools, limit_ctx, trace, skip_if_unchanged=True)
|
||||
from ouroboros.owner_wait import wait_after_tools
|
||||
|
||||
wait_after_tools(ctx, limit_ctx.messages, trace, limit_ctx.accumulated_usage,
|
||||
|
|
@ -197,6 +175,7 @@ def advance_explicit_acceptance(tools: Any, limit_ctx: Any, trace: dict,
|
|||
if not isinstance(request, dict):
|
||||
return
|
||||
tools._ctx._acceptance_request_pending = None
|
||||
tools._ctx._acceptance_pending_review_choice = "" # a new nomination starts a fresh wait/finish choice
|
||||
from ouroboros.loop_delivery import apply_delivery_subject_decision
|
||||
|
||||
subject = request.get("acceptance_subject")
|
||||
|
|
@ -692,7 +671,6 @@ def _finish_cyber_acceptance(ctx: _TaskAcceptanceContext, result: Any) -> bool:
|
|||
"status": ACCEPTANCE_ACCEPTED if clean else ACCEPTANCE_FINALIZED_UNACCEPTED,
|
||||
"reason": "clean_pass" if clean else "author_finish", "source": "task_acceptance_review",
|
||||
"author_disposition": author, "review_pending": pending,
|
||||
"rationale": "The author chose delivery; recorded critic outcomes and unfinished review work are unchanged.",
|
||||
})
|
||||
ctx.emit_progress("Task acceptance feedback remains advisory; Main chose delivery."
|
||||
+ (" Review is still running." if pending else ""))
|
||||
|
|
@ -735,7 +713,6 @@ def _finish_advisory_author(ctx: _TaskAcceptanceContext) -> bool:
|
|||
_loop()._set_acceptance_decision(ctx.llm_trace, {
|
||||
"status": ACCEPTANCE_FINALIZED_UNACCEPTED, "reason": "author_finish",
|
||||
"source": "task_acceptance_review", "author_disposition": author,
|
||||
"rationale": "The author finished after independent feedback; the current subject is author-accepted, not reviewer PASS.",
|
||||
"reviewer_signal": author["reviewer_signal"],
|
||||
"reviewer_binding_hash": feedback.get("binding_hash"),
|
||||
})
|
||||
|
|
@ -889,7 +866,7 @@ def _apply_task_acceptance_result(
|
|||
run["feedback_delivered"] = True
|
||||
break
|
||||
# The aggregate word is not an explanation: printing DEGRADED alone read
|
||||
# as "no valid quorum" while a capsule was in fact fed back for one more
|
||||
# as "no settled verdict" while a capsule was in fact fed back for one more
|
||||
# bounded pass. Name the pass being started and the recorded causes; a
|
||||
# wave that recorded none says THAT, so the verdict is a label beside a
|
||||
# stated absence rather than standing in for the reason.
|
||||
|
|
@ -917,13 +894,12 @@ def _apply_task_acceptance_result(
|
|||
"status": ACCEPTANCE_FINALIZED_UNACCEPTED,
|
||||
"reason": "review_degraded",
|
||||
"source": "task_acceptance_review",
|
||||
"rationale": "Acceptance reviewers did not reach a valid quorum.",
|
||||
"degraded_reasons": list(getattr(result, "degraded_reasons", []) or []),
|
||||
"open_obligations": [str(item.get("id")) for item in open_obligations],
|
||||
})
|
||||
# Show the slot failure causes beside the verdict, not only in task_results.
|
||||
ctx.emit_progress(
|
||||
"Task acceptance review: DEGRADED (no valid quorum; not recorded as PASS)."
|
||||
"Task acceptance review: DEGRADED (no settled verdict; not recorded as PASS)."
|
||||
+ _slot_cause_clause(result)
|
||||
)
|
||||
return False
|
||||
|
|
@ -1035,7 +1011,7 @@ def _record_acceptance_infra_failure(ctx: _TaskAcceptanceContext, exc: Exception
|
|||
"status": ACCEPTANCE_FINALIZED_UNACCEPTED,
|
||||
"reason": "infra_failure",
|
||||
"source": "task_acceptance_review",
|
||||
"rationale": "The mandatory host acceptance panel failed before a valid quorum.",
|
||||
"rationale": "The mandatory host acceptance panel failed before any reviewer answered.",
|
||||
"degraded_reasons": [f"{type(exc).__name__}: {safe_error}"],
|
||||
})
|
||||
ctx.emit_progress("Task acceptance review: DEGRADED after host review infrastructure failure.")
|
||||
|
|
@ -1441,6 +1417,10 @@ def _run_task_acceptance_review_once(
|
|||
packet_budget_chars=acceptance_packet_budget_chars(_acceptance_delivery_slots()),
|
||||
)
|
||||
try:
|
||||
from ouroboros.review_dispatch import reconcile_pending_acceptance_runs
|
||||
|
||||
reconcile_pending_acceptance_runs(
|
||||
llm_trace, drive_root=drive_root or tools._ctx.drive_root, usage_ctx=tools._ctx)
|
||||
from types import SimpleNamespace
|
||||
|
||||
from ouroboros.review_substrate import build_review_binding
|
||||
|
|
@ -1454,6 +1434,9 @@ def _run_task_acceptance_review_once(
|
|||
from ouroboros.loop_delivery import delivery_subject_hash
|
||||
|
||||
review_ctx.review_binding["subject_hash"] = delivery_subject_hash(tools._ctx, llm_trace, content)
|
||||
from ouroboros.loop_messages import owner_source_sha256
|
||||
|
||||
review_ctx.review_binding["owner_source_sha256"] = owner_source_sha256(tools._ctx) # the premises this panel judged
|
||||
if _loop()._task_acceptance_owner_generation_changed(tools._ctx):
|
||||
_loop()._supersede_task_acceptance_for_owner_followup(tools._ctx, llm_trace)
|
||||
return True
|
||||
|
|
@ -1466,6 +1449,11 @@ def _run_task_acceptance_review_once(
|
|||
seen_bindings, prior_run = _prior_acceptance_run(
|
||||
tools._ctx, llm_trace, binding_hash, paid_identity=paid_identity,
|
||||
)
|
||||
from ouroboros.acceptance_settlement import _deliver_under_running_panel
|
||||
|
||||
handled = _deliver_under_running_panel(review_ctx, prior_run)
|
||||
if handled is not None:
|
||||
return handled
|
||||
reused_result = None
|
||||
applied_before = bool(prior_run and (prior_run.get("applied_decision") or prior_run.get("feedback_delivered")))
|
||||
if prior_run is not None:
|
||||
|
|
@ -1515,26 +1503,21 @@ def _run_task_acceptance_review_once(
|
|||
passes_before_apply = int(
|
||||
getattr(tools._ctx, "_task_acceptance_improvement_passes", 0) or 0
|
||||
)
|
||||
if prior_run is not None and acceptance_run_pending(prior_run):
|
||||
from ouroboros.review_dispatch import collect_task_acceptance_run
|
||||
|
||||
panel_result = collect_task_acceptance_run(
|
||||
prior_run, drive_root=drive_root or tools._ctx.drive_root, usage_ctx=tools._ctx,
|
||||
)
|
||||
# Keep the paid operation's request; only its producer facts advance.
|
||||
prior_run.update({key: value for key, value in vars(panel_result).items()
|
||||
if key != "request"})
|
||||
else:
|
||||
panel_result = reused_result or _loop()._execute_task_acceptance_panel(review_ctx)
|
||||
panel_result = reused_result or _loop()._execute_task_acceptance_panel(review_ctx)
|
||||
run_record = prior_run if reused_result is not None else _record_host_acceptance_run(review_ctx, panel_result)
|
||||
if acceptance_run_pending(panel_result):
|
||||
tools._ctx._task_acceptance_pending = str(run_record.get("binding_hash") or "")
|
||||
run_record["enforcement_impact"] = "pending_feedback"
|
||||
from ouroboros.acceptance_settlement import remember_settlement_trace
|
||||
|
||||
remember_settlement_trace(tools._ctx, llm_trace, run_record)
|
||||
llm_trace["review_decision"].update({
|
||||
"eligibility": "review_in_flight", "operation_state": "in_flight",
|
||||
})
|
||||
emit_progress("Task acceptance review is running; Main can receive and answer messages.")
|
||||
if not review_enforcement_blocks("blocking"):
|
||||
from ouroboros.acceptance_settlement import acceptance_wait_chosen
|
||||
|
||||
if not acceptance_wait_chosen(tools._ctx):
|
||||
return _finish_cyber_acceptance(review_ctx, panel_result)
|
||||
return True
|
||||
tools._ctx._task_acceptance_pending = ""
|
||||
|
|
|
|||
|
|
@ -77,6 +77,9 @@ _DELIVERY_HOLD_CONTROLS = frozenset({
|
|||
_SKILL_ACTION_HOLD_CONTROL,
|
||||
_CHILD_ABSORPTION_HOLD_CONTROL,
|
||||
})
|
||||
# Header of the host's own rendered control block; identifies the transcript's
|
||||
# control history the way ``acceptance_observation`` marks the observation rows.
|
||||
_DELIVERY_CONTROL_MARKER = "[DELIVERY_FINALIZATION_CONTROL]"
|
||||
|
||||
|
||||
def _swarm_handoff_attempt(ctx: Any) -> Dict[str, Any]:
|
||||
|
|
@ -596,14 +599,15 @@ def _merge_finalization_trace(
|
|||
return llm_trace
|
||||
|
||||
|
||||
def _delivery_control_prompt(candidate: DeliveryCandidate, *, keep_allowed: bool) -> str:
|
||||
def _delivery_control_prompt(candidate: DeliveryCandidate, *, keep_allowed: bool,
|
||||
pending_review_choice: bool = False) -> str:
|
||||
keep_line = (
|
||||
"keep is allowed because no answer-invalidating evidence changed."
|
||||
if keep_allowed
|
||||
else "keep is NOT allowed because owner/tool/child/verification evidence changed."
|
||||
)
|
||||
return (
|
||||
"[DELIVERY_FINALIZATION_CONTROL]\n"
|
||||
f"{_DELIVERY_CONTROL_MARKER}\n"
|
||||
f"A complete answer candidate (revision {candidate.revision}, sha256 "
|
||||
f"{candidate.content_sha256[:12]}) is retained by the loop; do not replace it with a "
|
||||
f"service notice. {keep_line}\n"
|
||||
|
|
@ -616,6 +620,10 @@ def _delivery_control_prompt(candidate: DeliveryCandidate, *, keep_allowed: bool
|
|||
"\nEither form may include acceptance_subject with the latest observed "
|
||||
"owner_source_sha256, optional complete effective_criteria and material_tool_indices. "
|
||||
"Keep can retain answer text while explicitly changing its review subject."
|
||||
+ ('\nA paid acceptance panel on an earlier revision is still running. Either form may add '
|
||||
'"pending_review":"wait" (the default: hold this answer until its verdict arrives) or '
|
||||
'"pending_review":"finish" (deliver now; the verdict reaches you as advice when it settles).'
|
||||
if pending_review_choice else "")
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -645,29 +653,43 @@ def _arm_delivery_control(
|
|||
llm_trace: Dict[str, Any],
|
||||
*,
|
||||
control: str = "awaiting_control",
|
||||
skip_if_unchanged: bool = False,
|
||||
) -> None:
|
||||
candidate = getattr(tools._ctx, "_delivery_candidate", None)
|
||||
if not isinstance(candidate, _loop().DeliveryCandidate):
|
||||
return
|
||||
evidence_revision, evidence_fingerprint = _loop()._delivery_evidence_state(tools, ctx, llm_trace)
|
||||
candidate.finalization_control = control
|
||||
candidate.repair_attempted = False
|
||||
# The acceptance wake re-offers an EXISTING candidate's contract: its one
|
||||
# malformed-control repair stays spent. Every other arm starts a new episode.
|
||||
if not skip_if_unchanged:
|
||||
candidate.repair_attempted = False
|
||||
tools._ctx._delivery_control_required = True
|
||||
_loop()._append_or_merge_user_message(
|
||||
ctx.messages,
|
||||
_delivery_control_prompt(
|
||||
candidate,
|
||||
keep_allowed=_delivery_keep_allowed(
|
||||
candidate, evidence_revision, evidence_fingerprint,
|
||||
),
|
||||
),
|
||||
from ouroboros.acceptance_settlement import acceptance_choice_offered
|
||||
|
||||
control_prompt = _delivery_control_prompt(
|
||||
candidate,
|
||||
keep_allowed=_delivery_keep_allowed(candidate, evidence_revision, evidence_fingerprint),
|
||||
pending_review_choice=bool(getattr(tools._ctx, "_task_acceptance_pending", "")
|
||||
and acceptance_choice_offered()),
|
||||
)
|
||||
# ``skip_if_unchanged`` is the repeated re-offer (every acceptance wake shows
|
||||
# the keep contract): an unchanged candidate renders identical bytes, so the
|
||||
# transcript's control history is not repeated, the way
|
||||
# ``prepare_acceptance_observation`` skips an unchanged observation. Every
|
||||
# other caller arms because something changed and always appends. ``slot``
|
||||
# keeps an already-sent tail row byte-frozen rather than rewritten (#906).
|
||||
latest = next((row for row in reversed(ctx.messages)
|
||||
if _DELIVERY_CONTROL_MARKER in str(row.get("content") or "")), None)
|
||||
if not (skip_if_unchanged and latest is not None
|
||||
and control_prompt in str(latest.get("content") or "")):
|
||||
_loop()._append_or_merge_user_message(ctx.messages, control_prompt, slot=tools._ctx)
|
||||
candidate.control_episode_seen = True
|
||||
from ouroboros.loop_acceptance import capture_acceptance_observation, acceptance_observation_prompt
|
||||
|
||||
observed = capture_acceptance_observation(tools._ctx, llm_trace, getattr(ctx, "incoming_messages", None))
|
||||
if prompt := acceptance_observation_prompt(tools._ctx, observed):
|
||||
_loop()._append_or_merge_user_message(ctx.messages, prompt)
|
||||
_loop()._append_or_merge_user_message(ctx.messages, prompt, slot=tools._ctx)
|
||||
_loop()._publish_delivery_candidate(tools, candidate, llm_trace)
|
||||
|
||||
|
||||
|
|
@ -766,7 +788,9 @@ def _classify_parsed_delivery_control(
|
|||
if not isinstance(parsed, dict) or "delivery_control" not in parsed:
|
||||
return "none", "", exact_error
|
||||
selected = str(parsed.get("delivery_control") or "")
|
||||
keys = set(parsed) - {"acceptance_subject"}
|
||||
if "pending_review" in parsed and str(parsed.get("pending_review") or "").strip().lower() not in {"wait", "finish"}:
|
||||
return "invalid", "", 'pending_review must be "wait" or "finish"'
|
||||
keys = set(parsed) - {"acceptance_subject", "pending_review"}
|
||||
if selected == "keep" and keys == {"delivery_control"}:
|
||||
return "keep", "", ""
|
||||
if selected == "replace" and keys == {"delivery_control", "full_answer"}:
|
||||
|
|
@ -922,6 +946,11 @@ def _resolve_delivery_control(
|
|||
applied, subject_error = apply_delivery_subject_decision(tools, ctx, llm_trace, parsed["acceptance_subject"])
|
||||
if not applied:
|
||||
control_kind, error = "invalid", subject_error
|
||||
if control_kind in {"keep", "replace"} and isinstance(parsed, dict):
|
||||
# Recorded on every control answer (the classifier already refused any
|
||||
# other value), so an answer without the key always means "wait" rather
|
||||
# than inheriting an earlier round's choice.
|
||||
tools._ctx._acceptance_pending_review_choice = str(parsed.get("pending_review") or "wait").strip().lower()
|
||||
evidence_revision, evidence_fingerprint = _loop()._delivery_evidence_state(tools, ctx, llm_trace)
|
||||
valid = control_kind == "replace"
|
||||
if control_kind == "keep":
|
||||
|
|
|
|||
|
|
@ -701,8 +701,8 @@ def plan_review_disclosure(decision: Dict[str, Any], forced_reason: str = "") ->
|
|||
if decision.get("status") == "cycles_exhausted" and decision.get("enforcement") == "blocking":
|
||||
return (
|
||||
f"\n\n⚠️ Blocking plan review stayed open ({outcome or 'open'}) with the review-cycle "
|
||||
f"cap spent ({decision.get('cycles_paid')} paid cycle(s)); the task is finalized as "
|
||||
"blocked_with_evidence — the planned work must not be treated as done."
|
||||
f"cap spent ({decision.get('cycles_paid')} paid cycle(s)); the task ends blocked "
|
||||
"with its evidence recorded; the planned work must not be treated as done."
|
||||
)
|
||||
if decision.get("quorum_unreachable") and decision.get("enforcement") == "blocking":
|
||||
# B2b: the agent chose the honest blocked terminal while the reviewer quorum
|
||||
|
|
@ -712,8 +712,8 @@ def plan_review_disclosure(decision: Dict[str, Any], forced_reason: str = "") ->
|
|||
f"\n\n⚠️ Blocking plan review stayed open ({outcome or 'open'}) with its reviewer "
|
||||
"quorum structurally unreachable (typed window-exhausted reviewer lanes"
|
||||
+ (f"; earliest recorded reset {reset}" if reset else "")
|
||||
+ "); the task is finalized as blocked_with_evidence — the planned work must "
|
||||
"not be treated as done."
|
||||
+ "); the task ends blocked with its evidence recorded; the planned work "
|
||||
"must not be treated as done."
|
||||
)
|
||||
if decision.get("allow"):
|
||||
return (
|
||||
|
|
|
|||
|
|
@ -653,6 +653,53 @@ def routing_option_label(option: Any) -> str:
|
|||
OUTCOME_PHASE_HEADLINE = {"working": "Working", "done": "Done", "warn": "Done with warnings",
|
||||
"error": "Failed", "cancelled": "Cancelled"}
|
||||
|
||||
# One owner sentence per typed cause, for BOTH lifecycle writers and the card.
|
||||
# Keyed on the CODE only — never on (status × reason): the status word already
|
||||
# speaks, and a product of the two would be a matrix nobody maintains. A code
|
||||
# with no sentence stays raw (docs/DESIGN.md "Status and chips"), and the raw
|
||||
# code stays typed on the row. web/modules/log_events.js carries the twin;
|
||||
# web/tests/fixtures/outcome_phase_parity.json pins both.
|
||||
TASK_CAUSE_PHRASES = {
|
||||
# Acceptance-decision reasons. A clean accepted decision renders no clause,
|
||||
# so clean_pass and clean_pass_obligations_closed carry no sentence; an
|
||||
# accepted decision with a sentence here still states its cause.
|
||||
"previous_revision_accepted": "The reviewers approved the earlier version of this answer; it changed before they finished.",
|
||||
"author_finish": "The answer was delivered on Main's own judgement; the reviewers had not signed it off.",
|
||||
"review_degraded": "No reviewer verdict was established for this answer.",
|
||||
"infra_failure": "A review infrastructure failure prevented a settled verdict.",
|
||||
"dialogue_terminal": "The reviewers and Main could not agree, and both positions were kept.",
|
||||
"improvement_capsule": "The reviewers asked for one more pass and Main was given their notes.",
|
||||
"fence_reopen_failed": "The requested extra pass could not be started, so the answer stands as it was.",
|
||||
"review_cycles_exhausted": "The task used up its review rounds before the answer was signed off.",
|
||||
"open_obligations": "The answer was delivered with reviewer requests still open.",
|
||||
"improvement_window_closed": "There was no room left for another pass, so the answer stands as it was.",
|
||||
"capsule_spent": "The one allowed improvement pass was already used.",
|
||||
"reviewer_fail_no_capsule": "A reviewer rejected the answer and suggested nothing to change.",
|
||||
"no_actionable_changes": "The re-review was not clean and suggested nothing to change.",
|
||||
"identical_acceptance_refused": "Nothing had changed since the last review, so the recorded verdict stands.",
|
||||
"review_skipped_deadline_reserve": "There was not enough time left to review the answer.",
|
||||
"delivery_binding_superseded": "The answer or its evidence changed, so the earlier review no longer covered it.",
|
||||
"owner_followup": "A new message from you arrived, so the review was set aside for it.",
|
||||
"evidence_refresh": "The work changed after the review was frozen, so it no longer covered the answer.",
|
||||
"revision_unavailable_on_forced_rail": "The task had to stop, so the requested rework never happened.",
|
||||
"owner_hurry": "You asked me to hurry, so no further review was started.",
|
||||
"unspecified": "The answer was not signed off, and no cause was recorded.",
|
||||
# The rail that ended the task before an owed acceptance panel could run.
|
||||
"acceptance_bypassed_budget_exhausted": "The task ran out of budget before the answer could be reviewed.",
|
||||
"acceptance_bypassed_round_limit": "The task hit its round limit before the answer could be reviewed.",
|
||||
"acceptance_bypassed_deadline": "The task ran out of time before the answer could be reviewed.",
|
||||
"acceptance_bypassed_provider_unavailable": "The model provider was unavailable, so the answer was never reviewed.",
|
||||
"acceptance_bypassed_context_overflow": "The task outgrew its context before the answer could be reviewed.",
|
||||
"acceptance_bypassed_children_unabsorbed": "Some sub-tasks had not been folded in, so the answer was never reviewed.",
|
||||
# Execution reason codes, carried verbatim from the card's own old table.
|
||||
"plan_review_advisory": "Plan review never closed; the work continued under advisory enforcement",
|
||||
"host_child_status_suffix": "A child task had not settled when the answer was delivered",
|
||||
"invalid_delivery_control_after_repair": "The delivery control object was still malformed after repair",
|
||||
"budget_exhausted": "The task ran out of budget before it could finish cleanly",
|
||||
"delivery_control_degraded": "Delivery finished in a degraded control state",
|
||||
"delegated_custody_unreconciled": "Some delegated work was never reconciled.",
|
||||
}
|
||||
|
||||
|
||||
def outcome_phase(result: Dict[str, Any], event: Dict[str, Any]) -> str:
|
||||
"""The host mirror of the browser's terminality gate and severity fold.
|
||||
|
|
@ -986,20 +1033,18 @@ def _append_terminal_task_projection(
|
|||
phase = outcome_phase(effective, event)
|
||||
outcome = OUTCOME_PHASE_HEADLINE[phase]
|
||||
row_chat_id = int(event.get("chat_id") or task.get("chat_id") or 0)
|
||||
excerpt = _completion_excerpt(effective, chat_id=row_chat_id)
|
||||
details = f'Details: get_task_result(task_id="{tid}")'
|
||||
text = (
|
||||
f"{outcome}. role={role}; parent={parent_id or 'unknown'}; "
|
||||
f"root={root_id}; project={project_id or 'none'}."
|
||||
)
|
||||
if excerpt:
|
||||
text += f" {excerpt}"
|
||||
# The room IS the project and ``result_ref`` IS the pointer, so the row
|
||||
# says in words only what the model cannot read off the typed fields:
|
||||
# ``memory._format_chat_line`` renders the text and drops every other
|
||||
# key, leaving lineage as the one fact that must stay prose.
|
||||
text = (f"{outcome}. Root task {tid}." if is_root
|
||||
else f"{outcome}. {role} (child {tid} of {parent_id or 'unknown'}).")
|
||||
verdict = _completion_verdict(effective, event)
|
||||
if verdict:
|
||||
text += f" {verdict}"
|
||||
depth = (effective.get("swarm_efficiency") or {}).get("depth") if isinstance(effective.get("swarm_efficiency"), dict) else None
|
||||
if isinstance(depth, dict) and depth.get("requested_depth") is not None:
|
||||
text += f" Depth requested={depth['requested_depth']}, permitted={depth.get('permitted_depth')}, achieved={depth.get('achieved_depth')} ({depth.get('status')})."
|
||||
excerpt = _completion_excerpt(effective, chat_id=row_chat_id, salvage_only=True)
|
||||
if excerpt:
|
||||
text += f" {excerpt}"
|
||||
result_ref = {"kind": "task_result", "task_id": tid, "reader": "get_task_result"}
|
||||
row = {
|
||||
"ts": str(event.get("ts") or effective.get("ts") or utc_now_iso()),
|
||||
|
|
@ -1014,7 +1059,7 @@ def _append_terminal_task_projection(
|
|||
"outcome_authority": "canonical_task_result_after_finalization",
|
||||
"outcome_axes": effective.get("outcome_axes") or event.get("outcome_axes") or {},
|
||||
"reason_code": reason, "result_ref": result_ref,
|
||||
"text": f"{text} {details}",
|
||||
"text": text,
|
||||
}
|
||||
if isinstance(effective.get("model_execution"), dict):
|
||||
row["model_execution"] = dict(effective["model_execution"])
|
||||
|
|
@ -1086,7 +1131,8 @@ def _stop_receipt_reached_chat(result: Dict[str, Any], chat_id: Any) -> bool:
|
|||
return lineage is not None and str(lineage) == str(chat_id)
|
||||
|
||||
|
||||
def _completion_excerpt(result: Dict[str, Any], *, chat_id: Any = None) -> str:
|
||||
def _completion_excerpt(result: Dict[str, Any], *, chat_id: Any = None,
|
||||
salvage_only: bool = False) -> str:
|
||||
"""One plain-text excerpt for BOTH lifecycle writers (event + task_summary).
|
||||
|
||||
Markdown markers are stripped BEFORE whitespace flattening: the stripper's
|
||||
|
|
@ -1100,7 +1146,14 @@ def _completion_excerpt(result: Dict[str, Any], *, chat_id: Any = None) -> str:
|
|||
the caller's own pointer keeps owning the untruncated copy. ``chat_id`` is
|
||||
the row's destination: only there can the stop receipt already have
|
||||
published the same text, and only there does the label stand alone.
|
||||
|
||||
``salvage_only`` is how the durable rows ask for that ONE excerpt and
|
||||
nothing else: a cut of the model's own answer is already in the room the
|
||||
row lives in, while salvaged bytes exist nowhere else.
|
||||
"""
|
||||
salvaged = str(result.get("terminal_origin") or "") == TERMINAL_ORIGIN_HOST_SALVAGE
|
||||
if salvage_only and not salvaged:
|
||||
return ""
|
||||
body = ""
|
||||
for key in ("summary", "result", "error"):
|
||||
body = " ".join(strip_markdown(str(result.get(key) or "")).split())
|
||||
|
|
@ -1109,7 +1162,7 @@ def _completion_excerpt(result: Dict[str, Any], *, chat_id: Any = None) -> str:
|
|||
if not body:
|
||||
return ""
|
||||
excerpt = body if len(body) <= 240 else body[:239].rstrip() + "…"
|
||||
if str(result.get("terminal_origin") or "") != TERMINAL_ORIGIN_HOST_SALVAGE:
|
||||
if not salvaged:
|
||||
return excerpt
|
||||
if _stop_receipt_reached_chat(result, chat_id):
|
||||
return f"{SALVAGE_EXCERPT_LABEL}."
|
||||
|
|
@ -1148,11 +1201,15 @@ def _completion_verdict(result: Dict[str, Any], event: Dict[str, Any]) -> str:
|
|||
"""One TERMINATED host clause for BOTH lifecycle rows.
|
||||
|
||||
A host row must not present an unaccepted claim as the whole story: a
|
||||
non-accepted decision leads with its full upstream-bounded rationale;
|
||||
otherwise the execution reason stands. The Python twin of
|
||||
``taskReasonDetail``'s acceptance branch; callers add no punctuation.
|
||||
non-accepted decision speaks through the owner sentence of its own typed
|
||||
reason, otherwise the execution reason speaks. The stored reviewer
|
||||
rationale never reaches the row — it stays in the card, the task result and
|
||||
Logs, which is the complete text this pointer resolves to. The Python twin
|
||||
of ``taskReasonDetail``; callers add no punctuation.
|
||||
"""
|
||||
from ouroboros.outcomes import ACCEPTANCE_ACCEPTED, REASON_OWNER_REQUESTED_FINALIZATION
|
||||
from ouroboros.outcomes import (
|
||||
ACCEPTANCE_ACCEPTED, REASON_FINAL_MESSAGE, REASON_OWNER_REQUESTED_FINALIZATION,
|
||||
)
|
||||
|
||||
decision: Dict[str, Any] = {}
|
||||
veto: Dict[str, Any] = {}
|
||||
|
|
@ -1165,30 +1222,31 @@ def _completion_verdict(result: Dict[str, Any], event: Dict[str, Any]) -> str:
|
|||
if isinstance(holder, dict) and isinstance(holder.get("acceptance_decision"), dict):
|
||||
decision = holder["acceptance_decision"]
|
||||
status = str(decision.get("status") or "").strip()
|
||||
cause = str(decision.get("reason") or "")
|
||||
reason = str(result.get("reason_code") or event.get("reason_code") or "")
|
||||
# A healed debt is never restored here. The objective warning the overlay
|
||||
# froze keeps the headline at "Done with warnings" and the refresh may not
|
||||
# rewrite it, but that is the axis speaking about what was true at write
|
||||
# time; naming the code again would state a debt the same record shows as
|
||||
# empty. The current execution reason speaks when there is one, otherwise
|
||||
# the row states no cause and leaves the headline to the axis that owns it.
|
||||
reason, custody = _custody_debt_reason(reason, result, event)
|
||||
if (reason != REASON_OWNER_REQUESTED_FINALIZATION and status != ACCEPTANCE_ACCEPTED
|
||||
and status and outcome_phase(result, event) in {"done", "warn"}):
|
||||
clause = f"Acceptance: {status}"
|
||||
rationale = " ".join(strip_markdown(str(decision.get("rationale") or "")).split())
|
||||
if rationale:
|
||||
clause += " — " + rationale
|
||||
elif reason and reason != REASON_OWNER_REQUESTED_FINALIZATION:
|
||||
detail = veto.get("detail") if veto.get("reason") == reason else ""
|
||||
clause = f"Reason: {' '.join(strip_markdown(str(detail)).split()) if detail else reason}"
|
||||
if custody:
|
||||
clause += f" ({custody})"
|
||||
elif custody:
|
||||
clause = f"Reason: {custody}"
|
||||
else:
|
||||
if (reason != REASON_OWNER_REQUESTED_FINALIZATION and status
|
||||
and (status != ACCEPTANCE_ACCEPTED or cause in TASK_CAUSE_PHRASES)
|
||||
and outcome_phase(result, event) in {"done", "warn"}):
|
||||
clause = TASK_CAUSE_PHRASES.get(cause, cause)
|
||||
elif reason in {REASON_OWNER_REQUESTED_FINALIZATION, REASON_FINAL_MESSAGE}:
|
||||
return ""
|
||||
return clause if clause.endswith((".", "!", "?", "…")) else clause + "."
|
||||
else:
|
||||
# A healed debt is never restored here. The objective warning the
|
||||
# overlay froze keeps the headline and the refresh may not rewrite it,
|
||||
# but naming the code again would state a debt the same record shows as
|
||||
# empty. The debt is a warning BESIDE the rail cause, and a row with
|
||||
# neither states no cause and leaves the headline to its own axis.
|
||||
reason, custody = _custody_debt_reason(reason, result, event)
|
||||
detail = veto.get("detail") if veto.get("reason") == reason else ""
|
||||
clause = (" ".join(strip_markdown(str(detail)).split()) if detail
|
||||
else TASK_CAUSE_PHRASES.get(reason, reason))
|
||||
if clause and custody:
|
||||
clause += f" ({TASK_CAUSE_PHRASES.get(custody, custody)})"
|
||||
elif custody:
|
||||
clause = TASK_CAUSE_PHRASES.get(custody, custody)
|
||||
if not clause:
|
||||
return ""
|
||||
return clause if clause.endswith((".", "!", "?", "…", ")")) else clause + "."
|
||||
|
||||
|
||||
def _run_lives_in_its_project(
|
||||
|
|
@ -1270,14 +1328,15 @@ def enqueue_project_completion_summary(
|
|||
# Offering "Open the Project" would reproduce the reported defect —
|
||||
# a Main row leading into an empty room.
|
||||
return False
|
||||
excerpt = _completion_excerpt(result, chat_id=1)
|
||||
# Only the salvage excerpt survives here: a cut of the model's own
|
||||
# answer repeats bytes the Project already holds, while salvaged bytes
|
||||
# exist nowhere else. This writer's only pointer is the invitation, so
|
||||
# a salvage may never displace it — preserved bytes named with no way
|
||||
# to reach them are worse than the plain invitation.
|
||||
excerpt = _completion_excerpt(result, chat_id=1, salvage_only=True)
|
||||
verdict = _completion_verdict(result, task_done_event)
|
||||
lead = f"{verdict} " if verdict else ""
|
||||
# This writer's only pointer is the invitation below, so a salvage may
|
||||
# never displace it: preserved bytes named with no way to reach them are
|
||||
# worse than the plain invitation. Both salvage forms keep it, the
|
||||
# labelled excerpt and the label that stands alone.
|
||||
if excerpt.startswith(SALVAGE_EXCERPT_LABEL):
|
||||
if excerpt:
|
||||
excerpt = f"{excerpt} Open the Project for details."
|
||||
event = {
|
||||
"type": "send_message", "chat_id": 1, "task_id": tid,
|
||||
|
|
@ -1326,8 +1385,7 @@ def announce_project_started(
|
|||
)
|
||||
event = {
|
||||
"type": "send_message", "chat_id": 1, "task_id": tid,
|
||||
"text": (f"{snapshot['target_label']} · Started\n"
|
||||
"Work is running in this Project."),
|
||||
"text": f"{snapshot['target_label']} · Started",
|
||||
"role": "system", "system_type": "project_started",
|
||||
"delivery_id": f"project-start:{pid}",
|
||||
"progress_meta": {
|
||||
|
|
|
|||
|
|
@ -6,8 +6,9 @@ agent never drift:
|
|||
- ``gateway/projects.py`` turn-into-project conversion (reuses an admission name,
|
||||
or names inline as a race fallback);
|
||||
- ``ensure_project_scope`` (the agent self-creates + names a project);
|
||||
- ``admission_names`` (headless runs and chat promotion, no model call).
|
||||
A direct conversation turn is never named: it renders as an activity block.
|
||||
- ``admission_names`` (headless runs and chat promotion, no model call);
|
||||
- ``spawn_turn_namer`` (a direct Main turn, named once it starts working:
|
||||
the first non-addressing tool call triggers one bounded Light call).
|
||||
|
||||
Doctrine:
|
||||
- P5 LLM-first: the model COINS the name; post-processing is purely lexical
|
||||
|
|
@ -19,9 +20,12 @@ Doctrine:
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import contextvars
|
||||
import logging
|
||||
import pathlib
|
||||
import threading
|
||||
from dataclasses import replace
|
||||
from typing import Any, Dict, Optional, Sequence
|
||||
from typing import Any, Callable, Dict, Optional, Sequence
|
||||
|
||||
log = logging.getLogger("ouroboros.project_naming")
|
||||
|
||||
|
|
@ -269,6 +273,138 @@ async def llm_project_name_async(
|
|||
return fb
|
||||
|
||||
|
||||
def _refresh_root_cost_after_naming(drive_root: Any, task_id: str) -> None:
|
||||
"""Refresh a terminal root projection after the naming attempt settles."""
|
||||
try:
|
||||
from types import SimpleNamespace
|
||||
|
||||
from ouroboros.agent_task_pipeline import _set_root_post_task_checkpoint
|
||||
from ouroboros.task_results import load_task_result
|
||||
|
||||
current = load_task_result(drive_root, task_id) or {}
|
||||
refreshed = {**current, "id": task_id, "budget_drive_root": str(drive_root)}
|
||||
_set_root_post_task_checkpoint(
|
||||
SimpleNamespace(drive_root=pathlib.Path(drive_root)), refreshed, "refresh",
|
||||
)
|
||||
except Exception:
|
||||
log.debug("project naming cost refresh failed for %s", task_id, exc_info=True)
|
||||
|
||||
|
||||
def spawn_turn_namer(
|
||||
drive_root: Any, task_id: str, text: str, *, broadcast: Optional[Callable[[dict], None]] = None,
|
||||
) -> None:
|
||||
"""Coin an LLM title for a direct turn that started working, in a DAEMON thread.
|
||||
|
||||
Called once per Main turn from the first non-addressing tool-call frame (the
|
||||
seam that already stamps that frame as work: ``supervisor/log_addressing.py``),
|
||||
so a greeting that runs no tool costs no naming call while a turn that does
|
||||
real work gets a human title as its block becomes the task card. Writes the
|
||||
coined ``suggested_name`` onto the task result (turn-into-project then reuses
|
||||
it with zero extra call) and, via ``broadcast``, emits a ``task_named`` event so
|
||||
the live card shows the title. NEVER blocks the task. ``drive_root`` is captured at
|
||||
CALL time — NOT read from a mutable module global at thread-execution time — so a later
|
||||
context switch (or a test that swaps the supervisor drive) can't redirect this thread's
|
||||
write. Skips cleanly unless ``drive_root`` is a real directory (test safety: a stub /
|
||||
MagicMock drive must never materialise a stray path — chat_observed persists BEFORE the
|
||||
LLM call). Fail-soft."""
|
||||
from ouroboros.settings_integrity import copy_task_settings_context
|
||||
|
||||
body = " ".join(str(text or "").split())
|
||||
if not body:
|
||||
return
|
||||
try:
|
||||
if not pathlib.Path(str(drive_root)).is_dir():
|
||||
return
|
||||
except (OSError, TypeError, ValueError):
|
||||
return
|
||||
|
||||
def _work() -> None:
|
||||
try:
|
||||
# v6.58.0 (§3.4b): HARD total wall-clock bound. The transport timeout bounds
|
||||
# ONE attempt, but llm.chat's retry/fallback chain under a degraded provider
|
||||
# could stretch the whole call to tens of minutes (the incident where the
|
||||
# card was named 24 minutes late). A title is cosmetic: if it hasn't landed
|
||||
# within the transport budget + slack, drop it — the id/title heuristics and
|
||||
# the convert path's own bounded inline call (8s) already cover naming.
|
||||
_result: list[str] = []
|
||||
_detached = threading.Event()
|
||||
_finished = threading.Event()
|
||||
_refresh_lock = threading.Lock()
|
||||
_refreshed = False
|
||||
|
||||
def _refresh_detached_once() -> None:
|
||||
nonlocal _refreshed
|
||||
with _refresh_lock:
|
||||
if _refreshed:
|
||||
return
|
||||
_refreshed = True
|
||||
_refresh_root_cost_after_naming(drive_root, task_id)
|
||||
|
||||
def _call() -> None:
|
||||
try:
|
||||
_result.append(llm_project_name(body, drive_root=drive_root, task_id=task_id))
|
||||
except Exception:
|
||||
log.debug("turn namer inner call failed for %s", task_id, exc_info=True)
|
||||
finally:
|
||||
_finished.set()
|
||||
if _detached.is_set():
|
||||
_refresh_detached_once()
|
||||
|
||||
settings_context = contextvars.Context()
|
||||
copy_task_settings_context(settings_context)
|
||||
inner = threading.Thread(target=settings_context.run, args=(_call,), name=f"namer-call-{task_id}", daemon=True)
|
||||
inner.start()
|
||||
if not _finished.wait(timeout=max(0.0, _naming_timeout_sec() + 30.0)):
|
||||
_detached.set()
|
||||
# Close the race where settlement lands between wait() and
|
||||
# the detached marker. The once-guard covers both interleavings.
|
||||
if _finished.is_set():
|
||||
_refresh_detached_once()
|
||||
log.debug("turn namer exceeded its wall-clock bound for %s; skipped", task_id)
|
||||
return
|
||||
inner.join()
|
||||
if not _result:
|
||||
log.debug("turn namer exceeded its wall-clock bound for %s; skipped", task_id)
|
||||
return
|
||||
name = _result[0]
|
||||
if not name:
|
||||
return
|
||||
from ouroboros.task_results import (
|
||||
STATUS_RUNNING,
|
||||
load_task_result,
|
||||
write_task_result,
|
||||
)
|
||||
|
||||
# Persist suggested_name as same-status ENRICHMENT, not a RUNNING transition: a
|
||||
# fast task may already be terminal (completed/failed/cancelled) by the time this
|
||||
# daemon finishes, and write_task_result's monotonic guard DROPS a regressing
|
||||
# RUNNING write — which would silently lose the name the convert path reuses.
|
||||
# Writing under the current on-disk status lets the monotonic guard's same-status
|
||||
# enrichment carry the field through (and a benign drop only in the rare race where
|
||||
# the status advanced past our read — acceptable for a best-effort title).
|
||||
current = load_task_result(drive_root, task_id) or {}
|
||||
status = str(current.get("status") or "") or STATUS_RUNNING
|
||||
write_task_result(drive_root, task_id, status, suggested_name=name)
|
||||
# A cosmetic namer may settle concurrently with or after the ordinary
|
||||
# post-task worker. The shared refresh/checkpoint critical section
|
||||
# linearizes both cases without marking an unfinished phase complete.
|
||||
_refresh_root_cost_after_naming(drive_root, task_id)
|
||||
if broadcast is not None:
|
||||
try:
|
||||
broadcast({"type": "task_named", "task_id": task_id, "suggested_name": name})
|
||||
except Exception:
|
||||
log.debug("task_named broadcast failed for %s", task_id, exc_info=True)
|
||||
except Exception:
|
||||
log.debug("turn namer failed for %s", task_id, exc_info=True)
|
||||
|
||||
try:
|
||||
settings_context = contextvars.Context()
|
||||
copy_task_settings_context(settings_context)
|
||||
threading.Thread(target=settings_context.run, args=(_work,), name=f"namer-{task_id}", daemon=True).start()
|
||||
except Exception:
|
||||
log.debug("turn namer thread spawn failed for %s", task_id, exc_info=True)
|
||||
|
||||
|
||||
def admission_names(body: Dict[str, Any], description: str) -> tuple:
|
||||
"""The run's owner-facing name at admission: ``(title, suggested_name)``.
|
||||
|
||||
|
|
|
|||
|
|
@ -830,9 +830,15 @@ def task_presentation_snapshot(drive_root: Any, task_id: str, *, task: Any = Non
|
|||
if task_name:
|
||||
break
|
||||
task_name = task_name or "Task"
|
||||
label = f"{pname} › {task_name}" if pname else task_name
|
||||
# One name, said once: a task whose own name IS the project name renders
|
||||
# "Launch › Launch", which reads as two different things. Exact equality
|
||||
# only — a prefix rule would fold "Art" into "Arthur".
|
||||
if pname and task_name == pname:
|
||||
label = task_name
|
||||
return {"project_id": pid, "project_name": pname, "task_id": tid,
|
||||
"project_routable": registered, "task_name": task_name,
|
||||
"target_label": f"{pname} › {task_name}" if pname else task_name}
|
||||
"target_label": label}
|
||||
|
||||
|
||||
def create_project(
|
||||
|
|
|
|||
|
|
@ -995,11 +995,18 @@ def _settle_review_attempt(
|
|||
and str(getattr(actor, "operation_state", "") or "") == "not_dispatched")
|
||||
)
|
||||
released_wave: Dict[str, Any] = {}
|
||||
quorum_wave: Dict[str, Any] = {}
|
||||
if entry.released_early and entry.wave_key in _RELEASED_WAVES:
|
||||
roster = _RELEASED_WAVES[entry.wave_key]
|
||||
roster["slots"][str(getattr(slot, "slot_id", "") or "")] = str(actor.status or "settled")
|
||||
slot_id = str(getattr(slot, "slot_id", "") or "")
|
||||
roster["slots"][slot_id] = str(actor.status or "settled")
|
||||
if getattr(request, "surface", "") == "task_acceptance": # plan review's frame carries counts only
|
||||
roster.setdefault("verdicts", {})[slot_id] = _settled_slot_verdict(actor)
|
||||
if all(roster["slots"].values()):
|
||||
released_wave = _RELEASED_WAVES.pop(entry.wave_key)
|
||||
elif _released_quorum_reached(request, roster):
|
||||
roster["quorum_announced"] = True
|
||||
quorum_wave = copy.deepcopy(roster)
|
||||
if replayable and usage_ctx is not None and (late or explicit_retry):
|
||||
settled = getattr(usage_ctx, "_review_settled_attempts", None)
|
||||
if not isinstance(settled, dict):
|
||||
|
|
@ -1019,10 +1026,10 @@ def _settle_review_attempt(
|
|||
|
||||
announce_released_settlement(usage_ctx, request=request, task_id=task_id, slot=slot, actor=actor,
|
||||
settled_wave=dict(released_wave.get("slots") or {}), roster_size=int(released_wave.get("total") or 0))
|
||||
if getattr(request, "surface", "") == "task_acceptance" and released_wave:
|
||||
from ouroboros.loop_acceptance_review import announce_acceptance_settlement
|
||||
if getattr(request, "surface", "") == "task_acceptance" and (released_wave or quorum_wave):
|
||||
from ouroboros.acceptance_settlement import announce_acceptance_settlement
|
||||
|
||||
announce_acceptance_settlement(usage_ctx, request, released_wave)
|
||||
announce_acceptance_settlement(usage_ctx, request, released_wave or quorum_wave)
|
||||
if late and not pending_invocation and not custody_lost and usage_ctx is not None:
|
||||
try:
|
||||
from ouroboros.tools.review_helpers import emit_review_event
|
||||
|
|
@ -1046,9 +1053,55 @@ def _wave_key(request: Any) -> str:
|
|||
return "|".join(str(getattr(request, key, "") or "") for key in ("surface", "task_id", "retry_key"))
|
||||
|
||||
|
||||
def _released_quorum_reached(request: Any, roster: Dict[str, Any]) -> bool:
|
||||
"""Whether this wave has enough answered slots to be worth waking Main for.
|
||||
|
||||
The panel's own ``min_successful_slots`` is the quorum; slots that ANSWERED
|
||||
before the release (the drain collected an ok/empty actor, not a refusal or
|
||||
an expiry) are counted from the roster's registration. A wave
|
||||
announced at its quorum is never announced for that reason twice; the last
|
||||
straggler still announces through the completed-roster arm.
|
||||
"""
|
||||
if roster.get("quorum_announced"):
|
||||
return False
|
||||
policy = getattr(request, "policy", None) or {}
|
||||
try:
|
||||
quorum = max(1, int(policy.get("min_successful_slots") or 1))
|
||||
except (TypeError, ValueError):
|
||||
quorum = 1
|
||||
slots = roster.get("slots") or {}
|
||||
answered = sum(1 for status in slots.values() if status in {"ok", "empty"})
|
||||
early = set(roster.get("answered_before_release_ids") or ()) - set(slots)
|
||||
return len(early) + answered >= quorum
|
||||
|
||||
|
||||
def _settled_slot_verdict(actor: Any) -> Dict[str, str]:
|
||||
"""This ONE reviewer's own verdict, parsed where it settled.
|
||||
|
||||
Quorum, tier, contract demotion and dissent stay with the collecting call's
|
||||
reducer (``review_actor_aggregation``): a settlement thread that re-derived
|
||||
them would be a second aggregation authority. Only the reviewer's own words
|
||||
travel, so the wake can BE the advice instead of a pointer to it.
|
||||
"""
|
||||
from ouroboros.triad_review import parse_review_findings
|
||||
|
||||
try:
|
||||
parsed, findings, signal = parse_review_findings(str(getattr(actor, "raw_text", "") or ""))
|
||||
except Exception:
|
||||
log.debug("released acceptance verdict could not be parsed", exc_info=True)
|
||||
return {"verdict": "", "note": ""}
|
||||
from ouroboros.utils import truncate_review_artifact
|
||||
|
||||
note = str((parsed or {}).get("summary") or "") if isinstance(parsed, dict) else ""
|
||||
note = note or next((str(row.get("recommendation") or row.get("item") or "")
|
||||
for row in (findings or []) if isinstance(row, dict)), "")
|
||||
return {"verdict": str(signal or "").upper(), "note": truncate_review_artifact(" ".join(note.split()), limit=400)}
|
||||
|
||||
|
||||
def _register_released_roster(
|
||||
request: Any, slots: List[Any], slot_entries: Dict[str, Any], returned_ids: set,
|
||||
slot_deadlines: Dict[str, float], monotonic_now: Callable[[str], float],
|
||||
*, answered_before_release: Any = (),
|
||||
) -> set:
|
||||
"""Register the WHOLE released roster under ONE lock hold before any released row is
|
||||
minted: a slot settling at once then finds the complete roster and cannot split the
|
||||
|
|
@ -1066,6 +1119,10 @@ def _register_released_roster(
|
|||
if released_ids:
|
||||
# A collection re-releases the wave: merging keeps the recorded outcomes.
|
||||
roster = _RELEASED_WAVES.setdefault(_wave_key(request), {"slots": {}, "total": len(slots)})
|
||||
# Slot IDS, not a count: a re-released wave replays an already-settled
|
||||
# slot through the drain, and an id in the roster is never counted twice.
|
||||
roster["answered_before_release_ids"] = sorted(
|
||||
set(roster.get("answered_before_release_ids") or ()) | set(answered_before_release))
|
||||
for slot_id in released_ids:
|
||||
roster["slots"].setdefault(slot_id, "")
|
||||
return released_ids
|
||||
|
|
@ -1365,6 +1422,8 @@ def run_custodied_review_slots(
|
|||
returned_ids = {str(getattr(actor, "slot_id", "") or "") for actor in actors}
|
||||
released_ids = _register_released_roster(
|
||||
request, slots, slot_entries, returned_ids, slot_deadlines, monotonic_now,
|
||||
answered_before_release={str(getattr(actor, "slot_id", "") or "") for actor in actors
|
||||
if str(getattr(actor, "status", "") or "") in {"ok", "empty"}},
|
||||
) if drain_deadline is not None else set()
|
||||
for slot in slots:
|
||||
slot_id = str(getattr(slot, "slot_id", "") or "")
|
||||
|
|
|
|||
|
|
@ -176,6 +176,46 @@ def collect_task_acceptance_run(run: dict, *, drive_root: Any, usage_ctx: Any) -
|
|||
usage_ctx._review_frozen_rows = previous
|
||||
|
||||
|
||||
def reconcile_pending_acceptance_runs(
|
||||
llm_trace: dict, *, drive_root: Any, usage_ctx: Any,
|
||||
) -> int:
|
||||
"""Collect every already-paid acceptance panel still recorded as running, $0.
|
||||
|
||||
The dispatch barrier (``ReviewRequest.drain_deadline``) returns the host right
|
||||
after dispatch, so a panel whose subject was re-authored before it settled is
|
||||
left with ``pending_dispatch`` rows that nothing reads: the free-replay lookup
|
||||
matches only the CURRENT binding or paid identity, so verdicts the tree already
|
||||
bought were discarded. This is the acceptance twin of plan review's
|
||||
reconcile-before-supersede (I3): it sends nothing, pays nothing, samples no new
|
||||
evidence, and advances only producer facts on runs the tree already owns.
|
||||
Idempotent by ``acceptance_run_pending`` alone -- a settled, custody-lost or
|
||||
already-collected run is never collected again. Returns how many runs advanced.
|
||||
"""
|
||||
from ouroboros.loop_acceptance_review import acceptance_run_pending
|
||||
|
||||
advanced = 0
|
||||
for run in (llm_trace.get("review_runs") or []):
|
||||
# Agent-tool acceptance runs carry no barrier and drain synchronously.
|
||||
if not isinstance(run, dict) or run.get("authority") != "host_root":
|
||||
continue
|
||||
if not isinstance(run.get("request"), dict) or not run.get("slot_roster"):
|
||||
continue
|
||||
if not acceptance_run_pending(run):
|
||||
continue
|
||||
try:
|
||||
result = collect_task_acceptance_run(
|
||||
run, drive_root=drive_root, usage_ctx=usage_ctx,
|
||||
)
|
||||
except (OSError, TimeoutError, ValueError, KeyError) as exc:
|
||||
log.warning("acceptance run %s could not be reconciled: %s",
|
||||
str(run.get("panel_id") or "")[:16], exc)
|
||||
continue
|
||||
# Keep the paid operation's request; only its producer facts advance.
|
||||
run.update({key: value for key, value in vars(result).items() if key != "request"})
|
||||
advanced += not acceptance_run_pending(run)
|
||||
return advanced
|
||||
|
||||
|
||||
def task_acceptance_preclaim_refusal(ctx: Any) -> Any:
|
||||
"""Project every free refusal before assembly and again at dispatch."""
|
||||
from ouroboros.review_substrate import ReviewRunResult
|
||||
|
|
|
|||
|
|
@ -34,6 +34,15 @@ def _sub():
|
|||
|
||||
_TIER_ORDER = {OUTCOME_TIER_SOLVED: 0, OUTCOME_TIER_BEST_EFFORT: 1, OUTCOME_TIER_BLOCKED: 2}
|
||||
|
||||
# The reviewers' JSON keeps the identifier; the note the model READS (and may
|
||||
# echo to its human) says the assessment in words — the v6.61.4 token-parroting
|
||||
# class, where a ledger token becomes owner-facing prose by being quoted back.
|
||||
TIER_WORDS = {
|
||||
OUTCOME_TIER_SOLVED: "a verified solution",
|
||||
OUTCOME_TIER_BEST_EFFORT: "a partial result",
|
||||
OUTCOME_TIER_BLOCKED: "blocked, with the evidence recorded",
|
||||
}
|
||||
|
||||
|
||||
_CRITERION_STATUSES = frozenset({"supported", "missing", "partial", "rejected"})
|
||||
|
||||
|
|
@ -214,7 +223,7 @@ def _unresolved_evidence_ref_labels(run: Any) -> List[str]:
|
|||
return list(dict.fromkeys(labels))
|
||||
|
||||
|
||||
def panel_reason(run: Any) -> str:
|
||||
def panel_reason(run: Any, *, with_tier: bool = True) -> str:
|
||||
"""One honest reason line naming the REAL blocker (v6.74.0, A6); shared by
|
||||
the capsule header, the compact projection fallback, and progress lines.
|
||||
Accepts a ``ReviewRunResult`` or its dict/namespace record."""
|
||||
|
|
@ -224,6 +233,9 @@ def panel_reason(run: Any) -> str:
|
|||
run = SimpleNamespace(**run)
|
||||
aggregate = str(getattr(run, "aggregate_signal", "") or "UNKNOWN").upper()
|
||||
tier = aggregate_outcome_tier(run)
|
||||
# The tier is said in words (the reviewers' JSON keeps the identifier); the
|
||||
# capsule header states it once itself and asks for the reason alone.
|
||||
rated = f"rated {TIER_WORDS.get(tier, tier or 'unclassified')} — " if with_tier else ""
|
||||
if aggregate == "PASS":
|
||||
if task_acceptance_is_clean(run):
|
||||
return "clean acceptance"
|
||||
|
|
@ -234,11 +246,11 @@ def panel_reason(run: Any) -> str:
|
|||
if unresolved:
|
||||
more = f" (+{len(unresolved) - 3} more)" if len(unresolved) > 3 else ""
|
||||
return (
|
||||
f"tier={tier or 'unclassified'} — cited evidence does not resolve "
|
||||
f"{rated}cited evidence does not resolve "
|
||||
f"against the packet: {', '.join(unresolved[:3])}{more}"
|
||||
)
|
||||
return (
|
||||
f"tier={tier or 'unclassified'} — a PASS is not release-clean until "
|
||||
f"{rated}a PASS is not release-clean until "
|
||||
"every criterion is supported"
|
||||
)
|
||||
if aggregate == "FAIL":
|
||||
|
|
@ -270,8 +282,8 @@ def panel_reason(run: Any) -> str:
|
|||
break
|
||||
if named:
|
||||
compact = _sub().truncate_review_artifact(" ".join(named.split()), limit=300)
|
||||
return f"tier={tier or 'unclassified'} — {compact}"
|
||||
return f"tier={tier or 'unclassified'} — reviewer FAIL without a named finding"
|
||||
return f"{rated}{compact}"
|
||||
return f"{rated}reviewer FAIL without a named finding"
|
||||
reasons = [str(r) for r in (getattr(run, "degraded_reasons", None) or []) if str(r)]
|
||||
if len(reasons) > 4:
|
||||
return "; ".join(reasons[:4]) + f" ⚠️ OMISSION NOTE: +{len(reasons) - 4} more causes in the run record"
|
||||
|
|
@ -425,8 +437,8 @@ def build_improvement_capsule(
|
|||
# the agent sees WHAT failed instead of a bare ledger label.
|
||||
header = f"[Final improvement note] Review verdict: {aggregate_signal or 'UNKNOWN'}"
|
||||
if tier:
|
||||
header += f" (tier: {tier})"
|
||||
header += f" — {panel_reason(result)}."
|
||||
header += f" — rated {TIER_WORDS.get(tier, tier)}"
|
||||
header += f" — {panel_reason(result, with_tier=False)}."
|
||||
lines = [header]
|
||||
open_ids = [
|
||||
str(o.get("id"))
|
||||
|
|
|
|||
|
|
@ -102,9 +102,11 @@ log = logging.getLogger(__name__)
|
|||
# 300-line function gate; v6.70.0 added the ground-truth-probe contract).
|
||||
_PROMOTE_CHAT_DESCRIPTION = (
|
||||
"Promote real work out of this conversation into a supervised pooled task "
|
||||
"while the conversation remains available. Use it "
|
||||
"whenever a chat request needs tools/files/multi-step work rather than a "
|
||||
"conversational answer. Before framing the objective around an EXISTING artifact "
|
||||
"while the conversation remains available. Tools, files and several steps can "
|
||||
"stay in the conversation; promote when independent work is useful — its own "
|
||||
"queue slot, admission and reviews, steerable from chat — or when the owner "
|
||||
"explicitly asks for a separate task. "
|
||||
"Before framing the objective around an EXISTING artifact "
|
||||
"('check/fix/extend the X skill/file'), ground-truth its existence with one cheap probe "
|
||||
"first (skills: list_skills; files: list_files) — memory of past work is not evidence "
|
||||
"the referent still exists. Always give a short, human-readable task `title`. To "
|
||||
|
|
@ -335,10 +337,12 @@ def get_tools() -> List[ToolEntry]:
|
|||
}, _update_scratchpad),
|
||||
ToolEntry("send_user_message", {
|
||||
"name": "send_user_message",
|
||||
"description": "Send a separate reply to the owner during ongoing work, or reach out "
|
||||
"with an insight, a question, or an invitation to collaborate. "
|
||||
"The reply appears in the conversation and leaves the task running. "
|
||||
"Progress stays in the task card; the final answer is delivered automatically.",
|
||||
"description": "Send a separate reply to the owner while work continues: the first "
|
||||
"line of longer work (what I am about to do and why), or a mid-work "
|
||||
"insight, a question, or an invitation to collaborate. It appears in "
|
||||
"the conversation as a normal reply and leaves the work running; later "
|
||||
"progress stays in the card and the final answer is delivered "
|
||||
"automatically.",
|
||||
"parameters": {"type": "object", "properties": {
|
||||
"text": {"type": "string", "description": "Message text"},
|
||||
"reason": {"type": "string", "description": "Why you're reaching out (logged, not sent)"},
|
||||
|
|
@ -374,8 +378,7 @@ def get_tools() -> List[ToolEntry]:
|
|||
}, "required": ["action"]},
|
||||
}, _toggle_consciousness),
|
||||
ToolEntry("set_next_wakeup", {
|
||||
"name": "set_next_wakeup", "description": "Choose the consciousness wake-up interval in seconds: how long after a wake-up ends the next one starts (clamped into the owner's OUROBOROS_BG_WAKEUP_MIN/MAX bounds; a wake-up calling this sets its own next one; a pending wake-up keeps its time; stored for later when consciousness is off).",
|
||||
"parameters": {"type": "object", "properties": {"seconds": {"type": "integer", "description": "Seconds from the end of a wake-up to the next one"}}, "required": ["seconds"]},
|
||||
"name": "set_next_wakeup", "description": "Choose the consciousness wake-up interval in seconds: how long after a wake-up ends the next one starts (clamped into the owner's OUROBOROS_BG_WAKEUP_MIN/MAX bounds; a wake-up calling this sets its own next one; a pending wake-up keeps its time; stored for later when consciousness is off).", "parameters": {"type": "object", "properties": {"seconds": {"type": "integer", "description": "Seconds from the end of a wake-up to the next one"}}, "required": ["seconds"]},
|
||||
}, _set_next_wakeup),
|
||||
ToolEntry("switch_model", {
|
||||
"name": "switch_model",
|
||||
|
|
|
|||
|
|
@ -260,11 +260,12 @@ def _promote_chat_to_task(
|
|||
) -> str:
|
||||
"""Route real work out of the conversation lane into a supervised pooled task.
|
||||
|
||||
Option B of the multi-project chat plane (v6.32.0): the conversation stays
|
||||
in the fast in-process lane; ANY substantial work spawns a first-class
|
||||
pooled task with a live card. The decision is the model's own structural
|
||||
tool call (BIBLE P5 — no keyword routing). Follow-up owner messages reach
|
||||
the running task through its owner-mailbox.
|
||||
The conversation stays in the fast in-process lane and keeps its own tools;
|
||||
the model promotes when an independent task is useful (SYSTEM.md, Decision
|
||||
Loop), and the promoted work runs as a first-class pooled task with a live
|
||||
card. The decision is the model's own structural tool call (BIBLE P5 — no
|
||||
keyword routing). Follow-up owner messages reach the running task through
|
||||
its owner-mailbox.
|
||||
|
||||
``title`` is a short human name the model coins for the card AT CREATION
|
||||
(no extra request, owner P1) — reused as the project name if this task is
|
||||
|
|
|
|||
|
|
@ -225,7 +225,8 @@ need it.
|
|||
broad fallbacks, silent catches, or shims lacking a concrete reachable
|
||||
failure mode. Mid-task I ask: am I solving the class or patching symptoms, am
|
||||
I adding surface area, am I still within my human's stated scope?
|
||||
- For long work I emit concise progress — what I learned and the next step —
|
||||
- Before long work I send my human one message saying what I will check and
|
||||
why; progress after that is concise — what I learned and the next step —
|
||||
explaining the thought, not narrating tool calls. After a repeatable
|
||||
workflow I capture the recipe: trigger, authoritative files and logs,
|
||||
commands, validation, known false leads.
|
||||
|
|
@ -238,15 +239,14 @@ need it.
|
|||
|
||||
### Outcome honesty
|
||||
|
||||
Every task lands on one of three honest tiers: **solved** (verified against the
|
||||
task's own surface), **best_effort** (a real partial deliverable with
|
||||
unverified or incomplete parts explicitly marked), or **blocked_with_evidence**
|
||||
(what blocked me, the exact evidence, and the next action someone could take).
|
||||
When a deadline, budget, or round limit forces finalization, I extract the best
|
||||
verified result I have and mark the gaps — an honest best_effort is an expected
|
||||
outcome, not a failure; returning emptiness is the only true failure mode. I
|
||||
never inflate a tier: claiming solved without verification is worse than an
|
||||
honest best_effort.
|
||||
Every task ends in one of three honest states, and I say which plainly:
|
||||
solved and verified against the task's own surface; partly done, with the
|
||||
real partial result handed over and its unverified or missing parts marked;
|
||||
or blocked, with what blocked me, the exact evidence and the next action
|
||||
someone could take. When a deadline, budget or round limit forces me to
|
||||
finish, I extract the best verified result I have and mark the gaps. An
|
||||
honest partial result is an expected ending; returning nothing is the only
|
||||
real failure mode. I never claim more than I verified.
|
||||
|
||||
## Capability Acquisition
|
||||
|
||||
|
|
|
|||
|
|
@ -7,7 +7,7 @@ Split out of ``supervisor/events.py`` at the module-size boundary;
|
|||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from typing import Any, Dict
|
||||
from typing import Any, Callable, Dict, Optional
|
||||
|
||||
from ouroboros.contracts.chat_id_policy import HIDDEN_CHAT_ID, WEB_UI_CHAT_ID
|
||||
|
||||
|
|
@ -186,11 +186,16 @@ class TurnEventQueue:
|
|||
capture rule of DEVELOPMENT.md). Wraps the turn's real queue and stamps
|
||||
the turn chat onto its own still-unaddressed task-scoped payloads."""
|
||||
|
||||
def __init__(self, inner: Any, task_id: Any, chat_id: Any, initiator: Any = "") -> None:
|
||||
def __init__(self, inner: Any, task_id: Any, chat_id: Any, initiator: Any = "",
|
||||
on_first_work: Optional[Callable[[], None]] = None) -> None:
|
||||
self._inner = inner
|
||||
self._task_id = str(task_id or "")
|
||||
self._chat_id = int(chat_id or 0)
|
||||
self._initiator = str(initiator or "")
|
||||
# Fired once, on the first frame this proxy stamps as WORK (below):
|
||||
# the lane hangs the turn namer on it, so a turn is named exactly
|
||||
# when its block becomes a task card and never for a greeting.
|
||||
self._on_first_work = on_first_work
|
||||
|
||||
def stamp(self, item: Any) -> Any:
|
||||
if isinstance(item, dict):
|
||||
|
|
@ -200,8 +205,8 @@ class TurnEventQueue:
|
|||
data["chat_id"] = self._chat_id
|
||||
# The lane fact rides the same events by the same rule: a live
|
||||
# tool or progress frame names its direct turn on arrival, so
|
||||
# the chat block never wears managed chrome (a Task title, a
|
||||
# conversion control) in the window before the census lists it.
|
||||
# the header pill keeps the census verdict (Thinking…) beside
|
||||
# the block before the census lists the turn; chrome never reads it.
|
||||
data.setdefault("_is_direct_chat", True)
|
||||
# The turn's origin label (a consciousness wake-up) rides the
|
||||
# same events, so a tool-only wake is filed and labelled from
|
||||
|
|
@ -220,6 +225,12 @@ class TurnEventQueue:
|
|||
and not data.get("routing_action")
|
||||
):
|
||||
data.setdefault("cancelable", True)
|
||||
if self._on_first_work is not None:
|
||||
callback, self._on_first_work = self._on_first_work, None
|
||||
try:
|
||||
callback()
|
||||
except Exception:
|
||||
log.debug("first-work callback failed for %s", self._task_id, exc_info=True)
|
||||
return item
|
||||
|
||||
def put(self, item: Any, *args: Any, **kwargs: Any) -> Any:
|
||||
|
|
|
|||
|
|
@ -395,10 +395,12 @@ def _admit_chat_task(
|
|||
_pool()._report_binding_failure(task["id"], pid, exc, path="direct_project_turn")
|
||||
if not task["text"]:
|
||||
task["text"] = "(image attached)" if image_data else ""
|
||||
# A direct turn is not named: it renders as an activity block, never as
|
||||
# a titled task card, and joins a Project only through the model's own
|
||||
# scope tools (owner decision 14=A). Managed promotes keep their
|
||||
# admission names (worker_promotion._admitted_suggested_name).
|
||||
# A Main turn is named lazily: the turn queue below fires the namer on
|
||||
# the first non-addressing tool call (owner decision Q7=A, 16.09), so a
|
||||
# greeting costs no naming call and a working turn gets a title as its
|
||||
# block becomes the task card. A Project-room turn is named by its room.
|
||||
# Managed promotes keep their admission names
|
||||
# (worker_promotion._admitted_suggested_name).
|
||||
# A consciousness wake-up derives its level's disabled_tools and mode cap
|
||||
# here, before the contract reads them (consciousness_authority).
|
||||
apply_consciousness_authority(task)
|
||||
|
|
@ -444,9 +446,17 @@ def _execute_chat_task(admitted: Dict[str, Any]) -> bool:
|
|||
# straight to the agent's event queue DURING handle_task) and its
|
||||
# returned events can be consumed after this registry entry is gone:
|
||||
# stamp the authoritative chat identity before handing them off.
|
||||
on_first_work = None
|
||||
if not task.get("project_id"):
|
||||
from ouroboros.project_naming import spawn_turn_namer
|
||||
|
||||
on_first_work = lambda: spawn_turn_namer( # noqa: E731
|
||||
_pool().DRIVE_ROOT, str(task["id"]), task["text"], broadcast=_broadcast_task_named,
|
||||
)
|
||||
turn_queue = _TurnEventQueue(
|
||||
_pool().get_event_q(), task["id"], chat_id,
|
||||
initiator=str((admitted.get("task_metadata") or {}).get("initiator") or ""),
|
||||
on_first_work=on_first_work,
|
||||
)
|
||||
prev_queue = getattr(agent, "_event_queue", None)
|
||||
agent._event_queue = turn_queue
|
||||
|
|
|
|||
File diff suppressed because one or more lines are too long
|
|
@ -528,7 +528,7 @@ def test_a_revision_names_its_pass_and_causes_not_the_aggregate_word(monkeypatch
|
|||
|
||||
A DEGRADED wave that DID feed an improvement capsule back printed
|
||||
"Task acceptance review: DEGRADED - improvement note fed back", where
|
||||
DEGRADED elsewhere means "no valid quorum". The row now names the pass being
|
||||
DEGRADED elsewhere means "no settled verdict". The row now names the pass being
|
||||
started and the causes the wave actually recorded; the degraded terminal
|
||||
keeps the same causes through the one shared clause.
|
||||
"""
|
||||
|
|
@ -581,7 +581,7 @@ def test_the_degraded_terminal_keeps_its_wording_and_its_causes(monkeypatch, tmp
|
|||
|
||||
assert trace["acceptance_decision"]["reason"] == "review_degraded"
|
||||
assert emitted[-1] == (
|
||||
"Task acceptance review: DEGRADED (no valid quorum; not recorded as PASS)."
|
||||
"Task acceptance review: DEGRADED (no settled verdict; not recorded as PASS)."
|
||||
" Causes: s1 window_exhausted"
|
||||
)
|
||||
|
||||
|
|
|
|||
|
|
@ -10,6 +10,7 @@ from types import SimpleNamespace
|
|||
import pytest
|
||||
|
||||
from ouroboros import loop, review_substrate
|
||||
from ouroboros.loop_acceptance_review import acceptance_run_pending
|
||||
from ouroboros.review_records import ReviewSlot
|
||||
from ouroboros.tools.registry import ToolRegistry
|
||||
from tests.test_loop_acceptance_gate import _seed_acceptance_root
|
||||
|
|
@ -231,6 +232,446 @@ def test_new_criterion_same_answer_gets_one_new_panel_and_keeps_prior_request(fu
|
|||
assert host[0]["subject_hash"] != host[1]["subject_hash"]
|
||||
from ouroboros.task_results import project_task_acceptance_review_capacity
|
||||
assert project_task_acceptance_review_capacity(f.ctx, task_id=f.ctx.task_id)["claimed_cycles"] == 2
|
||||
# The superseded panel was paid for: its verdicts must have been read, not stranded.
|
||||
assert not acceptance_run_pending(host[0])
|
||||
assert host[0]["actors"][0]["parsed"]["verdict"] == "PASS"
|
||||
|
||||
|
||||
def test_reauthored_answer_collects_the_stranded_panel_before_paying_again(full_loop, monkeypatch):
|
||||
"""A re-authored answer moves the paid identity, so the free-replay lookup no
|
||||
longer sees the running panel. Its verdicts were bought; the host collects
|
||||
them at $0 before assembling evidence for, or refusing, anything new."""
|
||||
f = full_loop
|
||||
reauthored = ANSWER + " Budget: $12."
|
||||
recorded = []
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5)
|
||||
# The paid panel settles while Main is still working; nothing has
|
||||
# read its verdicts yet and the settlement wake is the only signal.
|
||||
f.release.set()
|
||||
with f.condition:
|
||||
assert f.condition.wait_for(lambda: f.settled_count >= 1, timeout=10)
|
||||
return {"content": "", "tool_calls": [call("send_user_message", {"text": "Still writing the report."}, "status")]}, 0.0
|
||||
if f.model_step == 3:
|
||||
pending = [r for r in f.ctx._execution_trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
recorded.append(pending[0]["actors"][0]["operation_id"])
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": reauthored}, "reauthored-review")]}, 0.0
|
||||
assert f.model_step < 8, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == reauthored
|
||||
# The re-authored subject still buys its own panel: no new refusal gate.
|
||||
assert len(f.review_sends) == len(set(f.review_sends)) == 2
|
||||
host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
assert len(host) == 2
|
||||
assert not acceptance_run_pending(host[0]), host[0]["actors"]
|
||||
# The EXACT recorded producer advanced, not a re-run.
|
||||
assert host[0]["actors"][0]["parsed"]["verdict"] == "PASS"
|
||||
assert host[0]["actors"][0]["operation_id"] == recorded[0]
|
||||
# The collected verdicts reached the next panel's dialogue history.
|
||||
history = f.review_requests[1].evidence["acceptance_dialogue_history"]
|
||||
assert history and history[0]["aggregate_signal"] == "PASS"
|
||||
# And the published projection, instead of a transport error on a row nobody read.
|
||||
from ouroboros.task_results import load_task_result, project_task_acceptance_review_capacity
|
||||
panels = {p["panel_id"]: p for p in
|
||||
load_task_result(f.ctx.drive_root, f.ctx.task_id)["review_projection"]["panels"]}
|
||||
collected = panels[host[0]["panel_id"]]
|
||||
assert collected["actors"][0]["transport_status"] == "success"
|
||||
assert collected["actors"][0]["parse_status"] == "valid"
|
||||
# Collection is free: exactly the two dispatched panels were ever claimed.
|
||||
assert project_task_acceptance_review_capacity(f.ctx, task_id=f.ctx.task_id)["claimed_cycles"] == 2
|
||||
|
||||
|
||||
def test_reauthored_answer_on_a_one_cycle_install_is_refused_after_its_panel_was_collected(full_loop, monkeypatch):
|
||||
"""The live incident shape (task 4525349b, OUROBOROS_REVIEW_MAX_CYCLES=1): the
|
||||
paid panel settles while Main is still working, Main re-authors, and the
|
||||
one-cycle cap refuses a second panel. The refusal is the product rule (owner
|
||||
14A) and stays; what the reconcile changes is that the paid verdicts are read
|
||||
BEFORE it — recorded as settled, present in the next evidence's dialogue
|
||||
history — and the decision carries no prose rationale claiming a quorum
|
||||
failure that never happened."""
|
||||
f = full_loop
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "1")
|
||||
reauthored = ANSWER + " Budget: $12."
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5)
|
||||
f.release.set()
|
||||
with f.condition:
|
||||
assert f.condition.wait_for(lambda: f.settled_count >= 1, timeout=10)
|
||||
return {"content": "", "tool_calls": [call("send_user_message", {"text": "Still writing the report."}, "status")]}, 0.0
|
||||
if f.model_step == 3:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": reauthored}, "reauthored-review")]}, 0.0
|
||||
assert f.model_step < 8, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == reauthored
|
||||
# One paid dispatch only: the cap refused the re-authored subject.
|
||||
assert len(f.review_sends) == 1
|
||||
host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
assert len(host) == 2
|
||||
# The paid panel was read at $0 before the refusal, not left as pending stubs.
|
||||
assert not acceptance_run_pending(host[0]), host[0]["actors"]
|
||||
assert host[0]["actors"][0]["parsed"]["verdict"] == "PASS"
|
||||
assert any("review_cycles_exhausted" in str(reason) for reason in host[1].get("degraded_reasons") or [])
|
||||
decision = trace.get("acceptance_decision") or {}
|
||||
assert decision.get("reason") == "review_degraded"
|
||||
assert "rationale" not in decision
|
||||
from ouroboros.task_results import project_task_acceptance_review_capacity
|
||||
assert project_task_acceptance_review_capacity(f.ctx, task_id=f.ctx.task_id)["claimed_cycles"] == 1
|
||||
|
||||
|
||||
def _terminal_record(trace):
|
||||
"""The host's own fold of the trace, as the terminal row and the card read it."""
|
||||
from ouroboros import outcomes
|
||||
|
||||
review = outcomes._review_axis(trace)
|
||||
return {"status": "completed", "reason_code": "final_message",
|
||||
"outcome_axes": {"execution": {"status": "ok"}, "review": review,
|
||||
"objective": outcomes._objective_axis(review)}}
|
||||
|
||||
|
||||
def test_a_rewritten_answer_delivers_under_the_running_panel_instead_of_buying_one(full_loop, monkeypatch):
|
||||
"""The live incident shape (task 4525349b, cap 1) under owner D4=A and fork 1=B:
|
||||
Main nominates, the reviewers are still reading when it rewrites the answer
|
||||
through the delivery control (round 7 of the incident). The rewrite is a
|
||||
DELIVERY, not a nomination: it buys nothing and is refused nothing; the host
|
||||
waits for the panel it already paid for, and when that panel PASSES the earlier
|
||||
revision the task is accepted on the reviewers' word and the row says so."""
|
||||
from ouroboros.project_dialogue import _completion_verdict, outcome_phase
|
||||
|
||||
f = full_loop
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "1")
|
||||
reauthored = ANSWER + " Budget: $12."
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5) and not f.release.is_set(), "the panel must still be running"
|
||||
observation = f.ctx._acceptance_observation
|
||||
return {"content": json.dumps({"delivery_control": "replace", "full_answer": reauthored,
|
||||
"acceptance_subject": {"owner_source_sha256": observation["owner_source_sha256"]}})}, 0.0
|
||||
assert f.model_step < 6, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == reauthored
|
||||
# The rewrite bought nothing and was refused nothing: one paid panel, no synthetic refusal run.
|
||||
assert len(f.review_sends) == 1
|
||||
host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
assert len(host) == 1 and not acceptance_run_pending(host[0]), host
|
||||
assert host[0]["actors"][0]["parsed"]["verdict"] == "PASS"
|
||||
assert not any("review_cycles_exhausted" in str(reason) for reason in host[0].get("degraded_reasons") or [])
|
||||
# The host waited for the panel it had paid for (blocking enforcement in this fixture).
|
||||
assert f.waits and f.waits[0]["reason"] == "review"
|
||||
assert any("holding the answer for its verdict" in line for line in f.progress), f.progress
|
||||
decision = trace["acceptance_decision"]
|
||||
assert decision["status"] == "accepted" and decision["reason"] == "previous_revision_accepted"
|
||||
assert decision["reviewer_signal"] == "PASS" and "rationale" not in decision
|
||||
from ouroboros.task_results import project_task_acceptance_review_capacity
|
||||
assert project_task_acceptance_review_capacity(f.ctx, task_id=f.ctx.task_id)["claimed_cycles"] == 1
|
||||
record = _terminal_record(trace)
|
||||
assert outcome_phase(record, {}) == "done", record["outcome_axes"]
|
||||
assert _completion_verdict(record, {}) == (
|
||||
"The reviewers approved the earlier version of this answer; it changed before they finished."
|
||||
)
|
||||
|
||||
|
||||
def test_a_rejected_earlier_revision_is_not_a_verdict_on_the_rewrite(full_loop, monkeypatch):
|
||||
"""Fork 1: a FAIL on the earlier revision must not paint the rewritten,
|
||||
unreviewed answer red. The settled negative verdict hands the delivery to the
|
||||
ordinary path: at cap 1 that is the typed capacity refusal, so the row says no
|
||||
verdict was established for THIS answer, and the task ends with warnings."""
|
||||
from ouroboros.project_dialogue import _completion_verdict, outcome_phase
|
||||
|
||||
f = full_loop
|
||||
f.reviewer_verdict = "FAIL"
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "1")
|
||||
reauthored = ANSWER + " Budget: $12."
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5) and not f.release.is_set()
|
||||
observation = f.ctx._acceptance_observation
|
||||
return {"content": json.dumps({"delivery_control": "replace", "full_answer": reauthored,
|
||||
"acceptance_subject": {"owner_source_sha256": observation["owner_source_sha256"]}})}, 0.0
|
||||
assert f.model_step < 6, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == reauthored and len(f.review_sends) == 1
|
||||
host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
assert [r.get("aggregate_signal") for r in host] == ["FAIL", "DEGRADED"], host
|
||||
assert host[0]["superseded_by_revision"] and not host[1].get("superseded_by_revision")
|
||||
assert trace["acceptance_decision"]["reason"] == "review_degraded"
|
||||
record = _terminal_record(trace)
|
||||
assert outcome_phase(record, {}) == "warn", record["outcome_axes"]
|
||||
assert _completion_verdict(record, {}) == "No reviewer verdict was established for this answer."
|
||||
|
||||
|
||||
def _advisory(monkeypatch):
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_ENFORCEMENT", "advisory")
|
||||
|
||||
|
||||
def test_a_conscious_finish_releases_the_answer_while_the_panel_runs(full_loop, monkeypatch):
|
||||
"""Owner D4=A point 3: under advisory enforcement Main chooses explicitly.
|
||||
``"pending_review":"finish"`` on the delivery control delivers now, without a
|
||||
park; the verdict reaches Main as advice when it settles."""
|
||||
f = full_loop
|
||||
_advisory(monkeypatch)
|
||||
f.ctx.owner_wait_callback = lambda *_a, **_kw: pytest.fail("a conscious finish must not park")
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "nominate")]}, 0.0
|
||||
assert f.model_step == 2 and f.entered.wait(5) and not f.release.is_set()
|
||||
control = json.loads(keep(f)["content"])
|
||||
assert '"pending_review":"finish"' in str(messages), "the choice is offered while the panel runs"
|
||||
return {"content": json.dumps({**control, "pending_review": "finish"})}, 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and len(f.review_sends) == 1 and f.waits == []
|
||||
assert trace["acceptance_decision"]["reason"] == "author_finish"
|
||||
assert trace["review_decision"]["review_pending"] is True
|
||||
from ouroboros.task_results import project_task_acceptance_review_capacity
|
||||
assert project_task_acceptance_review_capacity(f.ctx, task_id=f.ctx.task_id)["claimed_cycles"] == 1
|
||||
|
||||
|
||||
def test_a_panel_that_settles_after_the_loop_exited_is_attached_through_the_remembered_trace(full_loop, monkeypatch):
|
||||
"""Fable review round 2 (CRITICAL): the loop exit restores the context's
|
||||
``_execution_trace`` to its pre-loop value, so a wave settling after the
|
||||
turn ended found no trace and the late supplement never fired in production.
|
||||
The pending panel now remembers its trace; the settlement thread reads it
|
||||
back after the loop is gone and announces once in the task's room."""
|
||||
from ouroboros.task_results import load_task_result, write_task_result
|
||||
|
||||
f = full_loop
|
||||
_advisory(monkeypatch)
|
||||
f.ctx.owner_wait_callback = lambda *_a, **_kw: pytest.fail("a conscious finish must not park")
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "nominate")]}, 0.0
|
||||
assert f.model_step == 2 and f.entered.wait(5) and not f.release.is_set()
|
||||
control = json.loads(keep(f)["content"])
|
||||
return {"content": json.dumps({**control, "pending_review": "finish"})}, 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and trace["review_decision"]["review_pending"] is True
|
||||
assert getattr(f.ctx, "_execution_trace", None) is None, "the loop exit detached the live trace"
|
||||
# The pipeline seals the task before the straggler answers.
|
||||
write_task_result(f.ctx.drive_root, f.ctx.task_id, "completed", chat_id=1, result=ANSWER)
|
||||
f.release.set()
|
||||
with f.condition:
|
||||
assert f.condition.wait_for(lambda: f.settled_count >= 1, timeout=10)
|
||||
events = list(f.events.queue)
|
||||
rows = [e for e in events if e.get("system_type") == "acceptance_late_settlement"]
|
||||
assert len(rows) == 1, [e.get("type") for e in events]
|
||||
assert rows[0]["task_id"] == f.ctx.task_id and rows[0]["chat_id"] == 1
|
||||
assert rows[0]["text"].startswith("Reviewers later passed this answer. They reviewed the answer that was delivered.")
|
||||
assert "- acceptance-one: PASS" in rows[0]["text"]
|
||||
stored = load_task_result(f.ctx.drive_root, f.ctx.task_id)
|
||||
assert stored["status"] == "completed"
|
||||
actor = stored["review_projection"]["panels"][-1]["actors"][0]
|
||||
assert actor["transport_status"] == "success" and actor["parse_status"] == "valid"
|
||||
assert len(f.review_sends) == 1, "the supplement bought nothing"
|
||||
assert not getattr(f.ctx, "_acceptance_settlement_traces", {}), "a settled wave releases its remembered trace"
|
||||
|
||||
|
||||
def test_a_rejected_earlier_revision_buys_a_panel_on_the_rewrite_when_the_cap_allows(full_loop, monkeypatch):
|
||||
"""The other half of fork 1: after a FAIL on the earlier revision the ordinary
|
||||
path decides, and with review cycles left it buys a real panel on the
|
||||
rewritten bytes — the verdict that then accepts the task is about the
|
||||
delivered text, not the old one."""
|
||||
f = full_loop
|
||||
f.reviewer_verdict = "FAIL"
|
||||
reauthored = ANSWER + " Budget: $12."
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5) and not f.release.is_set()
|
||||
observation = f.ctx._acceptance_observation
|
||||
return {"content": json.dumps({"delivery_control": "replace", "full_answer": reauthored,
|
||||
"acceptance_subject": {"owner_source_sha256": observation["owner_source_sha256"]}})}, 0.0
|
||||
# The park released panel 1 (FAIL) before this round; the rewrite addresses the notes.
|
||||
f.reviewer_verdict = "PASS"
|
||||
assert f.model_step < 8, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == reauthored
|
||||
assert len(f.review_sends) == 2, "the rewrite got its own panel"
|
||||
host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
assert [r.get("aggregate_signal") for r in host] == ["FAIL", "PASS"], host
|
||||
assert host[0]["superseded_by_revision"] and not host[1].get("superseded_by_revision")
|
||||
assert f.review_requests[1].subject == reauthored
|
||||
assert trace["acceptance_decision"]["status"] == "accepted"
|
||||
assert trace["acceptance_decision"]["reason"] in {"clean_pass", "clean_pass_obligations_closed"}
|
||||
|
||||
|
||||
def test_an_older_fail_never_outvotes_the_pass_that_accepted_the_task(full_loop, monkeypatch):
|
||||
"""Astra review round 4: panel A rejects the first draft, Main re-nominates and
|
||||
panel B passes the second, Main rewrites once more under B. Both runs end up
|
||||
superseded; the decision names B. The review axis must read B alone — the
|
||||
old FAIL is audit evidence, not a vote against the accepted answer."""
|
||||
from ouroboros.project_dialogue import outcome_phase
|
||||
|
||||
f = full_loop
|
||||
f.reviewer_verdict = "FAIL"
|
||||
second = ANSWER + " Budget: $12."
|
||||
third = second + " Timeline: two weeks."
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5)
|
||||
f.release.set()
|
||||
with f.condition:
|
||||
assert f.condition.wait_for(lambda: f.settled_count >= 1, timeout=10)
|
||||
f.reviewer_verdict = "PASS"
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": second}, "second-review")]}, 0.0
|
||||
if f.model_step == 3:
|
||||
observation = f.ctx._acceptance_observation
|
||||
return {"content": json.dumps({"delivery_control": "replace", "full_answer": third,
|
||||
"acceptance_subject": {"owner_source_sha256": observation["owner_source_sha256"]}})}, 0.0
|
||||
assert f.model_step < 8, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == third and len(f.review_sends) == 2
|
||||
host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
assert [r.get("aggregate_signal") for r in host] == ["FAIL", "PASS"], host
|
||||
assert all(r.get("superseded_by_revision") for r in host)
|
||||
decision = trace["acceptance_decision"]
|
||||
assert decision["status"] == "accepted" and decision["reason"] == "previous_revision_accepted"
|
||||
assert decision["reviewed_panel_id"] == host[1]["panel_id"]
|
||||
record = _terminal_record(trace)
|
||||
assert record["outcome_axes"]["review"]["aggregate_signals"] == ["PASS"], record["outcome_axes"]["review"]
|
||||
assert outcome_phase(record, {}) == "done", record["outcome_axes"]
|
||||
# The verification ledger agrees: the older FAIL is superseded evidence, not a live failure.
|
||||
from ouroboros._outcome_receipts import review_run_ledger_status, select_current_review_runs
|
||||
selection = select_current_review_runs(trace["review_runs"], delivery_candidate=trace.get("delivery_candidate"),
|
||||
review_decision=trace.get("review_decision"))
|
||||
assert review_run_ledger_status(host[0], selection) == ("superseded", True)
|
||||
|
||||
|
||||
def test_an_owner_followup_acknowledged_through_the_control_sets_the_panel_aside(full_loop, monkeypatch):
|
||||
"""Astra review round 5: the owner changes the requirements while the panel
|
||||
runs; Main reads the message, acknowledges its source on the delivery control
|
||||
and rewrites. The rewrite is NOT a delivery under the old panel (it judged the
|
||||
old premises): the ordinary path buys a panel on the new answer and the old
|
||||
PASS never accepts it."""
|
||||
f = full_loop
|
||||
followup = "Also add a timeline section to the report."
|
||||
rewritten = ANSWER + " Timeline: two weeks."
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5) and not f.release.is_set()
|
||||
f.incoming.put(followup)
|
||||
return {"content": "", "tool_calls": [call("send_user_message", {"text": "Adding the timeline."}, "ack")]}, 0.0
|
||||
if f.model_step == 3:
|
||||
assert followup in str(messages), "the owner follow-up reached the model"
|
||||
observation = f.ctx._acceptance_observation
|
||||
return {"content": json.dumps({"delivery_control": "replace", "full_answer": rewritten,
|
||||
"acceptance_subject": {"owner_source_sha256": observation["owner_source_sha256"]}})}, 0.0
|
||||
assert f.model_step < 8, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == rewritten
|
||||
assert len(f.review_sends) == 2, ("the rewrite for the new premises got its own panel", f.progress)
|
||||
assert f.review_requests[1].subject == rewritten
|
||||
host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"]
|
||||
assert host[0]["superseded_by_revision"] and host[0]["owner_source_sha256"] != host[1]["owner_source_sha256"]
|
||||
assert trace["acceptance_decision"]["reason"] != "previous_revision_accepted"
|
||||
assert trace["acceptance_decision"]["status"] == "accepted"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("install", ["advisory_default", "blocking_finish", "cyber_pro"])
|
||||
def test_waiting_is_the_default_and_blocking_enforcement_never_offers_the_choice(full_loop, monkeypatch, install):
|
||||
"""Waiting needs no key; blocking enforcement waits whatever the model says and
|
||||
is never offered the choice; Cyber Pro keeps its own rule (Main's final response
|
||||
is its decision) and is not offered the choice either."""
|
||||
f = full_loop
|
||||
if install == "advisory_default":
|
||||
_advisory(monkeypatch)
|
||||
if install == "cyber_pro":
|
||||
_advisory(monkeypatch)
|
||||
monkeypatch.setattr("ouroboros.config.get_runtime_mode", lambda: "cyber_pro")
|
||||
f.ctx.owner_wait_callback = lambda *_a, **_kw: pytest.fail("Cyber Pro never parks on a review")
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "nominate")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5)
|
||||
control = json.loads(keep(f)["content"])
|
||||
if install == "blocking_finish":
|
||||
control["pending_review"] = "finish"
|
||||
return {"content": json.dumps(control)}, 0.0
|
||||
assert f.model_step < 6, f.progress
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and len(f.review_sends) == 1
|
||||
offered = any('"pending_review":"finish"' in str(inputs) for inputs in f.model_inputs)
|
||||
if install == "advisory_default":
|
||||
assert offered and f.waits and f.waits[0]["reason"] == "review"
|
||||
assert trace["acceptance_decision"]["status"] == "accepted"
|
||||
elif install == "blocking_finish":
|
||||
assert not offered and f.waits and f.waits[0]["reason"] == "review"
|
||||
assert trace["acceptance_decision"]["status"] == "accepted"
|
||||
else:
|
||||
assert not offered and f.waits == []
|
||||
assert trace["acceptance_decision"]["reason"] == "author_finish"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("pending_first", [False, True])
|
||||
|
|
|
|||
|
|
@ -107,3 +107,462 @@ def test_review_park_is_not_a_question_and_preserves_operation(tmp_path):
|
|||
def test_unknown_custody_is_not_an_active_wait():
|
||||
assert not acceptance_run_pending({"actors": [{"operation_state": "custody_lost", "late_result_pending": True}]})
|
||||
assert acceptance_run_pending({"actors": [{"operation_state": "pending_dispatch"}]})
|
||||
|
||||
|
||||
def _host_run(**overrides):
|
||||
run = {"authority": "host_root", "request": {"surface": "task_acceptance", "retry_key": "subject"},
|
||||
"slot_roster": [{"slot_id": "one", "model": "m", "route": "api_chat"}],
|
||||
"actors": [{"operation_state": "pending_dispatch"}]}
|
||||
run.update(overrides)
|
||||
return run
|
||||
|
||||
|
||||
def test_a_settled_acceptance_run_is_never_collected_twice(tmp_path, monkeypatch):
|
||||
"""``acceptance_run_pending`` is the whole idempotency guard: a settled,
|
||||
custody-lost or agent-tool run is never re-read, so no marker field exists."""
|
||||
from ouroboros import review_dispatch
|
||||
|
||||
monkeypatch.setattr(review_dispatch, "collect_task_acceptance_run",
|
||||
lambda *a, **k: pytest.fail("a run that is not pending was collected again"))
|
||||
trace = {"review_runs": [
|
||||
_host_run(actors=[{"operation_state": "settled", "parsed": {"verdict": "PASS"}}]),
|
||||
_host_run(actors=[{"operation_state": "custody_lost", "late_result_pending": True}]),
|
||||
_host_run(authority="agent_tool"),
|
||||
_host_run(slot_roster=[]),
|
||||
_host_run(request=None),
|
||||
]}
|
||||
advanced = review_dispatch.reconcile_pending_acceptance_runs(
|
||||
trace, drive_root=tmp_path, usage_ctx=SimpleNamespace(),
|
||||
)
|
||||
assert advanced == 0
|
||||
|
||||
|
||||
def test_an_uncollectable_pending_run_is_left_alone_and_never_raises(tmp_path, monkeypatch):
|
||||
"""Fail-soft like ``plan_review_collect.collect_before_gate``: the pass
|
||||
continues, the run stays pending, and nothing dispatches."""
|
||||
from ouroboros import review_dispatch
|
||||
|
||||
monkeypatch.setattr("ouroboros.review_substrate.run_review_request",
|
||||
lambda *a, **k: pytest.fail("a stranded run bought another review"))
|
||||
for error in (ValueError("recorded acceptance roster is unavailable"),
|
||||
KeyError("request"), OSError("custody store unavailable"), TimeoutError("slow")):
|
||||
def raising(*_a, _error=error, **_k):
|
||||
raise _error
|
||||
|
||||
monkeypatch.setattr(review_dispatch, "collect_task_acceptance_run", raising)
|
||||
pending = _host_run()
|
||||
advanced = review_dispatch.reconcile_pending_acceptance_runs(
|
||||
{"review_runs": [pending]}, drive_root=tmp_path, usage_ctx=SimpleNamespace(),
|
||||
)
|
||||
assert advanced == 0 and acceptance_run_pending(pending)
|
||||
|
||||
|
||||
def _mailbox_rows(root, task_id):
|
||||
from ouroboros.owner_mailbox import _mailbox_path
|
||||
|
||||
path = _mailbox_path(root, task_id)
|
||||
if not path.exists():
|
||||
return []
|
||||
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
|
||||
|
||||
|
||||
def test_the_settlement_wake_carries_the_reviewers_own_verdicts(tmp_path):
|
||||
"""A review is advice for its author (owner D4=A): the wake IS the advice,
|
||||
not a pointer to a collect verb the model reaches only by not moving on."""
|
||||
from ouroboros.acceptance_settlement import announce_acceptance_settlement
|
||||
|
||||
announce_acceptance_settlement(
|
||||
SimpleNamespace(drive_root=tmp_path),
|
||||
SimpleNamespace(retry_key="acceptance-subject-one", task_id="root"),
|
||||
{"slots": {"one": "ok", "two": ""}, "total": 2,
|
||||
"verdicts": {"one": {"verdict": "PASS", "note": "Budget section is complete."}}},
|
||||
)
|
||||
rows = _mailbox_rows(tmp_path, "root")
|
||||
assert len(rows) == 1 and rows[0]["provenance"] == "system"
|
||||
text = rows[0]["text"]
|
||||
assert "acceptance-subject-one" in text and "1 of 2 reviewer slot(s)" in text
|
||||
assert "advice for you, not a signature" in text
|
||||
assert "- one: PASS — Budget section is complete." in text and "- two: pending" in text
|
||||
assert "keep control" not in text
|
||||
|
||||
|
||||
class _SlotModel:
|
||||
"""One chat per roster slot, each held behind its own event."""
|
||||
|
||||
def __init__(self, gates, verdicts):
|
||||
self.gates, self.verdicts, self.calls = gates, verdicts, []
|
||||
|
||||
def chat(self, **kwargs):
|
||||
self.calls.append(kwargs)
|
||||
model = str(kwargs.get("model") or "")
|
||||
assert self.gates[model].wait(10), f"fixture did not release {model}"
|
||||
return ({"content": json.dumps({"verdict": self.verdicts[model], "findings": [],
|
||||
"summary": f"{model} says {self.verdicts[model]}"})},
|
||||
{"prompt_tokens": 5, "completion_tokens": 2})
|
||||
|
||||
|
||||
def _released_wave(tmp_path, ctx, *, slots, model, min_successful_slots=1, task_id="root",
|
||||
retry_key="acceptance-subject-one"):
|
||||
request = ReviewRequest(surface="task_acceptance", task_id=task_id, goal="goal",
|
||||
subject="complete result", evidence={"requirement": "exact"},
|
||||
policy={"min_successful_slots": min_successful_slots},
|
||||
retry_key=retry_key, drain_deadline=time.monotonic())
|
||||
return run_review_request(request, slots=slots, drive_root=tmp_path, usage_ctx=ctx, llm=model)
|
||||
|
||||
|
||||
def test_the_quorum_wake_carries_the_reviewers_verdicts_before_the_last_slot_settles(tmp_path, monkeypatch):
|
||||
"""Fable roast round 1 / owner D4=A: a two-slot wave with quorum 1 wakes Main
|
||||
when the first reviewer answers, then again when the straggler settles; the
|
||||
straggler's wake lists every slot."""
|
||||
settled = threading.Condition()
|
||||
count = {"n": 0}
|
||||
original_settle = __import__("ouroboros.review_custody", fromlist=["_settle_review_attempt"])._settle_review_attempt
|
||||
|
||||
def settle(*args, **kwargs):
|
||||
try:
|
||||
return original_settle(*args, **kwargs)
|
||||
finally:
|
||||
with settled:
|
||||
count["n"] += 1
|
||||
settled.notify_all()
|
||||
|
||||
monkeypatch.setattr("ouroboros.review_custody._settle_review_attempt", settle)
|
||||
gates = {"model/a": threading.Event(), "model/b": threading.Event()}
|
||||
model = _SlotModel(gates, {"model/a": "PASS", "model/b": "FAIL"})
|
||||
ctx = SimpleNamespace(task_id="root", task_attempt=1, drive_root=tmp_path, budget_drive_root=tmp_path,
|
||||
task_metadata={}, pending_events=[], event_queue=None)
|
||||
slots = [ReviewSlot(slot_id="a", model="model/a", effort="high", timeout_sec=20),
|
||||
ReviewSlot(slot_id="b", model="model/b", effort="high", timeout_sec=20)]
|
||||
try:
|
||||
first = _released_wave(tmp_path, ctx, slots=slots, model=model)
|
||||
assert acceptance_run_pending(first) and _mailbox_rows(tmp_path, "root") == []
|
||||
gates["model/a"].set()
|
||||
with settled:
|
||||
assert settled.wait_for(lambda: count["n"] >= 1, timeout=10)
|
||||
rows = _mailbox_rows(tmp_path, "root")
|
||||
assert len(rows) == 1, rows
|
||||
assert "1 of 2 reviewer slot(s)" in rows[0]["text"]
|
||||
assert "- a: PASS — model/a says PASS" in rows[0]["text"] and "- b: pending" in rows[0]["text"]
|
||||
gates["model/b"].set()
|
||||
with settled:
|
||||
assert settled.wait_for(lambda: count["n"] >= 2, timeout=10)
|
||||
rows = _mailbox_rows(tmp_path, "root")
|
||||
assert len(rows) == 2, rows
|
||||
assert "2 of 2 reviewer slot(s)" in rows[1]["text"]
|
||||
assert "- b: FAIL — model/b says FAIL" in rows[1]["text"]
|
||||
assert len(model.calls) == 2, "the wakes bought nothing"
|
||||
finally:
|
||||
for gate in gates.values():
|
||||
gate.set()
|
||||
with settled:
|
||||
settled.wait_for(lambda: count["n"] >= 2, timeout=10)
|
||||
|
||||
|
||||
def _terminal_ctx(tmp_path, *, task_id, retry_key="acceptance-subject-one"):
|
||||
"""A worker context whose task already ended, with the wave's run in its trace."""
|
||||
import queue
|
||||
|
||||
from ouroboros.task_results import write_task_result
|
||||
|
||||
write_task_result(tmp_path, task_id, "completed", chat_id=1, result="The delivered answer.")
|
||||
return SimpleNamespace(task_id=task_id, task_attempt=1, drive_root=tmp_path, budget_drive_root=tmp_path,
|
||||
task_metadata={}, pending_events=[], event_queue=queue.Queue(),
|
||||
_execution_trace={"review_runs": []}, _retry_key=retry_key)
|
||||
|
||||
|
||||
def test_a_panel_that_settles_after_the_task_ended_is_attached_and_announced_once(tmp_path, monkeypatch):
|
||||
"""Owner fork 2=A: one System row in the task's room, whatever the verdict;
|
||||
the projection reads the collected verdicts; nothing wakes a model."""
|
||||
from ouroboros.task_results import load_task_result
|
||||
|
||||
settled = threading.Event()
|
||||
original_settle = __import__("ouroboros.review_custody", fromlist=["_settle_review_attempt"])._settle_review_attempt
|
||||
|
||||
def settle(*args, **kwargs):
|
||||
try:
|
||||
return original_settle(*args, **kwargs)
|
||||
finally:
|
||||
settled.set()
|
||||
|
||||
monkeypatch.setattr("ouroboros.review_custody._settle_review_attempt", settle)
|
||||
gates = {"model/a": threading.Event()}
|
||||
model = _SlotModel(gates, {"model/a": "PASS"})
|
||||
ctx = _terminal_ctx(tmp_path, task_id="late-root")
|
||||
slot = ReviewSlot(slot_id="a", model="model/a", effort="high", timeout_sec=20)
|
||||
try:
|
||||
first = _released_wave(tmp_path, ctx, slots=[slot], model=model, task_id="late-root")
|
||||
assert acceptance_run_pending(first)
|
||||
run = {**json.loads(json.dumps(dataclasses.asdict(first))), "authority": "host_root",
|
||||
"panel_id": "panel_1", "binding_hash": "binding-one", "candidate_hash": "c1",
|
||||
"superseded_by_revision": True}
|
||||
ctx._execution_trace["review_runs"].append(run)
|
||||
gates["model/a"].set()
|
||||
assert settled.wait(10)
|
||||
events = []
|
||||
while not any(e.get("system_type") == "acceptance_late_settlement" for e in events):
|
||||
events.append(ctx.event_queue.get(timeout=10))
|
||||
finally:
|
||||
gates["model/a"].set()
|
||||
assert settled.wait(10)
|
||||
assert _mailbox_rows(tmp_path, "late-root") == [], "a terminal task has nobody to wake"
|
||||
rows = [e for e in events if e.get("system_type") == "acceptance_late_settlement"]
|
||||
# Exactly one chat row; the other queue entries are the typed operation facts
|
||||
# every settlement emits, never a second message and never a model turn.
|
||||
assert len(rows) == 1 and [e for e in events if e.get("type") == "send_message"] == rows, events
|
||||
event = rows[0]
|
||||
assert event["type"] == "send_message" and event["role"] == "system"
|
||||
assert event["chat_id"] == 1 and event["task_id"] == "late-root"
|
||||
assert event["text"].startswith("Reviewers later passed this answer. They reviewed the earlier version")
|
||||
assert "- a: PASS — model/a says PASS" in event["text"]
|
||||
assert event["delivery_id"] == "acceptance-late:acceptance-subject-one"
|
||||
assert not acceptance_run_pending(run) and run["actors"][0]["parsed"]["verdict"] == "PASS"
|
||||
stored = load_task_result(tmp_path, "late-root")
|
||||
assert stored["status"] == "completed", "the supplement never moves a terminal status"
|
||||
actor = stored["review_projection"]["panels"][0]["actors"][0]
|
||||
assert actor["transport_status"] == "success" and actor["parse_status"] == "valid"
|
||||
assert len(model.calls) == 1, "collection is free"
|
||||
# A second settlement of the same wave finds nothing to reconcile and announces nothing.
|
||||
from ouroboros.acceptance_settlement import attach_late_acceptance_settlement
|
||||
|
||||
assert attach_late_acceptance_settlement(
|
||||
ctx, SimpleNamespace(retry_key="acceptance-subject-one", task_id="late-root"),
|
||||
{"slots": {"a": "ok"}, "total": 1}, result=stored) is False
|
||||
assert ctx.event_queue.empty()
|
||||
|
||||
|
||||
def test_a_terminal_task_gets_one_row_at_completion_not_at_quorum(tmp_path, monkeypatch):
|
||||
"""Fable review round 2: a quorum settlement on an already-terminal task must
|
||||
not announce a half-settled wave (the straggler's verdict would never reach
|
||||
the chat, deduped behind the same delivery id). Nothing is reconciled until
|
||||
every slot answered, so the one row carries every reviewer."""
|
||||
settled = threading.Condition()
|
||||
count = {"n": 0}
|
||||
original_settle = __import__("ouroboros.review_custody", fromlist=["_settle_review_attempt"])._settle_review_attempt
|
||||
|
||||
def settle(*args, **kwargs):
|
||||
try:
|
||||
return original_settle(*args, **kwargs)
|
||||
finally:
|
||||
with settled:
|
||||
count["n"] += 1
|
||||
settled.notify_all()
|
||||
|
||||
monkeypatch.setattr("ouroboros.review_custody._settle_review_attempt", settle)
|
||||
gates = {"model/a": threading.Event(), "model/b": threading.Event()}
|
||||
model = _SlotModel(gates, {"model/a": "PASS", "model/b": "PASS"})
|
||||
ctx = _terminal_ctx(tmp_path, task_id="late-two")
|
||||
slots = [ReviewSlot(slot_id="a", model="model/a", effort="high", timeout_sec=20),
|
||||
ReviewSlot(slot_id="b", model="model/b", effort="high", timeout_sec=20)]
|
||||
try:
|
||||
first = _released_wave(tmp_path, ctx, slots=slots, model=model, task_id="late-two")
|
||||
run = {**json.loads(json.dumps(dataclasses.asdict(first))), "authority": "host_root",
|
||||
"panel_id": "panel_1", "binding_hash": "binding-one", "candidate_hash": "c1"}
|
||||
ctx._execution_trace["review_runs"].append(run)
|
||||
gates["model/a"].set()
|
||||
with settled:
|
||||
assert settled.wait_for(lambda: count["n"] >= 1, timeout=10)
|
||||
# The late row is enqueued inside the settle (before the wrapper counts), so this is deterministic.
|
||||
assert not [e for e in list(ctx.event_queue.queue) if e.get("system_type") == "acceptance_late_settlement"]
|
||||
gates["model/b"].set()
|
||||
with settled:
|
||||
assert settled.wait_for(lambda: count["n"] >= 2, timeout=10)
|
||||
events = []
|
||||
while not any(e.get("system_type") == "acceptance_late_settlement" for e in events):
|
||||
events.append(ctx.event_queue.get(timeout=10))
|
||||
finally:
|
||||
for gate in gates.values():
|
||||
gate.set()
|
||||
with settled:
|
||||
settled.wait_for(lambda: count["n"] >= 2, timeout=10)
|
||||
rows = [e for e in events if e.get("system_type") == "acceptance_late_settlement"]
|
||||
assert len(rows) == 1 and "- a: PASS" in rows[0]["text"] and "- b: PASS" in rows[0]["text"]
|
||||
assert "pending" not in rows[0]["text"] and _mailbox_rows(tmp_path, "late-two") == []
|
||||
|
||||
|
||||
def test_two_late_panels_of_one_task_each_announce_their_own_row(tmp_path, monkeypatch):
|
||||
"""Astra review round 4: a settlement reconciles only its own wave. Collecting
|
||||
every pending panel would let the first callback swallow the sibling's verdicts
|
||||
and leave that panel's own callback with nothing to announce."""
|
||||
settled = threading.Condition()
|
||||
count = {"n": 0}
|
||||
original_settle = __import__("ouroboros.review_custody", fromlist=["_settle_review_attempt"])._settle_review_attempt
|
||||
|
||||
def settle(*args, **kwargs):
|
||||
try:
|
||||
return original_settle(*args, **kwargs)
|
||||
finally:
|
||||
with settled:
|
||||
count["n"] += 1
|
||||
settled.notify_all()
|
||||
|
||||
monkeypatch.setattr("ouroboros.review_custody._settle_review_attempt", settle)
|
||||
# Both settlements are PUBLISHED (custody's lock block ran, the settled attempt
|
||||
# is collectible) before EITHER announcement runs, so the two announcements
|
||||
# race exactly as two panels settling together would; a broken barrier fails
|
||||
# the test instead of degrading it to a serial schedule.
|
||||
barrier = threading.Barrier(2, timeout=10)
|
||||
broken = []
|
||||
from ouroboros import acceptance_settlement as leaf
|
||||
original_announce = leaf.announce_acceptance_settlement
|
||||
|
||||
def announce(*args, **kwargs):
|
||||
try:
|
||||
barrier.wait()
|
||||
except threading.BrokenBarrierError:
|
||||
broken.append(True)
|
||||
return original_announce(*args, **kwargs)
|
||||
|
||||
monkeypatch.setattr(leaf, "announce_acceptance_settlement", announce)
|
||||
gates = {"model/a": threading.Event(), "model/b": threading.Event()}
|
||||
model = _SlotModel(gates, {"model/a": "PASS", "model/b": "FAIL"})
|
||||
ctx = _terminal_ctx(tmp_path, task_id="late-pair")
|
||||
try:
|
||||
for key, slot_model in (("wave-one", "model/a"), ("wave-two", "model/b")):
|
||||
slot = ReviewSlot(slot_id=key, model=slot_model, effort="high", timeout_sec=20)
|
||||
first = _released_wave(tmp_path, ctx, slots=[slot], model=model, task_id="late-pair", retry_key=key)
|
||||
ctx._execution_trace["review_runs"].append(
|
||||
{**json.loads(json.dumps(dataclasses.asdict(first))), "authority": "host_root",
|
||||
"panel_id": f"panel_{key}", "binding_hash": f"binding-{key}", "candidate_hash": f"c-{key}"})
|
||||
gates["model/a"].set()
|
||||
gates["model/b"].set()
|
||||
with settled:
|
||||
assert settled.wait_for(lambda: count["n"] >= 2, timeout=15)
|
||||
events = []
|
||||
deadline = time.monotonic() + 10
|
||||
while time.monotonic() < deadline and sum(1 for e in events if e.get("system_type") == "acceptance_late_settlement") < 2:
|
||||
try:
|
||||
events.append(ctx.event_queue.get(timeout=2))
|
||||
except Exception:
|
||||
break
|
||||
finally:
|
||||
for gate in gates.values():
|
||||
gate.set()
|
||||
with settled:
|
||||
settled.wait_for(lambda: count["n"] >= 2, timeout=10)
|
||||
rows = [e for e in events if e.get("system_type") == "acceptance_late_settlement"]
|
||||
assert not broken, "both announcements must have raced through the barrier"
|
||||
assert sorted(r["delivery_id"] for r in rows) == ["acceptance-late:wave-one", "acceptance-late:wave-two"], events
|
||||
assert any("passed" in r["text"] for r in rows) and any("rejected" in r["text"] for r in rows)
|
||||
|
||||
|
||||
def test_an_owner_followup_sets_the_running_panel_aside(tmp_path):
|
||||
"""Fable review round 4: after the owner changes the requirements, the answer
|
||||
Main writes for them is not a delivery under the panel that reviewed the old
|
||||
ones — the latch clears, so the ordinary path decides while the old panel's
|
||||
verdicts still arrive as advice."""
|
||||
from ouroboros import loop
|
||||
from tests.test_delivery_forced_finalization import _forced_test_context
|
||||
|
||||
_loop, registry, _ctx, trace = _forced_test_context(tmp_path)
|
||||
registry._ctx._task_acceptance_pending = "binding-one"
|
||||
trace["review_decision"] = {}
|
||||
loop._supersede_task_acceptance_for_owner_followup(registry._ctx, trace)
|
||||
assert registry._ctx._task_acceptance_pending == ""
|
||||
assert trace["acceptance_decision"]["reason"] == "owner_followup"
|
||||
|
||||
|
||||
def test_a_late_settlement_for_another_task_is_never_published(tmp_path):
|
||||
"""The worker may already be rebound: a wave whose retry key is not in the
|
||||
live trace returns False and writes nothing anywhere."""
|
||||
from ouroboros.acceptance_settlement import attach_late_acceptance_settlement
|
||||
from ouroboros.task_results import load_task_result
|
||||
|
||||
ctx = _terminal_ctx(tmp_path, task_id="other-root")
|
||||
ctx._execution_trace["review_runs"].append(_host_run(request={"surface": "task_acceptance", "retry_key": "another"}))
|
||||
before = json.dumps(load_task_result(tmp_path, "other-root"), sort_keys=True)
|
||||
assert attach_late_acceptance_settlement(
|
||||
ctx, SimpleNamespace(retry_key="acceptance-subject-one", task_id="other-root"),
|
||||
{"slots": {"a": "ok"}, "total": 1}, result={"chat_id": 1}) is False
|
||||
assert ctx.event_queue.empty() and _mailbox_rows(tmp_path, "other-root") == []
|
||||
assert json.dumps(load_task_result(tmp_path, "other-root"), sort_keys=True) == before
|
||||
|
||||
|
||||
def test_every_acceptance_wake_reoffers_a_changed_keep_contract(tmp_path, monkeypatch):
|
||||
"""A replacement candidate inherits ``control_episode_seen``; the contract for
|
||||
the NEW candidate must still be shown, while identical bytes are not repeated."""
|
||||
from ouroboros.loop_acceptance_review import wait_for_acceptance_feedback
|
||||
from tests.test_delivery_forced_finalization import _forced_test_context
|
||||
|
||||
loop, registry, ctx, trace = _forced_test_context(tmp_path)
|
||||
monkeypatch.setattr("ouroboros.owner_wait.wait_after_tools", lambda *_a, **_k: None)
|
||||
registry._ctx._task_acceptance_pending = "binding-one"
|
||||
blocks = lambda: str(ctx.messages).count("[DELIVERY_FINALIZATION_CONTROL]") # noqa: E731
|
||||
|
||||
first = loop._replace_delivery_candidate(registry, ctx, trace, "First complete answer.", control="candidate")
|
||||
assert first.control_episode_seen is False
|
||||
wait_for_acceptance_feedback(registry, ctx, trace, [], set())
|
||||
assert first.control_episode_seen is True and blocks() == 1
|
||||
# The same candidate renders identical bytes: a second wake adds no noise.
|
||||
wait_for_acceptance_feedback(registry, ctx, trace, [], set())
|
||||
assert blocks() == 1
|
||||
second = loop._replace_delivery_candidate(registry, ctx, trace, "Second complete answer.", control="candidate")
|
||||
assert second.control_episode_seen is True, "the inherited flag is what hid the contract"
|
||||
wait_for_acceptance_feedback(registry, ctx, trace, [], set())
|
||||
assert blocks() == 2 and second.content_sha256[:12] in str(ctx.messages)
|
||||
|
||||
|
||||
def test_the_acceptance_wake_keeps_the_one_repair_already_spent(tmp_path, monkeypatch):
|
||||
"""Scope review round 1: re-arming on every wake reset ``repair_attempted``,
|
||||
so a candidate could burn one malformed-control repair per wake instead of
|
||||
one per episode. The wake's re-offer preserves the spent repair; an ordinary
|
||||
arm (something changed) still opens a fresh episode."""
|
||||
from ouroboros.loop_acceptance_review import wait_for_acceptance_feedback
|
||||
from tests.test_delivery_forced_finalization import _forced_test_context
|
||||
|
||||
loop, registry, ctx, trace = _forced_test_context(tmp_path)
|
||||
monkeypatch.setattr("ouroboros.owner_wait.wait_after_tools", lambda *_a, **_k: None)
|
||||
registry._ctx._task_acceptance_pending = "binding-one"
|
||||
candidate = loop._replace_delivery_candidate(registry, ctx, trace, "Complete answer.", control="candidate")
|
||||
wait_for_acceptance_feedback(registry, ctx, trace, [], set())
|
||||
candidate.repair_attempted = True # the one repair was spent on a malformed control
|
||||
wait_for_acceptance_feedback(registry, ctx, trace, [], set())
|
||||
assert candidate.repair_attempted is True, "the wake re-offer must not refund the repair"
|
||||
loop._arm_delivery_control(registry, ctx, trace)
|
||||
assert candidate.repair_attempted is False, "an ordinary arm opens a new episode"
|
||||
|
||||
|
||||
def test_the_rearmed_contract_never_rewrites_an_already_sent_row(tmp_path):
|
||||
"""Issue #906: merging into a sent row discards the conversation cache. Every
|
||||
wake re-arms, so the control block must take the execution slot and append."""
|
||||
import copy as _copy
|
||||
|
||||
from ouroboros.transcript_prefix import observe_send
|
||||
from tests.test_delivery_forced_finalization import _forced_test_context
|
||||
|
||||
loop, registry, ctx, trace = _forced_test_context(tmp_path)
|
||||
loop._replace_delivery_candidate(registry, ctx, trace, "Complete answer.", control="candidate")
|
||||
ctx.messages.append({"role": "user", "content": "An owner follow-up that already went out."})
|
||||
observe_send(registry._ctx, ctx.messages, round_idx=1)
|
||||
sent = _copy.deepcopy(ctx.messages[-1])
|
||||
|
||||
loop._arm_delivery_control(registry, ctx, trace)
|
||||
|
||||
assert sent in ctx.messages, "an already-sent message was rewritten"
|
||||
control = [row for row in ctx.messages
|
||||
if "[DELIVERY_FINALIZATION_CONTROL]" in str(row.get("content") or "")]
|
||||
assert len(control) == 1 and control[0] is not ctx.messages[ctx.messages.index(sent)]
|
||||
# Only the acceptance wake's repeated re-offer is deduplicated. Every other
|
||||
# caller arms because something changed, so it always appends.
|
||||
loop._arm_delivery_control(registry, ctx, trace)
|
||||
assert str(ctx.messages).count("[DELIVERY_FINALIZATION_CONTROL]") == 2
|
||||
|
||||
|
||||
@pytest.mark.parametrize("choice", ["wait", "finish"])
|
||||
def test_pending_review_rides_beside_the_verb_and_is_recorded_on_every_answer(tmp_path, choice):
|
||||
"""WP-7: the optional wait/finish choice is a sibling of ``acceptance_subject``,
|
||||
never an extra key that invalidates the body, and every control answer records
|
||||
it (an answer without the key means wait)."""
|
||||
from tests.test_delivery_control_lineage import _start_control_episode
|
||||
|
||||
loop, registry, ctx, trace, candidate = _start_control_episode(tmp_path)
|
||||
loop._arm_delivery_control(registry, ctx, trace)
|
||||
status, text = loop._resolve_delivery_control(
|
||||
json.dumps({"delivery_control": "keep", "pending_review": choice}), registry, ctx, trace,
|
||||
)
|
||||
assert (status, text) == ("resolved", candidate.full_text)
|
||||
assert registry._ctx._acceptance_pending_review_choice == choice
|
||||
loop._arm_delivery_control(registry, ctx, trace)
|
||||
status, _text = loop._resolve_delivery_control(
|
||||
json.dumps({"delivery_control": "keep"}), registry, ctx, trace,
|
||||
)
|
||||
assert status == "resolved" and registry._ctx._acceptance_pending_review_choice == "wait"
|
||||
|
|
|
|||
|
|
@ -77,6 +77,7 @@ TERMINAL_WRITERS = {
|
|||
('ouroboros/mutation_attribution.py::record_terminal_mutation_candidates', 'status'): 'dynamic',
|
||||
('ouroboros/post_task_checkpoint.py::set_root_post_task_checkpoint', 'str(existing.get("status") or task.get("status") or STATUS_COMPLETED)'): 'terminal',
|
||||
('ouroboros/project_dialogue.py::_append_terminal_task_projection', 'status'): 'dynamic',
|
||||
('ouroboros/project_naming.py::spawn_turn_namer._work', 'status'): 'dynamic',
|
||||
('ouroboros/project_dialogue.py::persist_continuation_narrative', 'requested_status'): 'dynamic',
|
||||
# The locked field projector preserves the existing status, including a
|
||||
# terminal one; publishing review evidence never completes the task itself.
|
||||
|
|
|
|||
|
|
@ -454,7 +454,8 @@ def test_one_cancel_leaves_exactly_one_salvaged_paragraph_in_the_chat(tmp_path,
|
|||
)
|
||||
assert f"{SALVAGE_EXCERPT_LABEL}." in row["text"]
|
||||
assert salvage not in row["text"]
|
||||
assert 'get_task_result(task_id="stopped-one")' in row["text"]
|
||||
assert "get_task_result" not in row["text"]
|
||||
assert row["text"].startswith("Cancelled. Root task stopped-one.")
|
||||
|
||||
|
||||
@pytest.mark.serial
|
||||
|
|
@ -525,7 +526,7 @@ def test_cascade_receipt_dedups_the_actual_destination_and_preserves_main(qenv,
|
|||
if row.get("type") == "task_summary"
|
||||
)
|
||||
assert text not in terminal["text"] and SALVAGE_EXCERPT_LABEL in terminal["text"]
|
||||
assert 'get_task_result(task_id="settled-root")' in terminal["text"]
|
||||
assert "get_task_result" not in terminal["text"]
|
||||
queue.events.clear()
|
||||
assert enqueue_project_completion_summary(
|
||||
qenv.drive, {}, "settled-root", task, stored, {"status": "failed"},
|
||||
|
|
|
|||
|
|
@ -159,6 +159,13 @@ def test_ordinary_addressing_card_tracks_real_work_through_metrics_and_reload(
|
|||
else:
|
||||
card.wait_for(timeout=30000)
|
||||
assert card.count() == 1
|
||||
# A tool row or a tool error is content: the direct turn
|
||||
# wears the task card (chrome follows content, owner
|
||||
# decision 16.09) — a visible chip, a title, and the
|
||||
# Main conversion control on an unbound turn.
|
||||
assert card.locator("[data-live-phase]").is_visible()
|
||||
assert card.locator("[data-live-title]").inner_text().strip() != ""
|
||||
assert card.locator("[data-turn-into-project]").count() == 1
|
||||
page.screenshot(path=str(evidence / "live.png"), full_page=True, animations="disabled")
|
||||
page.reload(wait_until="domcontentloaded")
|
||||
page.get_by_text("The addressing attempt is complete.", exact=True).wait_for(timeout=30000)
|
||||
|
|
|
|||
|
|
@ -1,7 +1,9 @@
|
|||
"""Host facts behind the conversation activity block (WP-G).
|
||||
|
||||
The chat block keys its chrome on ``_is_direct_chat`` (a direct conversation
|
||||
turn vs a managed/Swarm root) and represents an addressing-only turn by the
|
||||
``_is_direct_chat`` (a direct conversation turn vs a managed/Swarm root) is a
|
||||
host fact for routing, the census kind, Stop custody and the header pill; the
|
||||
chat block's chrome follows the work it holds, never this fact. The block
|
||||
represents an addressing-only turn by the
|
||||
typed routing action, never by a client tool-name list. Live frames learn the
|
||||
fact from the activity census and the rebuilt ``task_done``; replay learns it
|
||||
from the terminal truth annotation and the authored summary row. These tests
|
||||
|
|
|
|||
|
|
@ -45,6 +45,8 @@ def _start_control_episode(
|
|||
'{"delivery_control":"replace","full_answer":7}',
|
||||
'{"delivery_control":"replace","full_answer":"one","full_answer":"two"}',
|
||||
'{"delivery_control":"keep","extra":true}',
|
||||
'{"delivery_control":"keep","pending_review":"later"}',
|
||||
'{"delivery_control":"replace","full_answer":"x","pending_review":true}',
|
||||
'```json\n{"delivery_control":"publish"}\n```',
|
||||
],
|
||||
)
|
||||
|
|
|
|||
|
|
@ -1012,7 +1012,7 @@ def test_main_history_admits_project_started_row_and_project_thread_excludes_it(
|
|||
"project_id": "launch",
|
||||
"project_name": "Launch 🚀",
|
||||
"target_label": "Launch 🚀 › Ship release",
|
||||
"text": "Launch 🚀 › Ship release · Started\nWork is running in this Project.",
|
||||
"text": "Launch 🚀 › Ship release · Started",
|
||||
},
|
||||
{
|
||||
"ts": "2026-08-21T00:00:02Z",
|
||||
|
|
|
|||
|
|
@ -126,7 +126,8 @@ def test_a_derived_name_is_not_reported_as_model_coined():
|
|||
"""Turn-into-project reuses the name slot; it must not claim authorship.
|
||||
|
||||
`suggested_name` is filled by admission naming (a caller title, or the
|
||||
request's first line) and by the agent's own scope tools; the conversion
|
||||
request's first line), by the lazy turn namer of a working Main turn and
|
||||
by the agent's own scope tools; the conversion
|
||||
cannot tell them apart, so its naming reason names the SLOT it read rather
|
||||
than a coiner that may not exist.
|
||||
"""
|
||||
|
|
|
|||
|
|
@ -64,14 +64,21 @@ def test_live_task_message_marker_uses_my_human_wording():
|
|||
|
||||
|
||||
def test_system_prompt_carries_outcome_honesty_and_capability_acquisition():
|
||||
"""v6.29.0 doctrine pins: the three-tier outcome lexicon and the capability-
|
||||
acquisition boldness clause must stay in SYSTEM.md. v6.60.0: the FINAL ANSWER
|
||||
marker doctrine deliberately MOVED to the per-task contract (answer_protocol) —
|
||||
the default prompt must NOT carry it (ordinary tasks never see the marker)."""
|
||||
"""v6.29.0 doctrine pins: the outcome-honesty doctrine and the capability-
|
||||
acquisition boldness clause must stay in SYSTEM.md. The doctrine now says
|
||||
the three endings in WORDS: the ledger identifiers belong to the reviewers'
|
||||
JSON contract, and a prompt that spells them is a prompt the model parrots
|
||||
back to its human. v6.60.0: the FINAL ANSWER marker doctrine deliberately
|
||||
MOVED to the per-task contract (answer_protocol) — the default prompt must
|
||||
NOT carry it (ordinary tasks never see the marker)."""
|
||||
import pathlib
|
||||
|
||||
text = (pathlib.Path(__file__).parent.parent / "prompts" / "SYSTEM.md").read_text(encoding="utf-8")
|
||||
assert "blocked_with_evidence" in text
|
||||
assert "### Outcome honesty" in text
|
||||
# Whitespace-normalized: the doctrine sentence is line-wrapped in the file.
|
||||
assert "the only real failure mode" in " ".join(text.split())
|
||||
assert "blocked_with_evidence" not in text
|
||||
assert "best_effort" not in text
|
||||
# Whitespace-normalized: a line-wrapped "FINAL\nANSWER" must not slip past.
|
||||
assert "FINAL ANSWER" not in " ".join(text.split())
|
||||
assert "## Capability Acquisition" in text
|
||||
|
|
|
|||
|
|
@ -481,6 +481,20 @@ def test_turn_event_queue_stamps_by_value_at_the_producer():
|
|||
zero = {"type": "log_event", "data": {"type": "x", "task_id": "turn1", "chat_id": 0}}
|
||||
assert proxy.stamp(zero)["data"]["chat_id"] == 0
|
||||
|
||||
# The first-work callback (the lane hangs the turn namer on it) fires
|
||||
# exactly once, on the first frame stamped as work — never on a receipt.
|
||||
fired = []
|
||||
named = _TurnEventQueue(SimpleNamespace(put=captured.append, put_nowait=captured.append), "turn2", 42,
|
||||
on_first_work=lambda: fired.append(1))
|
||||
named.put_nowait({"type": "log_event", "data": {
|
||||
"type": "tool_call_started", "task_id": "turn2", "tool": "promote_chat_to_task", "routing_action": "promote_chat_to_task"}})
|
||||
named.put_nowait({"type": "log_event", "data": {"type": "llm_round_error", "task_id": "turn2"}})
|
||||
assert fired == []
|
||||
named.put_nowait({"type": "log_event", "data": {"type": "tool_call_started", "task_id": "turn2", "tool": "read_file"}})
|
||||
named.put_nowait({"type": "log_event", "data": {"type": "tool_call_finished", "task_id": "turn2", "tool": "read_file"}})
|
||||
named.put_nowait({"type": "log_event", "data": {"type": "tool_call_started", "task_id": "turn2", "tool": "run_command"}})
|
||||
assert fired == [1]
|
||||
|
||||
# The execution half of the direct lane installs the proxy around
|
||||
# agent.handle_task (admission registers the turn; execution runs it).
|
||||
import inspect
|
||||
|
|
|
|||
|
|
@ -268,3 +268,41 @@ def test_native_post_task_retains_activity_and_delivers_answer_early(monkeypatch
|
|||
assert final["delivery_id"] == early["delivery_id"]
|
||||
assert done["type"] == "task_done" and done["task_id"] == task_id
|
||||
assert bus.empty()
|
||||
|
||||
|
||||
def test_first_working_tool_call_names_a_main_turn_once(monkeypatch, tmp_path):
|
||||
"""Owner decision Q7=A (16.09): a direct Main turn is named lazily, by the first
|
||||
non-addressing tool call, so a greeting costs no naming call and a working turn
|
||||
gets its title as its block becomes the task card; a Project-room turn is named
|
||||
by its room and spawns nothing."""
|
||||
from ouroboros import agent as agent_module, project_naming
|
||||
|
||||
_lane(monkeypatch, tmp_path)
|
||||
calls = []
|
||||
monkeypatch.setattr(project_naming, "spawn_turn_namer", lambda *a, **kw: calls.append((a, kw)))
|
||||
|
||||
class Actor:
|
||||
def __init__(self, **kwargs):
|
||||
self._owner_message_admission_lock = threading.Lock()
|
||||
|
||||
def handle_task(self, task):
|
||||
frames = self._event_queue
|
||||
frames.put_nowait({"type": "log_event", "data": {
|
||||
"type": "tool_call_started", "task_id": task["id"], "tool": "promote_chat_to_task",
|
||||
"routing_action": "promote_chat_to_task"}})
|
||||
assert calls == [], "an addressing call is a receipt, not work"
|
||||
for kind, tool in (("tool_call_started", "read_file"), ("tool_call_finished", "read_file"),
|
||||
("tool_call_started", "run_command")):
|
||||
frames.put_nowait({"type": "log_event", "data": {"type": kind, "task_id": task["id"], "tool": tool}})
|
||||
return []
|
||||
|
||||
monkeypatch.setattr(agent_module, "make_agent", Actor)
|
||||
workers.handle_chat_direct(1, "проверь, почему карточка задачи потеряла элементы")
|
||||
assert len(calls) == 1
|
||||
args, kwargs = calls[0]
|
||||
assert args[0] == tmp_path and args[2] == "проверь, почему карточка задачи потеряла элементы"
|
||||
assert callable(kwargs["broadcast"])
|
||||
|
||||
calls.clear()
|
||||
workers.handle_chat_direct(1, "и в комнате проекта", task_metadata={"project_id": "room"})
|
||||
assert calls == [], "a Project-room turn is named by its room"
|
||||
|
|
|
|||
|
|
@ -157,7 +157,7 @@ def test_missing_role_terminal_root_is_never_labeled_child(tmp_path):
|
|||
row = next(row for row in _chat_rows(tmp_path) if row.get("task_id") == "roleless-root")
|
||||
assert row["summary_kind"] == "terminal_root_projection"
|
||||
assert row["role"] == "root"
|
||||
assert "role=root" in row["text"]
|
||||
assert "Root task roleless-root." in row["text"]
|
||||
|
||||
|
||||
def test_terminal_child_projection_is_idempotent_and_honest_for_all_outcomes(tmp_path):
|
||||
|
|
@ -201,7 +201,9 @@ def test_terminal_child_projection_is_idempotent_and_honest_for_all_outcomes(tmp
|
|||
assert row["result_ref"] == {
|
||||
"kind": "task_result", "task_id": task_id, "reader": "get_task_result",
|
||||
}
|
||||
assert f'get_task_result(task_id="{task_id}")' in row["text"]
|
||||
# The reader is a typed field; a host row never spells a tool name.
|
||||
assert "get_task_result" not in row["text"]
|
||||
assert f"(child {task_id} of project-root)" in row["text"]
|
||||
|
||||
|
||||
def test_terminal_projection_dedup_does_not_lose_concurrent_chat_append(tmp_path):
|
||||
|
|
@ -472,10 +474,13 @@ def test_child_projection_enters_main_cognition_and_project_lineage_not_main_ui(
|
|||
project_context = "\n\n".join(
|
||||
build_recent_sections(Memory(tmp_path), env=None, thread_chat_id=project_chat)
|
||||
)
|
||||
assert "Reviewed exact SHA" in main_context
|
||||
assert "parent=root" in main_context
|
||||
assert "Reviewed exact SHA" in project_context
|
||||
assert "parent=root" in project_context
|
||||
# ``memory._format_chat_line`` renders the text and drops every typed
|
||||
# field, so lineage must stay in words. The child's own answer is not
|
||||
# repeated here: it is a turn of its own in the room this row lives in.
|
||||
assert "(child child-review of root)" in main_context
|
||||
assert "Reviewed exact SHA" not in main_context
|
||||
assert "(child child-review of root)" in project_context
|
||||
assert "Reviewed exact SHA" not in project_context
|
||||
|
||||
import asyncio
|
||||
|
||||
|
|
@ -500,7 +505,8 @@ def test_child_projection_enters_main_cognition_and_project_lineage_not_main_ui(
|
|||
{"chat_id": 1, "status": "completed"},
|
||||
)
|
||||
main_context = "\n\n".join(build_recent_sections(Memory(tmp_path), env=None))
|
||||
assert "Unscoped child truth" in main_context
|
||||
assert "researcher (child child-main of main-root)" in main_context
|
||||
assert "Unscoped child truth" not in main_context
|
||||
main_rows = json.loads(asyncio.run(endpoint(SimpleNamespace(
|
||||
query_params={"chat_id": "1"},
|
||||
))).body)["messages"]
|
||||
|
|
|
|||
|
|
@ -322,7 +322,11 @@ def test_cap_reached_returns_typed_exhausted_result_hold_and_event(harness, monk
|
|||
assert gate["status"] == "cycles_exhausted" and gate["allow"] is True and gate["closed"] is False
|
||||
decision = force_plan_decision(ctx, {}, enforcement="blocking")
|
||||
assert decision["required"] and decision["self_opened"] and decision["status"] == "cycles_exhausted"
|
||||
assert "blocked_with_evidence" in plan_review_disclosure(decision)
|
||||
# Owner-readable prose says the ending in words; the ledger identifier
|
||||
# stays on the typed objective axis asserted below.
|
||||
disclosure = plan_review_disclosure(decision)
|
||||
assert "the task ends blocked with its evidence recorded" in disclosure
|
||||
assert "blocked_with_evidence" not in disclosure
|
||||
events = []
|
||||
while not harness.events.empty():
|
||||
events.append(harness.events.get_nowait())
|
||||
|
|
|
|||
|
|
@ -69,7 +69,11 @@ def test_quorum_unreachable_releases_finalization_for_a_blocked_terminal(harness
|
|||
decision = force_plan_decision(ctx, {}, enforcement="blocking")
|
||||
assert decision["allow"] is True and decision["quorum_unreachable"] is True
|
||||
disclosure = plan_review_disclosure(decision)
|
||||
assert "blocked_with_evidence" in disclosure and "structurally unreachable" in disclosure
|
||||
# Owner-readable prose says the ending in words; the ledger identifier stays
|
||||
# on the typed objective axis asserted below.
|
||||
assert "the task ends blocked with its evidence recorded" in disclosure
|
||||
assert "blocked_with_evidence" not in disclosure
|
||||
assert "structurally unreachable" in disclosure
|
||||
assert "2030-01-01T00:00:00+00:00" in disclosure
|
||||
from ouroboros.outcomes import derive_loop_outcome
|
||||
|
||||
|
|
|
|||
|
|
@ -95,8 +95,11 @@ def test_completion_summary_event_text_is_plain_and_fully_normalized(
|
|||
ctx.DRIVE_ROOT, {"status": "completed"}, "root-project", root, result, done,
|
||||
) is True
|
||||
assert queued[0]["text"] == (
|
||||
f"Launch 🚀 › Ship release · Done\n{PLAIN_EXCERPT}"
|
||||
"Launch 🚀 › Ship release · Done\nOpen the Project for details."
|
||||
)
|
||||
# The model's own answer already lives in the room this row points at, so
|
||||
# Main states the outcome and the way in, never a cut of the same bytes.
|
||||
assert PLAIN_EXCERPT not in queued[0]["text"]
|
||||
for marker in ("#", "**", "`"):
|
||||
assert marker not in queued[0]["text"]
|
||||
# RO4 convergence: the producer already normalized, so the verbatim live
|
||||
|
|
@ -188,7 +191,7 @@ def test_history_normalizes_old_project_rows_on_read_without_rewriting_log(tmp_p
|
|||
"type": "project_started", "task_id": "root-project",
|
||||
"project_id": "launch", "project_name": "Launch",
|
||||
"target_label": "Launch › Ship",
|
||||
"text": "# Launch › Ship · Started\nWork is running in this Project.",
|
||||
"text": "# Launch › Ship · Started",
|
||||
},
|
||||
{
|
||||
"ts": "2026-08-21T00:00:02Z", "direction": "system", "chat_id": 1,
|
||||
|
|
@ -222,7 +225,7 @@ def test_history_normalizes_old_project_rows_on_read_without_rewriting_log(tmp_p
|
|||
by_type = {row.get("system_type"): row for row in payload["messages"] if row.get("role") == "system"}
|
||||
|
||||
started = by_type["project_started"]
|
||||
assert started["text"] == "Launch › Ship · Started\nWork is running in this Project."
|
||||
assert started["text"] == "Launch › Ship · Started"
|
||||
assert started["markdown"] is False
|
||||
|
||||
completion = by_type["project_completion_summary"]
|
||||
|
|
@ -370,9 +373,9 @@ def test_owner_requested_stop_is_done_on_the_host_row_too():
|
|||
|
||||
A4_DECISION = {
|
||||
"status": "finalized_unaccepted",
|
||||
"reason": "review_degraded",
|
||||
"rationale": "Acceptance reviewers did not reach a valid quorum.",
|
||||
}
|
||||
A4_CLAUSE = "Acceptance: finalized_unaccepted — Acceptance reviewers did not reach a valid quorum."
|
||||
|
||||
|
||||
def _a4_result(**overrides):
|
||||
|
|
@ -391,12 +394,16 @@ def _a4_result(**overrides):
|
|||
def test_host_verdict_states_an_unaccepted_acceptance_decision_in_its_own_words():
|
||||
"""S5-04: a warning caused by REVIEW used to be explained by the execution
|
||||
reason that happened to sit beside it (``Reason: final_message``), which
|
||||
named the delivery step rather than the cause."""
|
||||
from ouroboros.project_dialogue import _completion_verdict
|
||||
named the delivery step rather than the cause. The decision's own typed
|
||||
reason now speaks, through the one shared table."""
|
||||
from ouroboros.project_dialogue import TASK_CAUSE_PHRASES, _completion_verdict
|
||||
|
||||
assert _completion_verdict(_a4_result(), {}) == A4_CLAUSE
|
||||
# The stored rationale already ends in a period; the clause must not double it.
|
||||
assert not _completion_verdict(_a4_result(), {}).endswith("..")
|
||||
verdict = _completion_verdict(_a4_result(), {})
|
||||
assert verdict == TASK_CAUSE_PHRASES["review_degraded"]
|
||||
assert not verdict.endswith("..")
|
||||
# The stored reviewer rationale stays in the card, task_results and Logs:
|
||||
# the row carries one sentence, and never the machine words beside it.
|
||||
assert "quorum" not in verdict and "finalized_unaccepted" not in verdict
|
||||
|
||||
|
||||
def test_host_verdict_keeps_the_execution_reason_when_acceptance_was_reached():
|
||||
|
|
@ -406,7 +413,7 @@ def test_host_verdict_keeps_the_execution_reason_when_acceptance_was_reached():
|
|||
"execution": {"status": "ok"},
|
||||
"review": {"status": "pass", "acceptance_decision": {"status": "accepted"}},
|
||||
})
|
||||
assert _completion_verdict(accepted, {}) == "Reason: final_message."
|
||||
assert _completion_verdict(accepted, {}) == ""
|
||||
assert _completion_verdict({"status": "completed"}, {}) == ""
|
||||
# A hard failure explains itself by its execution reason, not by a decision.
|
||||
# The custody debt this row names is one the row STILL owes: since owner item
|
||||
|
|
@ -419,61 +426,52 @@ def test_host_verdict_keeps_the_execution_reason_when_acceptance_was_reached():
|
|||
outcome_axes={"execution": {"status": "failed"},
|
||||
"review": {"acceptance_decision": dict(A4_DECISION)}}),
|
||||
{},
|
||||
) == "Reason: delegated_custody_unreconciled."
|
||||
) == "Some delegated work was never reconciled."
|
||||
|
||||
|
||||
def test_host_verdict_flattens_and_strips_a_markdown_rationale():
|
||||
"""These are durable plain-text rows: the rationale is free owner-visible
|
||||
text up to 500 characters and may carry newlines and markdown markers."""
|
||||
from ouroboros.project_dialogue import _completion_verdict
|
||||
def test_the_stored_reviewer_rationale_never_reaches_the_row(tmp_path):
|
||||
"""The rationale is free reviewer text up to 500 characters, with newlines
|
||||
and markdown markers. It belongs where the full copy lives — the card, the
|
||||
task result and Logs — and the row carries the one table sentence."""
|
||||
from ouroboros.project_dialogue import TASK_CAUSE_PHRASES, _completion_verdict
|
||||
from ouroboros.task_results import write_task_result
|
||||
|
||||
rationale = "## Verdict\n\nThe **tests** never ran with `pytest`. " * 12
|
||||
noisy = _a4_result(outcome_axes={
|
||||
"execution": {"status": "ok"},
|
||||
"review": {"status": "degraded", "acceptance_decision": {
|
||||
"status": "revision_requested",
|
||||
"rationale": "## Verdict\n\nThe **tests** never ran with `pytest`",
|
||||
"status": "revision_requested", "reason": "evidence_refresh",
|
||||
"rationale": rationale,
|
||||
}},
|
||||
})
|
||||
verdict = _completion_verdict(noisy, {})
|
||||
assert verdict == "Acceptance: revision_requested — Verdict The tests never ran with pytest."
|
||||
for marker in ("#", "**", "`", "\n"):
|
||||
assert verdict == TASK_CAUSE_PHRASES["evidence_refresh"]
|
||||
for marker in ("#", "**", "`", "\n", "pytest"):
|
||||
assert marker not in verdict
|
||||
|
||||
# A rationale that already terminates itself keeps its own punctuation: a
|
||||
# question mark is as terminal as a period, and appending one would render
|
||||
# "…did the tests run?." to the owner.
|
||||
asking = _a4_result(outcome_axes={
|
||||
"execution": {"status": "ok"},
|
||||
"review": {"status": "degraded", "acceptance_decision": {
|
||||
"status": "revision_requested", "rationale": "Did the tests ever run?",
|
||||
}},
|
||||
})
|
||||
assert _completion_verdict(asking, {}) == "Acceptance: revision_requested — Did the tests ever run?"
|
||||
# The complete text stays resolvable: the decision the row points at keeps
|
||||
# the rationale the row no longer prints (BIBLE P1 — text or a pointer).
|
||||
fields = {key: value for key, value in noisy.items()
|
||||
if key not in {"task_id", "status"}}
|
||||
write_task_result(tmp_path, "root-project", "completed", **fields)
|
||||
stored = json.loads(
|
||||
(tmp_path / "task_results" / "root-project.json").read_text(encoding="utf-8")
|
||||
)
|
||||
decision = stored["outcome_axes"]["review"]["acceptance_decision"]
|
||||
assert decision["reason"] == "evidence_refresh"
|
||||
assert "never ran with" in decision["rationale"]
|
||||
|
||||
|
||||
def test_host_verdict_keeps_the_full_bounded_acceptance_rationale():
|
||||
from ouroboros.project_dialogue import _completion_verdict
|
||||
|
||||
rationale = ("Review evidence " + ("remains material and owner-visible. " * 9)).strip()
|
||||
result = _a4_result(outcome_axes={
|
||||
"execution": {"status": "ok"},
|
||||
"review": {"status": "degraded", "acceptance_decision": {
|
||||
"status": "revision_requested", "rationale": rationale,
|
||||
}},
|
||||
})
|
||||
|
||||
assert len(rationale) > 240
|
||||
assert _completion_verdict(result, {}) == f"Acceptance: revision_requested — {rationale}"
|
||||
|
||||
|
||||
def test_host_verdict_states_a_decision_without_a_rationale_alone():
|
||||
def test_host_verdict_states_no_cause_for_a_decision_without_a_typed_reason():
|
||||
"""A status word is not a cause, and the collapsed status is already the
|
||||
headline; a historical decision without a reason therefore says nothing."""
|
||||
from ouroboros.project_dialogue import _completion_verdict
|
||||
|
||||
bare = _a4_result(outcome_axes={
|
||||
"execution": {"status": "degraded"},
|
||||
"review": {"status": "degraded", "acceptance_decision": {"status": "revision_requested"}},
|
||||
})
|
||||
assert _completion_verdict(bare, {}) == "Acceptance: revision_requested."
|
||||
assert _completion_verdict(bare, {}) == ""
|
||||
|
||||
|
||||
def test_host_verdict_leads_both_lifecycle_rows(tmp_path, monkeypatch):
|
||||
|
|
@ -505,9 +503,12 @@ def test_host_verdict_leads_both_lifecycle_rows(tmp_path, monkeypatch):
|
|||
tmp_path, {"status": "completed"}, "root-project", root, result, done,
|
||||
) is True
|
||||
assert queued[0]["text"] == (
|
||||
f"Launch 🚀 › Ship release · Done with warnings\n{A4_CLAUSE} Release shipped."
|
||||
"Launch 🚀 › Ship release · Done with warnings\n"
|
||||
"No reviewer verdict was established for this answer. "
|
||||
"Open the Project for details."
|
||||
)
|
||||
assert "final_message" not in queued[0]["text"]
|
||||
assert "Release shipped." not in queued[0]["text"]
|
||||
|
||||
ordinary = _a4_result(
|
||||
reason_code="budget_exhausted",
|
||||
|
|
@ -519,7 +520,8 @@ def test_host_verdict_leads_both_lifecycle_rows(tmp_path, monkeypatch):
|
|||
) is True
|
||||
assert queued[1]["text"] == (
|
||||
"Launch 🚀 › Ship release · Done with warnings\n"
|
||||
"Reason: budget_exhausted. Release shipped."
|
||||
"The task ran out of budget before it could finish cleanly. "
|
||||
"Open the Project for details."
|
||||
)
|
||||
|
||||
assert append_terminal_task_projection(tmp_path, "root-project", root, result, done)
|
||||
|
|
@ -529,11 +531,18 @@ def test_host_verdict_leads_both_lifecycle_rows(tmp_path, monkeypatch):
|
|||
if line.strip()
|
||||
]
|
||||
projection = next(row for row in rows if row.get("summary_kind") == "terminal_root_projection")
|
||||
assert A4_CLAUSE in projection["text"]
|
||||
assert "Reason: final_message" not in projection["text"]
|
||||
assert projection["text"] == (
|
||||
"Done with warnings. Root task root-project. "
|
||||
"No reviewer verdict was established for this answer."
|
||||
)
|
||||
assert "final_message" not in projection["text"]
|
||||
# The room is the project and result_ref is the reader, so neither the id
|
||||
# soup nor a tool name has to be spelled into owner-visible prose.
|
||||
assert "get_task_result" not in projection["text"]
|
||||
assert projection["reason_code"] == "final_message"
|
||||
assert projection["project_id"] == "launch"
|
||||
assert projection["result_ref"]["reader"] == "get_task_result"
|
||||
assert projection["outcome"] == "Done with warnings"
|
||||
assert projection["text"].endswith('Details: get_task_result(task_id="root-project")')
|
||||
|
||||
|
||||
def test_host_verdict_and_the_card_line_compose_the_same_sentence():
|
||||
|
|
@ -577,8 +586,12 @@ def test_terminal_row_reports_the_depth_request_only_when_one_exists(tmp_path):
|
|||
for line in (tmp_path / "logs" / "chat.jsonl").read_text(encoding="utf-8").splitlines()
|
||||
if line.strip()
|
||||
}
|
||||
assert "Depth requested=2, permitted=4, achieved=2 (achieved)." in rows["swarm-root"]["text"]
|
||||
# The depth facts stay on the task result, where a reader can compare them;
|
||||
# the row states the outcome and the lineage, not a nesting audit.
|
||||
assert rows["swarm-root"]["text"] == "Done. Root task swarm-root."
|
||||
assert "Depth" not in rows["swarm-root"]["text"]
|
||||
assert "Depth" not in rows["flat-root"]["text"]
|
||||
assert result["swarm_efficiency"]["depth"]["requested_depth"] == 2
|
||||
for marker in ("#", "**", "`"):
|
||||
assert marker not in rows["swarm-root"]["text"]
|
||||
|
||||
|
|
@ -615,8 +628,8 @@ def test_a_host_salvage_row_is_never_a_bare_headline_and_reason(tmp_path):
|
|||
f"{SALVAGE_EXCERPT_LABEL}: Applied Rewrote the atlas builder and reran the suite."
|
||||
in row["text"]
|
||||
)
|
||||
assert "Reason: context_overflow." in row["text"]
|
||||
assert row["text"].endswith('Details: get_task_result(task_id="salvaged-root")')
|
||||
assert "context_overflow." in row["text"]
|
||||
assert "get_task_result" not in row["text"]
|
||||
for marker in ("#", "**", "`"):
|
||||
assert marker not in row["text"]
|
||||
|
||||
|
|
|
|||
|
|
@ -1106,7 +1106,7 @@ def test_project_completion_enqueues_once_for_root_and_never_for_child_or_direct
|
|||
"type": "send_message",
|
||||
"chat_id": 1,
|
||||
"task_id": "root-project",
|
||||
"text": "Launch 🚀 › Ship release · Done\nRelease shipped.",
|
||||
"text": "Launch 🚀 › Ship release · Done\nOpen the Project for details.",
|
||||
"role": "system",
|
||||
"system_type": "project_completion_summary",
|
||||
"delivery_id": "project-completion:root-project",
|
||||
|
|
|
|||
|
|
@ -139,9 +139,7 @@ def test_project_started_row_rides_outbox_pins_main_and_dedupes_durably(tmp_path
|
|||
assert queued[0]["system_type"] == "project_started"
|
||||
assert queued[0]["role"] == "system"
|
||||
assert queued[0]["chat_id"] == 1
|
||||
assert queued[0]["text"] == (
|
||||
"Launch 🚀 › Ship release · Started\nWork is running in this Project."
|
||||
)
|
||||
assert queued[0]["text"] == "Launch 🚀 › Ship release · Started"
|
||||
assert queued[0]["progress_meta"] == {
|
||||
"project_id": "launch",
|
||||
"project_name": "Launch 🚀",
|
||||
|
|
|
|||
|
|
@ -3,6 +3,7 @@ from __future__ import annotations
|
|||
import asyncio
|
||||
import json
|
||||
import threading
|
||||
import time
|
||||
|
||||
from ouroboros.skill_loader import (
|
||||
SkillReviewState,
|
||||
|
|
@ -425,6 +426,27 @@ def test_cancellation_during_extension_reconcile_keeps_lifecycle_lane(tmp_path,
|
|||
asyncio.run(main())
|
||||
|
||||
|
||||
def _read_heartbeat(job_path) -> str:
|
||||
"""Read the heartbeat stamp through the same Windows race the writer tolerates.
|
||||
|
||||
The beat replaces ``review_job.json`` atomically and retries its own sharing
|
||||
violations (``utils.replace_atomic``); the mirror image is this poll opening
|
||||
the file while that replace is in flight, which windows-latest refuses with
|
||||
``PermissionError`` (winerror 5/32). POSIX never raises here, so this is one
|
||||
read there and a bounded retry on Windows, never a silent skip.
|
||||
"""
|
||||
delay = 0.01
|
||||
for attempt in range(20):
|
||||
try:
|
||||
return json.loads(job_path.read_text(encoding="utf-8"))["last_heartbeat_at"]
|
||||
except PermissionError:
|
||||
if attempt == 19:
|
||||
raise
|
||||
time.sleep(delay)
|
||||
delay = min(delay * 2, 0.1)
|
||||
raise AssertionError("unreachable")
|
||||
|
||||
|
||||
def test_heartbeat_continues_during_extension_reconcile(tmp_path, monkeypatch):
|
||||
from ouroboros.skill_review import SkillReviewOutcome
|
||||
from ouroboros.skill_review_runner import (
|
||||
|
|
@ -476,13 +498,13 @@ def test_heartbeat_continues_during_extension_reconcile(tmp_path, monkeypatch):
|
|||
try:
|
||||
await _wait_for_reconcile(task, reconcile_started)
|
||||
job_path = review_job_state_path(drive_root, "alpha")
|
||||
before = json.loads(job_path.read_text(encoding="utf-8"))["last_heartbeat_at"]
|
||||
before = _read_heartbeat(job_path)
|
||||
# A bounded wait, not 20 x 10 ms: on windows-latest the heartbeat's atomic replace can
|
||||
# lose a few rounds to this very poll holding the file open (sharing violation, logged
|
||||
# and retried by the beat), and the clock ticks at ~15.6 ms — 200 ms saw no change.
|
||||
for _ in range(60):
|
||||
await asyncio.sleep(0.05)
|
||||
after = json.loads(job_path.read_text(encoding="utf-8"))["last_heartbeat_at"]
|
||||
after = _read_heartbeat(job_path)
|
||||
if after != before:
|
||||
break
|
||||
assert after != before, "no heartbeat within 3 s while the reconcile blocks"
|
||||
|
|
|
|||
272
tests/test_task_card_conversion_browser.py
Normal file
272
tests/test_task_card_conversion_browser.py
Normal file
|
|
@ -0,0 +1,272 @@
|
|||
"""Converting a RUNNING direct Main turn into a project, end to end in a browser.
|
||||
|
||||
Owner decision 16.09: a direct conversation turn that does real work IS a full
|
||||
task card in Main, and it offers "Turn into project" WHILE IT RUNS — not only
|
||||
after it ends. The running half is what no unit test can certify: the click has
|
||||
to land while the turn is inside a model round, the durable binding has to be
|
||||
written under the live turn, and the turn's own final answer has to follow that
|
||||
binding into the project room instead of landing back in Main.
|
||||
|
||||
The turn is held inside its SECOND model round by a ``ModelGate`` (the same
|
||||
event-gated hold S26/S27 use), so "still running" is a fact of the HTTP boundary,
|
||||
never a sleep or a race: round one performs a real ``read_file`` (the work the
|
||||
card stands on), round two blocks until this test releases it, and its release
|
||||
produces the final answer whose delivery room is the whole point.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import uuid
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from ouroboros.projects_registry import project_binding_for_task, project_id_for_task
|
||||
from tests.system_e2e.harness import (
|
||||
ArtifactOracle, ModelGate, body_text, keyless_settings,
|
||||
start_server, wait_durable_result, wait_until,
|
||||
)
|
||||
from tests.test_chat_addressing_browser import ToolCallOnlyModel
|
||||
from tests.test_owner_wait_integration import wait_clone as clone_fixture
|
||||
from tests.ui_chat_viewport_smoke import _CAPTURE_TEST_SOCKET
|
||||
|
||||
wait_clone = clone_fixture
|
||||
pytestmark = [pytest.mark.serial, pytest.mark.ui_browser]
|
||||
|
||||
ANSWER = "The converted turn finished its work."
|
||||
# The turn namer and the project namer are TOOL-LESS model calls, which the stub
|
||||
# answers with its default text — so pinning it to something the answer does not
|
||||
# contain keeps "the project name" and "the answer" separable strings, and the
|
||||
# room assertions below cannot pass on an accidental substring.
|
||||
CARD_NAME = "Held work of a direct turn"
|
||||
|
||||
|
||||
def _out_rows(oracle, text):
|
||||
return [row for row in oracle._jsonl("logs/chat.jsonl")
|
||||
if row.get("direction") == "out" and text in str(row.get("text") or "")]
|
||||
|
||||
|
||||
_MAIN_CARD_FACTS = """id => {
|
||||
const card = document.querySelector(`#page-chat .chat-live-card[data-task-id="${id}"]`);
|
||||
return {
|
||||
card: !!card,
|
||||
converted: card?.dataset.projectCreated || '',
|
||||
bound: card?.dataset.projectBound || '',
|
||||
pointer: !!card?.querySelector('.chat-live-bound-pointer'),
|
||||
convert_buttons: document.querySelectorAll('#page-chat [data-turn-into-project]').length,
|
||||
};
|
||||
}"""
|
||||
|
||||
|
||||
def _task_rows(oracle, task_id):
|
||||
return [{"direction": row.get("direction"), "chat_id": row.get("chat_id"),
|
||||
"type": row.get("type", ""), "text": str(row.get("text") or "")[:80]}
|
||||
for row in oracle._jsonl("logs/chat.jsonl")
|
||||
if str(row.get("task_id") or "") == task_id]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("width", [1440, 390], ids=["desktop", "mobile"])
|
||||
def test_running_direct_turn_converts_to_project_and_its_answer_follows(
|
||||
wait_clone, tmp_path, monkeypatch, width,
|
||||
):
|
||||
from playwright.sync_api import sync_playwright
|
||||
from tests.system_e2e.harness import KeylessIsolatedServer
|
||||
|
||||
marker = "CONVERT_RUNNING_" + uuid.uuid4().hex
|
||||
# Round 1 does real work; round 2 (held) carries the tool RESULT message and
|
||||
# answers once released. The script is exactly those two rounds.
|
||||
steps = [
|
||||
{"tool": "read_file", "arguments": {"root": "system_repo", "path": "VERSION"}},
|
||||
{"final": ANSWER},
|
||||
]
|
||||
|
||||
def held_round(body):
|
||||
"""The turn's own round that already carries a tool result — i.e. the
|
||||
round AFTER the card's work exists. Naming/review calls carry no tools
|
||||
and are never held, so the hold cannot starve the conversion itself."""
|
||||
messages = [m for m in body.get("messages", []) if isinstance(m, dict)]
|
||||
return (bool(body.get("tools")) and marker in body_text(body)
|
||||
and any(str(m.get("role") or "") == "tool" for m in messages))
|
||||
|
||||
gate = ModelGate(held_round, timeout=300)
|
||||
home = tmp_path / "home"
|
||||
home.mkdir()
|
||||
original_env = KeylessIsolatedServer._env
|
||||
monkeypatch.setattr(KeylessIsolatedServer, "_env", lambda server: {
|
||||
**original_env(server), "HOME": str(home), "USERPROFILE": str(home),
|
||||
"XDG_CONFIG_HOME": str(home / ".config"),
|
||||
})
|
||||
evidence = Path(os.environ.get("OUROBOROS_BROWSER_EVIDENCE_OUT") or tmp_path / "evidence")
|
||||
evidence = evidence / f"convert-running-{width}"
|
||||
evidence.mkdir(parents=True, exist_ok=True)
|
||||
with ToolCallOnlyModel(steps, final_answer=CARD_NAME, gate=gate) as stub:
|
||||
server = start_server(wait_clone, tmp_path / "instance", keyless_settings(stub, OUROBOROS_MAX_WORKERS=1))
|
||||
oracle = ArtifactOracle(server.data_root)
|
||||
try:
|
||||
with sync_playwright() as pw:
|
||||
browser = pw.chromium.launch()
|
||||
page = browser.new_page(viewport={"width": width, "height": 900}, has_touch=width < 980)
|
||||
errors = []
|
||||
page.on("pageerror", lambda error: errors.append(str(error)))
|
||||
page.add_init_script(f"({_CAPTURE_TEST_SOCKET})()")
|
||||
try:
|
||||
page.goto(server.base_url, wait_until="domcontentloaded")
|
||||
page.wait_for_function("() => window.__testSockets?.[0]?.readyState === WebSocket.OPEN")
|
||||
page.locator("#chat-input").fill(marker)
|
||||
page.locator("#chat-send").click()
|
||||
|
||||
task = wait_until(lambda: next((row["task"] for row in oracle.events("task_received")
|
||||
if row.get("task", {}).get("_is_direct_chat") and marker in row["task"].get("text", "")), None), 90)
|
||||
assert task, "composer send never reached an ordinary direct turn"
|
||||
task_id = task["id"]
|
||||
assert not task.get("_ephemeral_turn"), task
|
||||
assert gate.arrived.wait(120), "the direct turn never reached its second model round"
|
||||
|
||||
# ---- the turn is provably RUNNING, inside a model round ----
|
||||
card = page.locator(f'.chat-live-card[data-task-id="{task_id}"]')
|
||||
card.wait_for(timeout=30000)
|
||||
convert = card.locator("[data-turn-into-project]")
|
||||
convert.wait_for(timeout=30000)
|
||||
assert page.locator(
|
||||
f'.chat-live-card[data-task-id="{task_id}"][data-finished="0"]'
|
||||
).count() == 1, "the conversion control is not on a RUNNING card"
|
||||
assert convert.count() == 1
|
||||
assert task_id not in oracle.running_ids(), "a direct turn must not take a pool worker"
|
||||
at_click = {
|
||||
"stored_status": str(oracle.task_result(task_id).get("status") or ""),
|
||||
"gate_held": gate.held, "gate_released": gate.release.is_set(),
|
||||
"binding": project_id_for_task(server.data_root, task_id),
|
||||
"card_finished": card.get_attribute("data-finished"),
|
||||
"title": card.locator("[data-live-title]").inner_text().strip(),
|
||||
}
|
||||
assert at_click["stored_status"] == "running", at_click
|
||||
assert at_click["gate_held"] == 1 and at_click["gate_released"] is False
|
||||
assert at_click["binding"] == "", "the turn is bound before the owner clicked"
|
||||
page.screenshot(path=str(evidence / "live-running.png"), full_page=True, animations="disabled")
|
||||
|
||||
# ---- the one click, and the server's own answer to it ----
|
||||
with page.expect_response(
|
||||
lambda response: "/api/projects/from-task" in response.url, timeout=60000,
|
||||
) as captured:
|
||||
convert.click()
|
||||
response = captured.value
|
||||
assert response.status == 200, (response.status, response.text())
|
||||
payload = response.json()
|
||||
assert payload["adopted"] is False, payload
|
||||
project = payload["project"]
|
||||
assert project.get("id") and project.get("chat_id"), project
|
||||
# Named from the work (owner P1: one click, no prompt), never
|
||||
# from the answer the turn has not even produced yet.
|
||||
assert str(project.get("name") or "").strip(), project
|
||||
assert ANSWER not in str(project["name"]), project
|
||||
project_chat = int(project["chat_id"])
|
||||
assert project_chat != 1, project
|
||||
|
||||
# The Main card became the project chip, and stopped offering
|
||||
# a second conversion of the same work.
|
||||
converted = page.locator(
|
||||
f'.chat-live-card[data-task-id="{task_id}"][data-project-created="1"]')
|
||||
converted.wait_for(timeout=30000)
|
||||
page.wait_for_selector(
|
||||
f'.chat-live-card[data-task-id="{task_id}"].is-project', timeout=30000)
|
||||
assert converted.get_attribute("data-project-id") == project["id"]
|
||||
assert converted.locator(".chat-live-project-name").inner_text().strip()
|
||||
assert page.locator(f'.chat-live-card[data-task-id="{task_id}"]'
|
||||
' [data-turn-into-project]').count() == 0
|
||||
page.screenshot(path=str(evidence / "live-converted.png"), full_page=True, animations="disabled")
|
||||
|
||||
# The durable binding exists while the turn is STILL held.
|
||||
assert gate.held == 1 and not gate.release.is_set()
|
||||
assert project_id_for_task(server.data_root, task_id) == project["id"]
|
||||
binding = project_binding_for_task(server.data_root, task_id)
|
||||
assert int(binding["project_chat_id"]) == project_chat, binding
|
||||
assert str(oracle.task_result(task_id).get("status") or "") == "running"
|
||||
|
||||
# ---- release: the turn's own answer follows the binding ----
|
||||
gate.release.set()
|
||||
stored = wait_durable_result(oracle, task_id, timeout=120)
|
||||
assert stored["status"] == "completed", stored
|
||||
assert ANSWER in str(stored.get("result") or ""), stored
|
||||
delivered = wait_until(lambda: _out_rows(oracle, ANSWER), 60) or []
|
||||
assert len(delivered) == 1, delivered
|
||||
final_row = delivered[0]
|
||||
assert final_row["chat_id"] == project_chat, final_row
|
||||
assert final_row["task_id"] == task_id, final_row
|
||||
assert not [row for row in _out_rows(oracle, ANSWER) if row["chat_id"] == 1], \
|
||||
"the final answer of a converted turn was also delivered to Main"
|
||||
# Main DOES get one durable row: a system pointer naming the
|
||||
# project that now holds the result. It carries the project's
|
||||
# name, never the answer.
|
||||
rows = _task_rows(oracle, task_id)
|
||||
main_rows = [row for row in rows if row["chat_id"] == 1]
|
||||
assert len(main_rows) == 1 and main_rows[0]["direction"] == "system", rows
|
||||
assert project["name"] in main_rows[0]["text"], main_rows
|
||||
assert ANSWER not in main_rows[0]["text"], main_rows
|
||||
assert [row for row in rows if row["chat_id"] == project_chat
|
||||
and row["direction"] == "out"], rows
|
||||
|
||||
# ---- reload: the Project room owns the work; Main offers no second conversion ----
|
||||
page.reload(wait_until="domcontentloaded")
|
||||
page.wait_for_function("() => window.__testSockets?.[0]?.readyState === WebSocket.OPEN")
|
||||
page.wait_for_selector("#page-chat")
|
||||
page.get_by_text(marker, exact=False).first.wait_for(timeout=30000)
|
||||
reloaded = wait_until(
|
||||
lambda: (lambda facts: facts if facts["convert_buttons"] == 0 else None)(
|
||||
page.evaluate(_MAIN_CARD_FACTS, task_id)),
|
||||
30) or page.evaluate(_MAIN_CARD_FACTS, task_id)
|
||||
assert reloaded["convert_buttons"] == 0, \
|
||||
f"Main still offers to convert work that already has a project: {reloaded}"
|
||||
# A card that survives in Main must carry its project identity
|
||||
# rather than a second conversion offer. Observed on this tree:
|
||||
# no Main card survives at all (see the gap note below).
|
||||
assert not reloaded["card"] or reloaded["converted"] == "1" or reloaded["bound"] == "1", \
|
||||
f"a surviving Main card does not name its project: {reloaded}"
|
||||
main_text = page.locator("#page-chat").inner_text()
|
||||
assert ANSWER not in main_text, "Main replayed an answer that belongs to the project"
|
||||
assert marker in main_text, "Main lost the owner message the turn started from"
|
||||
# OBSERVED GAP (evidence: receipt.json "reloaded_main_history").
|
||||
# The Main pointer row above IS written durably, but it is
|
||||
# persisted as type="task_summary", not as one of the two
|
||||
# lifecycle types room_membership() exempts — so on replay the
|
||||
# Main rule "entry_chat not in project_chat_ids and not bound"
|
||||
# drops it, because the task is now bound. Main's reloaded
|
||||
# history is therefore the owner's message ALONE: the converted
|
||||
# chip is a live-session artifact and does not survive a reload.
|
||||
# Asserted here as the honest current contract, with the full
|
||||
# history kept as evidence rather than pinning the gap shut.
|
||||
main_history = page.evaluate(
|
||||
"async () => (await (await fetch('/api/chat/history?chat_id=1')).json())")
|
||||
replayed = [row for row in (main_history.get("messages") or [])
|
||||
if marker not in str(row.get("text") or "")]
|
||||
assert not [row for row in replayed if ANSWER in str(row.get("text") or "")], replayed
|
||||
page.screenshot(path=str(evidence / "reloaded-main.png"), full_page=True, animations="disabled")
|
||||
page.evaluate(
|
||||
"project => window.dispatchEvent(new CustomEvent('ouro:open-project', {detail:{project}}))",
|
||||
project)
|
||||
page.wait_for_selector("#project-panel:not([hidden])")
|
||||
page.locator("#project-panel").get_by_text(ANSWER, exact=False).first.wait_for(timeout=30000)
|
||||
page.screenshot(path=str(evidence / "reloaded-project.png"), full_page=True, animations="disabled")
|
||||
assert not errors, errors
|
||||
(evidence / "receipt.json").write_text(json.dumps({
|
||||
"width": width, "task": task, "at_click": at_click,
|
||||
"from_task_response": payload, "binding": binding,
|
||||
"stored_result_status": stored["status"],
|
||||
"final_chat_row": final_row, "project_chat_id": project_chat,
|
||||
"reloaded_main": reloaded, "task_chat_rows": rows, "reloaded_main_history": main_history,
|
||||
"progress_chat_ids": sorted({int(row.get("chat_id") or 0) for row in
|
||||
oracle._jsonl("logs/progress.jsonl")
|
||||
if str(row.get("task_id") or "") == task_id}),
|
||||
"model_rounds": stub.kinds(), "errors": errors,
|
||||
}, ensure_ascii=False, indent=2, default=str), encoding="utf-8")
|
||||
except Exception:
|
||||
page.screenshot(path=str(evidence / "failure.png"), full_page=True, animations="disabled")
|
||||
(evidence / "failure-dom.html").write_text(page.content(), encoding="utf-8")
|
||||
raise
|
||||
finally:
|
||||
gate.release.set()
|
||||
browser.close()
|
||||
finally:
|
||||
gate.release.set()
|
||||
server.stop()
|
||||
assert server.proc.poll() is not None
|
||||
assert not gate.timed_out
|
||||
|
|
@ -229,3 +229,99 @@ def test_read_with_stub_root_leaks_no_cwd_dir(tmp_path, monkeypatch):
|
|||
assert tr.load_task_result(stub_root, "x") is None
|
||||
leaked = [p.name for p in pathlib.Path(".").iterdir() if "MagicMock" in p.name]
|
||||
assert leaked == [], f"read scan leaked mock-named paths: {leaked}"
|
||||
|
||||
|
||||
def test_turn_namer_persists_name_on_already_terminal_task(tmp_path, monkeypatch):
|
||||
"""v6.40.0 #1: the turn namer must persist ``suggested_name`` even when the task
|
||||
already raced to a terminal status — it enriches under the CURRENT status instead of a
|
||||
regressing RUNNING write (which the monotonic guard would drop, losing the convert-reuse
|
||||
name)."""
|
||||
import threading
|
||||
import time
|
||||
|
||||
from ouroboros import project_naming
|
||||
|
||||
tr.write_task_result(tmp_path, "t", tr.STATUS_COMPLETED, result="fast done")
|
||||
monkeypatch.setattr(project_naming, "llm_project_name", lambda *a, **k: "Nice Title")
|
||||
project_naming.spawn_turn_namer(tmp_path, "t", "build me a thing")
|
||||
for _ in range(100): # join the daemon namer thread (best-effort, bounded)
|
||||
if not any(th.name == "namer-t" for th in threading.enumerate()):
|
||||
break
|
||||
time.sleep(0.02)
|
||||
r = tr.load_task_result(tmp_path, "t")
|
||||
assert r["status"] == tr.STATUS_COMPLETED, "namer must NOT regress a terminal task to running"
|
||||
assert r.get("suggested_name") == "Nice Title", "suggested_name must survive on a terminal task"
|
||||
|
||||
|
||||
def test_turn_namer_late_settlement_refreshes_cost_without_late_name(tmp_path, monkeypatch):
|
||||
"""A provider thread outliving the cosmetic deadline still closes accounting only."""
|
||||
import threading
|
||||
import time
|
||||
|
||||
from ouroboros import agent_task_pipeline, project_naming, usage_accounting
|
||||
|
||||
tr.write_task_result(
|
||||
tmp_path,
|
||||
"late-root",
|
||||
tr.STATUS_COMPLETED,
|
||||
root_task_id="late-root",
|
||||
accounted_upper_bound_usd=0.0,
|
||||
cost_final=True,
|
||||
root_phase_checkpoint={"post_task_synthesis": "completed"},
|
||||
)
|
||||
entered = threading.Event()
|
||||
release = threading.Event()
|
||||
|
||||
def late_paid_name(*_args, **_kwargs):
|
||||
entered.set()
|
||||
assert release.wait(2)
|
||||
attempt = usage_accounting.reserve_attempt(usage_accounting.AttemptRequest(
|
||||
model="openai/gpt-5.2",
|
||||
provider="openai",
|
||||
reservation_usd=0.25,
|
||||
drive_root=tmp_path,
|
||||
task_id="late-root",
|
||||
root_task_id="late-root",
|
||||
global_limit_usd=5.0,
|
||||
root_limit_usd=5.0,
|
||||
))
|
||||
usage_accounting.mark_dispatched(attempt)
|
||||
usage_accounting.settle_attempt(attempt, {}, cost_usd=0.25, cost_final=True)
|
||||
return "Too Late"
|
||||
|
||||
monkeypatch.setattr(project_naming, "llm_project_name", late_paid_name)
|
||||
monkeypatch.setattr(project_naming, "_naming_timeout_sec", lambda: -29.98)
|
||||
broadcasts = []
|
||||
project_naming.spawn_turn_namer(
|
||||
tmp_path, "late-root", "build a thing", broadcast=broadcasts.append,
|
||||
)
|
||||
assert entered.wait(1)
|
||||
for _ in range(100):
|
||||
if not any(th.name == "namer-late-root" for th in threading.enumerate()):
|
||||
break
|
||||
time.sleep(0.01)
|
||||
assert not any(th.name == "namer-late-root" for th in threading.enumerate())
|
||||
|
||||
release.set()
|
||||
# The subject is that the late settlement's refresh LANDS, not how fast: the
|
||||
# detached thread runs reserve → dispatch → settle → cost reconstruction →
|
||||
# locked task-result write, each a real lock cycle with fsyncs, and the
|
||||
# Windows full-test leg took longer than a 2 s poll twice out of five runs
|
||||
# (d0bb839e, 5fbdabd3) while green on the others. A refresh that never lands
|
||||
# still fails here — after a bound wide enough for a loaded runner.
|
||||
started = time.monotonic()
|
||||
for _ in range(2000):
|
||||
stored = tr.load_task_result(tmp_path, "late-root")
|
||||
if stored.get("accounted_upper_bound_usd") == 0.25:
|
||||
break
|
||||
time.sleep(0.01)
|
||||
stored = tr.load_task_result(tmp_path, "late-root")
|
||||
assert stored["accounted_upper_bound_usd"] == 0.25, f"no refresh after {time.monotonic() - started:.1f}s"
|
||||
assert stored["accounted_upper_bound_usd_with_children"] == 0.25
|
||||
assert stored["cost_final"] is True
|
||||
assert stored.get("suggested_name") is None
|
||||
assert broadcasts == []
|
||||
assert agent_task_pipeline._root_post_task_already_completed(
|
||||
type("Env", (), {"drive_root": tmp_path})(),
|
||||
{"id": "late-root", "root_task_id": "late-root"},
|
||||
)
|
||||
|
|
|
|||
|
|
@ -235,7 +235,7 @@ def test_project_completion_host_salvage_labels_its_bytes_and_points(tmp_path, m
|
|||
assert len(queued) == 1
|
||||
text = queued[0]["text"]
|
||||
assert raw not in text
|
||||
assert f"Reason: provider_unavailable. {SALVAGE_EXCERPT_LABEL}: RAW PATCH" in text
|
||||
assert f"provider_unavailable. {SALVAGE_EXCERPT_LABEL}: RAW PATCH" in text
|
||||
# The excerpt form keeps the pointer too: this writer has no other one.
|
||||
assert text.endswith(" Open the Project for details.")
|
||||
labelled = text.split(f"{SALVAGE_EXCERPT_LABEL}: ", 1)[1]
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
"""The Reason line names the cause the record actually holds (owner item, spam B).
|
||||
"""The cause line names the cause the record actually holds (owner item, spam B).
|
||||
|
||||
``_apply_terminal_custody_outcome`` stamps ``delegated_custody_unreconciled`` as
|
||||
the row's reason_code while a delegated run is still unreconciled. The debt then
|
||||
|
|
@ -17,7 +17,10 @@ lines and would cross the 1000-line target here.
|
|||
from __future__ import annotations
|
||||
|
||||
from ouroboros.outcomes import WARN_DELEGATED_CUSTODY_UNRECONCILED
|
||||
from ouroboros.project_dialogue import _completion_verdict
|
||||
from ouroboros.project_dialogue import TASK_CAUSE_PHRASES, _completion_verdict
|
||||
|
||||
# The one owner sentence for the debt; the raw code stays typed on the row.
|
||||
CUSTODY_SENTENCE = TASK_CAUSE_PHRASES[WARN_DELEGATED_CUSTODY_UNRECONCILED]
|
||||
|
||||
|
||||
def _row(**fields) -> dict:
|
||||
|
|
@ -31,7 +34,7 @@ def test_a_healed_debt_yields_the_execution_reason_instead() -> None:
|
|||
delegated_runs_unreconciled=[],
|
||||
outcome_axes={"execution": {"status": "degraded", "reason_code": "tool_failure"}},
|
||||
)
|
||||
assert _completion_verdict(healed, {}) == "Reason: tool_failure."
|
||||
assert _completion_verdict(healed, {}) == "tool_failure."
|
||||
|
||||
|
||||
def test_an_open_debt_is_still_named() -> None:
|
||||
|
|
@ -39,9 +42,7 @@ def test_an_open_debt_is_still_named() -> None:
|
|||
delegated_runs_unreconciled=["run-a1"],
|
||||
outcome_axes={"execution": {"status": "ok"}},
|
||||
)
|
||||
assert _completion_verdict(open_debt, {}) == (
|
||||
f"Reason: {WARN_DELEGATED_CUSTODY_UNRECONCILED}."
|
||||
)
|
||||
assert _completion_verdict(open_debt, {}) == CUSTODY_SENTENCE
|
||||
|
||||
|
||||
def test_both_real_are_stated_in_one_line() -> None:
|
||||
|
|
@ -49,9 +50,7 @@ def test_both_real_are_stated_in_one_line() -> None:
|
|||
delegated_runs_unreconciled=["run-a1", "run-b2"],
|
||||
outcome_axes={"execution": {"status": "failed", "reason_code": "provider_unavailable"}},
|
||||
)
|
||||
assert _completion_verdict(both, {}) == (
|
||||
f"Reason: provider_unavailable ({WARN_DELEGATED_CUSTODY_UNRECONCILED})."
|
||||
)
|
||||
assert _completion_verdict(both, {}) == f"provider_unavailable ({CUSTODY_SENTENCE})"
|
||||
|
||||
|
||||
def test_a_healed_debt_with_no_execution_cause_states_nothing() -> None:
|
||||
|
|
@ -69,12 +68,12 @@ def test_the_debt_may_arrive_on_the_event_instead_of_the_result() -> None:
|
|||
"delegated_runs_unreconciled": ["run-a1"],
|
||||
"outcome_axes": {"execution": {"status": "ok"}},
|
||||
}
|
||||
assert _completion_verdict({}, event) == f"Reason: {WARN_DELEGATED_CUSTODY_UNRECONCILED}."
|
||||
assert _completion_verdict({}, event) == CUSTODY_SENTENCE
|
||||
|
||||
|
||||
def test_every_other_reason_code_passes_through_untouched() -> None:
|
||||
plain = {"status": "failed", "reason_code": "provider_unavailable"}
|
||||
assert _completion_verdict(plain, {}) == "Reason: provider_unavailable."
|
||||
assert _completion_verdict(plain, {}) == "provider_unavailable."
|
||||
assert _completion_verdict({"status": "completed"}, {}) == ""
|
||||
|
||||
|
||||
|
|
@ -103,7 +102,7 @@ def test_a_healed_debt_is_never_resurrected_by_its_own_frozen_warning() -> None:
|
|||
# While the debt is real the row names it, headline and cause agreeing.
|
||||
owed = _row(outcome_axes=axes, delegated_runs_unreconciled=["run-a1"])
|
||||
assert OUTCOME_PHASE_HEADLINE[outcome_phase(owed, {})] == "Done with warnings"
|
||||
assert _completion_verdict(owed, {}) == f"Reason: {WARN_DELEGATED_CUSTODY_UNRECONCILED}."
|
||||
assert _completion_verdict(owed, {}) == CUSTODY_SENTENCE
|
||||
|
||||
# A row whose axes healed too reads as clean and also states nothing.
|
||||
clean = _row(outcome_axes={"execution": {"status": "ok"}},
|
||||
|
|
@ -115,7 +114,7 @@ def test_a_healed_debt_is_never_resurrected_by_its_own_frozen_warning() -> None:
|
|||
railed = _row(outcome_axes=custody_debt_axes(
|
||||
{"execution": {"status": "failed", "reason_code": "provider_unavailable"}}),
|
||||
delegated_runs_unreconciled=[])
|
||||
assert _completion_verdict(railed, {}) == "Reason: provider_unavailable."
|
||||
assert _completion_verdict(railed, {}) == "provider_unavailable."
|
||||
|
||||
|
||||
def test_the_event_and_the_row_render_one_reason_line_over_the_shared_fixture() -> None:
|
||||
|
|
@ -139,7 +138,7 @@ def test_the_event_and_the_row_render_one_reason_line_over_the_shared_fixture()
|
|||
/ "web" / "tests" / "fixtures" / "outcome_phase_parity.json")
|
||||
cases = json.loads(fixture.read_text(encoding="utf-8"))["cases"]
|
||||
asserted = [case for case in cases if case.get("acceptance_clause")]
|
||||
assert len(asserted) >= 5, "the fixture lost its Reason-line cases"
|
||||
assert len(asserted) >= 5, "the fixture lost its cause-line cases"
|
||||
for case in asserted:
|
||||
record, clause = case["record"], case["acceptance_clause"]
|
||||
assert _completion_verdict(record, {}) == clause, case["name"]
|
||||
|
|
|
|||
|
|
@ -39,7 +39,7 @@ def test_ui_project_completion_pointer_keeps_project_history_scoped(direct_serve
|
|||
task_id = "tower-root-1"
|
||||
append_jsonl(logs / "chat.jsonl", {
|
||||
"ts": "2026-08-22T10:00:00+00:00", "direction": "out", "chat_id": 1,
|
||||
"user_id": 1, "text": f"{target_label} · Completed\nRelease shipped.",
|
||||
"user_id": 1, "text": f"{target_label} · Completed\nOpen the Project for details.",
|
||||
"type": "project_completion_summary", "task_id": task_id,
|
||||
"project_id": project["id"], "project_name": project["name"],
|
||||
"target_label": target_label, "status": "completed",
|
||||
|
|
|
|||
|
|
@ -111,13 +111,16 @@ def test_bound_direct_task_header_and_review_cost_survive_reopen(
|
|||
|
||||
def assert_running(amount):
|
||||
card = page.locator(card_selector)
|
||||
# The subject is a DIRECT turn, so the block wears the compact
|
||||
# chrome (DESIGN.md "Conversation activity block"): no
|
||||
# Working/Done chip, a running indicator while the turn runs.
|
||||
# The phase therefore survives as the record's fact the block
|
||||
# logic reads, not as a visible chip.
|
||||
expect(card).to_have_attribute("data-direct", "1", timeout=10_000)
|
||||
expect(card.locator("[data-live-phase]")).to_have_attribute("data-phase", "working")
|
||||
# The subject is a DIRECT turn with narration rows: content, so
|
||||
# the block wears the task-card chrome (DESIGN.md "Conversation
|
||||
# activity block", owner decision 16.09) — a visible Working
|
||||
# chip, the coined title, a running indicator while it runs —
|
||||
# and, inside a Project panel, no conversion control.
|
||||
expect(card.locator("[data-live-phase]")).to_have_attribute("data-phase", "working", timeout=10_000)
|
||||
expect(card.locator("[data-live-phase]")).to_be_visible()
|
||||
expect(card.locator("[data-live-phase]")).to_have_text("Working")
|
||||
expect(card.locator("[data-live-title]")).to_have_text("Analyze greeting context")
|
||||
expect(card.locator("[data-turn-into-project]")).to_have_count(0)
|
||||
expect(card.locator("[data-live-typing]")).to_be_visible()
|
||||
expect(card.locator("[data-live-meta]")).not_to_contain_text("Activity unconfirmed")
|
||||
expect(card.locator("[data-live-meta]")).to_contain_text(f"up to ${amount}")
|
||||
|
|
@ -158,6 +161,8 @@ def test_bound_direct_task_header_and_review_cost_survive_reopen(
|
|||
card = page.locator(card_selector)
|
||||
expect(card).to_have_attribute("data-finished", "1", timeout=15_000)
|
||||
expect(card.locator("[data-live-phase]")).to_have_attribute("data-phase", "done")
|
||||
expect(card.locator("[data-live-phase]")).to_be_visible()
|
||||
expect(card.locator("[data-live-phase]")).to_have_text("Done")
|
||||
expect(card.locator("[data-live-typing]")).not_to_be_visible()
|
||||
page.evaluate(
|
||||
"row => window.__ouroWs.emit('log', {chat_id: row.chat_id, data: row})",
|
||||
|
|
|
|||
|
|
@ -907,7 +907,7 @@ def test_only_task_acceptance_fail_with_correction_rail_is_a_valid_veto(tmp_path
|
|||
)
|
||||
assert tier_veto.aggregate_signal == "FAIL"
|
||||
assert tier_veto.actors[2]["signal"] == "FAIL"
|
||||
assert "best_effort" in build_improvement_capsule(tier_veto)
|
||||
assert "rated a partial result" in build_improvement_capsule(tier_veto) # the tier in words, never the identifier
|
||||
|
||||
unanimous_minimal_fail = run_review_request(
|
||||
request,
|
||||
|
|
|
|||
|
|
@ -396,7 +396,13 @@ def test_capsule_leads_with_verdict_blocker_rails_and_three_moves():
|
|||
rails_line="money: $1.00 spent; time: 10 min left; review passes: 1 done",
|
||||
open_obligations=[{"id": "ob-1"}, {"id": "ob-2"}],
|
||||
)
|
||||
assert "Review verdict: FAIL (tier: best_effort)" in capsule
|
||||
# The note the model READS says the assessment in words: a ledger token in
|
||||
# this stream is text the model parrots back to its human (v6.61.4).
|
||||
assert "Review verdict: FAIL — rated a partial result" in capsule
|
||||
assert "(tier: best_effort)" not in capsule
|
||||
# The whole header is identifier-free: panel_reason used to restate `tier=…`.
|
||||
header = capsule.splitlines()[0]
|
||||
assert "tier=" not in header and "best_effort" not in header and "blocked_with_evidence" not in header, header
|
||||
assert "Open blocking obligation(s) (2): ob-1, ob-2." in capsule
|
||||
assert "Remaining headroom — money: $1.00 spent" in capsule
|
||||
assert "FIX" in capsule and "REBUT" in capsule and "DECLARE UNREACHABLE" in capsule
|
||||
|
|
|
|||
|
|
@ -758,23 +758,32 @@ export function createChatInstance({
|
|||
// mutation, no sticky flag; the completion note and receipt rows are not content.
|
||||
function blockVisible(record) {
|
||||
if (!record || record.isSubagent) return true;
|
||||
const id = record.groupId;
|
||||
return !!(record.modelWaiting || record.cancelPendingPolicy || record.reviewAnchor)
|
||||
return blockHasWork(record)
|
||||
|| !!(record.modelWaiting || record.cancelPendingPolicy || record.reviewAnchor)
|
||||
|| stopEligible(record)
|
||||
|| record.reviewController?.groups.size > 0
|
||||
|| [...subagentChildParents.values()].some((info) => info.parentId === id)
|
||||
|| record.items.some((item) => !item.receipt && !String(item.dedupeKey || '').startsWith('task_done|'))
|
||||
|| record.toolErrors > 0
|
||||
|| (record.finished && record.phaseEl?.dataset?.phase !== 'done');
|
||||
}
|
||||
|
||||
// The host's direct-turn fact (census kind, rebuilt task_done, history
|
||||
// rows) selects the block's chrome; it never decides presence.
|
||||
// The host's lane fact (census kind, rebuilt task_done, history rows) is
|
||||
// kept on the record for the header pill only: a direct turn keeps the
|
||||
// census verdict (Thinking…) beside its block. It never chooses chrome.
|
||||
function noteDirectTurn(record, direct) {
|
||||
if (!record || typeof direct !== 'boolean' || record.direct === direct) return;
|
||||
record.direct = direct;
|
||||
if (direct && !record.suggestedName && !record.lastHumanHeadline) record.titleEl.textContent = '';
|
||||
ensureLiveCardVisible(record);
|
||||
if (record && typeof direct === 'boolean') record.direct = direct;
|
||||
}
|
||||
|
||||
// The work the block stands on — the presence facts minus open attention and
|
||||
// minus a bare terminal outcome. It selects the chrome (owner decision 16.09):
|
||||
// a block with work is the task card whatever lane produced it (a title, the
|
||||
// conversion control in Main); a block that exists only for open attention or
|
||||
// a non-Done ending keeps no title placeholder and offers no conversion. The
|
||||
// host's lane fact (`_is_direct_chat`) keeps its host jobs and never chooses chrome.
|
||||
function blockHasWork(record) {
|
||||
const id = record.groupId;
|
||||
// No lane is always shown: the retired bg-consciousness card kind is gone.
|
||||
return record.reviewController?.groups.size > 0
|
||||
|| [...subagentChildParents.values()].some((info) => info.parentId === id)
|
||||
|| record.items.some((item) => !item.receipt && !String(item.dedupeKey || '').startsWith('task_done|'))
|
||||
|| record.toolErrors > 0;
|
||||
}
|
||||
|
||||
// Tool accounting from a metrics or terminal fact: the meta counts and, with
|
||||
|
|
@ -883,16 +892,20 @@ export function createChatInstance({
|
|||
}
|
||||
}
|
||||
|
||||
// The block's chrome from the record's facts: a direct conversation turn
|
||||
// renders the compact activity block (no title placeholder, no status
|
||||
// chip, no conversion); a managed/Swarm root keeps the task card with
|
||||
// "Turn into project" unless its origin is already bound to a Project —
|
||||
// the one binding fact the /api/state sweep in app.js reads too.
|
||||
// The conversion control from the record's facts: a block with work in Main
|
||||
// offers "Turn into project" unless its origin is already bound to a Project
|
||||
// (the one binding fact the /api/state sweep in app.js reads too); the title
|
||||
// writers apply the same work predicate.
|
||||
function syncBlockChrome(record) {
|
||||
if (record.isSubagent || record.root.dataset.projectCreated === '1') return;
|
||||
const direct = record.direct ? '1' : '0';
|
||||
if (record.root.dataset.direct !== direct) record.root.dataset.direct = direct;
|
||||
const wanted = isMain && !record.direct && record.root.dataset.projectBound !== '1'
|
||||
const work = blockHasWork(record);
|
||||
// The first row of work lands after the title writers ran for its frame:
|
||||
// an empty title takes the placeholder here, the writers own it from then on.
|
||||
if (work && !record.titleEl.textContent) {
|
||||
record.titleEl.textContent = record.suggestedName || record.lastHumanHeadline
|
||||
|| (record.finished ? 'Task activity' : 'Working...');
|
||||
}
|
||||
const wanted = isMain && work && record.root.dataset.projectBound !== '1'
|
||||
&& !(window.__ouroTaskBindings || {})[record.groupId];
|
||||
if (wanted === Boolean(record.turnProjectBtn)) return;
|
||||
if (!wanted) {
|
||||
|
|
@ -1291,8 +1304,8 @@ export function createChatInstance({
|
|||
// The proactively-coined LLM name; becomes the card title when set.
|
||||
suggestedName: '',
|
||||
reviewOwnerDetailObserved: false,
|
||||
// Host fact: a direct conversation turn (census kind, rebuilt
|
||||
// task_done, history rows) vs a managed/Swarm root.
|
||||
// Host fact for the header pill: a direct conversation turn (census
|
||||
// kind, rebuilt task_done, history rows) vs a managed/Swarm root.
|
||||
direct: String(activeDirectActivities.get(normalizedGroupId)?.kind || 'managed_task') !== 'managed_task',
|
||||
};
|
||||
const reviewDisclosure = reviewDisclosureByTask.get(normalizedGroupId) || {
|
||||
|
|
@ -1480,7 +1493,7 @@ export function createChatInstance({
|
|||
// P1: last bounded activity projection (remembered even while
|
||||
// the collapsed line is suppressed on unnamed root cards) + sticky cost.
|
||||
clearStickyCardState(record);
|
||||
record.titleEl.textContent = record.direct ? '' : 'Working...';
|
||||
record.titleEl.textContent = record.suggestedName || (blockHasWork(record) ? 'Working...' : '');
|
||||
setLiveCardPhase(record, 'working');
|
||||
record.countEl.hidden = true;
|
||||
record.countEl.textContent = '0 notes';
|
||||
|
|
@ -1675,10 +1688,11 @@ export function createChatInstance({
|
|||
const desiredPhase = desiredLiveCardPhase(record, activePhase);
|
||||
setLiveCardPhase(record, desiredPhase.phase, desiredPhase.text, desiredPhase.className);
|
||||
// A coined project name takes the title slot (the activity headline stays in the
|
||||
// timeline); a child's title is its lineage identity; a direct block carries only
|
||||
// a real narration headline; otherwise the activity headline.
|
||||
// timeline); a child's title is its lineage identity; a block without work
|
||||
// (open attention, a bare non-Done ending) carries no title; otherwise the
|
||||
// activity headline.
|
||||
const title = record.suggestedName || (record.isSubagent ? childTitle(record)
|
||||
: record.direct ? record.lastHumanHeadline
|
||||
: !blockHasWork(record) ? ''
|
||||
: (record.finished ? record.lastHumanHeadline || 'Task activity' : activeHeadline));
|
||||
if (record.titleEl.textContent !== title) record.titleEl.textContent = title;
|
||||
// The collapsed line is a compact presentation projection, while the
|
||||
|
|
@ -1784,7 +1798,7 @@ export function createChatInstance({
|
|||
if (record.isSubagent) record.titleEl.textContent = childTitle(record);
|
||||
else if (!record.suggestedName && !record.lastHumanHeadline
|
||||
&& record.titleEl.textContent !== presentation.headline) {
|
||||
record.titleEl.textContent = record.direct ? '' : 'Task activity';
|
||||
record.titleEl.textContent = blockHasWork(record) ? 'Task activity' : '';
|
||||
}
|
||||
settleLiveCard(record, activePhase, wasFinished);
|
||||
ensureLiveCardVisible(record);
|
||||
|
|
@ -2113,9 +2127,9 @@ export function createChatInstance({
|
|||
if (childInfo) return Boolean(changed || queued);
|
||||
const subagentChanged = updateSubagentCardFromEvent(evt, rawTs);
|
||||
// The host stamps the lane on the turn's own frames (task_done always,
|
||||
// a direct turn's tool frames too), so chrome never waits for a census;
|
||||
// the host-attested Stop marker rides a direct turn's tool frames the
|
||||
// way it rides its narration rows, so a tool-only turn offers Stop.
|
||||
// a direct turn's tool frames too), so the header pill never waits for
|
||||
// a census; the host-attested Stop marker rides a direct turn's tool
|
||||
// frames the way it rides its narration rows, so a tool-only turn offers Stop.
|
||||
if (typeof evt._is_direct_chat === 'boolean') noteDirectTurn(liveCardRecords.get(taskId), evt._is_direct_chat);
|
||||
if (evt.cancelable === true) markTaskCancelable(taskId);
|
||||
if (eventType === 'task_done' && summary.terminal) {
|
||||
|
|
@ -3392,7 +3406,7 @@ export function createChatInstance({
|
|||
|
||||
// Mounted, unfinished, not waiting: the blocks that host their own running
|
||||
// indicator. Only a managed root among them drives the header's Working…;
|
||||
// a direct block keeps the census verdict (Thinking…).
|
||||
// a direct turn's block keeps the census verdict (Thinking…).
|
||||
const foregroundCards = () => Array.from(liveCardRecords.values()).filter((r) => isForegroundLiveCard(r) && !r.modelWaiting);
|
||||
function hasActiveLiveCard() {
|
||||
return foregroundCards().some((r) => !r.direct);
|
||||
|
|
|
|||
|
|
@ -404,21 +404,51 @@ export function taskStoppedWithSummary(evt) {
|
|||
return String(evt?.reason_code || '') === 'owner_requested_finalization';
|
||||
}
|
||||
|
||||
// The typed degradation causes a card can state in the owner's words. The record
|
||||
// keeps the machine code (Logs, task detail, benchmark ledgers); only the card
|
||||
// speaks. An UNKNOWN code stays raw on purpose: a reason we have no sentence for
|
||||
// must read as itself rather than as a wrong sentence.
|
||||
const TASK_REASON_PHRASES = {
|
||||
plan_review_advisory: 'plan review never closed; the work continued under advisory enforcement',
|
||||
host_child_status_suffix: 'a child task had not settled when the answer was delivered',
|
||||
invalid_delivery_control_after_repair: 'the delivery control object was still malformed after repair',
|
||||
budget_exhausted: 'the task ran out of budget before it could finish cleanly',
|
||||
delivery_control_degraded: 'delivery finished in a degraded control state',
|
||||
// The typed causes a card can state in the owner's words, keyed on the CODE
|
||||
// alone. The record keeps the machine code (Logs, task detail, benchmark
|
||||
// ledgers); only the card speaks. An UNKNOWN code stays raw on purpose: a
|
||||
// reason we have no sentence for must read as itself rather than as a wrong
|
||||
// sentence. The byte-identical twin of project_dialogue.TASK_CAUSE_PHRASES;
|
||||
// web/tests/fixtures/outcome_phase_parity.json pins both.
|
||||
const TASK_CAUSE_PHRASES = {
|
||||
previous_revision_accepted: "The reviewers approved the earlier version of this answer; it changed before they finished.",
|
||||
author_finish: "The answer was delivered on Main's own judgement; the reviewers had not signed it off.",
|
||||
review_degraded: "No reviewer verdict was established for this answer.",
|
||||
infra_failure: "A review infrastructure failure prevented a settled verdict.",
|
||||
dialogue_terminal: "The reviewers and Main could not agree, and both positions were kept.",
|
||||
improvement_capsule: "The reviewers asked for one more pass and Main was given their notes.",
|
||||
fence_reopen_failed: "The requested extra pass could not be started, so the answer stands as it was.",
|
||||
review_cycles_exhausted: "The task used up its review rounds before the answer was signed off.",
|
||||
open_obligations: "The answer was delivered with reviewer requests still open.",
|
||||
improvement_window_closed: "There was no room left for another pass, so the answer stands as it was.",
|
||||
capsule_spent: "The one allowed improvement pass was already used.",
|
||||
reviewer_fail_no_capsule: "A reviewer rejected the answer and suggested nothing to change.",
|
||||
no_actionable_changes: "The re-review was not clean and suggested nothing to change.",
|
||||
identical_acceptance_refused: "Nothing had changed since the last review, so the recorded verdict stands.",
|
||||
review_skipped_deadline_reserve: "There was not enough time left to review the answer.",
|
||||
delivery_binding_superseded: "The answer or its evidence changed, so the earlier review no longer covered it.",
|
||||
owner_followup: "A new message from you arrived, so the review was set aside for it.",
|
||||
evidence_refresh: "The work changed after the review was frozen, so it no longer covered the answer.",
|
||||
revision_unavailable_on_forced_rail: "The task had to stop, so the requested rework never happened.",
|
||||
owner_hurry: "You asked me to hurry, so no further review was started.",
|
||||
unspecified: "The answer was not signed off, and no cause was recorded.",
|
||||
acceptance_bypassed_budget_exhausted: "The task ran out of budget before the answer could be reviewed.",
|
||||
acceptance_bypassed_round_limit: "The task hit its round limit before the answer could be reviewed.",
|
||||
acceptance_bypassed_deadline: "The task ran out of time before the answer could be reviewed.",
|
||||
acceptance_bypassed_provider_unavailable: "The model provider was unavailable, so the answer was never reviewed.",
|
||||
acceptance_bypassed_context_overflow: "The task outgrew its context before the answer could be reviewed.",
|
||||
acceptance_bypassed_children_unabsorbed: "Some sub-tasks had not been folded in, so the answer was never reviewed.",
|
||||
plan_review_advisory: "Plan review never closed; the work continued under advisory enforcement",
|
||||
host_child_status_suffix: "A child task had not settled when the answer was delivered",
|
||||
invalid_delivery_control_after_repair: "The delivery control object was still malformed after repair",
|
||||
budget_exhausted: "The task ran out of budget before it could finish cleanly",
|
||||
delivery_control_degraded: "Delivery finished in a degraded control state",
|
||||
delegated_custody_unreconciled: "Some delegated work was never reconciled.",
|
||||
};
|
||||
|
||||
export function taskReasonPhrase(code) {
|
||||
const raw = String(code || '');
|
||||
return TASK_REASON_PHRASES[raw] || raw;
|
||||
return TASK_CAUSE_PHRASES[raw] || raw;
|
||||
}
|
||||
|
||||
// The custody overlay stamps this code as the row's reason_code while a
|
||||
|
|
@ -456,9 +486,13 @@ export function taskReasonDetail(evt) {
|
|||
const decision = record.outcome_axes?.review?.acceptance_decision
|
||||
?? record.review_status?.acceptance_decision;
|
||||
const severity = taskOutcomeSeverity(evt);
|
||||
if (severity !== 'error' && severity !== 'cancelled' && decision?.status && decision.status !== 'accepted') {
|
||||
const rationale = String(decision.rationale || '').split(/\s+/).filter(Boolean).join(' ');
|
||||
return `Acceptance: ${decision.status}${rationale ? ` — ${rationale}` : ''}`;
|
||||
const decisionCause = String(decision?.reason || '');
|
||||
if (severity !== 'error' && severity !== 'cancelled' && decision?.status
|
||||
&& (decision.status !== 'accepted' || Object.hasOwn(TASK_CAUSE_PHRASES, decisionCause))) {
|
||||
// The decision's own typed reason speaks (an accepted decision only when
|
||||
// it has a sentence); the stored reviewer rationale stays in the card
|
||||
// body, the task result and Logs.
|
||||
return taskReasonPhrase(decisionCause);
|
||||
}
|
||||
if (!evt?.reason_code || evt.reason_code === 'final_message') return '';
|
||||
// A healed debt is never restored: naming it again would state a debt the
|
||||
|
|
@ -466,12 +500,12 @@ export function taskReasonDetail(evt) {
|
|||
// there is one, otherwise the row states no cause and leaves the headline
|
||||
// to the frozen outcome axis that owns it.
|
||||
const [reason, custody] = custodyDebtReason(record);
|
||||
if (!reason) return custody ? `Reason: ${custody}` : '';
|
||||
if (!reason) return taskReasonPhrase(custody);
|
||||
const receiptVeto = record.outcome_axes?.objective?.receipt_veto;
|
||||
const cause = receiptVeto?.reason === reason && receiptVeto.detail
|
||||
? String(receiptVeto.detail).split(/\s+/).filter(Boolean).join(' ')
|
||||
: taskReasonPhrase(reason);
|
||||
return `Reason: ${cause}${custody ? ` (${custody})` : ''}`;
|
||||
return `${cause}${custody ? ` (${taskReasonPhrase(custody)})` : ''}`;
|
||||
}
|
||||
|
||||
// S3 (HQ1): the ONE shared projection of a typed owner_hurry event for the
|
||||
|
|
@ -1133,7 +1167,7 @@ function summarizeChatLiveEventView(evt) {
|
|||
const resultText = describeText(evt.result || '', 320, { markdown: true });
|
||||
const traceText = describeText(evt.trace_summary || '', 320);
|
||||
const errorText = describeText(evt.error || '', 220);
|
||||
const reasonDetail = evt.reason_code ? `Reason: ${taskReasonPhrase(evt.reason_code)}` : '';
|
||||
const reasonDetail = evt.reason_code ? taskReasonPhrase(evt.reason_code) : '';
|
||||
const detailParts = [
|
||||
progressText.full,
|
||||
resultText.full ? `[RESULT]\n${resultText.full}` : '',
|
||||
|
|
|
|||
|
|
@ -99,7 +99,8 @@ export function createModelWaitController({ getRecord, onDomWrite = (fn) => fn()
|
|||
if (record && !record.finished) {
|
||||
record.modelWaiting = waiting;
|
||||
record.root.dataset.modelWaiting = waiting ? '1' : '0';
|
||||
if (waiting && !record.suggestedName && !record.lastHumanHeadline && !record.direct) record.titleEl.textContent = 'Task';
|
||||
// A block without work carries no title placeholder while it waits: the chrome
|
||||
// follows the work it stands on, never the lane (and no lane is always shown).
|
||||
const phase = desiredLiveCardPhase(record);
|
||||
setLiveCardPhase(record, phase.phase, phase.text, phase.className);
|
||||
}
|
||||
|
|
|
|||
|
|
@ -308,20 +308,12 @@ body.resizing-panels { user-select: none; }
|
|||
}
|
||||
|
||||
|
||||
/* A direct conversation turn renders the same component as a compact activity
|
||||
block (docs/DESIGN.md "Conversation activity block"): no Task/Working/Done
|
||||
chip while it runs or once it is Done (Cancelling…/Finalizing…/Waiting and a
|
||||
non-Done outcome stay visible), no title placeholder, the typing dots as its
|
||||
running indicator; the collapsed header carries the count, cost and duration. */
|
||||
.chat-live-card[data-direct="1"] {
|
||||
border-color: var(--accent-08);
|
||||
}
|
||||
|
||||
.chat-live-card[data-direct="1"] > .chat-live-summary-button [data-live-phase]:is(.done, .working:not(.cancelling):not(.finalizing):not(.waiting)) {
|
||||
display: none;
|
||||
}
|
||||
|
||||
.chat-live-card[data-direct="1"] > .chat-live-summary-button .chat-live-title:empty {
|
||||
/* A block that exists only for open attention (a model wait, a pending or
|
||||
host-offered Stop) or a bare non-Done ending carries no title placeholder
|
||||
(docs/DESIGN.md "Conversation activity block"); its chip says the state it is
|
||||
in (Waiting…, Cancelling…, Failed). The first row of work — a tool call,
|
||||
narration, a child, a review, an error — gives it a title, whatever lane ran it. */
|
||||
.chat-live-card > .chat-live-summary-button .chat-live-title:empty {
|
||||
display: none;
|
||||
}
|
||||
|
||||
|
|
@ -1845,6 +1837,14 @@ body.resizing-panels { user-select: none; }
|
|||
min-height: 1.45em;
|
||||
}
|
||||
|
||||
/* A block without work carries no title placeholder: the collapsed geometry
|
||||
reserves nothing for an empty title (review round 1: the clamp and the
|
||||
reserved line above outranked the general empty-title rule). */
|
||||
.chat-live-card:not([data-expanded="1"]) > .chat-live-summary-button .chat-live-title:empty {
|
||||
display: none;
|
||||
min-height: 0;
|
||||
}
|
||||
|
||||
/* Activity grows with useful content. Details retain the complete source. */
|
||||
.chat-live-activity {
|
||||
color: var(--text-meta);
|
||||
|
|
|
|||
|
|
@ -8,7 +8,7 @@
|
|||
// block leaves with its wait and its resolved episode cannot reopen it; the
|
||||
// header keeps the census verdict beside a block.
|
||||
import assert from 'node:assert/strict';
|
||||
import test from 'node:test';
|
||||
import test, { after } from 'node:test';
|
||||
import { createChatInstance } from '../modules/chat.js';
|
||||
import { summarizeChatLiveEvent } from '../modules/log_events.js';
|
||||
import { ElementStub, installDom, restoreDom, walkCard } from './chat_dom_fixture.js';
|
||||
|
|
@ -16,6 +16,7 @@ import { ElementStub, installDom, restoreDom, walkCard } from './chat_dom_fixtur
|
|||
// The flat fixture's querySelector does not descend; the status badge and
|
||||
// card internals need a real descendant lookup.
|
||||
const originalQuery = ElementStub.prototype.querySelector;
|
||||
after(() => { ElementStub.prototype.querySelector = originalQuery; });
|
||||
ElementStub.prototype.querySelector = function (selector) {
|
||||
const direct = originalQuery.call(this, selector);
|
||||
if (direct) return direct;
|
||||
|
|
@ -108,17 +109,17 @@ test('a wait with zero tools opens a block with the controls and leaves with the
|
|||
} finally { f.close(); }
|
||||
});
|
||||
|
||||
test('a tool frame stamped with the lane fact mints a direct block before any census lists the turn', () => {
|
||||
test('a tool frame mints the task card before any census lists the turn: a tool row is content', () => {
|
||||
const f = fixture();
|
||||
try {
|
||||
f.log({ type: 'tool_call_started', tool: 'read_file', tool_call_id: 'c1', args: { path: 'README.md' }, _is_direct_chat: true });
|
||||
assert.ok(f.card(), 'the first tool call mints the block');
|
||||
assert.equal(f.card().dataset.direct, '1', 'the frame carries the lane, the census is not awaited');
|
||||
assert.equal(f.card().querySelector('[data-turn-into-project]'), null, 'no conversion control in the pre-census window');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'), 'conversion is offered in Main from the first content row');
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, 'Working...', 'the placeholder title of a running task card');
|
||||
f.log({ type: 'tool_call_finished', tool: 'read_file', tool_call_id: 'c1', args: { path: 'README.md' }, duration_sec: 0.3, _is_direct_chat: true });
|
||||
f.census(direct());
|
||||
assert.equal(f.card().dataset.direct, '1');
|
||||
assert.equal(f.card().querySelector('[data-turn-into-project]'), null);
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, 'Working...', 'the census lane fact does not change chrome');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'));
|
||||
} finally { f.close(); }
|
||||
});
|
||||
|
||||
|
|
@ -130,7 +131,7 @@ test('a tool-only direct turn offers Stop from its stamped tool frame, without a
|
|||
f.log({ type: 'tool_call_started', tool: 'read_file', tool_call_id: 'c1', args: { path: 'README.md' }, _is_direct_chat: true, cancelable: true });
|
||||
assert.ok(f.card(), 'the tool row mints the block');
|
||||
assert.ok(f.card().querySelector('[data-cancel-run]'), 'the host marker on the work frame offers Stop');
|
||||
assert.equal(f.card().dataset.direct, '1');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'), 'a tool row is work: conversion is offered');
|
||||
f.census(direct());
|
||||
assert.ok(f.card().querySelector('[data-cancel-run]'));
|
||||
f.log({ ...final, type: 'task_done', status: 'completed', _is_direct_chat: true });
|
||||
|
|
@ -138,7 +139,7 @@ test('a tool-only direct turn offers Stop from its stamped tool frame, without a
|
|||
} finally { f.close(); }
|
||||
});
|
||||
|
||||
test('a direct turn with two successful tools shows two compact rows live and the summary row after a reload', async () => {
|
||||
test('a direct turn with two successful tools is the task card: two compact rows live and the summary row after a reload', async () => {
|
||||
const f = fixture();
|
||||
try {
|
||||
f.census(direct());
|
||||
|
|
@ -147,9 +148,8 @@ test('a direct turn with two successful tools shows two compact rows live and th
|
|||
f.log({ type: 'tool_call_started', tool: 'web_search', tool_call_id: 'c2', args: { query: 'ouroboros' } });
|
||||
f.log({ type: 'tool_call_finished', tool: 'web_search', tool_call_id: 'c2', args: { query: 'ouroboros' }, duration_sec: 1.2 });
|
||||
assert.ok(f.card(), 'the first tool call mints the block live');
|
||||
assert.equal(f.card().dataset.direct, '1', 'direct chrome from the census fact');
|
||||
assert.equal(f.card().querySelector('[data-turn-into-project]'), null, 'no conversion on a direct block');
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, '', 'no placeholder title');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'), 'a working direct turn is offered conversion in Main');
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, 'Working...', 'the running placeholder title');
|
||||
assert.equal(f.rows().length, 2, 'start and finish of one call share a row');
|
||||
assert.match(f.rows()[0].innerHTML, /read_file · README\.md/);
|
||||
assert.match(f.rows()[1].innerHTML, /web_search · ouroboros/);
|
||||
|
|
@ -173,8 +173,8 @@ test('a direct turn with two successful tools shows two compact rows live and th
|
|||
try {
|
||||
await g.instance.refreshHistory({ revision: 1 });
|
||||
assert.ok(g.card(), 'the same turn shows the block after a reload');
|
||||
assert.equal(g.card().dataset.direct, '1', 'replay reads the same host fact from the summary row');
|
||||
assert.equal(g.card().querySelector('[data-turn-into-project]'), null);
|
||||
assert.ok(g.card().querySelector('[data-turn-into-project]'), 'conversion survives a reload');
|
||||
assert.notEqual(g.card().querySelector('[data-live-title]').textContent, '', 'a finished task card carries a title');
|
||||
// Per-tool rows are live-only: replay carries the summary row and the completion note.
|
||||
assert.equal(g.rows().length, 2);
|
||||
assert.ok(g.rows().some((n) => /2 tool calls/.test(n.innerHTML)));
|
||||
|
|
@ -213,7 +213,6 @@ test('a managed Swarm root keeps the task card with Turn into project; an origin
|
|||
f.census(managed());
|
||||
f.emit('chat', { task_id: TASK, role: 'assistant', is_progress: true, content: 'Planning the swarm.' });
|
||||
assert.ok(f.card());
|
||||
assert.equal(f.card().dataset.direct, '0');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'));
|
||||
assert.equal(f.status(), 'Working...');
|
||||
globalThis.window.__ouroTaskBindings = { 'bound-root': { project_id: 'p1', chat_id: 7 } };
|
||||
|
|
@ -291,6 +290,8 @@ test('a zero-tool turn that ended failed keeps its block live and after a reload
|
|||
f.log({ ...failed, type: 'task_done', status: 'failed' });
|
||||
assert.ok(f.card(), 'a terminal outcome other than Done is its own reason to exist');
|
||||
assert.equal(phaseOf(f.card()), 'error');
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, '', 'a failed greeting keeps no title placeholder');
|
||||
assert.equal(f.card().querySelector('[data-turn-into-project]'), null, 'and offers no conversion: it did no work');
|
||||
assert.equal(f.rows().length, 1, 'only the terminal note, which the predicate never counts as content');
|
||||
assert.match(f.rows()[0].innerHTML, />Failed</);
|
||||
} finally { f.close(); }
|
||||
|
|
@ -304,6 +305,7 @@ test('a zero-tool turn that ended failed keeps its block live and after a reload
|
|||
await g.instance.refreshHistory({ revision: 1 });
|
||||
assert.ok(g.card(), 'presence is the same on reload, with no replay-only force branch');
|
||||
assert.equal(phaseOf(g.card()), 'error');
|
||||
assert.equal(g.card().querySelector('[data-turn-into-project]'), null);
|
||||
} finally { g.close(); }
|
||||
});
|
||||
|
||||
|
|
@ -315,6 +317,7 @@ test('an acceptance review row keeps a block for a zero-tool turn', () => {
|
|||
{ surface: 'task_acceptance', panel_id: 'p1', aggregate_signal: 'PASS', reason: 'Answer matches the ask.' },
|
||||
] } });
|
||||
assert.ok(f.card(), 'the review is content the block stands on');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'), 'a review group is work: conversion is offered');
|
||||
assert.equal(f.card().dataset.finished, '1');
|
||||
assert.match(f.card().querySelector('[data-live-review-summary]')?.textContent || '', /Reviews 1/);
|
||||
} finally { f.close(); }
|
||||
|
|
@ -404,12 +407,63 @@ test('Stop stays reachable on a census-vouched root, managed card and direct blo
|
|||
const stop = f.card('live-root').querySelector('[data-cancel-run]');
|
||||
assert.ok(stop, `the census restores Stop on a ${kind} root`);
|
||||
assert.equal(stop.textContent, 'Stop…');
|
||||
assert.equal(f.card('live-root').dataset.direct, kind === 'managed_task' ? '0' : '1');
|
||||
assert.ok(f.card('live-root').querySelector('[data-turn-into-project]'), `narration is work on a ${kind} root`);
|
||||
assert.doesNotMatch(f.meta('live-root'), /unconfirmed|unavailable/);
|
||||
} finally { f.close(); }
|
||||
}
|
||||
});
|
||||
|
||||
// Owner decision 16.09 (Q1=A): chrome follows content, not the lane. A block
|
||||
// that exists only for open attention is compact; its first content row makes
|
||||
// it the task card. The header pill keeps reading the lane fact (Thinking…).
|
||||
test('a wait-only block carries no title and no conversion until its first row of work', () => {
|
||||
const f = fixture();
|
||||
try {
|
||||
f.census(direct());
|
||||
f.emit('chat', wait({ role: 'system' }));
|
||||
const card = f.card();
|
||||
assert.equal(card.querySelector('[data-live-title]').textContent, '', 'no placeholder title without work');
|
||||
assert.equal(card.querySelector('[data-turn-into-project]'), null, 'nothing to convert yet');
|
||||
f.emit('chat', wait({ role: 'system', revision: 2, state: 'resolved', resolution: 'quota_restored' }));
|
||||
assert.equal(f.card(), null, 'the block leaves with its attention');
|
||||
f.log({ type: 'tool_call_started', tool: 'read_file', tool_call_id: 'c1', args: { path: 'README.md' } });
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, 'Working...');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'));
|
||||
assert.equal(f.status(), 'Thinking...', 'the header keeps the census verdict for a direct turn');
|
||||
} finally { f.close(); }
|
||||
});
|
||||
|
||||
test('a managed root waiting for access before its first row shows its admission name and no conversion until work arrives', () => {
|
||||
const f = fixture();
|
||||
try {
|
||||
f.census(managed());
|
||||
f.emit('task_named', { task_id: TASK, suggested_name: 'Ship release' });
|
||||
f.emit('chat', wait({ role: 'system' }));
|
||||
const card = f.card();
|
||||
assert.equal(card.querySelector('[data-live-title]').textContent, 'Ship release', 'the admission name is still the title');
|
||||
assert.equal(card.querySelector('[data-turn-into-project]'), null);
|
||||
f.emit('chat', { task_id: TASK, role: 'assistant', is_progress: true, content: 'Working on it.' });
|
||||
assert.ok(card.querySelector('[data-turn-into-project]'));
|
||||
assert.equal(card.querySelector('[data-live-title]').textContent, 'Ship release');
|
||||
} finally { f.close(); }
|
||||
});
|
||||
|
||||
test('a coined name titles a direct task card live and stays after its final', () => {
|
||||
const f = fixture();
|
||||
try {
|
||||
f.census(direct());
|
||||
f.log({ type: 'tool_call_started', tool: 'read_file', tool_call_id: 'c1', args: { path: 'README.md' } });
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, 'Working...');
|
||||
f.emit('task_named', { task_id: TASK, suggested_name: 'Проверка карточки' });
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, 'Проверка карточки');
|
||||
f.emit('chat', { ...final, tool_calls: 1 });
|
||||
assert.equal(f.card().dataset.finished, '1');
|
||||
assert.equal(f.card().querySelector('[data-live-title]').textContent, 'Проверка карточки');
|
||||
assert.equal(f.card().querySelector('[data-live-phase]').textContent, 'Done', 'the Done chip is the card\'s own');
|
||||
assert.ok(f.card().querySelector('[data-turn-into-project]'));
|
||||
} finally { f.close(); }
|
||||
});
|
||||
|
||||
// Promote-refusal placement (Q1=A): a refused addressing call inside a live
|
||||
// turn is a FAILED call — an error row, content the block stands on — while
|
||||
// the owner message carries the host's sentence (`cause`) as its receipt.
|
||||
|
|
|
|||
|
|
@ -3,6 +3,11 @@ import test from 'node:test';
|
|||
import { createChatInstance } from '../modules/chat.js';
|
||||
import { installDom, restoreDom, walkCard } from './chat_dom_fixture.js';
|
||||
|
||||
// The stub's querySelector reads direct children only; the conversion button
|
||||
// lives inside the card's actions row, so find it by walking the subtree.
|
||||
const convertButton = (node) => (node?.dataset && Object.hasOwn(node.dataset, 'turnIntoProject') ? node
|
||||
: (node?.children || []).map(convertButton).find(Boolean) || null);
|
||||
|
||||
// A typing frame no longer writes the client live-set: liveness is a projection
|
||||
// of the /api/state census. A PARTIAL census listing is the census-shaped
|
||||
// equivalent of the old typing-frame write — it inserts the activity and
|
||||
|
|
@ -1073,8 +1078,9 @@ test('history rebuild keeps a lineage-known branch nested, never appended top-le
|
|||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Direct-turn tool work, typed conclusions and accounting render the compact
|
||||
// activity block: no conversion, no title placeholder, the same card component.
|
||||
// Direct-turn tool work, typed conclusions and accounting render the same task
|
||||
// card as a managed root (chrome follows content, owner decision 16.09); Cancel
|
||||
// still needs the host's marker.
|
||||
// ---------------------------------------------------------------------------
|
||||
test('a direct turn renders tool work as an activity block and needs host authority for Cancel', async () => {
|
||||
const { prior, mount } = installDom(async () => ({ ok: true, json: async () => ({ active_direct_turns: [] }) }));
|
||||
|
|
@ -1090,7 +1096,7 @@ test('a direct turn renders tool work as an activity block and needs host author
|
|||
};
|
||||
let instance;
|
||||
try {
|
||||
// Main chat: the only surface that offers "Turn into project" — to managed roots.
|
||||
// Main chat: the only surface that offers "Turn into project" — to content blocks of either lane.
|
||||
instance = createChatInstance({
|
||||
ws, state: { activePage: 'chat', projectChatIds: new Set(), unreadCount: 0 },
|
||||
updateUnreadBadge() {}, stateSnapshots, chatId: 1, idPrefix: 'chat', mountEl: mount,
|
||||
|
|
@ -1107,9 +1113,8 @@ test('a direct turn renders tool work as an activity block and needs host author
|
|||
} });
|
||||
const card = walkCard(messages, 'eph-1');
|
||||
assert.ok(card, 'real tool work reveals the activity block');
|
||||
assert.equal(card.dataset.direct, '1', 'the block wears the direct chrome');
|
||||
assert.equal(card.querySelector('[data-turn-into-project]'), null, 'a direct turn is never offered conversion');
|
||||
assert.equal(card.querySelector('[data-live-title]').textContent, '', 'no Task/Working placeholder title');
|
||||
assert.ok(convertButton(card), 'a working direct turn is offered conversion in Main');
|
||||
assert.equal(card.querySelector('[data-live-title]').textContent, 'Working...', 'the running placeholder title');
|
||||
assert.equal(card.querySelector('[data-cancel-run]'), null, 'no host cancelable marker: no Cancel');
|
||||
handlers.get('chat')({
|
||||
chat_id: 1, role: 'assistant', is_progress: true,
|
||||
|
|
@ -1140,7 +1145,7 @@ test('a direct turn renders tool work as an activity block and needs host author
|
|||
} });
|
||||
assert.equal(card.dataset.finished, '1');
|
||||
assert.match(card.querySelector('[data-live-meta]').innerHTML, /\$2\.70/);
|
||||
assert.equal(card.querySelector('[data-turn-into-project]'), null);
|
||||
assert.ok(convertButton(card), 'conversion stays on the finished card');
|
||||
// A direct turn without tool work or progress stays a plain answer.
|
||||
handlers.get('log')({ chat_id: 1, data: {
|
||||
type: 'task_started', task_id: 'eph-2', ts: '2026-09-05T11:00:00Z',
|
||||
|
|
@ -1204,8 +1209,7 @@ test(`history replay of a direct turn preserves ${execution}`, async () => {
|
|||
assert.equal(card.querySelector('[data-live-phase]').dataset.phase, phase);
|
||||
assert.doesNotMatch(card.querySelector('[data-live-meta]').innerHTML, /\$0(?:\.00)?(?:\s|<|$)/);
|
||||
if (execution === 'ok') assert.match(card.querySelector('[data-live-meta]').innerHTML, /\$0\.75/);
|
||||
assert.equal(card.dataset.direct, '1', 'replay reads the same host fact');
|
||||
assert.equal(card.querySelector('[data-turn-into-project]'), null);
|
||||
assert.ok(convertButton(card), 'replayed content is offered conversion in Main');
|
||||
assert.equal(card.querySelector('[data-cancel-run]'), null);
|
||||
assert.equal(messages.children.filter((n) => /resets on Monday/.test(n.innerHTML)).length, 1);
|
||||
} finally {
|
||||
|
|
@ -1399,3 +1403,54 @@ test('native terminal replay retains the actual narration as title', async () =>
|
|||
assert.equal(card.querySelector('[data-live-title]').textContent, 'still working');
|
||||
} finally { instance?.destroy(); restoreDom(prior); }
|
||||
});
|
||||
|
||||
|
||||
test('a late acceptance settlement row never rewrites the finished card', async () => {
|
||||
// Owner fork 2=A (2026-09-16): a reviewer panel that settles after the task
|
||||
// ended adds ONE System row to the task's room. It carries the task id like
|
||||
// every task-scoped notice, so the regression to pin is the card: its Done
|
||||
// chip, title and counts stay exactly as the terminal row left them.
|
||||
const rows = [
|
||||
{ task_id: 'late-panel', is_progress: true, text: '💬 checking the budget', ts: '2026-09-16T00:00:00Z', task_terminal_status: 'completed' },
|
||||
{ task_id: 'late-panel', role: 'assistant', text: 'The report is ready.', ts: '2026-09-16T00:01:00Z' },
|
||||
{ task_id: 'late-panel', role: 'system', system_type: 'task_summary', text: 'Done.', ts: '2026-09-16T00:02:00Z',
|
||||
tool_calls: 3, rounds: 2, outcome_final: true, outcome_phase: 'done', outcome_axes: { execution: { status: 'ok' } } },
|
||||
];
|
||||
const { prior, mount } = installDom(async (url) => ({ ok: true, json: async () =>
|
||||
String(url).startsWith('/api/chat/history') ? { messages: rows } : { active_direct_turns: [] } }));
|
||||
const handlers = new Map();
|
||||
const ws = { on(type, fn) { handlers.set(type, fn); return () => handlers.delete(type); }, isConnected: () => true, send() {} };
|
||||
let instance;
|
||||
try {
|
||||
instance = createChatInstance({ ws, state: { activePage: 'chat', projectChatIds: new Set(), unreadCount: 0 },
|
||||
updateUnreadBadge() {}, stateSnapshots: { begin: () => ({ generation: 1, requestedAt: Date.now() }),
|
||||
isCurrent: () => true, apply() {} }, chatId: 1, idPrefix: 'chat', mountEl: mount });
|
||||
await instance.refreshHistory({ revision: 1 });
|
||||
const messages = globalThis.document.byId.get('chat-messages');
|
||||
const card = walkCard(messages, 'late-panel');
|
||||
const before = {
|
||||
phase: card.querySelector('[data-live-phase]')?.textContent,
|
||||
hidden: card.querySelector('[data-live-phase]')?.hidden,
|
||||
title: card.querySelector('[data-live-title]')?.textContent,
|
||||
finished: card.dataset.finished,
|
||||
html: card.innerHTML,
|
||||
};
|
||||
assert.equal(before.phase, 'Done');
|
||||
const bubbles = () => messages.children.filter((node) => node.classList.contains('chat-bubble')
|
||||
&& node.classList.contains('system') && !node.classList.contains('typing-bubble'));
|
||||
const systemBefore = bubbles().length;
|
||||
const text = 'Reviewers later passed this answer. They reviewed the earlier version, which was rewritten before delivery.\n- triad_one: PASS — Budget section is complete.';
|
||||
handlers.get('chat')({ chat_id: 1, task_id: 'late-panel', role: 'system', system_type: 'acceptance_late_settlement',
|
||||
content: text, ts: '2026-09-16T00:05:00Z' });
|
||||
const after = walkCard(messages, 'late-panel');
|
||||
assert.equal(after, card, 'the finished card is the same node');
|
||||
assert.equal(after.querySelector('[data-live-phase]')?.textContent, before.phase);
|
||||
assert.equal(after.querySelector('[data-live-phase]')?.hidden, before.hidden);
|
||||
assert.equal(after.querySelector('[data-live-title]')?.textContent, before.title);
|
||||
assert.equal(after.dataset.finished, before.finished);
|
||||
assert.equal(after.innerHTML, before.html, 'the late row is not folded into the card');
|
||||
const added = bubbles();
|
||||
assert.equal(added.length, systemBefore + 1, 'exactly one new System row');
|
||||
assert.match(added[added.length - 1].innerHTML, /Reviewers later passed this answer/);
|
||||
} finally { instance?.destroy(); restoreDom(prior); }
|
||||
});
|
||||
|
|
|
|||
|
|
@ -249,7 +249,7 @@ const PLAIN_ROW = {
|
|||
role: 'system',
|
||||
system_type: 'project_completion_summary',
|
||||
markdown: false,
|
||||
content: 'Launch › Ship · Completed\nPlain excerpt line.',
|
||||
content: 'Launch › Ship · Completed\nOpen the Project for details.',
|
||||
project_id: 'launch',
|
||||
project_name: 'Launch',
|
||||
ts: '2026-08-31T00:00:00Z',
|
||||
|
|
@ -267,7 +267,7 @@ test('plain project row renders escaped text with Open Project and no markdown m
|
|||
// Escaped plain text: the raw newline survives (pre-wrap), so the row
|
||||
// did NOT pass through the markdown renderer (whose no-parser fallback
|
||||
// rewrites \n to <br>) and produced no heading elements.
|
||||
assert.match(bubble.innerHTML, /Launch › Ship · Completed\nPlain excerpt line\./);
|
||||
assert.match(bubble.innerHTML, /Launch › Ship · Completed\nOpen the Project for details\./);
|
||||
assert.doesNotMatch(bubble.innerHTML, /<br>|<h1|<h2|md-h1|md-h2/);
|
||||
// Bug report #9: no enhancement pass — Mermaid/Chart/KaTeX/code-copy
|
||||
// only ever activate behind enhanceChatMarkdown's enhanced stamp.
|
||||
|
|
|
|||
264
web/tests/fixtures/outcome_phase_parity.json
vendored
264
web/tests/fixtures/outcome_phase_parity.json
vendored
|
|
@ -1,5 +1,5 @@
|
|||
{
|
||||
"purpose": "One status-word family across the browser and the host. Every case is a terminal-or-not task record read by log_events.js (taskDoneIsTerminal + taskTerminalPhase + taskPresentation) and by ouroboros.project_dialogue (outcome_phase + OUTCOME_PHASE_HEADLINE + _completion_verdict). A new axis, reason or acceptance status is added to both sides in the same commit, with a row here.",
|
||||
"purpose": "One status-word family across the browser and the host. Every case is a terminal-or-not task record read by log_events.js (taskDoneIsTerminal + taskTerminalPhase + taskPresentation) and by ouroboros.project_dialogue (outcome_phase + OUTCOME_PHASE_HEADLINE + _completion_verdict). Every typed cause with an owner sentence carries a case here, so the two TASK_CAUSE_PHRASES twins cannot drift apart unnoticed. A new axis, reason or acceptance status is added to both sides in the same commit, with a row here.",
|
||||
"cases": [
|
||||
{
|
||||
"name": "definite pre-branch publication failure stays Failed beside degraded review",
|
||||
|
|
@ -20,7 +20,7 @@
|
|||
},
|
||||
"phase": "error",
|
||||
"headline": "Failed",
|
||||
"acceptance_clause": "Reason: PR not created; publication stopped before creating a submission branch."
|
||||
"acceptance_clause": "PR not created; publication stopped before creating a submission branch."
|
||||
},
|
||||
{
|
||||
"name": "clean completed root",
|
||||
|
|
@ -93,7 +93,7 @@
|
|||
"acceptance_clause": ""
|
||||
},
|
||||
{
|
||||
"name": "degraded review with an unaccepted acceptance decision",
|
||||
"name": "an unaccepted decision speaks through its typed reason, never through the stored rationale",
|
||||
"record": {
|
||||
"status": "completed",
|
||||
"reason_code": "final_message",
|
||||
|
|
@ -103,6 +103,7 @@
|
|||
"status": "degraded",
|
||||
"acceptance_decision": {
|
||||
"status": "finalized_unaccepted",
|
||||
"reason": "review_degraded",
|
||||
"rationale": "Acceptance reviewers did not reach a valid quorum."
|
||||
}
|
||||
}
|
||||
|
|
@ -110,7 +111,7 @@
|
|||
},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "Acceptance: finalized_unaccepted — Acceptance reviewers did not reach a valid quorum."
|
||||
"acceptance_clause": "No reviewer verdict was established for this answer."
|
||||
},
|
||||
{
|
||||
"name": "an accepted decision leaves the execution reason standing",
|
||||
|
|
@ -127,7 +128,7 @@
|
|||
"acceptance_clause": ""
|
||||
},
|
||||
{
|
||||
"name": "a revision request without a rationale states the status alone",
|
||||
"name": "a decision carrying no typed reason states no cause at all",
|
||||
"record": {
|
||||
"status": "completed",
|
||||
"outcome_axes": {
|
||||
|
|
@ -137,7 +138,7 @@
|
|||
},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "Acceptance: revision_requested."
|
||||
"acceptance_clause": ""
|
||||
},
|
||||
{
|
||||
"name": "a hard failure explains itself by its execution reason, not by the decision",
|
||||
|
|
@ -173,7 +174,7 @@
|
|||
},
|
||||
"phase": "error",
|
||||
"headline": "Failed",
|
||||
"acceptance_clause": "Reason: provider_unavailable (delegated_custody_unreconciled)."
|
||||
"acceptance_clause": "provider_unavailable (Some delegated work was never reconciled.)"
|
||||
},
|
||||
{
|
||||
"name": "a healed custody debt names the current execution cause alone",
|
||||
|
|
@ -185,7 +186,7 @@
|
|||
},
|
||||
"phase": "error",
|
||||
"headline": "Failed",
|
||||
"acceptance_clause": "Reason: provider_unavailable."
|
||||
"acceptance_clause": "provider_unavailable."
|
||||
},
|
||||
{
|
||||
"name": "a record carrying no debt list states nothing about the debt",
|
||||
|
|
@ -196,7 +197,252 @@
|
|||
},
|
||||
"phase": "error",
|
||||
"headline": "Failed",
|
||||
"acceptance_clause": "Reason: provider_unavailable."
|
||||
"acceptance_clause": "provider_unavailable."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason author_finish says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "author_finish"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The answer was delivered on Main's own judgement; the reviewers had not signed it off."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason review_degraded says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "review_degraded"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "No reviewer verdict was established for this answer."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason infra_failure says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "infra_failure"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "A review infrastructure failure prevented a settled verdict."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason dialogue_terminal says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "dialogue_terminal"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The reviewers and Main could not agree, and both positions were kept."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason improvement_capsule says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "improvement_capsule"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The reviewers asked for one more pass and Main was given their notes."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason fence_reopen_failed says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "fence_reopen_failed"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The requested extra pass could not be started, so the answer stands as it was."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason review_cycles_exhausted says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "review_cycles_exhausted"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task used up its review rounds before the answer was signed off."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason open_obligations says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "open_obligations"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The answer was delivered with reviewer requests still open."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason improvement_window_closed says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "improvement_window_closed"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "There was no room left for another pass, so the answer stands as it was."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason capsule_spent says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "capsule_spent"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The one allowed improvement pass was already used."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason reviewer_fail_no_capsule says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "reviewer_fail_no_capsule"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "A reviewer rejected the answer and suggested nothing to change."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason no_actionable_changes says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "no_actionable_changes"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The re-review was not clean and suggested nothing to change."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason identical_acceptance_refused says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "identical_acceptance_refused"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "Nothing had changed since the last review, so the recorded verdict stands."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason review_skipped_deadline_reserve says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "review_skipped_deadline_reserve"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "There was not enough time left to review the answer."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason delivery_binding_superseded says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "delivery_binding_superseded"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The answer or its evidence changed, so the earlier review no longer covered it."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason owner_followup says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "owner_followup"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "A new message from you arrived, so the review was set aside for it."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason evidence_refresh says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "evidence_refresh"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The work changed after the review was frozen, so it no longer covered the answer."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason revision_unavailable_on_forced_rail says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "revision_unavailable_on_forced_rail"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task had to stop, so the requested rework never happened."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason owner_hurry says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "owner_hurry"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "You asked me to hurry, so no further review was started."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason unspecified says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "unspecified"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The answer was not signed off, and no cause was recorded."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason acceptance_bypassed_budget_exhausted says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "acceptance_bypassed_budget_exhausted"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task ran out of budget before the answer could be reviewed."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason acceptance_bypassed_round_limit says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "acceptance_bypassed_round_limit"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task hit its round limit before the answer could be reviewed."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason acceptance_bypassed_deadline says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "acceptance_bypassed_deadline"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task ran out of time before the answer could be reviewed."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason acceptance_bypassed_provider_unavailable says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "acceptance_bypassed_provider_unavailable"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The model provider was unavailable, so the answer was never reviewed."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason acceptance_bypassed_context_overflow says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "acceptance_bypassed_context_overflow"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task outgrew its context before the answer could be reviewed."
|
||||
},
|
||||
{
|
||||
"name": "acceptance reason acceptance_bypassed_children_unabsorbed says its own sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "acceptance_bypassed_children_unabsorbed"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "Some sub-tasks had not been folded in, so the answer was never reviewed."
|
||||
},
|
||||
{
|
||||
"name": "execution reason budget_exhausted says its own sentence",
|
||||
"record": {"status": "completed", "reason_code": "budget_exhausted", "outcome_axes": {"execution": {"status": "degraded"}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task ran out of budget before it could finish cleanly."
|
||||
},
|
||||
{
|
||||
"name": "execution reason delivery_control_degraded says its own sentence",
|
||||
"record": {"status": "completed", "reason_code": "delivery_control_degraded", "outcome_axes": {"execution": {"status": "degraded"}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "Delivery finished in a degraded control state."
|
||||
},
|
||||
{
|
||||
"name": "execution reason host_child_status_suffix says its own sentence",
|
||||
"record": {"status": "completed", "reason_code": "host_child_status_suffix", "outcome_axes": {"execution": {"status": "degraded"}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "A child task had not settled when the answer was delivered."
|
||||
},
|
||||
{
|
||||
"name": "execution reason invalid_delivery_control_after_repair says its own sentence",
|
||||
"record": {"status": "completed", "reason_code": "invalid_delivery_control_after_repair", "outcome_axes": {"execution": {"status": "degraded"}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The delivery control object was still malformed after repair."
|
||||
},
|
||||
{
|
||||
"name": "execution reason plan_review_advisory says its own sentence",
|
||||
"record": {"status": "completed", "reason_code": "plan_review_advisory", "outcome_axes": {"execution": {"status": "degraded"}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "Plan review never closed; the work continued under advisory enforcement."
|
||||
},
|
||||
{
|
||||
"name": "an open delegated-custody debt says its own sentence",
|
||||
"record": {"status": "completed", "reason_code": "delegated_custody_unreconciled", "delegated_runs_unreconciled": ["run-a1"], "outcome_axes": {"execution": {"status": "degraded"}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "Some delegated work was never reconciled."
|
||||
},
|
||||
{
|
||||
"name": "spent review cycles stay Done with warnings, never Failed",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "review_cycles_exhausted"}}}, "reason_code": "final_message"},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "The task used up its review rounds before the answer was signed off."
|
||||
},
|
||||
{
|
||||
"name": "an unknown acceptance reason stays raw rather than borrowing a wrong sentence",
|
||||
"record": {"status": "completed", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "degraded", "acceptance_decision": {"status": "finalized_unaccepted", "reason": "some_future_reason"}}}},
|
||||
"phase": "warn",
|
||||
"headline": "Done with warnings",
|
||||
"acceptance_clause": "some_future_reason."
|
||||
},
|
||||
{
|
||||
"name": "reviewers approved the earlier revision: Done, and the row says which revision",
|
||||
"record": {"status": "completed", "reason_code": "final_message", "outcome_axes": {"execution": {"status": "ok"}, "review": {"status": "pass", "acceptance_decision": {"status": "accepted", "reason": "previous_revision_accepted"}}, "objective": {"status": "pass", "source": "task_acceptance_review", "review_status": "pass"}}},
|
||||
"phase": "done",
|
||||
"headline": "Done",
|
||||
"acceptance_clause": "The reviewers approved the earlier version of this answer; it changed before they finished."
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -10,18 +10,20 @@ import {
|
|||
// keeps the machine code; the card says what actually happened.
|
||||
|
||||
test('a typed cause is stated in the owner\'s words', () => {
|
||||
// No 'Reason:' / 'Acceptance:' prefix: a label naming an internal machine
|
||||
// concept in front of an owner sentence is the same leak in a politer font.
|
||||
assert.equal(
|
||||
taskReasonDetail({ reason_code: 'plan_review_advisory' }),
|
||||
'Reason: plan review never closed; the work continued under advisory enforcement',
|
||||
'Plan review never closed; the work continued under advisory enforcement',
|
||||
);
|
||||
assert.equal(
|
||||
taskReasonDetail({ reason_code: 'delivery_control_degraded' }),
|
||||
'Reason: delivery finished in a degraded control state',
|
||||
'Delivery finished in a degraded control state',
|
||||
);
|
||||
});
|
||||
|
||||
test('an unknown cause stays raw rather than becoming a wrong sentence', () => {
|
||||
assert.equal(taskReasonDetail({ reason_code: 'some_future_code' }), 'Reason: some_future_code');
|
||||
assert.equal(taskReasonDetail({ reason_code: 'some_future_code' }), 'some_future_code');
|
||||
assert.equal(taskReasonPhrase('some_future_code'), 'some_future_code');
|
||||
});
|
||||
|
||||
|
|
@ -52,7 +54,60 @@ test('every typed cause the loop can record has a sentence', () => {
|
|||
for (const code of literals) {
|
||||
assert.notEqual(
|
||||
taskReasonPhrase(code), code,
|
||||
`no owner-facing sentence for degraded_reason "${code}" — add one to TASK_REASON_PHRASES`,
|
||||
`no owner-facing sentence for degraded_reason "${code}" — add one to TASK_CAUSE_PHRASES`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('an accepted decision with a sentence still states its cause', () => {
|
||||
// Owner fork 1=B (2026-09-16): reviewers who approved the earlier revision
|
||||
// accept the task, and the row says which revision they approved. A clean
|
||||
// accepted decision keeps rendering nothing.
|
||||
const accepted = (reason) => taskReasonDetail({
|
||||
status: 'completed', reason_code: 'final_message',
|
||||
outcome_axes: { execution: { status: 'ok' },
|
||||
review: { status: 'pass', acceptance_decision: { status: 'accepted', reason } } },
|
||||
});
|
||||
assert.equal(accepted('previous_revision_accepted'),
|
||||
'The reviewers approved the earlier version of this answer; it changed before they finished.');
|
||||
assert.equal(accepted('clean_pass'), '');
|
||||
assert.equal(accepted(''), '');
|
||||
});
|
||||
|
||||
test('every acceptance reason the host can record has a sentence', () => {
|
||||
// The second half of the same gate. Acceptance reasons are written as
|
||||
// `"reason": "<code>"` inside the four acceptance/finalization/settlement
|
||||
// leaves; the bypass family is a dict of literals in outcomes.py, four more
|
||||
// arrive through named constants there and the settlement leaf names its
|
||||
// own reason as a constant, so all three shapes are read explicitly.
|
||||
const pkg = new URL('../../ouroboros/', import.meta.url);
|
||||
const read = (name) => readFileSync(new URL(name, pkg), 'utf8');
|
||||
const decisions = [
|
||||
'loop_acceptance_review.py', 'loop_acceptance.py', 'loop_forced_finalization.py', 'acceptance_settlement.py',
|
||||
].map(read).join('\n');
|
||||
const outcomes = read('outcomes.py');
|
||||
const acceptance = new Set([
|
||||
...[...decisions.matchAll(/"reason":\s*(?:\n\s*)?"([a-z_]+)"/g)].map((m) => m[1]),
|
||||
...[...outcomes.matchAll(/"(acceptance_bypassed_[a-z_]+)"/g)].map((m) => m[1]),
|
||||
...[...outcomes.matchAll(
|
||||
/^REASON_(?:REVIEW_CYCLES_EXHAUSTED|IDENTICAL_ACCEPTANCE_REFUSED|ACCEPTANCE_REVIEW_SKIPPED_DEADLINE_RESERVE|ACCEPTANCE_SKIPPED_OWNER_HURRY) = "([a-z_]+)"$/gm,
|
||||
)].map((m) => m[1]),
|
||||
...[...read('acceptance_settlement.py').matchAll(/^REASON_[A-Z_]+ = "([a-z_]+)"$/gm)].map((m) => m[1]),
|
||||
]);
|
||||
assert.ok(acceptance.has('previous_revision_accepted'), 'the settlement leaf is scanned');
|
||||
// A CLEAN accepted decision renders no clause, the owner stop carries its
|
||||
// own marker instead, and queue_inspection_failed is a `{status, reason}`
|
||||
// probe shape rather than an acceptance reason.
|
||||
const exempt = new Set([
|
||||
'clean_pass', 'clean_pass_obligations_closed', 'queue_inspection_failed',
|
||||
'acceptance_bypassed_owner_requested_finalization',
|
||||
]);
|
||||
assert.ok(acceptance.size >= 20, `expected the acceptance vocabulary, saw ${acceptance.size}`);
|
||||
for (const code of acceptance) {
|
||||
if (exempt.has(code)) continue;
|
||||
assert.notEqual(
|
||||
taskReasonPhrase(code), code,
|
||||
`no owner-facing sentence for acceptance reason "${code}" — add one to TASK_CAUSE_PHRASES`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
|
@ -90,6 +145,7 @@ const A4 = {
|
|||
status: 'degraded',
|
||||
acceptance_decision: {
|
||||
status: 'finalized_unaccepted',
|
||||
reason: 'review_degraded',
|
||||
rationale: 'Acceptance reviewers did not reach a valid quorum.',
|
||||
},
|
||||
},
|
||||
|
|
@ -99,9 +155,12 @@ const A4 = {
|
|||
test('an unaccepted decision explains the warning in its own words', () => {
|
||||
assert.equal(
|
||||
taskReasonDetail(A4),
|
||||
'Acceptance: finalized_unaccepted — Acceptance reviewers did not reach a valid quorum.',
|
||||
'No reviewer verdict was established for this answer.',
|
||||
);
|
||||
assert.doesNotMatch(taskReasonDetail(A4), /final_message/);
|
||||
// The stored reviewer rationale belongs to the card body, the task result
|
||||
// and Logs; the row states one sentence and never the machine words.
|
||||
assert.doesNotMatch(taskReasonDetail(A4), /quorum|finalized_unaccepted/);
|
||||
});
|
||||
|
||||
test('an accepted decision omits neutral final_message and preserves substantive reasons', () => {
|
||||
|
|
@ -113,21 +172,26 @@ test('an accepted decision omits neutral final_message and preserves substantive
|
|||
},
|
||||
};
|
||||
assert.equal(taskReasonDetail(accepted), '');
|
||||
assert.equal(taskReasonDetail({ ...accepted, reason_code: 'custom_reason' }), 'Reason: custom_reason');
|
||||
assert.equal(taskReasonDetail({ ...accepted, reason_code: 'custom_reason' }), 'custom_reason');
|
||||
});
|
||||
|
||||
test('a decision without a rationale states its status alone', () => {
|
||||
test('a decision carrying no typed reason states no cause at all', () => {
|
||||
// A status word is not a cause, and the collapsed status is already the
|
||||
// headline; a historical record without a reason therefore says nothing.
|
||||
const record = {
|
||||
outcome_axes: { review: { acceptance_decision: { status: 'revision_requested' } } },
|
||||
status: 'completed',
|
||||
};
|
||||
assert.equal(taskReasonDetail(record), 'Acceptance: revision_requested');
|
||||
assert.equal(taskReasonDetail(record), '');
|
||||
});
|
||||
|
||||
test('a decision with no reason code still reaches the acceptance branch', () => {
|
||||
// The old single early return swallowed this frame before the branch.
|
||||
const record = { status: 'completed', review_status: { acceptance_decision: { status: 'revision_requested' } } };
|
||||
assert.equal(taskReasonDetail(record), 'Acceptance: revision_requested');
|
||||
const record = {
|
||||
status: 'completed',
|
||||
review_status: { acceptance_decision: { status: 'revision_requested', reason: 'owner_followup' } },
|
||||
};
|
||||
assert.equal(taskReasonDetail(record), 'A new message from you arrived, so the review was set aside for it.');
|
||||
});
|
||||
|
||||
test('a hard failure keeps explaining itself by its execution reason', () => {
|
||||
|
|
@ -140,17 +204,28 @@ test('a hard failure keeps explaining itself by its execution reason', () => {
|
|||
...A4, status: 'failed', reason_code: 'delegated_custody_unreconciled',
|
||||
delegated_runs_unreconciled: ['run-a1'],
|
||||
};
|
||||
assert.equal(taskReasonDetail(failed), 'Reason: delegated_custody_unreconciled');
|
||||
assert.equal(taskReasonDetail(failed), 'Some delegated work was never reconciled.');
|
||||
});
|
||||
|
||||
test('a multi-line rationale is flattened into one sentence', () => {
|
||||
test('a stored rationale never reaches the row, however it is written', () => {
|
||||
// The rationale used to be flattened into the line; it is free reviewer
|
||||
// text up to 500 characters and belongs where the full copy lives.
|
||||
const noisy = {
|
||||
status: 'completed',
|
||||
outcome_axes: {
|
||||
review: { acceptance_decision: { status: 'revision_requested', rationale: 'Two\n\nlines here.' } },
|
||||
review: {
|
||||
acceptance_decision: {
|
||||
status: 'revision_requested', reason: 'evidence_refresh',
|
||||
rationale: 'Two\n\nlines here.',
|
||||
},
|
||||
},
|
||||
},
|
||||
};
|
||||
assert.equal(taskReasonDetail(noisy), 'Acceptance: revision_requested — Two lines here.');
|
||||
assert.equal(
|
||||
taskReasonDetail(noisy),
|
||||
'The work changed after the review was frozen, so it no longer covered the answer.',
|
||||
);
|
||||
assert.doesNotMatch(taskReasonDetail(noisy), /lines/);
|
||||
});
|
||||
|
||||
// The custody overlay stamps `delegated_custody_unreconciled` as the row's
|
||||
|
|
@ -167,7 +242,7 @@ test('a healed custody debt yields the current execution reason on the card', ()
|
|||
reason_code: 'delegated_custody_unreconciled',
|
||||
delegated_runs_unreconciled: [],
|
||||
outcome_axes: { execution: { status: 'degraded', reason_code: 'tool_failure' } },
|
||||
}), 'Reason: tool_failure');
|
||||
}), 'tool_failure');
|
||||
});
|
||||
|
||||
test('an open custody debt is still named on the card', () => {
|
||||
|
|
@ -176,7 +251,7 @@ test('an open custody debt is still named on the card', () => {
|
|||
reason_code: 'delegated_custody_unreconciled',
|
||||
delegated_runs_unreconciled: ['run-a1'],
|
||||
outcome_axes: { execution: { status: 'ok' } },
|
||||
}), 'Reason: delegated_custody_unreconciled');
|
||||
}), 'Some delegated work was never reconciled.');
|
||||
});
|
||||
|
||||
test('a real debt beside a real execution cause is one card line', () => {
|
||||
|
|
@ -185,7 +260,7 @@ test('a real debt beside a real execution cause is one card line', () => {
|
|||
reason_code: 'delegated_custody_unreconciled',
|
||||
delegated_runs_unreconciled: ['run-a1', 'run-b2'],
|
||||
outcome_axes: { execution: { status: 'failed', reason_code: 'provider_unavailable' } },
|
||||
}), 'Reason: provider_unavailable (delegated_custody_unreconciled)');
|
||||
}), 'provider_unavailable (Some delegated work was never reconciled.)');
|
||||
});
|
||||
|
||||
// A LIVE task_done event carries the row's own debt list too
|
||||
|
|
@ -200,7 +275,7 @@ test('a live event carrying an open debt list names the debt through the list',
|
|||
reason_code: 'delegated_custody_unreconciled',
|
||||
delegated_runs_unreconciled: ['run-a1'],
|
||||
outcome_axes: { execution: { status: 'ok', reason_code: 'tool_failure' } },
|
||||
}), 'Reason: tool_failure (delegated_custody_unreconciled)');
|
||||
}), 'tool_failure (Some delegated work was never reconciled.)');
|
||||
});
|
||||
|
||||
test('a record carrying no debt list states nothing about the debt', () => {
|
||||
|
|
@ -208,7 +283,7 @@ test('a record carrying no debt list states nothing about the debt', () => {
|
|||
status: 'failed',
|
||||
reason_code: 'delegated_custody_unreconciled',
|
||||
outcome_axes: { execution: { status: 'failed', reason_code: 'provider_unavailable' } },
|
||||
}), 'Reason: provider_unavailable');
|
||||
}), 'provider_unavailable');
|
||||
assert.equal(taskReasonDetail({
|
||||
status: 'completed',
|
||||
reason_code: 'delegated_custody_unreconciled',
|
||||
|
|
|
|||
|
|
@ -64,7 +64,10 @@ test('live task_done and replay/log task truth have phase and headline parity',
|
|||
assert.deepEqual({ phase: replay.phase, headline: replay.headline }, expected, `${name}/replay`);
|
||||
assert.doesNotMatch(`${live.headline} ${replay.headline}`, /Issue|Notice|delegated_custody_unreconciled/);
|
||||
if (payload.reason_code) {
|
||||
assert.match(live.body, /Reason: delegated_custody_unreconciled/);
|
||||
// The card body says the cause in words; the raw code stays in the
|
||||
// record half (Logs meta), which is where a machine code belongs.
|
||||
assert.match(live.body, /Some delegated work was never reconciled\./);
|
||||
assert.doesNotMatch(live.body, /delegated_custody_unreconciled/);
|
||||
assert.ok(replay.meta.includes('delegated_custody_unreconciled'));
|
||||
}
|
||||
}
|
||||
|
|
@ -156,7 +159,8 @@ test('failed child remains a compact local fact without owner-alarm semantics',
|
|||
// Identity only; the chip carries `Failed` (DESIGN.md §4), the headline never does.
|
||||
assert.equal(child.headline, 'researcher');
|
||||
assert.doesNotMatch(child.body, /delegated_custody_unreconciled/);
|
||||
assert.match(child.fullBody, /Reason: delegated_custody_unreconciled/);
|
||||
assert.match(child.fullBody, /Some delegated work was never reconciled\./);
|
||||
assert.doesNotMatch(child.fullBody, /Reason: /, 'no machine label in front of the owner sentence');
|
||||
assert.equal('ownerAlarm' in child, false);
|
||||
assert.equal('notification' in child, false);
|
||||
const adapter = chatSource.slice(
|
||||
|
|
@ -406,6 +410,7 @@ test('a review-caused warning names the acceptance decision on the card and in L
|
|||
status: 'degraded',
|
||||
acceptance_decision: {
|
||||
status: 'finalized_unaccepted',
|
||||
reason: 'review_degraded',
|
||||
rationale: 'Acceptance reviewers did not reach a valid quorum.',
|
||||
},
|
||||
},
|
||||
|
|
@ -414,8 +419,10 @@ test('a review-caused warning names the acceptance decision on the card and in L
|
|||
const live = summarizeChatLiveEvent(evt);
|
||||
const replay = summarizeLogEvent(evt);
|
||||
assert.deepEqual({ phase: live.phase, headline: live.headline }, { phase: 'warn', headline: 'Done with warnings' });
|
||||
assert.match(live.body, /Acceptance: finalized_unaccepted — Acceptance reviewers did not reach a valid quorum\./);
|
||||
assert.match(live.body, /No reviewer verdict was established for this answer\./);
|
||||
assert.doesNotMatch(live.body, /final_message/);
|
||||
// The raw code lives on in the record half, never in the card body.
|
||||
assert.doesNotMatch(live.body, /finalized_unaccepted/);
|
||||
assert.ok(replay.meta.includes('review degraded'));
|
||||
assert.ok(replay.meta.includes('acceptance finalized_unaccepted'));
|
||||
});
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue