diff --git a/README.md b/README.md index 02a0a9300..5b1fa6761 100644 --- a/README.md +++ b/README.md @@ -12,7 +12,7 @@ [![Linux](https://img.shields.io/badge/Linux-x86__64-orange.svg)](https://ouroboros-agent.ai/install/#linux) [![Windows](https://img.shields.io/badge/Windows-x64-blue.svg)][download-windows-x64] [![OuroborosHub](https://img.shields.io/badge/OuroborosHub-skills%20marketplace-8A2BE2.svg)](https://github.com/razzant/OuroborosHub) -[![Version 7.0.0](https://img.shields.io/badge/version-7.0.0-green.svg)](VERSION) +[![Version 7.1.0](https://img.shields.io/badge/version-7.1.0-green.svg)](VERSION) Ouroboros is an open-source, general-purpose AI agent whose identity, durable memory, and history continue across tasks and restarts. It works on external projects, coordinates a live swarm of specialist agents, and can rewrite the implementation it runs on, including its code, architecture, prompts, tools, and dependencies. Reflection can also change how it understands itself without severing that continuity. @@ -64,13 +64,13 @@ The desktop packages already contain an optional CLI installer. On macOS, after -[download-macos-arm64]: https://github.com/razzant/ouroboros/releases/download/v7.0.0/Ouroboros-7.0.0.dmg -[download-windows-x64]: https://github.com/razzant/ouroboros/releases/download/v7.0.0/Ouroboros-7.0.0-windows-x64.zip -[download-linux-deb-amd64]: https://github.com/razzant/ouroboros/releases/download/v7.0.0/ouroboros_7.0.0_amd64.deb -[download-linux-rpm-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.0.0/ouroboros-7.0.0-1.x86_64.rpm -[download-linux-rpm-red80-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.0.0/ouroboros-7.0.0-1.red80.x86_64.rpm -[download-linux-appimage-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.0.0/Ouroboros-7.0.0-linux-x86_64.AppImage -[download-linux-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.0.0/Ouroboros-7.0.0-linux-x86_64.tar.gz +[download-macos-arm64]: https://github.com/razzant/ouroboros/releases/download/v7.1.0/Ouroboros-7.1.0.dmg +[download-windows-x64]: https://github.com/razzant/ouroboros/releases/download/v7.1.0/Ouroboros-7.1.0-windows-x64.zip +[download-linux-deb-amd64]: https://github.com/razzant/ouroboros/releases/download/v7.1.0/ouroboros_7.1.0_amd64.deb +[download-linux-rpm-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.1.0/ouroboros-7.1.0-1.x86_64.rpm +[download-linux-rpm-red80-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.1.0/ouroboros-7.1.0-1.red80.x86_64.rpm +[download-linux-appimage-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.1.0/Ouroboros-7.1.0-linux-x86_64.AppImage +[download-linux-x86_64]: https://github.com/razzant/ouroboros/releases/download/v7.1.0/Ouroboros-7.1.0-linux-x86_64.tar.gz Ouroboros bundles [Claudexor](https://github.com/razzant/claudexor) as its local execution layer for delegated coding and hosted-agent review. Ouroboros owns the task, memory, review, and final integration, while Claudexor runs the selected connected coding harness and returns durable execution evidence. [Explore Claudexor](https://claudexor.ai/). @@ -449,6 +449,7 @@ and the reason. | Version | Date | Description | |---------|------|-------------| +| 7.1.0 | 2026-09-17 | **feat: memory that understands people, Background Consciousness as an ordinary Main turn, and a chat that leads with narration.** Memory keeps an understanding of people in front of the mind through resident summaries, one memory shared by every room and honest consolidation (#910). Background Consciousness becomes an ordinary Main turn on an alarm clock (#988, closing the class behind #654). Chat blocks lead with narration while tool calls fold into one evidence row (#1007); task cards follow their work and treat acceptance review as advice to Ouroboros (#970); question pointers say whether an answer is wanted and show the recorded answer (#1025); the unified UI control system, source picker and model chooser name models and accounts truthfully (#768, #862, #890). Planning is asynchronous and Swarm tasks are admitted directly (#851); delegated work survives durable cancellation (#850); a task's own reviewer is never swept as its delegation (#1008); semantic duplicate vetoes leave subagent admission (#887); Presence profiles work in an owner-selected folder (#941). Prompt caches extend through an append-only acceptance observation and a shared Codex cache key per install and model (#929, #1015). Update letters include merged branch changes (#1003), the owned Claudexor daemon latches failed starts with a typed diagnosis, and the managed runtime advances to Claudexor 3.12.1 (#891, #962), alongside Windows, platform-installation and UI-smoke repairs. | | 7.0.0 | 2026-09-08 | **v7: a modular runtime with full ordinary-conversation tools, complete delegated inputs and visible owner dialogue.** Main and Project conversations retain their chosen tools and working folders; required owner waits preserve the live browser without occupying pooled execution capacity. Publish admission is visible in its intended chat, and GitHub/image/plan failures retain their actual causes. Planning advice stays optional; source-backed evidence, separate host notices and exact-workload test reuse preserve honest review without new approval machinery. Integrates the subscription model and account-role controls, owned startup/restart custody and cross-platform repairs, with Claudexor pinned to 3.10.1. | | 6.114.0 | 2026-09-01 | **feat: capability preservation as a first-class invariant, shape-first OpenRouter failover, and a green 3-OS matrix.** Capability preservation joins the immune system as a first-class invariant (#447, stages 1-3: golden capability suite, seam pins that actually differentiate, executable item-21 triggers, binaries judged by magic bytes). Same-model OpenRouter provider failover is now derived from the shape of the replayed reasoning artifact instead of a hard-coded model-family roster (#468): readable text/summary artifacts stay failover-eligible for every family, a bare `response_id` no longer pins, sealed artifacts (encrypted/signed/redacted, not vouched by the anthropic/gemini roster — `openai/` excluded on 2026-07 field evidence) keep the continuity pin on both the dispatch and reroute paths, and the pin reason plus the refusing upstream provider reach durable telemetry. Delegation and chat truth converge: delegated-run questions ride the escalation hierarchy (#204), the routing picker dispatches real routes (#198), typed owner quiz cards and the escalation channel (Q-2a/Q-2b), live-first owner deliveries, honest chat-activity conclusions with budget-pause resume, review findings in the task-card Reviews checkpoint, persistent stable-target registrations, a hold on the live delegated leaf when a nanny round dies unknown, and healed cross-generation custody disclosures. Rich chat stack (markdown, media, galleries, links — Andrei Kaznacheev), Chat viewport stability, Dashboard/Updates one-verdict action surface, Settings effect honesty, skill Repair affordance after preflight failure, node/npm launches through an execution-probed runtime ladder, the `ultra` reasoning-effort tier (@ndrew1337), and an O(delta) in-lock ledger read with a guarded torn-tail append boundary (@deebosh). Claudexor runtime pinned to 3.9.5 with delayed-startup reconciliation, a foreground quota-refresh bridge, and preserved delegated text fragments; Updates restart state survives same-SHA reconnects. CI: the serial service-log suites no longer race the child, web source pins are CRLF-safe on Windows, and the size-ratchet manifest is exact. This tag also carries the untagged 6.113.5 (generation-safe browser cleanup after timed-out tools, #440 audit fix-forward). | | 6.113.5 | 2026-08-31 | **fix: browser cleanup after a timed-out tool is generation-safe (integrates community PR #429 by @mikemikimike, closes #409).** A stateful-tool timeout now retires the whole browser generation: the shared state slot is replaced with a fresh object, the abandoned worker keeps writing only into its retired one, and the close is queued on the retiring executor so it always runs on the owning worker thread — including the already-settled race — with the cognitive lease closing after that cleanup. A late infrastructure-error retry that observes a replaced generation closes only its own retired session, so it can no longer cross-thread-kill the next command's browser. Closes the follow-up findings of the #440 post-merge audit; the never-settling-worker session leak stays a disclosed residual of in-process Playwright. | @@ -459,8 +460,7 @@ and the reason. | 6.113.0 | 2026-08-29 | **feat: delegation by construction — the nanny charter, typed $0 terminals, truthful executor cards, and honest route health (Claudexor runtime 3.9.0).** An `agent_session` child now IS work on its harness: the host pre-starts the physical leaf through the same configured `delegate_start` wrapper BEFORE the nanny's first LLM round and never waits — her first round arrives with a live `configured_session_started` receipt, waiting is her own `delegate_wait` decision, and children may run beside the leaf. A definite refusal to start (typed pre-POST, dispatch-blocked, engine-rejected — with a custody-handle guard that always prefers a model episode over a false terminal) ends the task typed and unrun at $0; ambiguity always wakes the model, and durable zero-run/unknown-evidence fences outrank blocked terminals. Zero-run receipts narrow to incomplete\|unknown, actor cleanliness requires a SUCCEEDED delegated run (children are evidence, never a completion path), unreadable custody projects typed unknown all the way into the finalization nudge, and only acts of delegation reset the economics baseline (the reminder-storm class is dead). route_health stops refusing on aggregate doctor status — admission belongs to the engine; the owner's enabled toggle stays a typed `route_disabled`. Acceptance sees substrate facts as visibility with zero gates. The executor chip tells the run truth for the whole lifecycle (dispatched → counted `N ok, M failed` → evidence honesty), all-failed can never render clean, actual_substrate reaches the wire, and the terminal evidence frame survives chat-0/A2A routing. The pinned Claudexor runtime moves to 3.9.0: per-vendor quota pacing with typed Retry-After floors (a poll 429 is never a quota fact), honest foreground cooldowns, first-429 short-circuit, cached accounts default, and the cursor delegation belt (live-E2E proven). This tag also carries the untagged 6.111.0 and 6.112.0 (P13 Emergence) below, and heals the branch's latent size-ratchet debt root-cause (settings_integrity extraction, cybergym module splits, regenerated manifest). | | 6.112.0 | 2026-08-28 | **feat: Principle 13 (Emergence) — designs must get better as intelligence grows.** BIBLE.md gains a new constitutional principle: code hardcodes the floor — truth, custody, budgets, authority, acceptance — never the ceiling; strategy belongs to the mind, patterns that worked are examples to record rather than laws to enforce, and every design faces the stronger-mind test: when the model gets smarter, does this get better on its own, or does it have to be torn out first? Pointed clarifications close the readings that used to license freezing today's shape: P2 defines the class by the invariant rather than the incident, P5 names how work is shaped (decomposition, roles, ordering, delegation) as behavior belonging to the LLM, and P7 distinguishes unused machinery (premature) from unused freedom (headroom). DEVELOPMENT.md adds the operational lens — the invariant question and the stronger-mind question, with a symmetric proof burden against both fossilizing the current case and speculating a framework. | | 6.110.0 | 2026-08-22 | **feat: preserve work-order authority across oversized delegation and recovery.** Complete external work orders remain byte-complete within the host serializer budget; when an order needs bounded continuation, Ouroboros asks the same actor for an exact readable canonical range and keeps incomplete coverage as typed `cannot_verify` evidence. Partial input can no longer authorize PASS, a destructive rewrite, or replacement of the full contract, while valid complete work continues through the existing flow. | -| 6.109.0 | 2026-08-21 | **feat: live task cost and ready-on-open agent accounts.** Running root-task heartbeats now project the existing physical-attempt ledger into one non-final subtree total, so compact Chat and Activity cards advance live without a second timer, endpoint, or client-side sum while preserving reserved, unresolved, and unmetered disclosure (PR #288). Opening Agents now wakes only an already-provisioned stale Claudexor home through the existing owner-action endpoint after a side-effect-free status read; background polling, first-time installs, foreign homes, and repair states remain untouched (PR #289). Fail-closed staged-binary review fixtures now inject exact Git tree-read errors instead of assuming loose object storage, removing the macOS stable-CI race without changing production behavior. | -Older releases are preserved in this repository's history. Older 6.x rows (including 6.110.1, 6.108.1, 6.106.0, 6.101.1, 6.97.2, 6.105.0, 6.97.1, 6.97.0, 6.96.1, 6.96.0, 6.95.0, 6.94.0, 6.93.0, 6.92.1, 6.92.0, 6.91.1, 6.90.3, 6.91.0, 6.90.2, 6.90.0, 6.87.5, 6.87.4, 6.87.3, 6.87.2, 6.84.0, 6.87.1, 6.83.0, 6.86.1, 6.81.1, 6.76.0, 6.75.0, 6.74.5, 6.74.4, 6.74.1, 6.74.0, 6.73.2, 6.73.1, 6.73.0, 6.72.0, 6.71.2, 6.71.1, 6.71.0, 6.70.0, 6.69.0, 6.68.0, 6.67.0, 6.66.0, 6.65.4, 6.65.3, 6.65.2, 6.65.1, 6.65.0, 6.64.3, 6.64.2, 6.64.1, 6.64.0, 6.63.0, 6.62.0, 6.61.4, 6.61.3, 6.61.1, 6.61.0, 6.60.0, 6.59.0, 6.58.0, 6.57.0, 6.56.0, 6.55.0, 6.54.4, 6.54.2, 6.54.1, 6.54.0, 6.53.4, 6.53.0, 6.51.0), the 5.2.0 through 5.33.0-rc.6 rows, and former `4.0.0` rows are rolled off to respect the P9 changelog cap; their full bodies remain in this file's git history, with release tags where present for historical versions. +Older releases are preserved in this repository's history. Older 6.x rows (including 6.109.0, 6.110.1, 6.108.1, 6.106.0, 6.101.1, 6.97.2, 6.105.0, 6.97.1, 6.97.0, 6.96.1, 6.96.0, 6.95.0, 6.94.0, 6.93.0, 6.92.1, 6.92.0, 6.91.1, 6.90.3, 6.91.0, 6.90.2, 6.90.0, 6.87.5, 6.87.4, 6.87.3, 6.87.2, 6.84.0, 6.87.1, 6.83.0, 6.86.1, 6.81.1, 6.76.0, 6.75.0, 6.74.5, 6.74.4, 6.74.1, 6.74.0, 6.73.2, 6.73.1, 6.73.0, 6.72.0, 6.71.2, 6.71.1, 6.71.0, 6.70.0, 6.69.0, 6.68.0, 6.67.0, 6.66.0, 6.65.4, 6.65.3, 6.65.2, 6.65.1, 6.65.0, 6.64.3, 6.64.2, 6.64.1, 6.64.0, 6.63.0, 6.62.0, 6.61.4, 6.61.3, 6.61.1, 6.61.0, 6.60.0, 6.59.0, 6.58.0, 6.57.0, 6.56.0, 6.55.0, 6.54.4, 6.54.2, 6.54.1, 6.54.0, 6.53.4, 6.53.0, 6.51.0), the 5.2.0 through 5.33.0-rc.6 rows, and former `4.0.0` rows are rolled off to respect the P9 changelog cap; their full bodies remain in this file's git history, with release tags where present for historical versions. --- diff --git a/VERSION b/VERSION index 66ce77b7e..a3fcc7121 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -7.0.0 +7.1.0 diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index fc0441455..1adffde76 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -1,4 +1,4 @@ -# Ouroboros v7.0.0 — Architecture & Reference +# Ouroboros v7.1.0 — Architecture & Reference This is the present-tense operational map of Ouroboros (BIBLE P6), in three layers: structure (what exists and where), operation (files, env keys, state paths, endpoints, flows), and rationale; it is NOT a changelog, and version history lives in README.md, git tags, and the commit log. Every important WHY stays in this book at least briefly, while mechanism detail lives in the module docstring the map points to by name, and rationale must be self-contained — future maintainers should not need old commits to understand why a guard, review gate, or lifecycle exists. The chapters below are the book: each owns one section of the map, and a change replaces the description of the node it touched. diff --git a/docs/architecture/03-web-ui-pages-and-buttons.md b/docs/architecture/03-web-ui-pages-and-buttons.md index e10ae452c..feba8f0b1 100644 --- a/docs/architecture/03-web-ui-pages-and-buttons.md +++ b/docs/architecture/03-web-ui-pages-and-buttons.md @@ -38,7 +38,7 @@ Shared frontend primitives keep pages from acquiring competing contracts — fro Messages, media bubbles and task-card roots order by raw numeric timestamps (ties keep arrival order; the typing indicator stays last). Application-controlled height mutations use a stable-viewport seam: within 48 CSS pixels of the live edge the transcript follows the bottom, otherwise it restores the visible anchor; remote delivery while the reader is away coalesces into one instance-local activity marker cleared at the bottom. Card text is selectable and a selecting drag never toggles the card; scroll position is remembered per chat instance; band anatomy and the quiet `Reviews N` count: `docs/DESIGN.md`. -History reconciliation is one synchronous two-pass transaction over existing keyed records; recent refresh, reconnect and archive prepend preserve the actual timeline/Review DOM nodes, disclosure, focus and selection. Historical narration can enrich a finished child without reviving Working or Stop; source identities distinguish equal timestamps and identical text while timestamps still order presentation; a task's current result outranks its historical terminal observation. The live-card growth trigger requests a fresh history read after 200 additional cards, and only completed, unrepresented cards captured before that request may leave the view (reading/focus/selection or retained-page membership protect them), so live accumulation stays bounded without replacing a reader's card or erasing events that arrived during the fetch. +History reconciliation is one synchronous two-pass transaction over existing keyed records; recent refresh, reconnect and archive prepend preserve the actual timeline/Review DOM nodes, disclosure, focus and selection. A child keeps one keyed current lifecycle row across live delivery and replay; older pages cannot regress its status, and each narration record remains separately visible. Historical narration can enrich a finished child without reviving Working or Stop; source identities distinguish equal timestamps and identical text while timestamps still order presentation; a task's current result outranks its historical terminal observation. The live-card growth trigger requests a fresh history read after 200 additional cards, and only completed, unrepresented cards captured before that request may leave the view (reading/focus/selection or retained-page membership protect them), so live accumulation stays bounded without replacing a reader's card or erasing events that arrived during the fetch. `chat_history.js` retains three ordinary rendered pages plus temporarily protected reading/focus/selection pages, holding page descriptors only — content stays with the row/card owners, which release distant page bodies and their media/listener resources while descriptors keep exact return navigation. An empty scan is not a page, there is no newer control, and the per-room scroll stash carries descriptors plus a physical reading anchor, never a second history copy; recent live refresh stays separate from the frozen archive range, and reaching the old physical beginning does not claim every page is loaded. diff --git a/docs/architecture/06-agent-core.md b/docs/architecture/06-agent-core.md index 8cbc46c07..1a4140715 100644 --- a/docs/architecture/06-agent-core.md +++ b/docs/architecture/06-agent-core.md @@ -30,7 +30,7 @@ Disclosed cancel-lifecycle residuals (deliberate): a cascade over a tree with no Host-enforced task acceptance is a root-owned completion coach, not the P3 commit gate. `off` disables it; `auto` and `required` both review observable effects and typed deliverables/criteria. In `auto`, an explicit `task_acceptance_review` request also qualifies, read-only research included, while queue membership, ordinary conversation, exploration or cognitive-memory updates alone do not, and no prose or tool-count classifier decides their meaning; `required` retains its non-direct-root criterion. Child reviews remain advisory evidence superseded by the root decision. -The explicit call nominates the complete ready result and returns `deferred_to_host_acceptance`, `authoritative=false`; after the whole tool-result block the host advances the same acceptance operation ordinary final delivery uses, and early feedback does not seal the task. Main authors the effective criteria, and the paid subject is result bytes, criteria and material effects, so a status question preserves a running review while a new criterion can buy review of unchanged text. A reviewer panel is advice for its author, never a signature on bytes the reviewers did not read: the wave wakes the original Main through the mailbox at its quorum and again when the last slot settles, each wake carrying each reviewer's own verdict (`acceptance_settlement.announce_acceptance_settlement`). A final answer that is not an explicit nomination neither buys a second panel nor is refused one while the panel this turn released still runs: Main waits (the default, and the only option under blocking enforcement; Cyber Pro keeps its own rule — Main's final response is its decision) or consciously finishes through the `pending_review` key of the delivery control; a panel that settled PASS on the earlier revision accepts the task on the reviewers' word (`previous_revision_accepted`, the owner row says so), and any other settled verdict hands delivery to the ordinary path with the collected verdicts in its dialogue history. Before a new panel is assembled and before any capacity refusal, every recorded still-pending panel of the same root is collected at $0 over its recorded request and roster (`review_dispatch.reconcile_pending_acceptance_runs`), so a subject re-authored mid-flight cannot discard verdicts the tree already bought; the collected verdicts enter the next panel's dialogue history outside the hashed material, so reading them mints no paid binding. Only actual final delivery seals ingress; stop, missing custody and unfinished work keep their observed outcomes. +The explicit call nominates the complete ready result and returns `deferred_to_host_acceptance`, `authoritative=false`; after the whole tool-result block the host advances the same acceptance operation ordinary final delivery uses, and early feedback does not seal the task. Main authors the effective criteria, and the paid subject is result bytes, criteria and material effects, so a status question preserves a running review while a new criterion can buy review of unchanged text. A reviewer panel is advice for its author, never a signature on bytes the reviewers did not read: the wave wakes the original Main through the mailbox at its quorum and again when the last slot settles, each wake carrying each reviewer's own verdict (`acceptance_settlement.announce_acceptance_settlement`). Only a still-pending panel may park the turn; `acceptance_settlement.awaited_panel_has_settled` lets the next round, including control repair, run once the panel is settled. A final answer that is not an explicit nomination neither buys a second panel nor is refused one while the panel this turn released still runs: Main waits (the default, and the only option under blocking enforcement; Cyber Pro keeps its own rule — Main's final response is its decision) or consciously finishes through the `pending_review` key of the delivery control; a panel that settled PASS on the earlier revision accepts the task on the reviewers' word (`previous_revision_accepted`, the owner row says so), and any other settled verdict hands delivery to the ordinary path with the collected verdicts in its dialogue history. Before a new panel is assembled and before any capacity refusal, every recorded still-pending panel of the same root is collected at $0 over its recorded request and roster (`review_dispatch.reconcile_pending_acceptance_runs`), so a subject re-authored mid-flight cannot discard verdicts the tree already bought; the collected verdicts enter the next panel's dialogue history outside the hashed material, so reading them mints no paid binding. Only actual final delivery seals ingress; stop, missing custody and unfinished work keep their observed outcomes. A completed reviewer from an older plan wave is attached as a historical supplement through the locked task-result writer and exact producer CAS: it never rewrites the original verdict, aggregate, closure, author dispositions or current-wave pointer, settles its historical cost without another cycle, and reaches a terminal parent without a new model turn. Task acceptance has the same twin: a panel that settles after its task is terminal is collected at $0 over the recorded operation, republished on the task's review projection with a host-composed `late_settlement` note (the verdict, the revision it covered, settled after the terminal; honest about a reviewer whose physical outcome is still unknown) and announced once in the task's room as a System row stamped `card_row="reviews"` (`acceptance_settlement.attach_late_acceptance_settlement`, deduped by delivery id) — a timeline item of the card, read by the next turn from chat history; no model turn starts. diff --git a/docs/install/index.html b/docs/install/index.html index 1bf17982b..2badf3a02 100644 --- a/docs/install/index.html +++ b/docs/install/index.html @@ -46,24 +46,24 @@

macOS 12+ · Apple silicon only

macOS

- Download for macOS (.dmg) + Download for macOS (.dmg)

Open the DMG, drag Ouroboros.app to Applications, then launch it from Applications.

Windows x64

Windows

- Download for Windows (.zip) + Download for Windows (.zip)

Extract the ZIP, open the Ouroboros folder, and run Ouroboros.exe.

Linux x86_64

Linux

- Debian / Ubuntu / Astra (.deb) - Fedora / RHEL (.rpm) - RED OS 8 (.rpm) - Portable AppImage - tar.gz archive + Debian / Ubuntu / Astra (.deb) + Fedora / RHEL (.rpm) + RED OS 8 (.rpm) + Portable AppImage + tar.gz archive

Prefer the package for your distribution. Use the AppImage or archive on other distributions; Git must already be installed.

@@ -74,7 +74,7 @@

macOS quick start

    -
  1. Click Download for macOS (.dmg).
  2. +
  3. Click Download for macOS (.dmg).
  4. Open the DMG and drag Ouroboros.app onto the Applications shortcut.
  5. Open Ouroboros from Applications. If Gatekeeper asks, right-click the app and choose Open.
diff --git a/ouroboros/acceptance_settlement.py b/ouroboros/acceptance_settlement.py index 1e9e08fa5..6e22f9040 100644 --- a/ouroboros/acceptance_settlement.py +++ b/ouroboros/acceptance_settlement.py @@ -162,29 +162,44 @@ def announce_acceptance_settlement(usage_ctx: Any, request: Any, wave: Dict[str, def panel_awaiting_this_turn(tools_ctx: Any, llm_trace: Dict[str, Any]) -> Optional[Dict[str, Any]]: - """The panel THIS turn released, found by the binding the host recorded when - it went pending — the one identity a re-authored answer cannot move. + """Read the panel this turn released without changing delivery authority. - A panel bound under older owner input is not that panel: once Main has - acknowledged a newer owner source (the owner's words changed the premises, - owner rule 4=A), the latch clears and the ordinary path decides, while the - old panel keeps its custody and its verdicts still arrive as advice. + The binding recorded when it went pending is the physical identity a + re-authored answer cannot move; current owner-source validation belongs + to delivery, never to the wait's settlement check. """ - from ouroboros.loop_messages import owner_source_sha256 - binding = str(getattr(tools_ctx, "_task_acceptance_pending", "") or "") if not binding: return None - run = next((run for run in reversed(llm_trace.get("review_runs") or []) - if isinstance(run, dict) and run.get("authority") == "host_root" - and str(run.get("binding_hash") or "") == binding), None) + return next((run for run in reversed(llm_trace.get("review_runs") or []) + if isinstance(run, dict) and run.get("authority") == "host_root" + and str(run.get("binding_hash") or "") == binding), None) + + +def awaited_panel_has_settled(tools_ctx: Any, llm_trace: Dict[str, Any]) -> bool: + """Whether the panel this turn waits for has already settled (a $0 look). + + Settled means its verdicts already woke Main and sit in the transcript: the + only thing left is the next model round — a control repair, for one — so + the loop must run it instead of parking behind a settlement that will never + arrive again (the keyless E2E lane hung that way until the task deadline). + The run record is reconciled at $0 first, exactly as delivery does, because + the trace learns of a settlement only through that collection. + """ + from ouroboros.loop_acceptance_review import acceptance_run_pending + from ouroboros.review_dispatch import reconcile_pending_acceptance_runs + + run = panel_awaiting_this_turn(tools_ctx, llm_trace) if run is None: - return None - reviewed_source = str(run.get("owner_source_sha256") or "") - if reviewed_source and reviewed_source != str(owner_source_sha256(tools_ctx) or ""): - tools_ctx._task_acceptance_pending = "" - return None - return run + return False + if acceptance_run_pending(run): + try: + reconcile_pending_acceptance_runs( + {"review_runs": [run]}, drive_root=pathlib.Path(tools_ctx.drive_root), + usage_ctx=tools_ctx) + except Exception: + log.debug("awaited acceptance panel could not be reconciled", exc_info=True) + return not acceptance_run_pending(run) def acceptance_choice_offered() -> bool: @@ -233,6 +248,7 @@ def _deliver_under_running_panel(ctx: Any, prior_run: Any) -> Optional[bool]: _end_acceptance_terminal, _finish_cyber_acceptance, _set_applied_host_acceptance_impact, acceptance_run_pending, ) + from ouroboros.loop_messages import owner_source_sha256 from ouroboros.outcomes import ACCEPTANCE_ACCEPTED tools_ctx = ctx.tools._ctx @@ -241,6 +257,12 @@ def _deliver_under_running_panel(ctx: Any, prior_run: Any) -> Optional[bool]: run = panel_awaiting_this_turn(tools_ctx, ctx.llm_trace) if run is None: return None + # Older owner premises cannot authorize this delivery (owner rule 4=A). + # The panel keeps its physical custody and still arrives as advice. + reviewed_source = str(run.get("owner_source_sha256") or "") + if reviewed_source and reviewed_source != str(owner_source_sha256(tools_ctx) or ""): + tools_ctx._task_acceptance_pending = "" + return None if acceptance_run_pending(run): if acceptance_wait_chosen(tools_ctx): ctx.emit_progress("Task acceptance review is still running; holding the answer for its verdict.") diff --git a/ouroboros/loop_acceptance_review.py b/ouroboros/loop_acceptance_review.py index 3e1f09025..0b6385ed9 100644 --- a/ouroboros/loop_acceptance_review.py +++ b/ouroboros/loop_acceptance_review.py @@ -159,6 +159,10 @@ def wait_for_acceptance_feedback(tools: Any, limit_ctx: Any, trace: dict, binding = getattr(ctx, "_task_acceptance_pending", "") if not binding: return + from ouroboros.acceptance_settlement import awaited_panel_has_settled + + if awaited_panel_has_settled(ctx, trace): + return # its verdicts already woke this turn; the next round runs, nothing settles again from ouroboros.loop_transport import _owner_signal_pending # A wake arriving during Main's request must reach the normal ingress drain. diff --git a/pyproject.toml b/pyproject.toml index c8964f8f2..892f46496 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "ouroboros" -version = "7.0.0" +version = "7.1.0" description = "Self-creating AI agent with constitution, background consciousness, and persistent identity" readme = "README.md" license = {text = "MIT"} diff --git a/site/install/index.html b/site/install/index.html index 1f3572499..d5abe836e 100644 --- a/site/install/index.html +++ b/site/install/index.html @@ -46,24 +46,24 @@

macOS 12+ · Apple silicon only

macOS

- Download for macOS (.dmg) + Download for macOS (.dmg)

Open the DMG, drag Ouroboros.app to Applications, then launch it from Applications.

Windows x64

Windows

- Download for Windows (.zip) + Download for Windows (.zip)

Extract the ZIP, open the Ouroboros folder, and run Ouroboros.exe.

Linux x86_64

Linux

Prefer the package for your distribution. Use the AppImage or archive on other distributions; Git must already be installed.

@@ -74,7 +74,7 @@

macOS quick start

    -
  1. Click Download for macOS (.dmg).
  2. +
  3. Click Download for macOS (.dmg).
  4. Open the DMG and drag Ouroboros.app onto the Applications shortcut.
  5. Open Ouroboros from Applications. If Gatekeeper asks, right-click the app and choose Open.
diff --git a/tests/system_e2e/test_system_scenarios_w3a.py b/tests/system_e2e/test_system_scenarios_w3a.py index 4e3a05075..98909e5e8 100644 --- a/tests/system_e2e/test_system_scenarios_w3a.py +++ b/tests/system_e2e/test_system_scenarios_w3a.py @@ -5,11 +5,19 @@ Every scenario asserts DURABLE artifacts (never an HTTP 200 alone, never a harness exit code) and synchronizes by durable-event polling: * S14 — PLAN REVIEW: a scripted task drives ``plan_task`` through a real - REVISE→ACCEPT cycle (cycle 1: every triad slot returns a blocking finding → - REVISE_PLAN; cycle 2: a CHANGED spec → all-clean → GREEN closed), the durable - chronicle is honest (``plan_review_state`` on the stored task row: paid-cycle - count, wave aggregates; immutable per-wave artifacts with the exact reviewer - outputs), and the shared owner cycle cap (``OUROBOROS_REVIEW_MAX_CYCLES``) is + REVISE→ACCEPT cycle over the ASYNCHRONOUS event route, which is the whole + point of the scenario — a fresh dispatch returns at the dispatch barrier with + an OPEN custody-pending wave (no verdict yet), the released reviewer slots + settle under process-local custody, the LAST settlement writes ONE system + frame into this task's mailbox, and the IDENTICAL envelope resubmitted after + that frame is the $0 collection that closes the wave and pays the cycle + exactly once. So the scripted agent dispatches, waits on its own task id + (the wait returns on ``owner_mailbox_pending``), collects, and only then + revises (cycle 1: every slot returns a blocking finding → REVISE_PLAN; + cycle 2: a CHANGED spec → all-clean → GREEN closed). The durable chronicle is + honest (``plan_review_state`` on the stored task row: paid-cycle count, wave + aggregates; immutable per-wave artifacts with the exact reviewer outputs), + and the shared owner cycle cap (``OUROBOROS_REVIEW_MAX_CYCLES``) is respected: a third paid cycle is refused with the typed ``PLAN_REVIEW_CYCLES_EXHAUSTED`` result at $0 (no reviewer dispatched) plus the durable ``review_cycles_exhausted`` escalation event. @@ -55,6 +63,7 @@ absorb and the update variations in wave 4. Still deferred: gateway/UI truth from __future__ import annotations import json +import re import subprocess import pytest @@ -67,6 +76,7 @@ from tests.system_e2e.harness import ( ArtifactOracle, ReviewScript, ScriptedStubModel, + body_text, classify_call, clone_repo, keyless_reviewer_slots, @@ -274,6 +284,92 @@ def _plan_step(spec: dict, note: str) -> dict: }} +# The ONE system frame plan_review_collect.announce_released_settlement writes +# into the task's mailbox when the last released slot of a wave settles, and the +# task identity the host states in every prompt (the scripted agent waits on its +# OWN id, exactly as the plan-review contract text tells the model to). +_SETTLED_FRAME_RE = re.compile( + r"Plan review wave ([0-9a-f?]+): \d+ of \d+ reviewer slot\(s\) settled") +_ROOT_TASK_ID_RE = re.compile(r'"root_task_id":\s*"([0-9a-f]{8,})"') +# Bounded so a wave that never settles fails LOUDLY inside the scenario +# (the E2E_SCRIPT_ERROR final answer below) instead of running the task out +# of the harness's own 600s wait with no explanation. +_WAIT_WINDOW_SEC = 60 +_WAIT_ROUNDS_MAX = 4 + + +class _Again: + """Marker a callable step returns to stay on the script one more round.""" + + __slots__ = ("step",) + + def __init__(self, step: dict) -> None: + self.step = step + + +class _HoldingStubModel(ScriptedStubModel): + """``ScriptedStubModel`` whose callable steps may HOLD the script index. + + The asynchronous plan-review route needs a step that repeats until a + precondition is visible in the transcript: the agent waits for the settled- + wave mailbox frame and only then resubmits the identical envelope. A blind + sleep inside the step is not an option — ``_answer`` runs under the model's + call lock, so sleeping there would also block the reviewer-slot calls whose + settlement it is waiting for. A callable step therefore returns ``_Again`` + to be served now AND decide again on the next round; everything else is + served and consumed exactly as the base class does, so ``script_consumed()`` + still means "every phase of the scenario ran". + """ + + def _next_step(self, body) -> dict | None: + index = self._script_index + step = super()._next_step(body) + if callable(step): + step = step(body) + if isinstance(step, _Again): + self._script_index = index + step = step.step + return step + + +def _after_wave_settled(wave_ordinal: int, then: dict): + """Hold the script until the host says plan-review wave ``wave_ordinal`` + settled, then emit ``then``. + + Until the frame is visible the step waits on this task's own id + (``wait_task`` returns early on ``owner_mailbox_pending``) — the route the + plan-review contract text names. Acting on an in-flight wave instead would + either re-read DEGRADED (a collection) or be refused as + ``PLAN_REVIEW_IN_FLIGHT`` (a revision), so the wait is what makes either + next move deterministic rather than a race with the reviewer slots. + """ + waits = {"rounds": 0} + + def step(body: dict) -> dict: + text = body_text(body) + settled = list(dict.fromkeys(_SETTLED_FRAME_RE.findall(text))) + if len(settled) >= wave_ordinal: + return then + waits["rounds"] += 1 + task_ids = _ROOT_TASK_ID_RE.findall(text) + if not task_ids or waits["rounds"] > _WAIT_ROUNDS_MAX: + return {"final": ( + "E2E_SCRIPT_ERROR: no settled-wave frame for plan-review wave " + f"{wave_ordinal} after {waits['rounds']} round(s); " + f"frames seen: {settled}; own task id visible: {bool(task_ids)}")} + return _Again({"tool": "wait_task", "arguments": { + "task_id": task_ids[-1], "timeout_sec": _WAIT_WINDOW_SEC}}) + + return step + + +def _collect_step(spec: dict, note: str, *, wave_ordinal: int): + """The $0 collection of wave ``wave_ordinal``: the IDENTICAL envelope, + resubmitted after the wave settled. It closes the recorded wave, aggregates + the verdicts and pays the cycle once, dispatching no second panel.""" + return _after_wave_settled(wave_ordinal, _plan_step(spec, note)) + + @pytest.mark.integration @pytest.mark.serial def test_s14_plan_review_revise_then_accept_cycle_with_honest_chronicle( @@ -281,8 +377,14 @@ def test_s14_plan_review_revise_then_accept_cycle_with_honest_chronicle( require_lane(LANE_MOCK) root = tmp_path_factory.mktemp("s11") review_script = ReviewScript({"plan_review": [W3A_PLAN_RED] * 3}) - stub = ScriptedStubModel( - [_plan_step(S11_SPEC_V1, "cycle 1"), _plan_step(S11_SPEC_V2, "cycle 2 — revised")], + # Dispatch -> wait for the settled-wave frame -> collect ($0), twice: the + # honest shape of the asynchronous event route (a fresh call returns at the + # dispatch barrier, the identical resubmission is the collection). + stub = _HoldingStubModel( + [_plan_step(S11_SPEC_V1, "cycle 1"), + _collect_step(S11_SPEC_V1, "cycle 1", wave_ordinal=1), + _plan_step(S11_SPEC_V2, "cycle 2 — revised"), + _collect_step(S11_SPEC_V2, "cycle 2 — revised", wave_ordinal=2)], review_script=review_script, ) with stub: @@ -307,26 +409,41 @@ def test_s14_plan_review_revise_then_accept_cycle_with_honest_chronicle( assert waves[-1].get("closed") is True, waves[-1] # The immutable per-wave artifacts carry the exact reviewer wave - # bytes: one REVISE_PLAN wave, one GREEN wave. The task artifact - # store lives under the SERVER data root (task_results/artifacts/), - # not the task's forked drive. + # bytes. The asynchronous route snapshots each cycle TWICE and both + # snapshots are evidence: the OPEN barrier wave recorded at dispatch + # (custody pending, unpaid, no verdict yet) and the wave the $0 + # collection closed. The verdict chronicle is the collected pair: + # one REVISE_PLAN wave, one GREEN wave. The task artifact store + # lives under the SERVER data root (task_results/artifacts/), not + # the task's forked drive; filenames end in a content hash, so the + # snapshots are keyed by their own facts, never by sort order. artifacts = sorted( (oracle.data_root / "task_results" / "artifacts").rglob("plan-review-wave-*.json")) - assert len(artifacts) == 2, artifacts - aggregates = [] - for path in artifacts: - payload = json.loads(path.read_text(encoding="utf-8")) - aggregates.append(str(payload.get("aggregate") or "")) + payloads = [json.loads(path.read_text(encoding="utf-8")) for path in artifacts] + assert len(payloads) == 4, artifacts + def by_cycle(payload: dict) -> int: + return int(payload.get("cycle_index") or 0) + + barrier = sorted((p for p in payloads if p.get("custody_pending")), key=by_cycle) + collected = sorted((p for p in payloads if not p.get("custody_pending")), key=by_cycle) + assert [p.get("aggregate") for p in barrier] == ["DEGRADED"] * 2, barrier + assert [p.get("paid") for p in barrier] == [False] * 2, barrier + assert [p.get("cycle_index") for p in collected] == [1, 2], collected + assert [p.get("paid") for p in collected] == [True] * 2, collected + aggregates = [str(p.get("aggregate") or "") for p in collected] assert aggregates == ["REVISE_PLAN", "GREEN"], aggregates # The REVISE wave chronicled the scripted finding honestly. - revise_payload = json.loads(artifacts[0].read_text(encoding="utf-8")) - revise_blob = json.dumps(revise_payload) + revise_blob = json.dumps(collected[0]) assert "checkable marker" in revise_blob, revise_blob[:2000] # Exactly two paid waves of three slots each hit the model; the - # scripted red round was fully served. + # scripted red round was fully served. FOUR plan_task calls bought + # only those two panels: each cycle is one dispatch plus one $0 + # collection of the same envelope, never a second panel. assert stub.kinds().count("plan_review") == 6, stub.kinds() + plan_rows = _tool_rows(oracle.task_drive(task_id), "plan_task") + assert len(plan_rows) == 4, plan_rows review_script.assert_consumed() assert stub.script_consumed() finally: @@ -341,10 +458,15 @@ def test_s14_plan_review_cycle_cap_refuses_third_paid_cycle(e2e_clone, tmp_path_ root = tmp_path_factory.mktemp("s11cap") review_script = ReviewScript({"plan_review": [W3A_PLAN_RED] * 6}) - stub = ScriptedStubModel( + # The cap is enforced at DISPATCH, so this scenario never collects — but it + # still waits for each panel to settle before submitting the next envelope. + # A revised envelope sent while the previous panel is still in flight is + # refused as PLAN_REVIEW_IN_FLIGHT, which would mask the cap refusal under + # a different typed reason. + stub = _HoldingStubModel( [_plan_step(S11_SPEC_V1, "cycle 1"), - _plan_step(S11_SPEC_V2, "cycle 2 — still open"), - _plan_step(S11_SPEC_V3, "cycle 3 — must be refused by the cap")], + _after_wave_settled(1, _plan_step(S11_SPEC_V2, "cycle 2 — still open")), + _after_wave_settled(2, _plan_step(S11_SPEC_V3, "cycle 3 — must be refused by the cap"))], review_script=review_script, ) with stub: @@ -702,6 +824,28 @@ S14_ANSWER_V1 = "Final answer: the summary is drafted (first pass)." S14_ANSWER_V2 = "Final answer: the summary is complete. W3A_DONE" +_OWNER_SOURCE_RE = re.compile(r'"owner_source_sha256": "([0-9a-f]{64})"') + + +def _control_step(delivery_control: str, full_answer: str | None = None): + """A scripted answer to the host's delivery control. + + Acceptance is asynchronous (owner D4=A): the panel settles in custody and its + verdicts wake Main with the keep/replace control re-offered, so the scripted + agent answers that control instead of a bare final. The exact owner source + selector is read from the transcript's last ``[ACCEPTANCE_SUBJECT_OBSERVATION]`` + (the wave-2 dynamic-argument contract of ``scripted_completion``).""" + def step(body: dict) -> dict: + found = _OWNER_SOURCE_RE.findall(body_text(body)) + control = {"delivery_control": delivery_control} + if full_answer is not None: + control["full_answer"] = full_answer + if found: + control["acceptance_subject"] = {"owner_source_sha256": found[-1]} + return {"final": json.dumps(control)} + return step + + def _s14_settings(stub) -> dict: return keyless_settings( stub, @@ -718,8 +862,11 @@ def test_s17_acceptance_reject_rework_accept(e2e_clone, tmp_path_factory): review_script = ReviewScript({ "acceptance": [W3A_ACCEPT_REJECT] * 3 + [W3A_ACCEPT_PASS] * 3, }) + # V1 is nominated; the REJECT wave wakes Main with the control re-offered and + # the rework answers it as a typed replace (a second paid panel); the PASS + # wave wakes Main again and a typed keep delivers under the accepted panel. stub = ScriptedStubModel( - [{"final": S14_ANSWER_V1}, {"final": S14_ANSWER_V2}], + [{"final": S14_ANSWER_V1}, _control_step("replace", S14_ANSWER_V2), _control_step("keep")], review_script=review_script, ) with stub: @@ -761,8 +908,10 @@ def test_s17_acceptance_identical_rework_is_free_replay_refusal( require_lane(LANE_MOCK) root = tmp_path_factory.mktemp("s14b") review_script = ReviewScript({"acceptance": [W3A_ACCEPT_REJECT] * 3}) + # Keep collects the REJECT verdict; the following improvement round + # resubmits the identical answer, which must replay at $0 without a new panel. stub = ScriptedStubModel( - [{"final": S14_ANSWER_V1}, {"final": S14_ANSWER_V1}], # rework changes NOTHING + [{"final": S14_ANSWER_V1}, _control_step("keep"), {"final": S14_ANSWER_V1}], review_script=review_script, ) with stub: diff --git a/tests/system_e2e/test_system_scenarios_w5.py b/tests/system_e2e/test_system_scenarios_w5.py index 4506c03a9..b380ef24c 100644 --- a/tests/system_e2e/test_system_scenarios_w5.py +++ b/tests/system_e2e/test_system_scenarios_w5.py @@ -131,6 +131,7 @@ _BUILDER_ROW = { "subagent_id": "cx-builder", "recommended_use": "Delegated builder for the system_e2e mutation scenarios.", "route": {"kind": "agent_session", "target_id": "fake-harness=mock-model"}, + "access": "workspace_write", "effort": "low", } S24_MARKER = "S24_PARENT_FINAL_e2e_w5" diff --git a/tests/test_acceptance_repair_after_settlement.py b/tests/test_acceptance_repair_after_settlement.py new file mode 100644 index 000000000..9b95c219c --- /dev/null +++ b/tests/test_acceptance_repair_after_settlement.py @@ -0,0 +1,67 @@ +"""The control-repair round after a settled acceptance panel runs; it never parks again.""" +from __future__ import annotations + +import copy +import json + +import pytest + +from ouroboros import loop +from tests.test_acceptance_async_loop import ANSWER, call, keep +from tests.test_acceptance_async_loop import full_loop as _full_loop + +full_loop = _full_loop # noqa: F811 - pytest fixture re-export + + +@pytest.mark.parametrize("owner_followup", [False, True], ids=["same_owner", "owner_followup"]) +def test_prose_after_the_settled_verdict_wake_gets_its_repair_round_not_a_second_park( + full_loop, monkeypatch, owner_followup, +): + """The wake after a settled panel re-offers the keep/replace control; a prose + answer there is a malformed control, so the loop queues its ONE repair round. + That round must run: the panel has already settled, so parking again would + wait for a settlement that never comes (the keyless E2E lane hung this way + until the task deadline). One park while the panel ran, then the repair.""" + f = full_loop + f.reviewer_verdict = "FAIL" + monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "1") + reauthored = ANSWER + " Budget: $12." + followup = "Also show the budget as an explicit dollar amount." + + def park(ctx, checkpoint): + f.park(ctx, checkpoint) + if owner_followup and len(f.waits) == 1: + # The owner's new input reaches the same Main wake as the settled + # verdict. That panel no longer covers the current owner corpus. + f.incoming.put(followup) + + f.ctx.owner_wait_callback = park + + def main(_llm, messages, *_a, **_kw): + f.model_inputs.append(copy.deepcopy(messages)) + f.model_step += 1 + if f.model_step == 1: + return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0 + if f.model_step == 2: + assert f.entered.wait(5) and not f.release.is_set() + return keep(f), 0.0 # held under the running panel: the one legitimate park + if f.model_step == 3: + assert f.settled.is_set() and "FAIL" in str(messages) + if owner_followup: + assert followup in str(messages) + return {"content": reauthored}, 0.0 # prose where the control object was due + if f.model_step == 4: + assert "[DELIVERY_CONTROL_REPAIR]" in str(messages[-1].get("content")) + observation = f.ctx._acceptance_observation + return {"content": json.dumps({"delivery_control": "replace", "full_answer": reauthored, + "acceptance_subject": {"owner_source_sha256": observation["owner_source_sha256"]}})}, 0.0 + assert f.model_step < 7, f.progress + return keep(f), 0.0 + + monkeypatch.setattr(loop, "call_llm_with_retry", main) + result, _usage, trace = f.run() + assert result == reauthored and len(f.review_sends) == 1 + assert [wait.get("reason") for wait in f.waits] == ["review"], f.waits + assert f.model_step == 4, f.progress + host = [r for r in trace["review_runs"] if r.get("authority") == "host_root"] + assert [r.get("aggregate_signal") for r in host] == ["FAIL", "DEGRADED"], host diff --git a/tests/test_account_catalog_browser.py b/tests/test_account_catalog_browser.py index 6630e3079..353e7a441 100644 --- a/tests/test_account_catalog_browser.py +++ b/tests/test_account_catalog_browser.py @@ -215,8 +215,10 @@ def test_partial_account_refresh_keeps_draft_and_only_failed_account_history(acc assert "not in discovery" in suggestions(page, field)[DRAFT_MODEL] account.select_option("work") assert "not checked" in suggestions(page, field)[DRAFT_MODEL] - assert row.locator("[data-subagent-status]").inner_text() == "Draft · Not checked" - assert "model list could not be read" in row.locator("[data-subagent-meta]").inner_text() + status = row.locator("[data-subagent-status]") + assert status.inner_text() == "Draft · Not checked" + # The unread account catalog is disclosed on the row's own status sentence. + assert "model list could not be read" in status.get_attribute("title") account.select_option("") roles.capture(page, "account-catalog-partial-" + consumer.lower().replace(" ", "-")) assert_saved(ui, consumer, DRAFT_MODEL, "") diff --git a/tests/test_author_ui_kit_browser.py b/tests/test_author_ui_kit_browser.py index 58621fbf3..55fa3d99c 100644 --- a/tests/test_author_ui_kit_browser.py +++ b/tests/test_author_ui_kit_browser.py @@ -195,6 +195,9 @@ def test_author_kit_authenticated_mount_and_lifetime(author_kit_server, tmp_path assert frame.locator('[role="status"]').inner_text() == "Enter a title." assert frame.locator('[role="status"]').get_attribute("data-tone") == "danger" frame.get_by_label("Title", exact=True).fill("Personal") + # select_option assigns the value without focusing the control. + # Exercise the native focus/blur cycle as part of choosing a view. + frame.get_by_label("View", exact=True).focus() frame.get_by_label("View", exact=True).select_option("Grid") frame.get_by_label("Enabled", exact=True).uncheck() frame.get_by_role("button", name="Preview", exact=True).click() @@ -216,6 +219,7 @@ def test_author_kit_authenticated_mount_and_lifetime(author_kit_server, tmp_path for style_page, name in zip(style_pages, ("spa", "wizard")): style_page.goto(server["url"] + "/style-" + name) style_page.locator('[name="title"]').wait_for() + style_page.get_by_label("View", exact=True).focus() style_page.get_by_label("View", exact=True).select_option("Grid") style_page.get_by_label("Enabled", exact=True).uncheck() style_page.evaluate("document.activeElement.blur()") diff --git a/tests/test_cyber_review_provenance.py b/tests/test_cyber_review_provenance.py index e52d3f577..1226aa76e 100644 --- a/tests/test_cyber_review_provenance.py +++ b/tests/test_cyber_review_provenance.py @@ -1,6 +1,7 @@ """Recovered first-review and exact-author evidence regressions from the preserved Ouroboros candidate.""" import copy import json +import threading from types import SimpleNamespace import pytest from tests.test_loop_acceptance_gate import _seed_acceptance_root @@ -28,13 +29,27 @@ def test_plan_first_dispatch_then_author_repeat_keeps_exact_raw_wave(harness, mo from ouroboros.tools import plan_review_runtime transport = AccountedFakeLLM(h.drive, reply=raw) + release = threading.Event() + chat = transport.chat + + def held_chat(**kwargs): + assert release.wait(20), "test did not release the physical reviewer calls" + return chat(**kwargs) + + monkeypatch.setattr(transport, "chat", held_chat) monkeypatch.setattr(plan_review_runtime, "LLMClient", lambda: transport) ctx = h.make_ctx() - barrier = _control(_call(ctx)) - assert barrier == {"outcome": "DEGRADED", "closed": False} - open_wave = _state(h)["waves"][-1] - assert open_wave["custody_pending"] is True and open_wave["paid"] is False - assert {actor["operation_state"] for actor in open_wave["actors"]} == {"pending_dispatch"} + try: + barrier = _control(_call(ctx)) + assert barrier == {"outcome": "DEGRADED", "closed": False} + open_wave = _state(h)["waves"][-1] + # Pin the pre-dispatch phase: a fast real send may otherwise settle a + # slot before the barrier snapshot and correctly make the wave paid. + assert open_wave["custody_pending"] is True and open_wave["paid"] is False + assert {actor["operation_state"] for actor in open_wave["actors"]} == {"pending_dispatch"} + assert not transport.calls + finally: + release.set() assert _wait_until(lambda: len(_mailbox_entries(h.drive, ctx.task_id)) == 1) dispatched = len(transport.calls) first = _control(_call(ctx)) # the identical envelope collects the settled wave diff --git a/tests/test_executor_card_browser.py b/tests/test_executor_card_browser.py index c733fcbb5..3e9e44058 100644 --- a/tests/test_executor_card_browser.py +++ b/tests/test_executor_card_browser.py @@ -13,7 +13,8 @@ def progress(task, text, **extra): @pytest.mark.parametrize('width', [390, 1440]) -def test_observed_executor_role_activity_and_project_pointer_live_replay(subscription_ui, width): +@pytest.mark.parametrize('linux_metrics', [False, True], ids=['native-font', 'linux-metrics']) +def test_observed_executor_role_activity_and_project_pointer_live_replay(subscription_ui, width, linux_metrics): ui = subscription_ui page = ui['page'] page.set_viewport_size({'width':width,'height':800}) @@ -64,6 +65,13 @@ def test_observed_executor_role_activity_and_project_pointer_live_replay(subscri assert activity.evaluate('e=>getComputedStyle(e).webkitLineClamp') == '1' pointer=page.locator('#executor-proof .project-work-pointer') assert 'UI coherence' in pointer.inner_text() + if linux_metrics: + # DejaVu Sans renders these two labels at 137px / 81px. Exercise that + # real Linux geometry on every host without bundling another font. + page.add_style_tag(content=''' + #executor-proof .project-work-coverage { min-width: 137px; } + #executor-proof .chat-panel-statusbar .status-badge { min-width: 81px; } + ''') # One-line pointer: the label ellipsizes instead of wrapping the status bar open. label=pointer.locator('.project-work-pointer-label') assert label.evaluate('e=>[getComputedStyle(e).whiteSpace,getComputedStyle(e).textOverflow]')==['nowrap','ellipsis'] @@ -85,6 +93,6 @@ def test_observed_executor_role_activity_and_project_pointer_live_replay(subscri page.evaluate('''frame=>{for(const fn of executorProof.handlers.get('chat')||[])fn({...frame,chat_id:101})}''',terminal) page.wait_for_function("document.querySelector('#executor-proof [data-task-id=child]').textContent.includes('Observed: Cursor Grok 4.6')") assert 'Coordinator: gpt-6-astra' in card.inner_text() - capture(page,f'executor-card-{width}') + capture(page,f'executor-card-{width}-{"linux-metrics" if linux_metrics else "native-font"}') page.evaluate('executorProof.instance.destroy()') assert page.locator('#executor-proof .project-work-pointer').count()==0 diff --git a/tests/test_extension_plugin_api.py b/tests/test_extension_plugin_api.py index 751073c73..34a1366de 100644 --- a/tests/test_extension_plugin_api.py +++ b/tests/test_extension_plugin_api.py @@ -230,7 +230,10 @@ def test_unload_does_not_deadlock_with_inflight_get_settings(tmp_path): unload_thread.start() time.sleep(0.1) release_reader.set() - unload_thread.join(timeout=1.0) + # The assertion is "no deadlock", so the deadline only has to be shorter + # than a hang; one second is a wall-clock guess that a loaded Windows CI + # runner (xdist workers sharing two cores) misses without any deadlock. + unload_thread.join(timeout=15.0) assert unload_done.is_set() assert extension_loader.snapshot()["extensions"] == [] diff --git a/tests/test_model_chooser_browser.py b/tests/test_model_chooser_browser.py index 8558e97de..90441077c 100644 --- a/tests/test_model_chooser_browser.py +++ b/tests/test_model_chooser_browser.py @@ -20,6 +20,9 @@ def catalog(ui): def select_model_field(ui, consumer): roles.configure_mixed(ui) + # The catalog above is unprefixed, so every consumer is put on the same + # OpenRouter API lane the suggestions belong to. + lane = roles.api_lane(ui) if consumer in ['Scope', 'Deep']: slots = ui['fixture']['preview']['reviewer_slots'] if consumer == 'Scope': slots['scope'] = [{'slot_id': 'scope_1', 'route': {'kind': 'api_chat', 'target_id': 'owner/model'}}] @@ -33,15 +36,16 @@ def select_model_field(ui, consumer): page.locator('[data-settings-tab="models"]').click() if consumer == 'Fallback': page.locator('[data-model-add]').click() group = page.locator('[data-model-role="main"]') if consumer == 'Models' else page.locator('[data-model-role-group="fallback"] .model-role-row').first - group.locator('[data-model-role-source]').select_option('openrouter') + group.locator('[data-model-role-source]').select_option(lane) return page, group.locator('[data-model-role-model]') if consumer == "Actor": row = page.locator('[data-subagent-row]').nth(1) + row.locator('[data-subagent-field="route"]').select_option(lane) return page, row.locator('[data-subagent-field="model"]') if consumer == 'Scope': return page, page.locator('[data-slot-id="scope_1"] [data-slot-custom-api]') if consumer == 'Deep': return page, page.locator('[data-deep-review-api-model]') row = page.locator('[data-slot-id="triad_1"]') if consumer == "Triad" else page.locator('[data-advisory-row]') - row.locator('[data-slot-route], [data-advisory-route]').select_option('api') + row.locator('[data-slot-route], [data-advisory-route]').select_option(lane) return page, row.locator('[data-slot-custom-api], [data-advisory-api-model]') diff --git a/tests/test_model_wait_browser.py b/tests/test_model_wait_browser.py index 9305b470a..3dfafae44 100644 --- a/tests/test_model_wait_browser.py +++ b/tests/test_model_wait_browser.py @@ -13,12 +13,15 @@ capture = setup_browser.capture pytestmark = [pytest.mark.ui_browser, pytest.mark.serial] TASK = "analysis-task" +# The wait picker offers an API lane per provider whose credential is stored. +API_LANE = "api:openai" @pytest.fixture def waiting_ui(subscription_ui): ui = subscription_ui page = ui["page"] + ui["settings"]["OPENAI_API_KEY"] = "***set***" rows = { key: {"wait_id": key, "revision": 1, "task_attempt": 1, "role": role, "model": "claudexor::codex=gpt-test", "source": "codex", @@ -162,7 +165,7 @@ def test_multiple_waits_toggle_and_exact_role_switch_wait_for_application(waitin page.wait_for_selector('[data-wait-id="light-wait"] [data-wait-change]:not([disabled])') assert not light.locator('[data-wait-auto]').is_checked() light.locator('[data-wait-change]').click() - light.locator('[data-model-role-source]').select_option('openai') + light.locator('[data-model-role-source]').select_option(API_LANE) light.locator('[data-model-role-model]').fill('owner-model') assert not light.locator('[data-wait-persist]').is_checked() light.locator('[data-wait-apply]').click() @@ -320,7 +323,7 @@ def test_wait_reconnect_preserves_unsubmitted_form_without_context_or_animation( field = row.locator('[data-model-role-model]') field.wait_for() assert row.locator('.model-role-details').count() == 0 - row.locator('[data-model-role-source]').select_option('openai') + row.locator('[data-model-role-source]').select_option(API_LANE) assert row.locator('.model-role-details').count() == 0 field.fill('unfinished-owner-model') row.locator('[data-wait-persist]').check() diff --git a/tests/test_model_wait_recovery_browser.py b/tests/test_model_wait_recovery_browser.py index 7489e46fd..847885ca5 100644 --- a/tests/test_model_wait_recovery_browser.py +++ b/tests/test_model_wait_recovery_browser.py @@ -9,7 +9,7 @@ from ouroboros import model_wait, owner_mailbox from ouroboros.task_results import load_task_result from tests.test_llm_claudexor import setup as subscription_transport from tests.test_model_wait import live_wait as wait_fixture -from tests.test_model_wait_browser import TASK, waiting_ui as waiting_fixture +from tests.test_model_wait_browser import API_LANE, TASK, waiting_ui as waiting_fixture from tests.test_subscription_setup_browser import subscription_ui as ui_fixture, capture setup = subscription_transport @@ -97,7 +97,7 @@ def test_fallback_local_change_discloses_task_only_and_keeps_other_roles(waiting waiter = page.locator('[data-wait-id="light-wait"]') waiter.locator('[data-wait-role]').filter(has_text="Fallback 2").wait_for() waiter.locator('[data-wait-change]').click() - waiter.locator('[data-model-role-source]').select_option("openai") + waiter.locator('[data-model-role-source]').select_option(API_LANE) waiter.locator('[data-model-role-model]').fill("replacement") local = waiter.locator('[data-model-local]') persist = waiter.locator('[data-wait-persist]') diff --git a/tests/test_processing_preferences_browser.py b/tests/test_processing_preferences_browser.py index 24ce0731e..d6de10621 100644 --- a/tests/test_processing_preferences_browser.py +++ b/tests/test_processing_preferences_browser.py @@ -80,7 +80,7 @@ def test_onboarding_processing_draft_returns_and_reaches_finish(subscription_ui) main.locator('summary').click() main.locator('[data-model-role-processing]').select_option('standard') page.locator('#next-btn').click() - page.wait_for_selector('#reviewer-slots-section') + page.wait_for_selector('[data-collapse="reviewers"]') page.locator('#back-btn').click() page.wait_for_selector('[data-global-processing]') assert page.locator('[data-global-processing]').input_value() == 'fast' diff --git a/tests/test_reference_book_migration.py b/tests/test_reference_book_migration.py index 87d59050b..8cac87bee 100644 --- a/tests/test_reference_book_migration.py +++ b/tests/test_reference_book_migration.py @@ -155,7 +155,11 @@ def test_entrypoint_keeps_its_h1_line_and_carries_only_the_membership(book_id): rel = BOOK_ENTRYPOINTS[book_id] raw = (REPO / rel).read_bytes() lines = raw.decode("utf-8").split("\n") - assert lines[0] == OLD_MONOLITHS[book_id]["h1"] + # The architecture h1 is a release version carrier (`release_sync` rewrites + # its version token on every bump): the whole line stays pinned, and only + # that token follows VERSION instead of the split-time 7.0.0. + version = (REPO / "VERSION").read_text(encoding="utf-8").strip() + assert lines[0] == OLD_MONOLITHS[book_id]["h1"].replace("v7.0.0", f"v{version}") book = load_reference_book(REPO, book_id) # One `## Chapters` heading and nothing else: the entrypoint orients, the # chapters carry the book. diff --git a/tests/test_subscription_role_routes_browser.py b/tests/test_subscription_role_routes_browser.py index f47467d40..d314c4713 100644 --- a/tests/test_subscription_role_routes_browser.py +++ b/tests/test_subscription_role_routes_browser.py @@ -9,6 +9,15 @@ from tests import test_subscription_setup_browser as setup_browser pytestmark = [pytest.mark.ui_browser, pytest.mark.serial] subscription_ui = setup_browser.subscription_ui capture = setup_browser.capture +# Every route picker offers an API lane per provider whose credential is +# stored, so a fixture that selects one advertises that provider's key first. +API_KEYS = {"openrouter": "OPENROUTER_API_KEY", "openai": "OPENAI_API_KEY"} + + +def api_lane(ui, provider="openrouter"): + """Store `provider`'s key in the served settings and return its route choice.""" + ui["settings"][API_KEYS[provider]] = "***set***" + return f"api:{provider}" @pytest.fixture @@ -49,10 +58,7 @@ def configure_mixed(ui): "advisory": {"enabled": True, "route": {"kind": "api_chat", "target_id": target, "profile_id": "personal"}}, "deep_review": {"subagent_id": "native"}, } - # An actor's saved route is not provider readiness: offer OpenAI through - # the same configured-credential fact the real source picker reads. - ui["settings"].update(OPENAI_API_KEY="fixture-openai-credential", - OUROBOROS_SUBAGENTS=actors, OUROBOROS_REVIEWER_SLOTS=json.dumps(slots)) + ui["settings"].update(OUROBOROS_SUBAGENTS=actors, OUROBOROS_REVIEWER_SLOTS=json.dumps(slots)) ui["fixture"]["preview"]["reviewer_slots"] = slots ui["fixture"]["catalog"]["model_sources"] = [ {"id": "opaque-source", "label": "Codex", "credentialHarness": "codex"}, @@ -75,6 +81,7 @@ def open_agents(ui): def test_reviewer_source_roundtrip_restores_its_own_model_and_account(role_ui): configure_mixed(role_ui) + lane = api_lane(role_ui, 'openai') page = open_agents(role_ui) for selector, model_field, account_field in [ ('[data-slot-id="triad_1"]', '[data-slot-custom-api]', '[data-slot-profile]'), @@ -82,13 +89,15 @@ def test_reviewer_source_roundtrip_restores_its_own_model_and_account(role_ui): ]: row = page.locator(selector) route = row.locator('[data-slot-route], [data-advisory-route]') - route.select_option('api:openai') + route.select_option(lane) + # The chooser holds the model alone; the source select names the provider. row.locator(model_field).fill('other-choice') route.select_option('subscription:opaque-source') assert row.locator(model_field).input_value() == 'gpt-test' assert row.locator(account_field).input_value() == 'personal' - route.select_option('api:openai') + route.select_option(lane) assert row.locator(model_field).input_value() == 'other-choice' + assert route.input_value() == lane route.select_option('subscription:opaque-source') page.locator('[data-advisory-row]').scroll_into_view_if_needed() capture(page, "reviewer-source-roundtrip-restored") @@ -96,7 +105,7 @@ def test_reviewer_source_roundtrip_restores_its_own_model_and_account(role_ui): with page.expect_response('**/api/reviewer-slots'): page.locator('#btn-reload-settings').click() page.get_by_role('button', name='Discard and continue', exact=True).click() - page.locator('[data-advisory-route]').select_option('api:openai') + page.locator('[data-advisory-route]').select_option(lane) assert page.locator('[data-advisory-api-model]').input_value() == '' @@ -167,9 +176,10 @@ def test_subscription_accounts_roundtrip_existing_editors(role_ui, width): def test_source_switch_and_catalog_refresh_keep_focus_and_draft(role_ui): ui = role_ui configure_mixed(ui) + lane = api_lane(ui, 'openai') page = open_agents(ui) triad = page.locator('[data-slot-id="triad_1"]') - triad.locator('[data-slot-route]').select_option("api:openai") + triad.locator('[data-slot-route]').select_option(lane) assert triad.locator('[data-slot-profile]').count() == 0 triad.locator('[data-slot-custom-api]').fill("openai::gpt-api") triad.locator('[data-slot-route]').select_option("subscription:opaque-source") diff --git a/tests/test_subscription_wait_integration_browser.py b/tests/test_subscription_wait_integration_browser.py index 81e0adf4a..414821fa5 100644 --- a/tests/test_subscription_wait_integration_browser.py +++ b/tests/test_subscription_wait_integration_browser.py @@ -63,6 +63,8 @@ def integrated_wait(subscription_ui, live_wait, monkeypatch): config.SETTINGS_PATH.write_text(json.dumps(settings), encoding="utf-8") original_settings = config.SETTINGS_PATH.read_bytes() ui["settings"].update(settings) + # The wait picker offers an API lane per provider whose credential is stored. + ui["settings"]["OPENAI_API_KEY"] = "***set***" source_entered, source_release = threading.Event(), threading.Event() catalog_entered, catalog_release = threading.Event(), threading.Event() @@ -265,7 +267,7 @@ def test_browser_wait_controls_apply_once_and_only_to_light(integrated_wait, per assert not row.locator("[data-wait-auto]").is_checked() row.locator("[data-wait-change]").click() - row.locator("[data-model-role-source]").select_option("openai") + row.locator("[data-model-role-source]").select_option("api:openai") row.locator("[data-model-role-model]").fill("owner-model") assert not row.locator("[data-wait-persist]").is_checked() if persist_role: diff --git a/tests/test_ui_smoke_large_artifacts.py b/tests/test_ui_smoke_large_artifacts.py index 2fe6f998a..3ca7b5249 100644 --- a/tests/test_ui_smoke_large_artifacts.py +++ b/tests/test_ui_smoke_large_artifacts.py @@ -99,12 +99,29 @@ def test_large_attachment_returns_through_real_document_handler_and_download( "name": "send_file", "arguments": json.dumps({"file_path": str(staged[0]), "caption": "Complete large dataset"}), }, }]} + finish = "tool_calls" if message.get("tool_calls") else "stop" payload = {"id": "mock-large-file", "object": "chat.completion", - "choices": [{"message": message, "finish_reason": "tool_calls" if message.get("tool_calls") else "stop"}], + "model": request.get("model") or "mock-model", + "choices": [{"index": 0, "message": message, "finish_reason": finish}], "usage": {"prompt_tokens": 1, "completion_tokens": 1, "total_tokens": 2}} - data = json.dumps(payload).encode() + content_type = "application/json" + if request.get("stream"): + # The main loop streams every completion: answer in SSE frames with + # the terminal framing the assembler requires. + content_type = "text/event-stream" + delta = dict(message) + if delta.get("tool_calls"): + delta["tool_calls"] = [dict(call, index=index) + for index, call in enumerate(delta["tool_calls"])] + common = {"id": payload["id"], "model": payload["model"], "object": "chat.completion.chunk"} + frames = [{**common, "choices": [{"index": 0, "delta": delta, "finish_reason": finish}]}, + {**common, "choices": [], "usage": payload["usage"]}] + data = ("".join("data: " + json.dumps(frame) + "\n\n" for frame in frames) + + "data: [DONE]\n\n").encode() + else: + data = json.dumps(payload).encode() handler.send_response(200) - handler.send_header("Content-Type", "application/json") + handler.send_header("Content-Type", content_type) handler.send_header("Content-Length", str(len(data))) handler.end_headers() handler.wfile.write(data) diff --git a/tests/test_ui_smoke_playwright.py b/tests/test_ui_smoke_playwright.py index f74307345..b277a82ba 100644 --- a/tests/test_ui_smoke_playwright.py +++ b/tests/test_ui_smoke_playwright.py @@ -1050,11 +1050,13 @@ def test_ui_smoke_collapsed_activity_line_named_vs_unnamed( assert 0.9 <= bands["running-act"]["activity"]["lines"] <= 1.2, bands assert bands["unnamed-act"]["activity"]["display"] == "none", bands # Useful root activity sizes naturally up to three lines; empty activity - # reserves no band on either running or finished cards. + # reserves no band on either running or finished cards. An uncoined turn + # promotes its only note to the title, so the collapsed line stays empty + # while the block still stands on that row of work. _emit_ws_frame(page, { "type": "chat", "role": "assistant", "is_progress": True, - "chat_id": 1, "task_id": "done-empty", "suggested_name": "Quick task", - "content": "", "ts": "2026-07-29T10:00:03+00:00", + "chat_id": 1, "task_id": "done-empty", + "content": "Quick task", "ts": "2026-07-29T10:00:03+00:00", }) done_empty = page.locator('.chat-live-card[data-task-id="done-empty"]') done_empty.wait_for(state="attached", timeout=30_000) @@ -1313,8 +1315,11 @@ def test_ui_smoke_chat_chronology_reconnect_and_plain_answer_marker(direct_serve timeout=30_000, ) state = page.evaluate( + # The paged-history control heads the list and carries no + # timestamp of its own: it is chrome, not a transcript row. """() => [...document.querySelector('#chat-messages').children] .filter((node) => !node.classList.contains('typing-bubble') + && !node.classList.contains('chat-load-older') && !node.textContent.includes('Reconnected')) .map((node) => ({ text: node.textContent, @@ -1323,6 +1328,7 @@ def test_ui_smoke_chat_chronology_reconnect_and_plain_answer_marker(direct_serve taskId: node.dataset.taskId || '', }))""" ) + assert page.locator("#chat-messages > .chat-load-older").count() == 1 assert [item["card"] for item in state] == [ False, True, True, False, True, False, False, ] @@ -1725,6 +1731,8 @@ def test_ui_smoke_direct_mode_nests_subagent_child_cards(direct_server_with_data assert "compared output" in expanded_text assert "done" in expanded_text.lower() assert "Scheduled subagent child1" not in expanded_text + assert "Subagent child1 running" not in expanded_text + assert child.locator('[data-live-line-key="terminal-subagent-lifecycle-child1"]').count() == 1 assert child_summary.get_attribute("aria-expanded") == "true" assert child.locator("[data-live-timeline]").first.get_attribute("id") assert result_toggle.get_attribute("aria-controls") @@ -1765,9 +1773,15 @@ def test_ui_smoke_direct_mode_nests_subagent_child_cards(direct_server_with_data replay_progress.locator(".chat-live-line-toggle").click() assert child_activity_early in replay_progress.inner_text() assert child_activity_tail in replay_progress.inner_text() + assert "Scheduled subagent child1" not in replay_child.inner_text() + assert "Subagent child1 running" not in replay_child.inner_text() + assert replay_child.locator('[data-live-line-key="terminal-subagent-lifecycle-child1"]').count() == 1 page.wait_for_timeout(900) # cover the routine background history sync assert replay_child.locator('.chat-live-line-repeat:not([hidden])').count() == 0 page.screenshot(path=str(data_dir.parent / "review-truth-child-reconnect.png"), full_page=True) + replay_progress.locator(".chat-live-line-toggle").click() + replay_child.scroll_into_view_if_needed() + page.screenshot(path=str(data_dir.parent / "lifecycle-current-status.png"), full_page=True) assert page.locator(".chat-bubble.progress").count() == 0 assert page.locator(".chat-bubble", has_text="Final child answer should stay inside the child card.").count() == 0 @@ -3036,11 +3050,8 @@ def test_ui_owner_context_mode_and_scope_slot_save(direct_server_with_data): assert after["context_mode_auto_low"] is False # 2. A scope slot saves with no window question anywhere. - # 6.2: the scope route is a review-lane row. D-10 moved the lanes - # out of Models into their own Agents tab. The seeded row is - # already an API row, so retyping its model id is the whole edit — - # the grouped combobox offers `api:` values, never a - # bare `api`, so nothing is selected here. + # Select the fixture's configured provider, then edit its model; + # the grouped combobox uses provider-specific API choices. page.click('[data-nav-page="settings"]') page.wait_for_selector("#s-context-mode", state="attached", timeout=30_000) page.locator('[data-settings-tab="agents"]').click() @@ -3049,7 +3060,8 @@ def test_ui_owner_context_mode_and_scope_slot_save(direct_server_with_data): '#reviewer-scope-rows .reviewer-slot-row [data-slot-route]' ).first scope_route.wait_for(state="visible", timeout=30_000) - assert str(scope_route.input_value()).startswith("api"), scope_route.input_value() + scope_route.select_option("api:openai-compatible") + assert scope_route.input_value() == "api:openai-compatible" custom_input = page.locator( '#reviewer-scope-rows .reviewer-slot-row [data-slot-custom-api]' ).first diff --git a/tests/test_ui_smoke_settings_drafts.py b/tests/test_ui_smoke_settings_drafts.py index 60f228e4a..5b1257e2f 100644 --- a/tests/test_ui_smoke_settings_drafts.py +++ b/tests/test_ui_smoke_settings_drafts.py @@ -211,6 +211,8 @@ def test_initial_settings_document_survives_early_edit_while_enrichment_waits(su ui, pending = subscription_ui, {"reviewers": [], "status": [], "catalog": []} page = ui["page"] ui["settings"]["OUROBOROS_MODEL"] = "claudexor::opaque-source=gpt-test" + # An API lane is offered per provider whose credential is stored. + ui["settings"]["OPENROUTER_API_KEY"] = "***set***" ui["fixture"]["catalog"]["model_sources"] = [ {"id": "opaque-source", "label": "Managed models", "credentialHarness": "codex"}, ] @@ -225,7 +227,7 @@ def test_initial_settings_document_survives_early_edit_while_enrichment_waits(su assert pending["status"], "the status read must still be pending" assert page.locator('#btn-save-settings').is_enabled(), "the known document can be saved before enrichment" source = main.locator('[data-model-role-source]') - source.select_option("openrouter") + source.select_option("api:openrouter") model = main.locator('[data-model-role-model]') model.fill("owner-kept-model") model.evaluate("element => { window.__earlyModel = element; element.setSelectionRange(4, 4); }") @@ -245,7 +247,7 @@ def test_initial_settings_document_survives_early_edit_while_enrichment_waits(su page.wait_for_function("""() => document.querySelector('[data-model-role="main"] [data-model-role-source]') .querySelector('option[value="subscription:opaque-source"]')""") expect(model).to_have_value("owner-kept-model") - assert source.input_value() == "openrouter" + assert source.input_value() == "api:openrouter" assert model.evaluate("element => element === window.__earlyModel && element.selectionStart === 4") expect(page.locator('#settings-unsaved-indicator')).to_have_class(re.compile('is-visible')) expect(page.locator('#btn-save-settings')).to_be_enabled() diff --git a/tests/ui_chat_viewport_smoke.py b/tests/ui_chat_viewport_smoke.py index 84ec2fddc..9dc8b0199 100644 --- a/tests/ui_chat_viewport_smoke.py +++ b/tests/ui_chat_viewport_smoke.py @@ -125,6 +125,14 @@ def run_chat_viewport_smoke( def hold_first_route(routes): return lambda route: routes.append(route) if not routes else route.fallback() + def incomplete_activity_census(route): + # These WS-only tasks do not exist on the fixture server. Its empty + # roster cannot disprove them; reconciliation cases opt in below. + response = route.fetch() + payload = response.json() + payload["active_chat_activities_complete"] = False + route.fulfill(response=response, json=payload) + def visible_card_anchor(page): return page.evaluate( """() => { @@ -152,6 +160,7 @@ def run_chat_viewport_smoke( page = browser.new_page(viewport={"width": 1280, "height": 760}) try: page.add_init_script(f"({_CAPTURE_TEST_SOCKET})()") + page.route("**/api/state", incomplete_activity_census) page.goto(url, wait_until="domcontentloaded", timeout=30_000) page.add_style_tag( content="#chat-messages, #chat-messages * { overflow-anchor: none !important; }" @@ -552,6 +561,24 @@ def run_chat_viewport_smoke( "ts": "2026-08-03T10:04:00+00:00", } _emit_ws_frame(page, late_child_frame) + # Observe this mount's real height change before testing the + # viewport; elapsed animation frames alone do not establish it. + # A missing or zero-height child still fails. + try: + page.wait_for_function("""minimum => { + const parent = document.querySelector('.chat-live-card[data-task-id="vp-parent"]'); + const child = parent?.querySelector(':scope > .chat-subagents > [data-task-id="vp-late-child"]'); + return child && child.getBoundingClientRect().height > 0 + && parent.getBoundingClientRect().height > minimum; + }""", arg=parent_before_mount + 30, timeout=10_000) + except PlaywrightError as exc: + geometry = parent.evaluate("""(card, before) => { + const child = card.querySelector('[data-task-id="vp-late-child"]'); + return {before, after: card.getBoundingClientRect().height, + expanded: card.dataset.expanded, childHeight: child?.getBoundingClientRect().height, + childParent: child?.parentElement?.dataset.subagentsFor}; + }""", parent_before_mount) + raise AssertionError(f"Late child did not grow its parent: {geometry}") from exc assert parent.evaluate("card => card.getBoundingClientRect().height") > parent_before_mount + 30 assert abs(card_top(page, anchor_id) - anchor_before) <= 6 parent.evaluate("""card => { window.__subagentNoopMutations = []; window.__subagentNoopObserver = new MutationObserver(records => window.__subagentNoopMutations.push(...records)); window.__subagentNoopObserver.observe(card, {attributes: true, attributeOldValue: true, childList: true, characterData: true, subtree: true}); }""") @@ -717,6 +744,7 @@ def run_chat_viewport_smoke( page.evaluate(_SETTLE_TWO_FRAMES) assert_noop_read(page, noop_top) page.unroute("**/api/state") + page.route("**/api/state", incomplete_activity_census) # A production-shaped review reference hydrates asynchronously; # both the fetch result and its review DOM reconcile stay anchored. diff --git a/uv.lock b/uv.lock index 78fca4d27..890e2ed5b 100644 --- a/uv.lock +++ b/uv.lock @@ -1535,7 +1535,7 @@ wheels = [ [[package]] name = "ouroboros" -version = "7.0.0" +version = "7.1.0" source = { editable = "." } dependencies = [ { name = "croniter" }, diff --git a/web/modules/api_types.js b/web/modules/api_types.js index 59bb13cf2..2560f2b61 100644 --- a/web/modules/api_types.js +++ b/web/modules/api_types.js @@ -1459,7 +1459,7 @@ export const MAX_QUIZ_OPTIONS = 6; // REFUSES a longer comment (it is delivered verbatim, never truncated), so // the card must not offer to send one. export const MAX_DECISION_COMMENT = 2000; -export const GATEWAY_CONTRACT_VERSION = '7.0.0'; +export const GATEWAY_CONTRACT_VERSION = '7.1.0'; /** * @typedef {Object} ChatHistoryPosition diff --git a/web/modules/chat_history_replay.js b/web/modules/chat_history_replay.js index b70e0dc5e..c0d596a46 100644 --- a/web/modules/chat_history_replay.js +++ b/web/modules/chat_history_replay.js @@ -22,15 +22,16 @@ export function mergeHistoricalTimelineItem(record, summary, row, ts) { if (summary.visible === false || !(summary.headline || summary.body)) return false; const identity = String(row?.history_id || ''); if (!identity) return false; - // A terminal projection is still one logical completion; narration consists - // of independently addressed source records, even when text and time match. - const key = summary.terminal ? summary.dedupeKey : `history:${identity}`; + // A child's lifecycle is one evolving status, just as it is live. Its + // authored progress keeps every source record, even when text and time match. + const evolving = summary.terminal || String(summary.dedupeKey || '').startsWith('subagent-lifecycle:'); + const key = evolving ? summary.dedupeKey : `history:${identity}`; let item = record.items.find((entry) => entry.dedupeKey === key); - if (!item && !summary.terminal) { + if (!item && !evolving) { item = record.items.find((entry) => !entry.historyId && entry.dedupeKey === summary.dedupeKey && entry.sourceTs === row.ts); } - if (item && summary.terminal) { + if (item && evolving) { const incomingTime = Date.parse(row.ts), existingTime = Date.parse(item.sourceTs || ''); if (incomingTime < existingTime || incomingTime === existingTime && compareHistoryPosition(row.history_position, item.historyPosition) < 0) return false; @@ -57,9 +58,9 @@ export function mergeHistoricalTimelineItem(record, summary, row, ts) { body: summary.body || '', fullBody: summary.fullBody || summary.body || '', fullRef: summary.fullRef || '', truncated: summary.truncated || false, ts: ts || '', sourceTs: row.ts || '', count: 1, dedupeKey: key, - ...(summary.terminal ? { sourceHistoryId: identity } : { historyId: identity }), + ...(evolving ? { sourceHistoryId: identity } : { historyId: identity }), historyPosition: row.history_position, - lineKey: summary.terminal ? `terminal-${String(key).replace(/[^A-Za-z0-9_-]/g, '-')}` + lineKey: evolving ? `terminal-${String(key).replace(/[^A-Za-z0-9_-]/g, '-')}` : `history-${identity.replace(/[^A-Za-z0-9_-]/g, '-')}`, }); } diff --git a/web/modules/chat_render_batch.js b/web/modules/chat_render_batch.js index d419b9b4b..d53b912dc 100644 --- a/web/modules/chat_render_batch.js +++ b/web/modules/chat_render_batch.js @@ -156,6 +156,32 @@ export function createHistoryResyncScheduler({ * Equal timestamps preserve arrival order; timestamp-free nodes append. * (Moved verbatim from chat.js — that module sits at its byte ceiling.) */ +// A focused text control keeps its own caret; Chromium mirrors it into the +// document Selection, so clearing and rebuilding document ranges around a +// timeline move collapses that caret (a typed wait-picker draft lost its +// selection on every reconnect). While such a control is focused, the document +// ranges are that mirror: leave them alone and restore the control's caret. +function textControlCaret(active) { + const tag = String(active?.tagName || active?.nodeName || '').toUpperCase(); + if (tag !== 'INPUT' && tag !== 'TEXTAREA') return null; + try { + const { selectionStart, selectionEnd, selectionDirection } = active; + if (selectionStart == null || selectionEnd == null) return null; + return { selectionStart, selectionEnd, selectionDirection: selectionDirection || 'none' }; + } catch { + return null; // input types without a caret (number, email, ...) throw on read + } +} + +function restoreTextControlCaret(active, caret) { + if (!caret || typeof active?.setSelectionRange !== 'function') return; + try { + active.setSelectionRange(caret.selectionStart, caret.selectionEnd, caret.selectionDirection); + } catch { + // The control changed type or lost its value between capture and restore. + } +} + export function insertTimelineNode(messages, node, typing = null) { const rawNodeTs = node?.dataset?.ts; const nodeTs = rawNodeTs == null || rawNodeTs === '' ? NaN : Number(rawNodeTs); @@ -176,7 +202,8 @@ export function insertTimelineNode(messages, node, typing = null) { if (node.parentNode === messages && node.nextElementSibling === target) return { before }; const doc = messages.ownerDocument; const active = doc?.activeElement; - const selection = doc?.getSelection?.(); + const caret = textControlCaret(active); + const selection = caret ? null : doc?.getSelection?.(); const ranges = Array.from({ length: selection?.rangeCount || 0 }, (_, index) => { const range = selection.getRangeAt(index); return [range.startContainer, range.startOffset, range.endContainer, range.endOffset]; @@ -184,6 +211,7 @@ export function insertTimelineNode(messages, node, typing = null) { if (target) messages.insertBefore(node, target); else messages.appendChild(node); if (active && node.contains?.(active) && doc.activeElement !== active) active.focus({ preventScroll: true }); + if (caret) restoreTextControlCaret(active, caret); if (ranges.length && ranges.every(([start, , end]) => start.isConnected && end.isConnected)) { selection.removeAllRanges(); for (const [start, startOffset, end, endOffset] of ranges) { @@ -293,7 +321,8 @@ export function createLiveCardTimelineRenderer({ withStableViewport, buildTimeli const prevTop = el.scrollTop; const byKey = new Map(Array.from(el.children).map((node) => [node.dataset.liveLineKey, node])); const active = el.ownerDocument?.activeElement; - const selection = el.ownerDocument?.getSelection?.(); + const caret = textControlCaret(active); + const selection = caret ? null : el.ownerDocument?.getSelection?.(); const ranges = Array.from({ length: selection?.rangeCount || 0 }, (_, index) => { const range = selection.getRangeAt(index); return [range.startContainer, range.startOffset, range.endContainer, range.endOffset]; @@ -321,6 +350,7 @@ export function createLiveCardTimelineRenderer({ withStableViewport, buildTimeli } if (moved) { if (el.contains(active) && el.ownerDocument.activeElement !== active) active.focus({ preventScroll: true }); + if (caret) restoreTextControlCaret(active, caret); const intact = ranges.filter(([start, , end]) => start.isConnected && end.isConnected); if (intact.length) { selection.removeAllRanges(); @@ -567,6 +597,9 @@ export function updateLiveTimelineItem(record, summary, { ts, rawTs, syntheticKe // content again once it reports an error. receipt: Boolean(summary.receipt), ts: ts || it.ts, + // Replay compares the child's current status with older pages. + // Its source time advances with the live status, not with narration. + ...(syntheticKey.startsWith('subagent-lifecycle:') && rawTs ? { sourceTs: rawTs } : {}), }; if (Object.entries(patch).some(([key, value]) => it[key] !== value)) { Object.assign(it, patch); diff --git a/web/package-lock.json b/web/package-lock.json index d00c86812..230eed381 100644 --- a/web/package-lock.json +++ b/web/package-lock.json @@ -1,12 +1,12 @@ { "name": "ouroboros-web", - "version": "7.0.0", + "version": "7.1.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "ouroboros-web", - "version": "7.0.0", + "version": "7.1.0", "devDependencies": { "eslint": "10.10.0", "globals": "17.12.0" diff --git a/web/package.json b/web/package.json index 0d304443b..84357a358 100644 --- a/web/package.json +++ b/web/package.json @@ -1,6 +1,6 @@ { "name": "ouroboros-web", - "version": "7.0.0", + "version": "7.1.0", "private": true, "type": "module", "description": "Ouroboros browser UI package boundary", diff --git a/web/style.css b/web/style.css index fe898f305..b94bf4f5b 100644 --- a/web/style.css +++ b/web/style.css @@ -7199,7 +7199,8 @@ textarea.chat-input { text-overflow: ellipsis; } .project-work-coverage { font-size: var(--type-meta); color: var(--text-meta); } -.chat-panel-statusbar { flex-wrap: wrap; gap: var(--space-2); } +/* Compact horizontal spacing leaves room for wider native system-font metrics. */ +.chat-panel-statusbar { flex-wrap: wrap; gap: var(--space-2) var(--space-1); } .chat-live-executor-chip { max-width: 100%; white-space: normal; } .chat-live-meta-text, .chat-live-executor-chip .harness-identity-label { overflow-wrap: anywhere; } .chat-bubble .message:not(:has(> *)), diff --git a/web/tests/chat_chronology.test.js b/web/tests/chat_chronology.test.js index d875335c3..0b0fc8047 100644 --- a/web/tests/chat_chronology.test.js +++ b/web/tests/chat_chronology.test.js @@ -131,3 +131,30 @@ test('URL-less media history rows cannot finalize a live task card', () => { assert.equal(isNonTerminalMediaHistoryRow({ system_type: 'video', task_id: 'running' }), true); assert.equal(isNonTerminalMediaHistoryRow({ system_type: 'task_summary', task_id: 'done' }), false); }); + +test('a timeline move keeps the caret of a focused text control instead of clearing document ranges', () => { + // Chromium mirrors a focused input's caret into the document Selection: the + // range save/restore around a node move collapsed a typed draft's selection. + const calls = []; + const input = { + tagName: 'INPUT', selectionStart: 2, selectionEnd: 8, selectionDirection: 'forward', + setSelectionRange(start, end, direction) { calls.push(['set', start, end, direction]); this.selectionStart = start; this.selectionEnd = end; }, + focus() { calls.push(['focus']); }, + }; + const selection = { + rangeCount: 1, + getRangeAt: () => ({ startContainer: {}, startOffset: 0, endContainer: {}, endOffset: 0 }), + removeAllRanges() { calls.push(['removeAllRanges']); input.selectionStart = input.selectionEnd = 0; }, + addRange() { calls.push(['addRange']); }, + }; + const typing = makeNode('typing'); + const timeline = makeTimeline([makeNode('t3', 3), typing]); + timeline.ownerDocument = { activeElement: input, getSelection: () => selection, createRange: () => ({ setStart() {}, setEnd() {} }) }; + const moving = makeNode('t1', 1); + moving.contains = () => false; + insertTimelineNode(timeline, moving, typing); + assert.deepEqual(ids(timeline), ['t1', 't3', 'typing']); + assert.ok(!calls.some(([name]) => name === 'removeAllRanges'), calls); + assert.deepEqual([input.selectionStart, input.selectionEnd], [2, 8]); + assert.deepEqual(calls, [['set', 2, 8, 'forward']]); +}); diff --git a/web/tests/chat_history_replay.test.js b/web/tests/chat_history_replay.test.js index ae1407c58..c21bfbc4e 100644 --- a/web/tests/chat_history_replay.test.js +++ b/web/tests/chat_history_replay.test.js @@ -2,6 +2,70 @@ import assert from 'node:assert/strict'; import test from 'node:test'; import { mergeHistoricalTimelineItem, compareHistoryPosition } from '../modules/chat_history_replay.js'; import { createChatHistoryPager } from '../modules/chat_history.js'; +import { updateLiveTimelineItem } from '../modules/chat_render_batch.js'; + +test('one child lifecycle survives chronological replay and older pages without losing narration', () => { + const narration = 'Searching evidence. '.repeat(60) + 'COMPLETE_NARRATION_END'; + const lifecycleKey = 'subagent-lifecycle:child'; + const frames = [ + { headline: 'Scheduled', phase: 'queued', dedupeKey: lifecycleKey }, + { headline: 'Searching evidence', body: narration, dedupeKey: 'subagent-progress:child' }, + { headline: 'Running', phase: 'working', dedupeKey: lifecycleKey }, + { headline: 'Searching evidence', body: narration, dedupeKey: 'subagent-progress:child' }, + { headline: 'Completed', phase: 'done', terminal: true, dedupeKey: lifecycleKey }, + ]; + const row = index => ({ history_id: `progress:${index}`, + history_position: { source: 'progress', offset: index }, ts: `2026-09-12T12:00:0${index}Z` }); + for (const order of [[0, 1, 2, 3, 4], [4, 3, 2, 1, 0], [2, 1, 0, 4, 3]]) { + const record = { items: [], finished: true }; + for (const index of order) mergeHistoricalTimelineItem(record, frames[index], row(index), String(index)); + const status = record.items.filter(item => item.dedupeKey === lifecycleKey); + assert.equal(status.length, 1); + assert.equal(status[0].headline, 'Completed'); + assert.equal(status[0].sourceHistoryId, 'progress:4'); + const voices = record.items.filter(item => item.historyId); + assert.deepEqual(voices.map(item => item.historyId), ['progress:1', 'progress:3']); + assert.deepEqual(voices.map(item => item.fullBody), [narration, narration]); + assert.equal(record.items.length, 3); + const before = JSON.stringify(record.items); + for (const index of order) assert.equal(mergeHistoricalTimelineItem(record, frames[index], row(index), String(index)), false); + assert.equal(JSON.stringify(record.items), before); + assert.equal(record.finished, true); + } +}); + +test('an older lifecycle page cannot regress a status already updated live', () => { + const record = { items: [] }; + const key = 'subagent-lifecycle:child'; + const frame = (headline, second) => updateLiveTimelineItem(record, + { headline, phase: 'working', dedupeKey: key }, + { headline, ts: `12:00:0${second}`, rawTs: `2026-09-12T12:00:0${second}Z`, + syntheticKey: key, inPlaceByKey: true }); + frame('Scheduled', 0); + const item = record.items[0]; + frame('Running', 3); + const old = { history_id: 'progress:1', history_position: { source: 'progress', offset: 1 }, + ts: '2026-09-12T12:00:01Z' }; + assert.equal(mergeHistoricalTimelineItem(record, + { headline: 'Scheduled', phase: 'queued', dedupeKey: key }, old, '12:00:01'), false); + assert.equal(record.items[0], item); + assert.equal(item.headline, 'Running'); + assert.equal(record.items.length, 1); +}); + +test('equal-time child lifecycle rows use source order, not page arrival order', () => { + const record = { items: [] }; + const summary = { headline: 'Running', phase: 'working', dedupeKey: 'subagent-lifecycle:child' }; + const row = offset => ({ history_id: `progress:${offset}`, ts: '2026-09-12T12:00:00Z', + history_position: { source: 'progress', offset } }); + mergeHistoricalTimelineItem(record, summary, row(2), '12:00'); + assert.equal(mergeHistoricalTimelineItem(record, { ...summary, headline: 'Scheduled' }, row(1), '12:00'), false); + assert.equal(record.items[0].headline, 'Running'); + mergeHistoricalTimelineItem(record, { ...summary, headline: 'Completed', terminal: true }, row(3), '12:00'); + assert.equal(record.items.length, 1); + assert.equal(record.items[0].headline, 'Completed'); + assert.equal(mergeHistoricalTimelineItem(record, summary, row(2), '12:00'), false); +}); test('equal-time identical narration retains physical identities in source order', () => { const record = { items: [], finished: true };