Reserve chat composer space with a flex item for native WebKit. Combine queue and live activity identities through the existing task controls. Give a worker with its own early progress one readiness extension bounded to 300 seconds from spawn. Wait for the actual chat socket before the large-attachment browser test sends its files. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
87 KiB
3. Web UI Pages & Buttons
This chapter owns the browser-facing surface: the pages of the build-free single-page application, the shared frontend contracts they must not each reinvent, and the truth rules that keep Chat, Files, Skills, Widgets, Dashboard and Settings showing the facts the runtime actually recorded. It exists because the same application is served to the desktop shell, ordinary browsers, Docker deployments, the Telegram mini app, phones and the Linux browser fallback, so a presentation decision made for one surface has to be stated for all of them.
The Web UI is a build-free vanilla-JavaScript SPA (web/index.html, shared CSS, web/modules/*) — no TypeScript or bundler step, so the running interface stays inspectable and editable by Ouroboros without regenerating opaque artifacts. web/app.js owns top-level page, Project-panel and mobile-navigation state; feature modules own their presentation; every long-lived UI acquisition carries a disposer bound to the lifecycle that created it (DEVELOPMENT "UI resources carry a disposer"). The same SPA serves every surface: the Linux browser fallback changes presentation only — no second API, onboarding contract, runtime identity or state owner — and the browser stays the owner's application outside process custody (§1). One shared WebSocket serves the whole application; Projects open no independent sockets, and REST remains the recovery and durable-read path (protocol and connection order: the WebSocket protocol subsection of §4).
The desktop shell exposes a small MainApi JS bridge (window.pywebview.api, launcher.py): three native confirmation methods (runtime mode, reviewed-skill auto-grant, skill key grant), download_file_to_downloads, open_file_with_default_app, open_external_url (absolute http(s)/mailto) and save_bytes_to_downloads for live base64 payloads; the loopback-file methods share one guard — loopback host, exact server port, and a path allowlist of /api/files/download, /api/extensions/... and /api/tasks/.... Because the embedded WebView has no new-window or download delegate, ui_helpers.js installs a shell-only link interceptor in BOTH top-level documents (SPA and framed onboarding wizard) when the bridge is present, routing each URL class to the matching bridge method; methods are feature-detected per call, because the packaged launcher updates only on reinstall while the served frontend updates with the managed repo, and a missing method degrades to copy-link-plus-toast or the file-helper fallback chain. The authenticated Telegram proxy's X-Ouroboros-Telegram-MiniApp: 1 presentation marker makes server_web.make_index_page include the asynchronous Telegram SDK and host hint in the main document; an unavailable SDK reports a click failure and never replays it after loading, and the marker grants no authentication authority.
Navigation and shared UI contracts
Primary navigation exposes Chat (Main), a collapsible Projects group, Files, Skills, Widgets, Dashboard and Settings; About is a Settings sub-tab. syncNavigationState() is the single presentation state machine for active page, active Project, Projects expansion, mobile drawer and panel backdrop, so independent toggles cannot leave multiple rows active or a hidden surface looking selected. Sidebar and Project panel widths are owner-local UI preferences, not runtime settings. The application has exactly one client-side route: a #<page> fragment for the injected page ids, honoured once on load and validated against the existing #page-<name> section (an unknown fragment is ignored rather than painting a blank surface), and never written back on navigation — the desktop shell and the Linux browser fallback have no address bar, and the Telegram mini app always loads at / — so a browser gains a shareable /#widgets link while no other surface changes. Projects are panels, not routes, and the sidebar is the document's one <nav> landmark.
Each active or deleting Project has a sidebar row with pointer- and keyboard-operable open/rename/delete; the backend owns the 80-character name limit and lifecycle truth. Unread Projects sort ahead of read, then by durable activity; a deleting Project becomes non-openable and stays visibly transitional until the server publishes authoritative registry state. On narrow screens navigation is an explicit drawer and the Project chat a full-width overlay; no gesture-only navigation layer competes with message scroll, text selection or the software keyboard.
Shared frontend primitives keep pages from acquiring competing contracts — frontend work must not reimplement supervisor, review, marketplace, extension and provider semantics per page:
web/ui.css← shared values and the explicit field/button/status/popup classes;index.htmlandonboarding_template.htmlload it before their page sheets,style.csskeeps shell composition and tab chrome, and page-specific layout does not duplicate the palette (DESIGN §8 names the migrated families/regions).page_header.js/page_icons.js← page headers, tab-strip markup andbindTabStrip(selected class, ARIA, roving keyboard focus, reveal within the strip; pages keep panel/loading ownership; silentselect()never calls the callback); navigation/header icons.ui_interactions.js←bindDialogFocus,bindMenuand geometry-onlybindPopoverPosition: callers mount/portal and remove their own markup, measured--ui-popup-*values drive the fixed, viewport-bounded.ui-popup, every binding returns teardown, and there is no global overlay registry.scroll_fade.js←bindScrollFadeprojects actually hidden content onto the scroll body's edge attributes, keeping a fully visible first/last edge readable, and owns its observers, listener and pending frame.api_client.js/api_types.js← browser API calls with typed error propagation; mirrors of the browser-facing contract shapes.ui_primitives.js← self-containedrenderSafeField,collectSafeFieldValues,escapeHtmlAttr,normalizeToneandsetInlineStatuswith no shell/API or document-global initialization;ui_helpers.js/utils.jsre-export them, so author pages need no second field renderer.ui_helpers.js← host-bridge, shared row/badge helpers and keyboard menu-lock behavior; the design-system action button and wrapping action-row composition for system chat and routing rows (createSystemMessageAction,createSystemMessageActions); installs its Alt guard on both top-level documents.skill_card_renderer.js← installed-skill cards.hub_sync.js← the one catalog×listing hub-card verdict (install/installed/update/adopt/wait_pr/none plus badges) for the OuroborosHub tab and My-skills badges, joining the catalog with/api/extensionsby canonical name.client_surface.js← the send-time sending-surface snapshot (raw observables, no device taxonomy) thatchat.jsspreads into each chat frame.chat_markdown.js← the chat rich-markdown renderer: marked+DOMPurify, a chat-local URL policy (external http(s)/mailto plus the canonical/api/files/downloadform), KaTeX (single-$unsupported), lazy mermaid, bounded```chartrendering, and an enhance/destroy lifecycle whose disposerschat.jsinvokes on bubble removal (vendored versions:web/vendor/VENDOR-MANIFEST.md).log_events.js← event classification, shared technical outcome reducers and the one factual task-presentation projection for Chat and Logs: task truth maps only toWorking/Done/Done with warnings/Failed/Cancelled; it owns no actions, incidents or notifications, compact headlines never expose raw reason codes, and a task-authored message is visible in the receiving task's timeline (task_message_injected, sender by value) and in the sender's (task_message_routed, written or refused with the host reason).toast.js,masonry.js,widget_frame.js,widget_job.js, CSS tokens ← notifications, layout, framed-widget bootstrap/lifecycle, bounded widget request/job policy.task_control_menu.js← the shared task stop/hurry dropdown, verbatim on Chat live cards and the Activity tab: "Wrap up" / "Hurry up" / "Stop now" (frozen owner wording); a host-attested budget-paused member swaps the working pair for "Resume" (POST /api/tasks/{id}/resume, refusals surfaced verbatim). Eligibility gates differ per surface; actions, endpoint bindings (stop_policymapping, stable per-taskrequest_idretry), locking and typed refusals do not. Choosing an action executes it immediately, dismissing continues the run, a pending cancel leaves only the hard escalation, and "Hurry up" acknowledges via local toast, never a chat message.
confirm_dialog.js::openConfirmDialog is the one browser-dialog authority: confirm mode resolves a strict boolean, input mode {confirmed, value} (empty on cancellation), alert mode one acknowledgement button; Cancel, Close, backdrop, Escape and supersession all resolve as non-confirmation. Native window.prompt/confirm/alert are forbidden in web/modules: inconsistent across shells, event-loop blocking, and window.prompt silently returns null in the macOS PyWebView shell. Critical controls act only on the exact confirmed result — Panic's confirm-and-send is one testable operation. Confirm and New Project dialogs use bindDialogFocus without sharing result semantics: callers bind after mounting and dispose before removal, restoration does not take focus from a different active surface, and menus close before their actions open a dialog. model_chooser.js supplies the shared editable suggestions for model roles and actor/reviewer route editors using popup geometry alone, so menu focus behavior cannot interfere with native text editing: the input keeps its draft/caret/composition through discovery updates, Arrow keys highlight and Enter or pointer selection assigns, Escape/blur dismiss without assignment; fixed source/account choices remain native selects, and unknown saved or arbitrary API model ids stay editable without an inventory allowlist.
Chat and Projects
Timeline ownership and ordering
web/modules/chat.js owns the canonical message timeline, drafts, attachment staging, runtime controls, budget projection, routing annotations, task and child cards, WebSocket subscriptions, thread routing, dedupe/insertion order, unread state and reconnect reconciliation; its per-instance chat_media.js controller owns delivered-media presentation and resources and binds the composer's paste/drag-drop file targets (bindComposerFileTargets), handing files to chat.js's stager. Every ordinary message has one canonical durable chat row; Project views are lenses over those rows and task bindings, so conversion creates no second message, unread event or cost record. Routing acknowledgements live in a compact sidecar keyed by client_message_id and update the existing owner message without a synthetic assistant bubble; a successful annotation carries the confirmed destination Project address for the shared open-or-focus action, and opening a pointer never toggles an already open room closed.
Messages, media bubbles and task-card roots order by raw numeric timestamps (ties keep arrival order; the typing indicator stays last). Application-controlled height mutations use a stable-viewport seam: within 48 CSS pixels of the live edge the transcript follows the bottom, otherwise it restores the visible anchor; remote delivery while the reader is away coalesces into one instance-local activity marker cleared at the bottom. Card text is selectable and a selecting drag never toggles the card; scroll position is remembered per chat instance; band anatomy and the quiet Reviews N count: docs/DESIGN.md.
History reconciliation is one synchronous two-pass transaction over existing keyed records; recent refresh, reconnect and archive prepend preserve the actual timeline/Review DOM nodes, disclosure, focus and selection. A child keeps one keyed current lifecycle row across live delivery and replay; older pages cannot regress its status, and each narration record remains separately visible. Historical narration can enrich a finished child without reviving Working or Stop; source identities distinguish equal timestamps and identical text while timestamps still order presentation; a task's current result outranks its historical terminal observation. The live-card growth trigger requests a fresh history read after 200 additional cards, and only completed, unrepresented cards captured before that request may leave the view (reading/focus/selection or retained-page membership protect them), so live accumulation stays bounded without replacing a reader's card or erasing events that arrived during the fetch.
chat_history.js retains three ordinary rendered pages plus temporarily protected reading/focus/selection pages, holding page descriptors only — content stays with the row/card owners, which release distant page bodies and their media/listener resources while descriptors keep exact return navigation. An empty scan is not a page, there is no newer control, and the per-room scroll stash carries descriptors plus a physical reading anchor, never a second history copy; recent live refresh stays separate from the frozen archive range, and reaching the old physical beginning does not claim every page is loaded.
Main rows and host-stamped card rows
Main receives ordinary main-thread dialogue, required Project question pointers and the two host-stamped Project lifecycle rows — project_started and the terminal project_completion_summary, each with the shared "Open Project" action; all other Project traffic stays in the Project thread. Lifecycle rows are plain dashboard text: the producer strips markdown once before durable write and live send, history normalizes older rows on read (chat.jsonl is never rewritten), and the renderer escapes any system row without markdown: true (except skill_review's dedicated renderer). An origin-addressed notice (task_not_started, task_start_unconfirmed, steer_not_delivered) stays in its issuing chat on replay even when the target is bound to another Project (project_dialogue.room_membership uses ORIGIN_ADDRESSED_NOTICE_TYPES). A steering refusal for an owner act with no owner-message receipt is one factual System row, never model-authored speech.
A required Project quiz additionally projects one keyed System pointer into Main, without a second durable message or answer form. Its task/quiz identity opens the existing Project question; task detail restores a question outside bounded history using the stored labels and optional aligned option details. Recorded answer/expiry updates the same pointer. The pointer row is complete for display: project_dialogue.project_question_pointer (history, the live delivery in message_bus.send_quiz, and the required_question of the activity census) carries the question, the option labels, the recorded answer and the wait facts (owner_wait_projection: the task's wait record when it names this quiz, a record that moved on to another quiz, the closed bound), so Main paints it from the row alone and reads task detail only to open the original form outside the loaded Project history. question_presentation.js owns the lifecycle/action/preview projection shared with the quiz header, and history attaches the same wait facts to the Project room's quiz rows, because a wait the owner resumed by ordinary input leaves no quiz_state frame behind; project_dialogue.QUESTION_STATUS is the Python twin, and question_presentation_parity.json pins both sides on the rows Python emits. Waiting needs positive evidence; resumed and finished tasks stay separately answerable. Freshness is the ordinary history reconciliation plus the quiz_state frame, and lifecycle only moves forward in the chat instance's observation of each question: a settled state never reopens, an unavailable row keeps what is known, and once a live frame closed a wait, a snapshot that still waits (an older history row, a detail read begun before the frame) cannot reopen it. A successful answer HTTP response must contain a valid recorded answer before the UI settles; missing or malformed confirmation never substitutes the local draft, and a late answer's toast says where it went (forwarded). active_chat_activities.required_question recovers current waits through the existing stat-keyed result memo, shared with finalizing.
A host fact about a task is a row of that task's card, never a standalone bubble beside it: the producer stamps the placement (card_row, timeline or reviews, with card_row_id as the row's stable identity across live delivery, outbox replay and history) and the browser attaches the row to the task's card record as one timeline item (reviews also refreshes the card's Reviews group). Only a row whose task has no card record in the page renders as the standalone System row it always was, with the same words; the untyped terminal host notice and the origin-addressed notices carry no placement fact and stay ordinary rows by design.
Finality metadata is row-specific. Durable task_summary rows carry the SAME phase the browser paints: project_dialogue.outcome_phase mirrors log_events.js — the terminality gate plus the severity fold over normalized axes — stamped as outcome_phase beside the status word from OUTCOME_PHASE_HEADLINE, so there is no second status-word family. The pre-finalization authored row carries the phase with outcome_final=false but no host verdict clause; only terminal task_summary rows append the host verdict clause (_completion_verdict, the host twin of taskReasonDetail's acceptance branch) when one exists. The Main project_completion_summary puts the shared headline and any verdict directly in its plain text and carries no outcome_phase. The start row carries neither outcome finality nor a verdict. A verdict is one owner sentence from the shared cause table (project_dialogue.TASK_CAUSE_PHRASES, keyed by the typed reason code; a code with no sentence stays raw): the execution reason on ordinary terminal summaries, or — when the host's acceptance decision is anything but accepted — the decision's reason; the stored reviewer rationale never enters the row (it stays on the task result and in Logs), and the sentence leads the Main completion row so a host-authored row never presents an unaccepted claim as the whole story. Older status words stay unrewritten, and the Telegram skill reads the stamped phase instead of deriving its own. Project panels accept only their own registered chat_id; the complete Project chat-id set is separate from the sidebar's bounded summary, and a projects_changed frame adds a new chat id synchronously before the asynchronous state refresh, so an early frame cannot be misclassified as Main.
Composer, attachments and delivered media
The Main composer exposes one-shot Swarm planning, the owner Nano/Low/Max context choice, file attachment and Send. Swarm places a structural force_plan fact on the next ordinary message and disarms after that send — never inferred from keywords. The host admits one managed root in the sending room without a model call, carrying the verbatim owner objective, origin, surface and attachments; the root may choose its Project through ensure_project_scope (§6 In-task project scoping), and a promote it makes before any plan wave carries its planning obligation to the new root (§6 Task lifecycle). Nano/Low/Max uses the dedicated owner endpoint: a derived Low fit never masquerades as an owner-selected posture, and selecting Max may invoke the exact-route capability acknowledgement when provider metadata cannot prove the window. Project panels omit global Restart, Panic, evolution, consciousness, review and budget controls — those belong to the one Ouroboros process, not a Project thread.
Attachments stage via paperclip, paste or drag-and-drop with no count or size rejection; files stream through upload and canonical staging while the composer and task authority show bounded previews. Staging is local, upload begins immediately before Send, and offline handling is the WebSocket queue contract (§4). Structured attachment metadata lets the gateway expose a native image to a vision-capable route and stage the complete file set through the shared artifact substrate.
Simple text messages may appear immediately as pending local bubbles and reconcile against their echoed client_message_id. Delivered documents capture immutable task-owned bytes and carry a verified file_ref plus canonical download URL, rebuilding from history without persisting base64. Files above the 50 MiB inline boundary are file-backed (desktop saves go through the launcher helper); photos and videos keep base64 only in the live frame — supported media is stored under the canonical task artifact root with a content-addressed URL so replay and rebuilds restore the same bubble. Structured links rows render at most twelve independently revalidated HTTP(S) buttons, identically live and replayed. The supervisor transport owns durable media copies so all producers converge on the canonical data root; failed persistence stays an honest caption row and cannot finalize a task card.
Direct turns and the activity block
Ordinary Main and Project conversations use complete native execution even while other work runs: each turn owns its agent and creates no supervisor queue record. The thread-safe process-local supervisor.active_activity.DirectActivityRegistry retains the private actor handle from preparation through result delivery; Stop, Hurry, quiz and steering resolve that exact actor, while active_direct_turns and typing frames expose presentation fields only. The managed-update writer drain waits for the registered executions, post-task work included; custody maintenance and restart census include those live IDs; and the consciousness alarm clock does not start a wake-up while an owner direct turn is live (a wake ALREADY running keeps running beside a new owner turn).
The producer carries _is_direct_chat into the durable result and terminal event, and the supervisor's rebuilt task_done, the authored task_summary row, the history terminal-truth annotation (gateway/history.py) and the activity census kind all carry that one fact, so late or replayed ordinary Project completions stay in their room without a Main task-completion summary, a chat block reads the same host truth live, on reload and on reconnect, and a direct turn waiting for a model after its answer is never relabelled managed_task. The authored summary row also carries the turn's typed_routing_action, replayed as addressing_only. Every worker progress frame carries the typed narration fact (ouroboros/agent.py::_emit_progress: True only from the model-narration producer ouroboros/loop_messages.py::_emit_round_progress, False from every host note such as a checkpoint, model fallback or review verdict), whitelisted into history (gateway/history.py::_PROGRESS_META_FIELDS) and replayed, so only narration reaches the card title and collapsed activity line while a host note stays a visible timeline row; a frame without the fact is legacy and read as narration.
A turn's activity block is in the transcript iff its record holds content, open attention or launched work — a model wait, a pending or host-offered Stop, a child, a review group, a content row, or a terminal outcome other than Done (web/modules/chat.js::blockVisible, one predicate re-read at every mutation, no sticky flag); the completion note is not content, so a greeting keeps no block. Successful tool calls fold into one evidence row per block (web/modules/chat_activity.js::toolEvidenceView, timeline key tools|<task>, phase calling/warn/result, per-tool counts on Expand; a failed or timed-out call keeps its own error row and is counted), derived live from the call frames and replaced by the host metrics when they arrive, so a tool-using turn shows its block from its first call without a row per call. The block's chrome follows the work it stands on (web/modules/chat.js::blockHasWork: the presence facts minus open attention and minus a bare non-Done ending), never the lane: a block with work is the task card whether a managed root or a direct turn produced it, a block that exists only for open attention or a non-Done ending of a turn that did no work carries no title placeholder and no conversion until its first row of work, the chip always says the block's state, and _is_direct_chat keeps its host jobs and, on the client, the header pill only.
An addressing call (promote_chat_to_task, route_to_project, steer_task, ensure_project_scope) is stamped by the host on its live tool-call frames (routing_action) and counted in the task metrics (routing_tool_calls), both from the one routing-verb table in tool_capabilities.py; the typed annotation on the owner message is its receipt, so the call is a receipt row that never makes the block exist, and a turn that ran only such calls without error draws no block live or on reload — no client tool-name list exists. A refused addressing call keeps the block by its error row, and its receipt line says the cause in the owner's words: the host's cause sentence rides the message_annotation frame and the replayed annotation, routingAnnotationText prints it verbatim, and «Choose a target» appears only with real options. An act the host itself issued and then refused adds one plain 📋 System row (task_not_started / task_start_unconfirmed) in the owner's chat, keyed to the never-started task, which replay neither mints nor finishes a card for. Metrics and authored task-summary rows carry the exact per-tool invocation and error counts from the same execution trace, so replay rebuilds the same evidence row (child cards replay none) and block presence is identical live, on reload and on reconnect — except a turn that moved itself into a Project, whose block and answer replay in the Project room. Historical rows missing those facts remain legacy, without an inferred cost or recovered outcome.
Liveness census and the chat header
GET /api/state exposes active_chat_activities: direct turns united with ROOT managed queue tasks as kind="managed_task" with phase queued, budget_paused (PENDING fenced by a budget-pause fact; nothing dispatches it until an explicit resume, so it must not masquerade as queued), working, or finalizing (RUNNING with an open post-task checkpoint), so a chat instance created after the task started hydrates from the queue authority rather than transient typing frames. active_chat_activities_complete authorizes absence-based reconciliation only when every required live-identity source was read successfully; partial snapshots still carry positive activity. That census is the only inserter into the client's live-activity set (finals and the census itself delete): a complete supervisor-ready snapshot deletes every id it does not list, whatever the entry's kind, and an incomplete one deletes nothing, so an unready supervisor loop or a failed live-identity source pauses deletion until the next complete snapshot. Hydration is a plain projection of that census through the page-wide snapshot sequencer — re-read every 3 seconds while the chat page is open and every 20 seconds otherwise — with no per-entry generation.
A typing frame is a submission receipt, not liveness: it retires the matching local Sending... by client_message_id and asks for a census read, and never inserts, revives or extends a live-set entry. kind is presentation — it selects the label (Thinking for a direct turn, Working/Queued/Paused for a managed root) and exempts nothing from deletion; child typing still arrives without a kind, and the Telegram transport's native typing consumer ignores it. The client status reducer derives the chat header only from connection state, that census, the owner's own unconfirmed sends and live managed task cards (a direct turn's block is not one: the header keeps the census verdict beside it, and a mounted unfinished block hosts its own running indicator instead of the typing bubble); a terminal failure stays a factual task result, never a reasonless header Attention, and a local Sending... is retired only by authoritative evidence (typing receipt, snapshot turn, durable routing receipt, replayed user row, turn conclusion, or offline-queue eviction) — never by the live user-row echo or a socket write. The census enumerates direct turns and managed ROOTS; descendants are carried by their own cards, so a running child that has typed but not yet emitted a progress row is absent from the header until that row mints its card, and a root the census does not enumerate (a Presence turn is never registered in the direct registry) is present through its card alone. The header badge has exactly one writer, the status reducer, with the panel-boot "Online" seed as the sole exception.
Owner-message continuity is journal-backed: each locally sent owner row is kept (bounded, with its routing annotation) until a fetched history response returns the same client_message_id, and a full rebuild re-renders unconfirmed journal rows after the server response, so a stale history snapshot cannot erase a message the owner just sent. Task finalization is honest: a root's early final answer carries task_phase="finalizing" (the same fact replay derives from the open post-task checkpoint), the card holds a sticky Finalizing… phase, and task_cost_finalized is bookkeeping that never resolves a card. If an observed managed root disappears from a queue-authoritative snapshot whose request began after the observation, the state-refresh fan-out removes activity/cancel authority and starts one single-flight durable task-detail read; request generations prevent an older response from undoing a newer projection.
Two invariants close the class of liveness the server cannot name — typing-fed header phantoms and cards shielded by a client-held live-set. First, the history window is lineage-closed: subagent lineage older than the window's progress recency floor is kept only while the child still runs or finalizes, or its parent is REPRESENTED by the response (a row proving its card: telemetry, its own message or summary, a folded review group or plan reference it owns — a typed photo/document/quiz delivery does NOT count), or the parent is alive; an anchored child represents its own children in turn, so a nested swarm is kept or dropped whole, and a FINAL row failing that test is emitted without lineage fields (ouroboros/gateway/history.py), so replay cannot mint an unfinishable parent card while a visible or live parent keeps its completed children and their executor receipts. Second, liveness reconciles over the CARD SET, not only the activity census: a card is shielded by an id the census actually lists and by nothing else; an unvouched connected root card is finished ONLY by proven durable terminal detail; a card without a current result stays unconfirmed until a complete fresh activity snapshot excludes it, then uses the retained typed terminal fact or a neutral Outcome unavailable anchor, without task phase or controls. A root proven terminal (its task_done or its terminal detail) settles the descendant cards a lost child terminal left open from each child's OWN durable result — one single-flight read per child through the ordinary child-terminal path, never a cascade from the parent's outcome, never from the card-set scan; until such a read settles it, a child whose terminal frames were all lost keeps a Working chip without live dots — a card-level residual the header no longer masks.
Task cards, errors and reason lines
Task activity collapses into a live task card per root instead of flooding the transcript; log_events.js keeps the live task card and grouped task cards on one reducer across Chat and Dashboard Logs. Non-terminal LLM/tool/checkpoint failures stay inspectable timeline facts but do not promote the card; only authoritative terminal task truth changes terminal status, and unknown Chat event names do not acquire severity from keyword substrings. Dashboard Logs files a row under Errors from the same typed projection that paints its phase pill (categorizeLogEvent reads the summarizeLogEvent phase — a failed task_done, an errored/killed/timed-out tool, an LLM call failure, a review lifecycle error): a replayed tools.jsonl row reads its is_error/signal facts exactly like the live frame, typed ok/level facts outrank a failure-shaped name, the name-substring test survives only as the disclosed non-expanding remainder for an unknown name carrying no typed fact, and a task group keeps Errors once earned. A delegated harness run's swarm_fanout row (host role="delegated_run") is labelled a delegated run, never a subagent. A failed child keeps a local Failed chip while its root continues. A card keeps a plain-text latest-activity line and an expandable timeline; server-truncated rows fetch the complete typed task record on demand. The compact card is a navigation projection, not the result authority; urgent toast/unread behavior stays confined to explicit task_incident facts, never inferred from severity. The ONE explanatory line under a terminal headline is taskReasonDetail: an owner-requested soft stop has none, a hard failure or cancellation names its execution reason, a non-accepted host acceptance decision states its reason as the owner sentence from the same cause table (no status word, no stored rationale), and everything else keeps the typed reason phrase (an unknown code stays raw); Logs meta additionally names review <status> and acceptance <status>.
Child cards and executor presentation
Subagents render as distinct child cards keyed by their actual child task ids; parents keep lineage references without duplicating the child's final answer. Nested children collapse by default: the headline names the role (a short task id disambiguates identical siblings; Logs keep the diagnostic model/id detail), status is carried by the chip, and Agent model or Coordinator is labelled separately in metadata beside the executor observation/receipt. A child keeps one useful activity line visible and a root allows three, without an empty activity band or a duplicated headline; full narration and Reviews retain independent disclosure. A collapsed child keeps that identity row at every card width (the narrow regime keeps its 160px title demand and wraps the side controls under it), and nested frames stay: each child box borders its whole branch with an opaque background, so ancestry is visible without translucent brightness accumulating at depth; a full history rebuild keeps branches inside their parents, minting a parent represented only by its final row as the nested card its lineage names. A harness pin the resolution refused shows {harness} · blocked on the child card (the route the pin named, projected onto the live frame beside the typed subagent_executor_unavailable terminal), never dispatched or no run yet. Reviews are not task lineage: a reviewer run's execution receipt belongs inside the real owning task card, and its harness or neutral API mark identifies the delivery channel without minting a child card or proving execution. Selected subagent_id/configured snapshot, requested route, effective engine route/model/account and terminal execution evidence stay separate facts — intent must not be redrawn as proof of where the run settled. Legacy lane/executor fields stay readable on historical cards only.
Live executor presentation travels as optional executor_observation on progress metadata: delegate_progress.emit (fed by tools/delegate.py from already-polled run detail) selects the latest typed timeline actor (harnessId, attemptId, event type, snapshot lastSeq) into a record bound by task_id, task_attempt (empty only when unknown), run_id, attempt_id, harness_id, phase and revision; an optional model names its model_source, today only requested and only when custody's requested harness matches that actor — the serving model is absent from the live timeline, and a final-attempt summary never fills the gap. subagent_messages.executor_observation_meta validates the same task-bound shape at Agent emission, supervisor delivery and history replay; a later plain coordinator note inherits no observation. log_events.executorChip labels it last update, keeps requested/observed source explicit and rejects older timestamp/revision observations at the sticky-chip seam; settled execution_evidence.harness_models appears separately as observed history; harness_presentation.js::executorIdentityMarkup is the one markup owner of the executor chip (Coordinator/Agent model and observed-model list), and chat.js assembles no second identity label. This adds no polling, persistent current-actor record, task authority or terminal evidence; execution_evidence and actual_substrate keep their settlement-only meaning.
Reviews projection
An Advisory task-author decision is presented once inside its task's acceptance group; an older panel never inherits a later author's hash, and the same keyed renderer preserves disclosure and reading position.
The task card's Reviews section is a read-only projection over independent domain authorities (skill history, typed plan-review state, task-acceptance evidence, repository-review records) that keep their own lifecycle, verdict and enforcement. A row is admitted only with stable review identity, an exact real presentation-owner task, typed domain state, and exact subject/candidate binding where required; incomplete review stays on its domain surface. Admitted today: task-bound Skill Review, plan review, task acceptance; advisory and commit review stay on their domain surfaces until a bounded exact projection exists. When review history is the only retained fact for an exact owner, Chat may render an inert owner anchor with Reviews — no task phase or liveness until canonical task activity arrives. The projection never synthesizes a task-wide verdict and changes no review routing, status, attention or enforcement. Skill references fold into groups and attempts by the bounded Chat-history reader; plan review and task acceptance hydrate through the task-detail seam (a canonical plan-state write appends one empty typed review_reference for a single-flight refresh; task-result plan_review_state stays authority). A review_reference row is addressed by the task's bound project chat, else by the chat its caller named, else by the hidden partition: Main is never its default, so a run with no owner-visible room keeps its review rows out of the owner's conversation (rows written before a bind are re-classified on read, never rewritten). Truncation uses the quota cause so Load older expands both overlays; duplicate Skill lifecycle acknowledgements are typed lifecycle_pointer rows that may enrich a present exact owner but never mint a card. Opening and closing Reviews belongs only to the user (docs/DESIGN.md). chat_id=0 remains the hidden Skill Review partition, never a Main surface. No review inbox, generic review endpoint, review ledger or second state machine.
Card cost
Card cost is sticky task-scope evidence. Only frames carrying task accounting status, finality, subtree, reservation or unknown-cost fields may update it; an unrelated per-call cost_usd delta is never relabelled as the task total. Compact cards render one amount — the accounted upper bound worded up to while the ledger is open, plain once final — preferring the complete subtree projection; a running root's heartbeat may carry a non-final aggregate from the physical-attempt ledger, so the amount advances without a second timer, endpoint or client-side sum, and a history window replays the same non-final projection (live_root_cost_projection, the one cost owner) on a running or finalizing root's latest in-window progress row so a reload between heartbeats shows the same ceiling — a subtree with no attributable rows stays absent, never zero. Precedence: unavailable, then pending, then final, newer evidence winning within a class; costless frames cannot erase a known value. Dashboard and task-detail accounting derive from the physical-attempt ledger, not the card; compact review rows copy or sum no money, and Skill wave attempt/slot money appears only inside the lazy exact-job detail, joined from the same ledger via physical_attempt_v1.
Stop and cancel controls
A Cancel action appears only on unfinished, unconverted root cards carrying the host-attested cancelable=true; pooled tasks and addressable native direct turns use the existing custody owner (§5), and card shape alone grants no control. The stop control is the shared dropdown above, sent with cascade:true (root plus live descendant subtree): the hard path answers only after teardown or a typed refusal; the graceful "Wrap up" path records a durable finalize intent and answers immediately with a typed 202 pending acknowledgement. Natural completion wins a race, a missing live task reconciles from the durable record, and a process that cannot be proven dead remains a visible refusal rather than being painted Cancelled.
Turning a card into a Project
A Main managed or Swarm root card may be turned into a Project: conversion creates or reuses the Project, names it owner-facing, binds the task and its canonical origin message, moves live work onto the Project lane and replaces the Main action with a calm Project pointer. Tasks already bound to a Project get no second conversion button, and a block that has done no work yet (open attention only, or a bare non-Done ending) has none to offer. A direct conversation turn's card is convertible like a managed root's: a turn outside the queue has no lane to mark, so the durable bind alone is its commit point, and its later rows route to the Project room per event (events_chat_delivery._bound_project_chat_id), as for a turn that called ensure_project_scope. The button is rendered from the record's facts (chat.js::syncBlockChrome) and the /api/state binding sweep in app.js reads the same task_bindings fact; the toast says Project opened when the server adopted the existing Project (adopted: true) and Project created otherwise.
The binding, origin-claim and adoption mechanism is owned by §6 Project binding by task and by origin; the UI relies on its outcomes. The one-click conversion (no explicit project id) ADOPTS the Project the task's owner message already has instead of minting a second one: the existing row comes back with adopted: true, the coined name is discarded, and a card whose button is still on screen after a sibling bound it answers with its own Project rather than an error. A fresh conversion binds that origin's live sibling roots and direct turns to the new Project (disclosed in events.jsonl as project_origin_siblings_bound, skipping a sibling whose own binding names another room), so — for work born from one owner message, while the binding store is readable and that Project is still active — a promoted root and the turn that promoted it cannot become two Projects for one message, and a second click mints nothing. A request naming a DIFFERENT project id is an explicit choice and keeps the 409 refusal. An ADOPTING conversion binds only the clicked card, so /api/state.task_bindings also carries the message's other LIVE task ids, marked origin_bound: their convert button closes and their pointer opens the project, while their live cards stay in the chat they were started from, because nothing durable bound them and their chat rows have not moved. The conversion answer always states adopted, true or false.
The pointer is the one shared project chip (ui_helpers.js::renderProjectChip): cards inside a Project panel and nested subagent cards never receive it, and clicking opens the panel or does nothing when already open — a pointer never closes what it points at. Naming reuses an already coined model title or falls back through the server naming path; the UI invents no second name authority. Project-SCOPED is not project-BOUND: a headless/CLI run carries a project_id for lease and memory without a durable binding — addressed to the Project thread at admission (§5), it never mints a Main card and needs no conversion button.
Project rooms
A Project chat has one compact navigation pointer in its status bar (project_work_pointer.js). projectWorkTarget selects the latest connected non-child card by timestamp and physical source position that is unfinished, or the latest represented root when all are finished; it creates no card or execution state. Clicking scrolls only that chat to the card and records reading intent; it never selects the recipient of the next message. The label is Working · / Latest task · plus the card's coined name or title, capped by projectWorkLabel to one line — never the card's full status headline. The pointer keeps the bar's original 180 px flex basis, so a default desktop panel renders one row while the status pill is short; a longer pill, a narrower panel or a phone wraps the bar to a second row, never a third. Unless historyWindow.complete is explicitly true the adjacent note says Loaded messages only; an empty represented set hides button and note, never claiming the Project has no work. No separate task pane, history fetch, poller or persisted pointer state exists.
The New Project dialog supports exactly one source: no folder, a fresh managed genesis workspace, an attached existing folder, or a cloned Git URL; the selected folder stays visible independently of the directory being browsed and is what submission uses, and its shared focus boundary leaves source/creation and non-confirming cancellation semantics with project_create.js. Attach uses a server-side directory browser so the flow works in web and Docker too; ordinary attached folders are supported directly, Git-specific operations require a Git worktree, and clone failures distinguish missing credentials. All three disclose that Project tasks receive read, write and shell access in the chosen folder; provenance is a durable historical fact, not recomputed from current git state.
Deleting a Project is lifecycle work, not filesystem deletion (§6 Project registry and lease): the server fences new admission, cancels and quiesces the Project task subtree, then tombstones the registry entry; the UI may acknowledge that deletion started but never claims completion early. Project id, canonical history, task bindings, memory, provenance and the working folder remain — deleting the row is not permission to erase the owner's repository or the agent's history.
Project unread state is the durable comparison visible_revision > project_seen_revision; only owner-visible assistant/result content, delivered media or a real incident advances the visible revision. Opening a panel does not clear unread by itself: the browser refreshes and paints the exact revision, verifies the panel is still visible and connected, then posts that revision as acknowledgement; the server clamps and max-merges it so a stale tab cannot move the cursor backwards or acknowledge future output.
A chat instance has an explicit resource lifecycle: destroy() marks it dead, disposes subscriptions, listeners, observers and timers, and removes the DOM last; late async work checks the destroyed flag. app.js keeps at most one live Project chat instance; closing or switching stashes scroll intent and destroys it. The narrow exception is client state the server cannot reconstruct (staged File objects or an upload in flight): such an instance is hidden and marked pending, reused on reopen, and returns to the destroy policy once that work settles; typed but unsent text survives separately in per-thread session storage. Hidden Project rooms therefore neither accumulate listeners nor acknowledge unseen revisions, and no data is discarded.
History reads and the SSE v2 transport
All five task-event sources (progress, chat, events, tools and supervisor) read their retained archive chains plus live files. Interactive Chat history uses GET /api/chat/history, preserving recent-window defaults, separate human/telemetry quotas and the legacy limit. An additive opaque cursor selects older physical ranges; page_cursor replays the same frozen portion and next_cursor advances backward; both streams retain their initial upper boundary through rotation and live append. Continuation resumes from its last processed position, including foreign rows, without rescanning the visited tail; an empty sparse Project page with has_more=true is not EOF. A cursor-bound source failure retains the exact input cursor for Retry; a recent read with an unavailable archive still projects its readable rows, reports the gap, and offers Retry with null cursors until the boundary is known. Changed room membership returns history_view_changed rather than mixing views. The shared JSONL chain helper owns metadata and inode rebinding. history.py uses one mapper for recent and older selections, preserving quiz answers, lineage, media, Reviews and current terminal authority; typed quiz-answer/terminal rows may carry replay evidence without a visible bubble; quota-deferred rows remain reachable, including children whose parents are on another page. window.complete means complete reachable history, independently of physical has_more; a successful individual page does not establish it. Display reads avoid materializing artifacts.
Task-event lineage discovery, task-list ordering and Main's newest-result selection share gateway/task_list_scan.py's compact process-local memo: dev/inode/size/mtime/ctime invalidate changed files, failed or torn reads are never cached, full selected results still pass the schema reader, and directory enumeration stays proportional to the number of results. A living event follower's retained per-file child bindings survive only failed reads; valid parent/role changes, schema refusals and deletion remove them. Nonfatal history_gap frames with lineage_incomplete distinguish retained bindings from unknown membership and do not invent a view change; fresh connections never treat client cursor paths as proof; these diagnostics consume no legacy history rank and acknowledge no log bytes. Main's dialogue projection reads twenty nonempty text rows from a live tail and at most two archives, retaining message ids and the 500-character text preview; its unvisited-row count is explicitly unknown.
POST /api/tasks/{id}/events is the read-only SSE v2 transport (TaskEventsRequest): JSON {v:2, wait, cursor} carries {v,seq,view,positions}, one byte offset per root/source across the complete archive+live chain. Events use physical root/source order, never timestamp ranks; each log row carries a stable root/source/byte-position event_id and the cursor after THAT row. Filter changes (task ids, roots, task_done suppression or a proven creation floor) emit cursor_replay and replay the new view. The CLI advances only after consuming a frame and deduplicates its most recent 4096 log identities; older replay duplicates remain possible. A cursor_checkpoint at wait expiry advances skipped bytes without incrementing delivery sequence or claiming task progress. Live partial lines wait; immutable malformed/partial lines emit history_gap and advance; unreadable required segments or shorter chains emit cursor_unavailable without resetting. The cursor rules: byte positions survive rotation (one source-bound metadata snapshot per pass over the shared JSONL chain helper, the original live generation rebound by inode after rotation, consumed prefixes stat'd without opening, at most the live file and one archive held); each 64 KiB-plus-final-line buffer is closed before network delivery; each source pins its logical chain EOF for one pass, even through rotation, so later appends belong to the next pass and cannot starve other sources or result/checkpoint handling (a row cut only by that boundary stays pending even if its file rotates before the next batch); failed or torn reads are never cached (the memo above). POST structure is validated by the derived validate_ingress contract before semantic cursor checks. Archives must remain immutable and retained: manual prefix removal masked by subsequent growth is outside the offset guarantee. Only explicit creation facts may skip older archives, and skipped bytes still count in positions: fresh host UUID allocations in HTTP task creation, tool scheduling and supervisor scheduling stamp created_at before preparation; supplied/restored ids, legacy rows and first-at-final results are not backfilled and conservatively retain a full cold scan. A result-directory failure stays explicit: SSE emits an error rather than an empty view, Main marks result omissions unknown with the source error, and the legacy list wrapper keeps its fail-soft warning. Each connection emits a fresh task-result projection and materializes the terminal result once; synthetic results have no log identity and are never deduplicated. Legacy GET and iter_task_events keep timestamp-sorted integer ranks, whose retroactive insertions may repeat or omit rows across reconnects (a from-zero replay recovers retained history); the CLI falls back to GET only on a first-connection HTTP 405. Subagent admission/custody dedupe remains the supervisor queue's transition (§5).
The agent-facing chat_history reader uses the same live-plus-rotated timeline and may narrow by exact provider, account, conversation, thread, actor and inclusive date bounds before the count/offset/text-search window — presence provenance is searchable as structured transport fact, not only flattened prose.
Files
Files is a full gateway-backed file manager, not a chat attachment picker: directory navigation with breadcrumbs and filtering, image and sandboxed PDF preview, text preview and editing with an explicit binary/unsupported state, file and directory creation, save, drag-and-drop upload, download, open in the default OS application, copy, move, paste and recursive delete. The desktop host bridge and the web fallback share one download contract. Unsaved text is guarded on selection, directory change, navigation and unload. Save exists only for a writable COMPLETE text read — a truncated preview stays read-only — and save failures and clipboard feedback land in a sibling status region, never over editable content. The backend is the path authority (root confinement and symlink policy: §4); UI path strings and disabled buttons are presentation only.
Skills and Widgets
Skills has three views: installed skills, ClawHub and OuroborosHub. The installed view merges extension truth with the serialized lifecycle queue, so install, update, review, dependency work, enable, disable, repair, uninstall and failure stay visible while queued or running. A failed primary read keeps a labelled previous list or an unavailable state, never a successful empty list, and ClawHub cannot derive Install from an unavailable installed-state read.
Installation, deterministic preflight, LLM review, owner grants, dependency readiness, extension loading, enablement and execution are separate lifecycle facts: a fresh review implies neither granted keys nor installed dependencies, and enabled=true does not override a blocked review or a load error. Owner attestation, where eligible, skips only the expensive LLM review; deterministic preflight and post-pass reconciliation still run. Repair and run is an ordinary managed development task visible in Chat, entered through normal owner-message ingress, in which, after review and prerequisites, the model enables and tests the selected installation; a later direct owner disable wins over that older request, and automatic repair-and-review carries no enable authority. Hub publication uses the selected-skill preflight and an ordinary managed task; the passive Installed projection neither runs Betterleaks nor claims publication readiness.
Widgets page
Widgets is a separate page because extension UI is an execution surface, not catalogue metadata. It renders only the UI tabs of reviewed live extensions, read from the dedicated passive projection GET /api/widgets (each card's revision is the owning skill's live payload content_hash, a change signature, not an ETag; the Skills page stays on /api/extensions). The list is reconciled by key against an order-independent change signature: an unchanged list moves no <article> node, and a changed card is stopped in order before its replacement mounts. The same reconcile runs on a visible extension_lifecycle event, on every WebSocket (re)connect and from the contextual Retry a failed list read shows beside its error while the last good cards stay; there is no list polling and no global Refresh/remount control, so retrying discovery never stops a kept-running program behind the owner's back.
Three render modes share one capability set (sandbox="allow-scripts allow-pointer-lock allow-downloads", allow="autoplay; fullscreen; clipboard-write", never allow-same-origin; author-facing contract: docs/CREATING_SKILLS.md). An extension-route iframe (kind: iframe) is the skill's own page with no bridge and must declare the route it loads; the frozen registration contract still accepts an omitted route, which paints a not-supported card rather than a widget. A declarative widget is rendered by host-owned code from a validated schema. A reviewed module widget runs in an opaque-origin srcdoc iframe under a document CSP that web/modules/widget_module.js builds from the page origin (default-src 'none', no connect-src; scripts only inline, blob: and the skill's module prefix, with 'wasm-unsafe-eval'). Framed declarations may set a bounded height (320–8,192 px); a module without one starts at the floor and reports its #root content height through the nonce-bound bridge, capped by an optional module-only max_height; legacy route iframes stay explicit-height-only because an opaque document cannot be measured, and geometry keys are rejected for declarative renders. Card order and the owner's per-card launch-policy override (widget_start_mode) are host UI state in /api/ui/preferences; the masonry (web/modules/masonry.js) packs cards in that order and writes only --masonry-* custom properties, so a reorder never moves a node or reloads a frame (disclosed residual: Tab order follows the DOM until a reload). Engineering rules: docs/DEVELOPMENT.md "Embedded surfaces declare geometry and refresh semantics", "UI resources carry a disposer", "Declarative widgets".
Framed cards start under a launch policy — the owner's override over the author's validated render.start over the kind default (module and route iframe → manual, declarative → auto), a pure function in web/modules/widget_card.js — with one primary Start/Stop control and a policy menu (Auto / Manual / Keep running). retain keeps a framed card mounted while Widgets is hidden and ends it on Stop, when its skill leaves the live list, on a changed revision and with the window, never with the server alone. A mounted widget owns its resources through one disposer: a module widget is stopped in order with acknowledgement (ouro-widget-disposed, or WIDGET_DISPOSE_ACK_TIMEOUT_MS — one second), with one mount in flight per card key so a remount waits for the pending stop; a route iframe disposes synchronously. Poll and WebSocket writers use monotonic progress per job, so an older response cannot rewind a newer event. Chart.js is bundled locally; rendering must not depend on a third-party CDN.
Declarative widgets
Declarative widgets support forms and actions, status/data/text/code/markdown, tables, tabs, charts, polls, jobs, streams, subscriptions, progress, media, files, maps, calendars, kanban, and composition through group, metric and callout. One recursive validator limits the tree to depth 8 and 256 nodes and reports the exact failing path; nested interactive components use an explicit id or stable tree path as identity. Within one fixed declaration mount, data updates reconcile nodes in place, preserving input identity, selection, composition, native popup state and live password text (retained form-value snapshots exclude passwords). Forms and actions carry their own keyed pending/result/error feedback, job settlement included. subscription.render remains transitively passive so an incoming event cannot smuggle a new active control tree past validation. Text, attributes, links, media routes and field values are escaped or constrained for their actual sink.
Module widget bridge
A module widget receives one parent-mediated I/O bridge on a per-mount nonce — the frame's only scriptable network path, since connect-src is closed — which keeps useful route I/O without giving reviewed skill JavaScript the SPA's cookies, DOM or broad API authority. OuroborosWidget.fetch (also the frame's fetch) is relayed by the parent, which accepts only the exact owning prefix under /api/extensions/<skill>/..., sends same-origin credentials, refuses to follow a redirect, and streams the answer back as a real Response over a ReadableStream — no default timeout (the author's init.signal or init.timeoutMs aborts), while declarative requests and the module source load keep a 25-second bound. The skill's namespaced WebSocket events arrive through OuroborosWidget.onEvent, filtered by the card's ws_prefix; out-of-process responses ride the same framed stream with no total or pre-header timer. Downloads reuse the common native/browser save owners, and external links go through ui_helpers.openExternalViaHostBridge on the same bridge (trusted anchor clicks, OuroborosWidget.openExternal, the frame's no-handle window.open): the host hands off to the native opener, the ready Telegram SDK or the browser before any asynchronous work, and a null noopener handle is not proof of failure. The out-of-process WS push (POST /ui/ws-message) admits a 60-message burst reserve per skill refilling one message per second and refuses the excess with a typed 429 (§12). The module source endpoint (GET /api/extensions/{skill}/module/{entry:path}) authorizes against the live loader registration only and serves the declared entry or any reviewed sibling .js/.mjs from the texts captured when the bundle registered — an edit after load is not served until the skill reloads — answering with Access-Control-Allow-Origin: * because the requesting frame is an opaque origin. The same nonce carries the fault channel: an in-frame script error, unhandled rejection or CSP violation is posted as one bounded, deduplicated ouro-widget-error message into the card's own status slot while the lifecycle state stays running, because the frame is still mounted; the frame also exposes data-widget-content-height and data-widget-frame-capped, so a widget pinned at its ceiling is distinguishable from one that painted nothing.
Author UI kit
Optional author controls reuse the SAME installed CSS and pure field/status source, not a second declarative renderer. ouroboros.server_web.read_author_kit_assets(request.app.state.repo_dir) reads the fixed web/ui.css and web/modules/ui_primitives.js through the serving-root web resolver and creates no endpoint, cache or state; docs/examples/author_ui_kit/ holds the two ordinary recipes (a module fetches the texts through its own GET route and OuroborosWidget.fetch; a route iframe embeds them under its own nonce CSP); neither changes auth, the bridge, the sandbox or the widget schema. Authors opt controls into the named classes and .ouro-ui and may override or omit the kit; there is no body reset, required conformance, hot-theme protocol or forced remount; retained frames keep their loaded copy, and a failed kit load is shown inside that author application. Tests: tests/test_author_ui_kit.py, tests/test_author_ui_kit_browser.py.
Out-of-process extension responses
An out-of-process extension response keeps its process and bundle contexts through asynchronous startup and teardown; blocking context entry/exit, process registration, pipe shutdown and termination run off the ASGI loop. runtime_limits.py (re-exported by config.py) owns the 64 KiB response chunk and the two-second post-response cleanup grace; neither is a response deadline. The reader bounds one metadata frame to 512 KiB and one body frame to its chunk plus flag, never a whole response. An abnormal child exit keeps its bounded, sanitized stderr and exit code without altering a body already delivered; static module assets come from the captured registration, without a child.
Dashboard
Dashboard groups Logs, Evolution, Costs, Updates and Activity under the common tab binder; each sub-tab owns its loading and refresh policy instead of running every expensive reader while hidden. Charting uses the bundled Chart.js (§3 Widgets).
Logs merges live WebSocket log frames with bounded REST backfill from the events, tools, progress and supervisor logs: chronological order, dedupe against live and reconnect overlap, bounded grouped task cards, raw record on demand, and shared category/severity/review presentation from log_events.js. A failed backfill names the unavailable history sources in a separate status row while live events continue. Clearing the visible panel does not delete the underlying logs.
Activity combines running and pending queue entries with the live activity census from the same refresh's /api/state: an identity already in the queue keeps its title, runtime and budget controls; a census-only direct turn shows its kind, phase, elapsed time and the shared task controls. A partial census supplies positive rows but cannot prove nothing is running. The view also shows the consciousness alarm clock (enabled, status line, autonomy level, next and last wake-up with its outcome, the rolling-24h allowance and how many of its tasks run) and scheduled work, with mechanical controls only where that surface is authoritative: typed cascade cancellation, start/stop for background consciousness, enable/disable/delete for owner-managed schedules. Each section reports its own failed read; a failed queue read neither erases the background/schedule sections nor claims an empty queue. A schedule reconciled from a skill manifest is read-only here — a direct edit would be overwritten by the skill lifecycle and falsely appear durable.
Costs is a projection of the physical-attempt ledger distinguishing confirmed, reserved, unresolved upper-bound, unknown/unmetered usage, open rows and finality; unavailable data renders as unavailable, not $0. Breakdowns by model, key, model category and task category are views over the same ledger. Total budget can hot-apply; the per-task value is a hard cost cap over the whole root tree of the next task, not an own-task soft warning (§7 Default settings), and raising a cap does not resume work that already finalized or paused.
Evolution shows the current evolution state and the consciousness alarm clock, campaign objective/progress, queue/failure/budget information and durable history. Starting a campaign uses the shared input dialog: cancellation starts nothing, confirmed empty input selects the backend's default autonomous objective, entered text becomes the objective. Light runtime mode disables self-modifying campaigns rather than presenting a control the backend will refuse.
Updates
Updates separates passive status from explicit mutation. Opening the page reads cached/local update and git state without touching the network; the passive read carries the last real check's timestamp and a minimal update_tx projection, so a re-opened panel sees an assisted resolution in progress. One exported pure verdict (updates.js::updateVerdict(status, phase), unit-pinned by web/tests/update_verdict.test.js) renders exactly ONE action button whose label is always the real next continuation, from Check for updates to Restart now (the degraded case where the automatic restart callback failed); "Up to date" is claimed only over an actual check result, a failed check keeps its actionable label, and unknown backend warning classes surface verbatim. The apply flow verifies the preflight (verifiedUpdatePlan) and confirms BEFORE applying, naming the path: a clean update restarts the server, a conflicting one starts the reviewed assisted task (model spend disclosed) whose progress lands in chat. The served-SHA decision remains the only page-reload authority; the boot-owned pending_boot_smoke and applying_replace phases keep the synthetic restarting state until a post-reconnect update_status_ready proves boot finalization returned. Typed apply-failure facts reach the owner rather than one string. Recovery holds the separately confirmed replace action, "Save recovery point" (moves only this installation's local ouroboros-stable fallback, never the official QA feed — §8) and ONE restore list where local tags label the commits they point at (/api/git/log tags carry their peeled sha); unavailable, divergent, dirty, unsafe, failed-check, rollback and restart-required stay visible as states, never as extra buttons. Passive Git status disables optional index locks so those reads cannot contend with the update writer, and reconnect preserves a pending preflight/apply.
The update letter — one short markdown paragraph Ouroboros writes about the update on offer — rides the same status payload as an additive letter key and refreshes on the same update_status_ready event, so it costs no endpoint, socket type or poll of its own. updates.js::updateLetterView(status, phase) projects it beside the verdict without touching it: a pending letter reads "What's new"; an applied one (HEAD equals the target, or the last check that described this HEAD found the target already inside it — a divergent install applies through a merge commit) reads "What changed in this version", because the letter outlives the update it describes; a superseded or relocated one keeps its text under the range it was written for; a failed write — including a range git could not read, never recorded as "nothing to say" — shows the last good letter, labelled by ITS range, with the reason, or the reason alone. It renders through the sanitizing chat markdown pipeline and adds no second action: unmanaged, unreadable and never-checked states and the restart phase hide it rather than grow a control. The sidebar update pill (update_status.js) is a pointer, not a second apply surface; visual rules: DESIGN.md §8.
The synchronous lock-owning apply executor publishes process-local stage observations through ouroboros/gateway/update_progress.py into the existing status payload (update_progress); update_progress_changed only invalidates that view. The observation names the real stages, survives an HTTP disconnect, and never authorizes mutation or recovery; a new process starts without it, while durable transaction recovery and the post-reconnect update_status_ready keep their authority. The panel shows actual stages without percentages, timers or an additional action.
Settings and onboarding
Settings has Accounts, Secrets, Models, Agents, Behavior, Advanced and About tabs — a sequence from connections to runtime detail. Accounts: managed subscriptions and their shared service banner, API providers, custom compatible endpoints, local runtime entry points, and the optional non-loopback network gate. Secrets: known provider/integration secrets, skill-requested keys and owner-defined custom keys, without returning stored values. Models: compact source/model/account role rows, ordered fallbacks, context assertions and effort lanes. Agents: task actors and review lanes, with delegation permissions, per-root and depth limits and subagent path roots; their accounts are managed in Accounts. Behavior: context, safety-supervisor coverage, task acceptance, self-evolution, prompt-cache posture. Advanced: process, timeout, local-model, integration, source-control and cleanup controls (worker count is process capacity, so it lives here). About reports application/runtime identity. Keys, defaults and per-key semantics are the §7 Default settings table, not this chapter. Models, Available subagents and Review lanes share ONE grouped source select owned by web/modules/route_editor_primitives.js (routeChoiceGroups, configuredApiProviders): the owner picks a source and the editor composes the stored id, so the provider prefixes (provider::model, claudexor::source=model, harness=model) are serialization only, never owner input.
The Settings client validates the whole current draft before Save; a local error keeps every value available for correction and sends no partial save (settings_controls.js keeps custom-key collection pure and dirty reads passive). Ordinary refresh and failed writes preserve current edits, leaving or explicitly reloading a dirty draft asks first, the write response distinguishes saved, unsaved and unknown outcomes, and there is no durable cross-page draft store or secret persistence.
Advanced also hosts the settings sections live extensions register, hydrated before they are shown: the client reads each declarative form/action component's own route with GET, and a 2xx JSON object of saved values pre-fills exactly the fields it names through renderSafeField (passwords are never hydrated), so Save rewrites current values instead of posting first options over them. A 404/405 means the skill serves no current values and the form renders from its declared defaults with Save enabled; a transport or server failure renders it with Save disabled and a note pointing at Reload Settings, because a blank form must never overwrite values the client could not read.
Each provider card has one compact Test action backed by POST /api/providers/test: request exactly {provider_id, overrides?}, response exactly {ok, error?}. Overrides are request-local — an omitted field reads the saved value, an explicitly edited empty field stays empty (on the compatible card it suppresses the legacy OpenAI fallback) — and draft credentials never mutate Settings, the process environment or LLMClient caches. The probe uses a configured route, then the provider's maintained main default, and only the generic OpenAI-compatible route performs bounded catalogue discovery when neither exists; it is one bounded, physically accounted call (ouroboros/llm_probe.py) with none of the normal-chat retry, fallback, tool, reasoning, web, cache or capability-learning paths. The card renders Testing…, Works or Not ready with one controlled short reason; the tooltip notes the request may incur provider charges.
Desktop onboarding and the blocking web overlay are the same served /onboarding page — same steps, same backend normalization; context mode stays a separate owner setting. Startup readiness is structural (§2): a configured managed-model Main route can satisfy the gate without an API key, while credential validity, entitlement, model availability and local-process health remain runtime status, not onboarding admission. Every host completes through the single POST /api/onboarding/complete transaction (§2), so completion is all-or-nothing and install-time agent defaults are part of the same save.
Agent accounts
Accounts is the owner-facing projection of Ouroboros's owned Claudexor daemon; the browser never receives its control token or interprets credentials. Status combines daemon/runtime readiness, login-capable harness discovery, credential profiles, honest vendor-live versus local-session verification, fresh quota windows and optional model discovery. One /v2/quota envelope supplies both windows and typed per-subject absences (quota, quota_absences), so a missing usage reading stays separate from login truth; only fresh snapshots are eligible for percentages or exhaustion, and a fully-used ratio without a valid future reset stays non-blocking and visibly unproven while an explicit active cooldown remains effective. API-key-only adapters gain no fake Login buttons merely by appearing in a broader execution catalogue.
Accounts group into one card per agent family: an aggregate status counting the accounts rotation can actually use (signed-in AND enabled), a fail-safe "Next up" badge naming who an unpinned run would take (unified accountPools first, legacy per-harness next_up second), and the add action. Rows are ONE type on both engine generations: on a UNIFIED engine (the server-stamped unified_accounts feature fact) every account — migrated default logins included, under the reserved <harness>-default registry ids — is a named row with an Enabled toggle and Remove; on a LEGACY engine the native pseudo-row keeps the same layout and only its ACTIONS differ — no Remove, no toggle, because that engine has no route for either and a dead button would claim an effect this process cannot have. verification=not_run is neutral unknown, and an explicit availability=unknown re-runs the shared Refresh rather than a new sign-in — an auth probe failure is never proof of logout. The legacy ''-keyed quota alias is granted only to the literal reserved <harness>-default id; pinned route health applies no alias and fails open to the engine's own typed refusal. Removing a named account preserves the complete deletion receipt: a refusal remains a refusal, and a vendor-owned / left-unchanged / OS-user disposition becomes a retained-credential warning, not a false sign-out claim. One service banner explains a daemon or runtime problem once, per facet: claudexor_accounts.py fans the catalog, account and quota reads out independently and stamps each classification into the payload's reads block, so one refusal never collapses its siblings into a global unreachable verdict; only a legacy payload without the stamp is read coarsely.
Connect is link-first and harness-agnostic: a typed disclosure renders the sign-in URL and any one-time code. Engines publish optional-without-default setupLogin per exact harness row (in_app → omitted setup transport, external_terminal → client_pty). Typed facts select every continuation, prose never does: credential_profile_required plus add_named_account selects the name-the-account face, a typed duplicate profile is idempotent, typed pre-job refusals prove setup custody absent/released, and the exact terminal_transport_* code (or durable job.nativeCommand.errorCode) offers the explicit external-terminal continuation. Only the engine's structural missing-vendor-binary first-create job lets the same owner action also consent to one hidden local install through the exact managed Claudexor CLI and one retry; success requires Claudexor's strict post-install proof (absolute installedBinary plus bounded installedVersion) — exit zero without proof is refused, and a second login refusal is returned rather than looped. A client_pty job exposes its labelled attach command immediately; there is no embedded terminal login surface. Terminal job state and the account row reconcile, so a stale verification read during login cannot claim failure after the account connected.
Account status refresh runs immediately and on visible page/tab activation; hidden pages do not pay for daemon round-trips. Entering Agents is an explicit owner action: after the fresh read an already-provisioned stale home restarts through the wake endpoint, while not_provisioned, foreign-owned and repair states stay behind Connect; background polling is read-only and never wakes the daemon. Job polling uses one request at a time, backs off on consecutive failures, and after ten stops with an honest unconfirmed state — lost contact proves neither failure nor settlement. One transition lock covers Start, Retry and Dismiss, and a new login begins only once release of the prior job is proven (loginReleaseProven): a 2xx cancel alone is not proof, and a network/server failure retains the card and job id because dropping it could orphan a live server job.
Review lanes and Available subagents
Review lanes edits one structured reviewer configuration (contract: §7 Reviewer slots). Each triad, scope, optional advisory or deep self-review row picks its reviewer from ONE flat select: Available-subagents roster rows lead as references, then the inline channels — API delivery or a coding-agent session — followed by model, optional credential profile and effort. The deep self-review block states its one difference from the advisory where the owner picks — an API model there is ONE packed review (Atlas + memory), not an inspection episode — and says so while the row is only synthesized from OUROBOROS_MODEL_DEEP_SELF_REVIEW. An untouched synthesized (or empty) placeholder is OMITTED from the save payload, so an unrelated save never writes the key's value into the setting; editing the row materializes it, and a blanked model box is refused typed at save. Saved choices that disappear from discovery stay visible as unavailable rather than silently changing; capability labels configure nothing, and server-returned limits plus the last effective execution disclose what a saved row actually ran as. The backend enforces the same contract independently of the client's pre-POST validation, and an unloaded view authors no replacement. A row pinned to an account discovery no longer lists keeps its pin, disclosed once above the rows and only on the word of a facet actually read (account pins answer to accounts, models to catalog; an unread or failed facet yields "not checked", not "not in discovery"). The standing note states the rule, never the situation: every review surface — commit, scope, plan, advisory, skill review and task acceptance — follows its configured rows and waits for subscription capacity rather than falling back to API spend. An agent-session row's meta line carries the live availability sentence the Available-subagents cards compute (sessionRouteVerdict): the row, never the model option, says whether an account can run the selected model.
Available subagents is the single task-actor editor. Its list-level Enabled flag and at most ten stable rows are the saved OUROBOROS_SUBAGENTS intent. The owner sees numbered compact cards (docs/DESIGN.md §6 row anatomy) and authors one prose field, Description (recommended_use), beside the structured API-model or Agent-session route, optional effort and optional managed-model/session account pin (empty pin = Claudexor's compatible-account rotation). Each card's status dot composes two axes — intent (saved / draft / generated) and availability (an API-model route reads Checked at start) — with the tone the worse of the two (subagent_status_primitives.sessionRouteVerdict / rowStatus). Identity is the stable internal subagent_id plus live route facts; parse drops a legacy display name, and a visual ordinal never becomes durable identity. Agent-session rows also select native Access: Full system access (default) or Working files. A saved lower choice is explicit; rows saved without the field default to full for new tasks, while running tasks keep their captured profile and read-only assignments stay read-only. The editor shares only neutral route/model/account/status primitives with Review lanes; Add and Duplicate reveal the new entry through the shared ui_helpers.revealNewRow. A fresh entry is not an error: its meta line carries a neutral hint until a save attempt (noteSaveAttempt; validate() stays pure), after which the offending card is tinted and names its error.
Saved intent and live evidence are separate axes: row status distinguishes saved/generated/draft intent from current availability and may show the last requested→effective run evidence, and a status/catalog/accounts failure annotates rows, never erases them. Settings GET may offer an unsaved migration/default candidate when no canonical value exists; the editor materializes it only on Save, and a late status or preview response may replace a still-clean generated baseline, never an owner-edited draft. Generic Settings save strictly validates and canonicalizes the materialized value before the serialized off-event-loop owner transaction; a running task keeps its immutable start snapshot, and the save response says changes apply from the next task.
Mutative subagents use Off, Auto and On: an explicit value applies to every acting surface, while Auto delegates the default to runtime mode — surface-aware in Light, which allows only children building outside the Ouroboros runtime (external workspace, genesis); why: §6 Safety and runtime mode, exact matrix ouroboros/config.py. Read-only children stay available. Owner-controlled, applies from the next task, independent of API vs coding-agent route.
Prompt Cache TTL is one global owner choice (OUROBOROS_PROMPT_CACHE_TTL, §7 Default settings) shared by every builder lane (task, review and safety), with the explicit tier applied only to cache markers on compatible Anthropic-family wire payloads, so builders cannot drift into conflicting cache horizons; the UI does not promise cache behavior on providers that manage it implicitly, and the shipped one-hour posture favors reuse across long waits and review cycles.
Settings save effects
Settings save classifies effects into three classes rather than claiming everything became live at once (exact key membership: settings_scales.py IMMEDIATE_SETTINGS / RESTART_REQUIRED_SETTINGS, applied by gateway/settings.py). Hot-apply — e.g. total budget, tool timeout, MCP configuration (a failed hot-reconfigure surfaces as a save warning); the retained soft/hard timeout keys are accepted only as deprecated audited no-ops and reported as retired. From the next task — e.g. models, credentials, reviewer/subagent configuration; a running task keeps its starting snapshot and the response says so, and a failed task-start settings reload is disclosed as the persisted, chat-visible task_start_settings_reload_failed event. Restart-required — e.g. worker count, bind host, the skills repo path; such a save offers Restart now over the owner /restart command. A changed skills repo path or runtime mode also reloads the server's extension loader at save time (a failed reload is a save warning), and a pooled worker does not wait for a respawn to see a newly published extension: before each task it adopts the server's published extension generation with one bounded reload per distinct generation (supervisor/worker_process.py _adopt_published_extensions; events extension_generation_adopted with its converged flag and extension_generation_adoption_failed) — a bounded attempt whose worker-side failure is caught and disclosed, not a convergence guarantee. Cyber can configure context, review and Supervisor through the same audited settings writer, including while a task is active; Access remains pending until restart. A saved choice never rewrites an existing task snapshot or physical request; configured review enforcement and Cyber's decision to continue remain separate facts. Every in-process owner-settings writer holds the shared settings_document_mutation() lock across its read/merge/write; the file lock is a write precondition, not a substitute (§7 Reading and writing the settings document). Loading reviewer and delegation settings waits at most the boundedStatusRefresh two-second foreground beat for Claudexor; a cold refresh continues and repaints when it lands. Backend failure and browser transport failure stay distinct, and absence of a successful status read is never evidence a runtime is healthy.
Visual verification policy
A visible change is exercised in at least one relevant real consumer flow and the rendered result is inspected with vision; a saved screenshot alone is not verification. Mobile, WebKit, additional browsers and special viewports are selected from the actual interaction risk — not a universal matrix (reviewer items: docs/CHECKLISTS.md). No visual-QA runner, endpoint, ledger, or mandatory device matrix is introduced by this policy.