Merge PR #729: streaming skill responses and owned external operations
Some checks are pending
CI / quick-test (push) Waiting to run
CI / benchmark-methodology (push) Waiting to run
CI / full-test (macos-latest) (push) Waiting to run
CI / full-test (ubuntu-latest) (push) Waiting to run
CI / full-test (windows-latest) (push) Waiting to run
CI / betterleaks-platform-smoke (macos-latest) (push) Waiting to run
CI / betterleaks-platform-smoke (ubuntu-latest) (push) Waiting to run
CI / betterleaks-platform-smoke (windows-latest) (push) Waiting to run
CI / integration-test (push) Waiting to run
CI / skill-smoke (macos-latest) (push) Waiting to run
CI / skill-smoke (ubuntu-latest) (push) Waiting to run
CI / docker-ui-smoke (push) Waiting to run
CI / system-e2e-mock (push) Waiting to run
CI / release-preflight (push) Blocked by required conditions
CI / vendor-package-smoke (push) Blocked by required conditions
CI / skill-smoke (windows-latest) (push) Waiting to run
CI / marker-guards (push) Waiting to run
CI / ui-smoke (push) Waiting to run
CI / docker-portable-test (push) Waiting to run
CI / e2e-live (SM1 x1 — largest subset feasible under the $30 cap) (push) Waiting to run
CI / build (dmg, macos-latest, macos-arm64, syft_1.50.0_darwin_arm64.tar.gz, syft, e32fdb9d47823fa633748a1efca2528fd77c37469ea93c9e40ab835da44e4cce) (push) Blocked by required conditions
CI / build (tar.gz, ubuntu-latest, linux-x86_64, syft_1.50.0_linux_amd64.tar.gz, syft, bf7b29ff57f06da30918266a0e1c2885a8f99784798d1bdb1628886aa015d788) (push) Blocked by required conditions
CI / build (zip, windows-latest, windows-x64, syft_1.50.0_windows_amd64.zip, syft.exe, 815ee6973ec5dff6a671d7f41b0e78835a8c45b91d5a39f4743ea1cee833d3be) (push) Blocked by required conditions
CI / release (push) Blocked by required conditions
Claudexor platform gate (API keys — subscription auth NOT covered) / fixture · macos-latest · exact managed runtime, fake harness, no model (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / fixture · ubuntu-latest · exact managed runtime, fake harness, no model (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / fixture · windows-latest · exact managed runtime, fake harness, no model (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · ubuntu-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · windows-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · codex · API key only, subscription NOT covered (push) Waiting to run

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
Anton Razzhigaev 2026-09-07 15:30:51 +00:00
commit a5e6b983e3
53 changed files with 4264 additions and 437 deletions

View file

@ -278,6 +278,7 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
├── skill_review_cycles.py ← Paid skill-review cycle counting, $0 replay, typed `review_cycles_exhausted`, accepted-rebuttal ledger; the shared cap SSOT is review_cycles.py
├── extension_loader.py ← Extension loading: in-process pure-Python via `PluginAPIImpl`, child-process proxies for isolated-dep/native extensions
├── extension_process_runner.py ← Extension child processes: scrubbed env, per-skill deps, timeouts, graceful host errors
├── extension_route_stream.py ← Portable stdio response frames and ASGI relay; child standard responses retain headers/HEAD/Range/background work, parent backpressure and exact-bundle cancellation use the existing runner process owner and supervised_futures
├── extension_ui_validation.py ← The one host-owned declarative-schema-v1 widget validator
├── extension_isolated_deps.py ← Legacy/forced in-process bridge for isolated-dependency extensions
├── extension_health.py ← Durable process-qualified per-skill health at `data/state/skills/<name>/health.json`; server observation is authoritative, worker observation is a handoff-qualified view
@ -741,7 +742,19 @@ Widgets is a separate page because extension UI is an execution surface, not cat
Declarative widgets support forms and actions, status/data/text/code/markdown, tables, tabs, charts, polls, jobs, streams, subscriptions, progress, media, files, maps, calendars, kanban, and composition through `group`, `metric`, and `callout`. One recursive validator limits the tree to depth 8 and 256 nodes and reports the exact failing path. Nested interactive components use an explicit id or stable tree path as identity; `subscription.render` remains transitively passive so an incoming event cannot smuggle a new active control tree past validation. Text, attributes, links, media routes, and field values are escaped or constrained for their actual sink.
Module widgets receive one parent-mediated I/O bridge on a per-mount nonce — the frame's only scriptable network path, since `connect-src` is closed: `OuroborosWidget.fetch` (also the frame's `fetch`) posts the request, the parent accepts only the exact owning prefix under `/api/extensions/<skill>/...`, issues it with same-origin credentials, refuses to follow a redirect, and streams the answer back as header/data/end frames from which the child rebuilds a real `Response` over a `ReadableStream` (binary by default; incremental reads work; no default timeout — the author's `init.signal` or `init.timeoutMs` aborts; declarative requests and the module source load keep a 25-second bound); the skill's namespaced WebSocket events are forwarded the same way (`OuroborosWidget.onEvent`, filtered by the card's `ws_prefix`). Disclosed limits: a route served by the out-of-process runner is buffered whole (about 380 KiB of body) and the out-of-process WS push is capped at 60 messages per minute per skill. The module source endpoint (`GET /api/extensions/{skill}/module/{entry:path}`) authorizes against the live loader registration only and serves the declared entry or any reviewed sibling `.js`/`.mjs` from the texts captured when the bundle registered (an edit after load is not served until the skill reloads), answering every response with `Access-Control-Allow-Origin: *` because the requesting frame is an opaque origin. This keeps useful route I/O without giving reviewed skill JavaScript the SPA's cookies, DOM, or broad API authority. The same nonce carries the frame's fault channel: an in-frame script error, an unhandled rejection or a CSP violation is posted as one `ouro-widget-error` message (bounded, deduplicated, clipped) and reaches the card's own status slot, while the lifecycle state stays running because the frame is still mounted; the frame also exposes `data-widget-content-height` and `data-widget-frame-capped` so a widget pinned at its ceiling is distinguishable from one that painted nothing. Chart.js is bundled locally; rendering must not depend on a third-party CDN. Framed cards start under a launch policy — the owner's override over the author's validated `render.start` over the kind default (module and route iframe → `manual`, declarative → `auto`), a pure function in `web/modules/widget_card.js` — with one primary Start/Stop control and a policy menu (Auto / Manual / Keep running); `retain` keeps a framed card mounted while Widgets is hidden and ends it on Stop, on its skill leaving the live list, on a changed `revision` (stopped in order and re-mounted) and with the window, never with the server alone. A mounted widget owns its resources through one disposer: for a module widget that is the ordered stop with acknowledgement — dispose message posted, bridge kept answering the child's hooks, then abort/unlisten/remove on `ouro-widget-disposed` or after `WIDGET_DISPOSE_ACK_TIMEOUT_MS` (one second) — with one settle promise and one mount in flight per card key, so a remount waits for the pending stop instead of racing it; a route iframe disposes synchronously. Poll and WebSocket writers use monotonic progress per job so an older response cannot rewind a newer event; job polling keeps its `job_id` across bounded retryable failures while explicit terminal states stay terminal; transient list failure preserves the last good widgets.
Module widgets receive one parent-mediated I/O bridge on a per-mount nonce — the frame's only scriptable network path, since `connect-src` is closed: `OuroborosWidget.fetch` (also the frame's `fetch`) posts the request, the parent accepts only the exact owning prefix under `/api/extensions/<skill>/...`, issues it with same-origin credentials, refuses to follow a redirect, and streams the answer back as header/data/end frames from which the child rebuilds a real `Response` over a `ReadableStream` (binary by default; incremental reads work; no default timeout — the author's `init.signal` or `init.timeoutMs` aborts; declarative requests and the module source load keep a 25-second bound); the skill's namespaced WebSocket events are forwarded the same way (`OuroborosWidget.onEvent`, filtered by the card's `ws_prefix`). The bridge asks for the next body chunk on consumer pull. Out-of-process responses use the same child-to-host framed stream: ordered raw headers, standard HEAD/Range handling, cancellation until completion and separate diagnostics for cleanup failure after delivery; there is no total or pre-header timer. The existing loaded bundle owns cancellation through supervised_futures, and only the dispatched instance is affected by unload. Module download calls and ordinary supported export links reuse the common native/browser save owners; route URLs stream without building a Blob. The out-of-process WS push (`POST /ui/ws-message`) admits a 60-message burst reserve per skill that refills one message per second, refusing the excess with a typed 429 (§12). The module source endpoint (`GET /api/extensions/{skill}/module/{entry:path}`) authorizes against the live loader registration only and serves the declared entry or any reviewed sibling `.js`/`.mjs` from the texts captured when the bundle registered (an edit after load is not served until the skill reloads), answering every response with `Access-Control-Allow-Origin: *` because the requesting frame is an opaque origin. This keeps useful route I/O without giving reviewed skill JavaScript the SPA's cookies, DOM, or broad API authority. The same nonce carries the frame's fault channel: an in-frame script error, an unhandled rejection or a CSP violation is posted as one `ouro-widget-error` message (bounded, deduplicated, clipped) and reaches the card's own status slot, while the lifecycle state stays running because the frame is still mounted; the frame also exposes `data-widget-content-height` and `data-widget-frame-capped` so a widget pinned at its ceiling is distinguishable from one that painted nothing. Chart.js is bundled locally; rendering must not depend on a third-party CDN. Framed cards start under a launch policy — the owner's override over the author's validated `render.start` over the kind default (module and route iframe → `manual`, declarative → `auto`), a pure function in `web/modules/widget_card.js` — with one primary Start/Stop control and a policy menu (Auto / Manual / Keep running); `retain` keeps a framed card mounted while Widgets is hidden and ends it on Stop, on its skill leaving the live list, on a changed `revision` (stopped in order and re-mounted) and with the window, never with the server alone. A mounted widget owns its resources through one disposer: for a module widget that is the ordered stop with acknowledgement — dispose message posted, bridge kept answering the child's hooks, then abort/unlisten/remove on `ouro-widget-disposed` or after `WIDGET_DISPOSE_ACK_TIMEOUT_MS` (one second) — with one settle promise and one mount in flight per card key, so a remount waits for the pending stop instead of racing it; a route iframe disposes synchronously. Poll and WebSocket writers use monotonic progress per job so an older response cannot rewind a newer event; job polling keeps its `job_id` across bounded retryable failures while explicit terminal states stay terminal; transient list failure preserves the last good widgets.
The extension response handle retains the existing process and bundle contexts
through asynchronous startup and teardown. Blocking context entry/exit, process
registration, pipe shutdown and termination run off the ASGI loop; cancellation
waits for its exact startup worker before closing any spawned child. Process facts
retain their existing thread-local publisher. `runtime_limits.py`, re-exported by `config.py`, owns the 64 KiB response
chunk and two-second post-response cleanup grace; neither is a response deadline.
The reader bounds one metadata frame to 512 KiB and one body frame to its chunk
plus flag before allocating its payload; this does not bound a whole response.
An abnormal child exit retains its bounded, sanitized stderr and exit code after
the existing drain, without changing a successfully delivered body.
Static module assets still come from the captured registration without a child.
### Dashboard
@ -917,6 +930,8 @@ Every `/api/files/*` operation resolves its requested path and refuses the opera
| GET | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/tools/schemas` | `gateway.host_service._api_tool_schemas` |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/allocate-internal` | `gateway.host_service._api_allocate_internal` |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/inject` | `gateway.host_service._api_chat_inject` |
| GET | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/operations/{operation_ref:path}` | `gateway.host_service._api_chat_operation` (the calling skill's own accepted message: pending, running with its task or turn, the durable answer, or the terminal task status) |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/cancel` | `gateway.host_service._api_chat_cancel` (the existing cancellation owner on work that message started; a typed outcome, never a cancellation that did not happen) |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/decision` | `gateway.host_service._api_chat_decision` (the `task_decision.answer_decision` ingress relayed for a transport skill) |
| POST | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/presence/turn` | `gateway.host_service._api_presence_turn` |
| GET | `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/presence/work/{work_ref}` | `gateway.host_service._api_presence_work` |
@ -981,7 +996,7 @@ Heartbeat and progress are different evidence: a heartbeat proves a process or l
No retry or new assignment may occupy a timed-out slot until the original process is provably dead. If kill and join cannot establish death, the reaper preserves a low-rank RUNNING result, keeps the slot marked `reaping`, emits a visible wedged receipt and restart hint, and performs no terminal write, `task_done`, retry, or respawn. One slot is sacrificed rather than letting a still-running process race a replacement and overwrite its result; the next supervisor generation reconciles the durable record after old-generation process custody has run.
A spawned or respawned slot is not capacity until its child confirms it. Both spawn paths install the slot `reaping` — the marker assignment, the crash detector and the reaper already honour — and hand it to one readiness seam (`supervisor/worker_pool_lifecycle.py`), which opens the slot only when the child's own `worker_ready` row names its pid and verifies the booted SHA in the same step. A child alive but silent past `WORKER_READY_WINDOW_SEC` (`ouroboros/config.py`, beside the spawn grace) is torn down and replaced through the ordinary respawn path — one transaction under the pool's lifecycle serializer, so a concurrent pool restart cannot install a fresh slot in the gap for the stale watcher's respawn to evict — at most `WORKER_READY_MAX_ATTEMPTS` consecutive times per slot before the slot is parked and the owner told; each outcome is a typed `supervisor.jsonl` row (`worker_sha_verify`, `worker_ready_timeout`, `worker_ready_released`). A child that dies during boot is released to the crash detector, which owns process death. The seam's own failure is never a parked wave: an exception inside the watcher releases every slot of that wave still booting (`worker_ready_released`, `reason=watcher_error`), degrading to the crash detector's ownership rather than holding capacity until restart, and its event reader treats a missing `events.jsonl` (not written yet, the rotator's rename-to-touch instant, removed by hand) as an empty read. Readiness is a contract distinct from process liveness (`is_alive`) and from the task idle rail: a child deadlocked on a lock inherited across fork is alive and holds no task, so only readiness can see it. The pool does not fork from the multi-threaded supervisor at all: Linux workers start by forkserver (one single-threaded parent), macOS and Windows by spawn.
A spawned or respawned slot is not capacity until its child confirms it. Both spawn paths install the slot `reaping` — the marker assignment, the crash detector and the reaper already honour — and hand it to one readiness seam (`supervisor/worker_pool_lifecycle.py`), which opens the slot only when the child's own `worker_ready` row names its pid and verifies the booted SHA in the same step. A child alive but silent past `WORKER_READY_WINDOW_SEC` (`ouroboros/runtime_limits.py`, re-exported by `config.py`, beside the spawn grace) is torn down and replaced through the ordinary respawn path — one transaction under the pool's lifecycle serializer, so a concurrent pool restart cannot install a fresh slot in the gap for the stale watcher's respawn to evict — at most `WORKER_READY_MAX_ATTEMPTS` consecutive times per slot before the slot is parked and the owner told; each outcome is a typed `supervisor.jsonl` row (`worker_sha_verify`, `worker_ready_timeout`, `worker_ready_released`). A child that dies during boot is released to the crash detector, which owns process death. The seam's own failure is never a parked wave: an exception inside the watcher releases every slot of that wave still booting (`worker_ready_released`, `reason=watcher_error`), degrading to the crash detector's ownership rather than holding capacity until restart, and its event reader treats a missing `events.jsonl` (not written yet, the rotator's rename-to-touch instant, removed by hand) as an empty read. Readiness is a contract distinct from process liveness (`is_alive`) and from the task idle rail: a child deadlocked on a lock inherited across fork is alive and holds no task, so only readiness can see it. The pool does not fork from the multi-threaded supervisor at all: Linux workers start by forkserver (one single-threaded parent), macOS and Windows by spawn.
Unexpected worker death is a three-way decision. An already-terminal durable result wins and is projected idempotently; a negative process exit code is terminal for every task, because replaying the same infrastructure signal usually repeats the failure and burns budget; only an otherwise-eligible non-signal crash retries within `QUEUE_MAX_RETRIES`. Repeated busy-worker or all-workers-dead failures trip the crash-storm fence, disable pooled admission, and surface recovery instead of cycling workers indefinitely. Direct chat stays available because it is not owned by the pooled scheduler.
@ -1877,15 +1892,6 @@ Add the field to the active frozen owner — `ouroboros/contracts/` for the pack
## 12. Host Service, Companion Processes, and Chat IDs
The Host Service is a loopback, authenticated callback boundary for reviewed skills (`ouroboros/gateway/host_service.py`, `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}`). Every request authenticates an opaque `x-skill-token` bound to the skill's content hash, executable review, enablement, and grants — secrets never enter the token, a payload edit stales it, and the client wrapper refuses stringification (`skill_token.py`). The frozen route family is exactly: `/identity`, `/tools/schemas`, `/chat/allocate-internal`, `/chat/inject`, `/chat/decision`, `/presence/turn`, `/presence/work/{work_ref}`, `/ui/ws-message`, and WS `/events`; permissions still decide which route works for a given skill. Review of a transport skill evaluates identity binding, attribution, polling bounds, panic cleanup, token confinement, and exfiltration — an owner-bound reviewed transport may be a first-class control surface, not a screen-only integration. External slash commands bind a separate positive-identity external owner slot (`supervisor/state.py`), so an unidentified transport can never bind commands and the local web owner can never lock out a real remote owner. `wait_for_response` on an injected chat message requires an A2A-allocated chat: only A2A chats have single-conversation semantics — on a human chat the first non-progress frame can be any concurrent task's answer.
Host chat attachment copies remain request-owned until enqueue succeeds. The
existing synchronous copy runs through `gateway._helpers.run_sync_to_completion`,
which waits for it even under HTTP cancellation before releasing the in-flight
slot. A request-local cleanup stack removes only newly created destinations on
cancellation or partial admission failure; accepted message attachments survive
later disconnects. The skill retains its original files throughout.
Chat uploads have one `gateway.files` storage owner for Host-confined paths and
completed multipart spools. It uses the artifact substrate's streaming hash and
atomic copy; borrowed spools promise descriptor identity, not an original pathname.
@ -1907,6 +1913,20 @@ directs the owner to the app. Delivery never starts a tunnel or publishes a new
public/token-bearing artifact URL, and this notice does not claim the bytes were
uploaded to Telegram.
The Host Service is a loopback, authenticated callback boundary for reviewed skills (`ouroboros/gateway/host_service.py`, `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}`). Every request authenticates an opaque `x-skill-token` bound to the skill's content hash, executable review, enablement, and grants — secrets never enter the token, a payload edit stales it, and the client wrapper refuses stringification (`skill_token.py`). The frozen route family is exactly: `/identity`, `/tools/schemas`, `/chat/allocate-internal`, `/chat/inject`, `/chat/operations/{operation_ref}`, `/chat/cancel`, `/chat/decision`, `/presence/turn`, `/presence/work/{work_ref}`, `/ui/ws-message`, and WS `/events`; permissions still decide which route works for a given skill. Review of a transport skill evaluates identity binding, attribution, polling bounds, panic cleanup, token confinement, and exfiltration — an owner-bound reviewed transport may be a first-class control surface, not a screen-only integration. External slash commands bind a separate positive-identity external owner slot (`supervisor/state.py`), so an unidentified transport can never bind commands and the local web owner can never lock out a real remote owner. `wait_for_response` remains restricted to A2A-allocated chats. Named messages wait for their exact operation; only the legacy unnamed path uses chat-wide response subscriptions. Admission is one limiter with two policies (`host_service._RateLimiter`): the WS relay lane (`/ui/ws-message`) is a token bucket — a 60-message burst reserve per skill refilling one message per second, so a burst no longer silences a widget for the rest of a minute — while every other lane keeps its 60-per-60-seconds sliding window. A refused relay is visible at the host and aggregated per burst: one warning when the reserve empties, then one warning plus one durable `host_service_ws_relay_dropped` row in `logs/events.jsonl` carrying the dropped count when the lane admits again or the idle bucket is swept; the 429 carries `retry_after_sec` and `dropped_in_burst`, the child's `send_ws_message` stays best-effort `None`, and the in-process broadcast path keeps no bound.
Host attachment copies use one request-local cleanup stack and the shared settled
HTTP-worker wait. Partial copy/refusal/cancellation removes only new unaccepted
copies. Named ingress retains those copies once its existing canonical write is
attempted, through an internal custody callback; a replay never adopts its duplicate copies.
Cancellation waits for that admission worker. Accepted bytes survive disconnects
and unknown write/queue outcomes. Such an error does not claim acceptance and
may retain an unused copy; no new reconciliation or transaction log is introduced.
Operation correlation (#667): a named injected message has `operation_ref=<chat_id>:<client_message_id>` on 202, 200, 504 and disconnect responses. `supervisor.message_bus.accept_local_message` serializes check, canonical inbound-row acceptance and enqueue; its existing `log_chat` writer must succeed before work is queued. The host-minted full origin is carried internally to `record_inbound_message`, which verifies the source and consumes it without writing a second row. Repeated same-id, same-text, same-skill delivery rejoins even before supervisor dequeue; changed content or source is refused with 409. This preserves the existing in-memory queue: a crash after acceptance can lose delivery and reads honestly as `lost`, never authorizing a second enqueue. Reads span retained canonical chat generations. Routing annotations and outbound task ids are discovery hints only: task reads and cancellation require the actual queue/task record's complete `origin_message_ref` to match the authenticated skill's canonical source. `DirectActivityRegistry` carries that same origin. Named response waits poll this exact operation and its retry-aware effective task result; chat ordering alone never proves a reply. Cancel enters the existing durable intent and cascade-custody owner only for work with that origin and the same installation root. A different or unavailable owner root is disclosed as cancel_unsupported before any intent is written. Unaddressable/foreign work remains `cancel_unsupported`; unresolved custody remains explicit and never becomes a false `cancelled`.
Successful extension children also carry a bounded `ws_relay_failures` count map from the existing PluginAPI transport owner through their result envelope and process facts. Missing transport, network errors and HTTP refusals remain visible without changing `send_ws_message -> None` or the producer's success. One host warning reports the aggregate; diagnostics contain no message bodies, URLs or credentials. A child that dies before its final envelope may lose this aggregate; the existing measured death/timeout facts remain authoritative. The separate streaming route implementation carries the same aggregate in completion X after body/background work.
Presence flow: `POST /presence/turn` requires the content-hash-bound `presence` permission, one `binding_id`, one exact transport event, and optionally staged files confined to the skill's state root; the binding resolves from `state/presence_bindings.json` with exact provider/account/conversation/thread origin verification; cross-process locks enforce the installation-wide cap and serialize one `conversation_key`; a stable event-derived task id makes transport retries idempotent; input and output join ordinary dialogue history with full transport/actor provenance. The one typed outcome is message/silent/tool_delivered/deferred — `deferred` only with a correlated `work_ref`, because an unanchored "deferred" would be an unanchored promise — and `GET /presence/work/{work_ref}` polls the bound late result without exposing the general task API. Promotion of Presence work into a managed task clears requested Project/workspace/source widening: a public conversation may promote long work but cannot choose new authority, and the cost ceiling plus return destination follow the promoted root by value. `presence_cancel_work` acts only on a `work_ref` whose stored binding and conversation match the current turn; owner chat and Background Consciousness may `initiate_presence` on an existing enabled binding.
Companion processes are host-supervised: reviewed manifest-declared descriptors enter durable custody, reconcile after lifecycle changes and restart, and stop on disable/unload/panic; `state/extension_generation.json` carries the opposite direction — the server's published live set, which a running task worker adopts at a task's start or at a dispatch miss, so an enable after boot is not invisible until the pool respawns. Worker-side changes write durable reconcile requests (`state/extension_reconcile/`) rather than spawning server-owned children; every reconcile state names the process that answered and whether that marker request was written, and the tool and review receipts pass both facts through. Health observations are process-qualified: aggregate and Skills UI health use the server observation as authority and expose the worker observation with its handoff outcome only as a qualifier, so failed handoff success cannot advance `last_known_good`; restart-budget exhaustion persists a terminal reason in that health state, cleared only by a later successful start. A companion's cwd is the reviewed payload directory, so a payload edit stales review before reload instead of silently mutating a live process. The live projection is `state/extension_companions.json`.

View file

@ -324,8 +324,13 @@ Out-of-process caveats: (1) `register()` and `on_unload` run for **each per-call
child** (every tool/route/WS dispatch and catalog), so `on_unload` fires per call,
not once per disable — keep it cheap and idempotent and put durable/once-per-session
teardown in a `companion_process` shutdown. (2) `send_ws_message` relays through the
loopback Host Service and is best-effort and rate-limited (~60/min per skill), so a
progress-heavy job should throttle updates or rely on poll-based status. (3) A
loopback Host Service and is best-effort: the relay lane holds a 60-message burst
reserve per skill that refills one message per second, the excess is refused (a 429
with `retry_after_sec`; the host records each refused burst once with its dropped
count), so a progress-heavy job should throttle updates or rely on poll-based
status. Successful children also return aggregate transport/HTTP refusal counts
through their normal result/process diagnostics; `send_ws_message` still returns
`None`, and acceptance does not prove browser delivery. (3) A
`companion_process` is spawned and supervised by the host **server** process: enabling
a companion skill from the agent's `toggle_skill` tool or via post-review auto-enable
records it in the worker process and writes a durable
@ -577,7 +582,7 @@ calls and runtime behaviours:
| `supervised_task` | The skill may register an in-process host-supervised async task. |
| `companion_process` | The skill may register a manifest-declared companion subprocess supervised by the host. |
| `subscribe_event` | The skill may subscribe to manifest-declared host event topics such as `chat.outbound` or `skill.lifecycle`. Chat topics require owner permission grants; `skill.lifecycle` does not. |
| `inject_chat` | The skill may request Host Service chat injection after an explicit owner permission grant: `POST /chat/inject` carries text, an inline image, or `attachments` (`[{path, name?, mime?}]` — regular files under the skill's own state root, at most 25 per message, which the host copies without the former 50 MiB upload cap into the shared `data/uploads` chat-upload store and stages for the task; a file-only message needs no text). The same grant lets the skill relay the owner's decision-card answer through `POST /chat/decision` (`{request_id, decision_id, option_index?, comment?}`, the `POST /api/decisions` contract). |
| `inject_chat` | The skill may request Host Service chat injection after an explicit owner permission grant: `POST /chat/inject` carries text, an inline image, or `attachments` (`[{path, name?, mime?}]` — regular files under the skill's own state root, at most 25 per message, which the host copies without the former 50 MiB upload cap into the shared `data/uploads` chat-upload store and stages for the task; a file-only message needs no text). The same grant lets the skill relay the owner's decision-card answer through `POST /chat/decision` (`{request_id, decision_id, option_index?, comment?}`, the `POST /api/decisions` contract). A message that carries a `client_message_id` becomes an addressable operation: the host answers with its `operation_ref` (`<chat_id>:<client_message_id>`) on 202, 200 and 504; a repeated delivery of the same message rejoins it instead of enqueueing again (a different message under a reused id is refused with 409); `GET /chat/operations/{operation_ref}` reports the skill's own accepted message (`pending`, `running` with its task or turn, the durable answer, a terminal task status, or `lost` after a host restart); and `POST /chat/cancel` (`{operation_ref, reason?}`) runs the existing cancellation owner on work that message started, answering `cancelled`, `already_terminal`, `unresolved` or `cancel_unsupported` — never a cancellation that did not happen. |
| `presence` | A reviewed transport skill may submit authenticated non-owner conversation events to the Host Service Presence boundary and poll only their correlated late work. Requires an explicit content-hash-bound owner grant. |
A missing permission causes the matching `register_*` call to raise
@ -873,6 +878,20 @@ ui_tab:
start: manual # auto | manual | retain — see "Launch policy" below
```
The manifest declaration is checked during preflight and review; it does not
create a live tab. With `permissions: [widget]` in your extension manifest,
register the same surface in `plugin.py`:
```python
def register(api):
api.register_ui_tab("editor", "Editor", render={
"kind": "module", "entry": "widget.js", "start": "manual",
})
```
`register_ui_tab` creates the tab on the Widgets page after the extension loads.
Keep `widget.js` beside `plugin.py`; it renders into the provided `#root` element.
The host fetches reviewed JS through `GET /api/extensions/<skill>/module/<entry>`,
embeds it in an opaque-origin iframe (`sandbox="allow-scripts allow-pointer-lock
allow-downloads"`, never `allow-same-origin`; see "What the frame may do"
@ -913,7 +932,7 @@ The frame has no scriptable network of its own: `connect-src` stays closed, so
`XMLHttpRequest`, `WebSocket`, `EventSource` and beacons are refused by the
document policy, and every request goes through the parent over one nonce-bound
message grammar (passive image, media and font loads from your own route prefix
are the one exception — "What the frame may do" below). Two calls cover it:
are the one exception — "What the frame may do" below). The bridge exposes:
- **`OuroborosWidget.fetch(url, init)`** (also installed as the frame's
`fetch`). `url` must resolve under `/api/extensions/<skill>/...`; anything
@ -925,7 +944,8 @@ are the one exception — "What the frame may do" below). Two calls cover it:
answers with a redirect rejects instead of being followed; it streams the
answer back, so you get a
real `Response`: `status`, `statusText`, **every** response header, and a
body that is binary by default — `.text()`, `.json()`, `.arrayBuffer()`,
body that is binary by default. Each next chunk is read only when your
consumer requests it — `.text()`, `.json()`, `.arrayBuffer()`,
`.blob()` and incremental `body.getReader()` reads all work. Server-sent
events are a plain streaming `GET` with `Accept: text/event-stream` read
through `body.getReader()` (there is no `EventSource` polyfill); NDJSON works
@ -943,13 +963,29 @@ are the one exception — "What the frame may do" below). Two calls cover it:
host strips its own namespace prefix. The first listener subscribes the frame,
the last unsubscribe stops delivery, and other skills' events never reach it.
Two limits are disclosed rather than hidden: a route served by the
out-of-process runner (isolated dependencies) is buffered whole before the frame
sees it and capped at about 380 KiB of body (the same ceiling as the
out-of-process module-bytes route below) — only an in-process route's `StreamingResponse` streams chunk by
chunk; and the out-of-process / companion WS push (`POST /ui/ws-message`) is
capped at 60 messages per 60 seconds per skill, so throttle progress events or
fall back to poll-based status for bursts.
- **`OuroborosWidget.download(name, source)`** saves an existing `Blob`, a
`data:` URL, or a URL under this skill's extension route prefix. It resolves
to the host's delivery result or rejects with a visible error. In the desktop
app it uses the same native Downloads owner as other file controls; in a
browser success means the download was started, not that disk writing was
confirmed. Large backend files should be passed as route URLs so the host
does not turn an HTTP stream into a Blob. Ordinary `<a download>` controls
using these routes, `data:` URLs or frame-created Blob URLs use this same path.
Out-of-process routes execute standard Starlette responses in the child and
stream their ordered headers and body to the host, including `FileResponse`
HEAD/Range behavior and background actions. There is no total or pre-header
request timer: finish, abort, disconnect or unloading that skill instance ends
its response. Slow consumption is backpressure, not a timeout. A body failure
breaks the stream; cleanup failure after the complete body is logged separately.
The incoming request body retains its 512 KiB cap, and one-shot tool/catalog/WS
results retain their existing caps and timeouts.
The out-of-process / companion WS push (`POST /ui/ws-message`) admits
a 60-message burst per skill and then one message per second (the excess gets a
429 with `retry_after_sec`, and the host logs each refused burst once with its
dropped count), so throttle sustained progress streams or fall back to
poll-based status.
#### What the frame may do
@ -987,15 +1023,12 @@ What that gives you, verified on Chromium and WebKit through
`data:` / `blob:` URLs — "Assets" below, including the CORS rule for fonts.
- **Clipboard write** (`navigator.clipboard.writeText`) from a user click; the
clipboard is never readable from the frame.
- **Downloads**: an `<a download>` or `blob:` link clicked by the owner
downloads in browsers (`allow-downloads`). The desktop shell's link
interceptor runs in the parent document only and the frame cannot reach the
shell bridge, so a download started inside the frame may be ignored there.
A download that must also work in the desktop shell stays host-side today:
serve the file from a skill route and let a declarative widget's `file`
component or a chat-delivered file offer it — both go through the host's
`downloadViaHostBridge` path. A module-frame download call over the bridge
is not built yet (disclosed).
- **Downloads**: module widgets use `OuroborosWidget.download` or ordinary
`<a download>` controls for their own route files, data URLs and frame-created
Blobs; the existing host save path supports both the desktop app and browsers.
A legacy `kind: iframe` route page has no module bridge: its downloads still
depend on the embedding engine. Use a module widget or a host-side declarative
`file` component when a native Downloads handoff is required.
- **Pointer lock** (`allow-pointer-lock`) and **fullscreen**
(`allowfullscreen` + `allow="fullscreen"`) for games and emulators. Both need
a user gesture and a focused window; feature-detect with
@ -1114,12 +1147,9 @@ review blockers. Reviewers judge the JavaScript that instantiates the module
and the module's provenance instead of its bytes.
Ship and load it through your own route: register a route that returns the
module bytes (an in-process handler may return a Starlette `Response` or
`FileResponse` of any size; an out-of-process handler's body is buffered by the
host and capped at about 380 KiB of body — `_RESULT_CAP` = 512 KiB in
`ouroboros/extension_process_runner.py` bounds the base64-encoded result — so a
larger module needs an in-process skill or the runtime-download path described
under assets below), then in the widget:
module bytes with a Starlette `Response` or `FileResponse`. The same response
runs in an isolated child for dependency-bearing skills and streams without the
old serialized-result body cap. Then in the widget:
```js
const bytes = await (await OuroborosWidget.fetch('/api/extensions/<skill>/core.wasm')).arrayBuffer();
@ -1136,8 +1166,8 @@ routes. The frame CSP admits this with `'wasm-unsafe-eval'` — there is no plai
#### Assets: fonts, audio, video, images
Widget assets are ordinary payload files and travel the same way as
WebAssembly: your own routes serve them (`register_route` returning the bytes;
an out-of-process handler answers about 380 KiB of body per response, as above), the
WebAssembly: your own routes serve them (`register_route` returning a standard
response, streamed from an isolated child when required), the
widget references them by `/api/extensions/<skill>/...` URL, and review
sees each non-text asset as a content-hash-bound descriptor. The module
endpoint stays JavaScript-only. Hub packages admit `.png .jpg .jpeg .gif .webp

View file

@ -1584,12 +1584,6 @@ or error otherwise, never failing on its own.
non-manifest history with last-5 retention (history is for recovery, not a
second deliverable list). The logical `root=deliverables` tool stays
read/list/search-only and is not granted to children.
- Host chat attachment admission owns its new upload copies until enqueue succeeds.
Keep copying under the shared settled HTTP-worker wait; cancelling the waiter
first settles copying, then cleans its unaccepted destinations. Accepted inputs
and the skill's original files survive cancellation; do not treat disconnect
as task cancellation.
- Large task files use `artifacts.stream_artifact_file` and atomic
`copy_artifact_file`; do not read complete datasets into a bytes object.
HTTP admission and materialization run their complete blocking operation off
@ -2243,7 +2237,7 @@ by "Provider Independence" above. Call-site imperatives:
the one import surface. Register the env key; do not scatter magic wait
numbers across call sites (`tests/test_timeout_policy.py`).
- Worker readiness is its own bound, not a tuning knob: `WORKER_READY_WINDOW_SEC`
and `WORKER_READY_MAX_ATTEMPTS` sit in `config.py` beside the spawn grace as
and `WORKER_READY_MAX_ATTEMPTS` sit in `runtime_limits.py` (re-exported by `config.py`) beside the spawn grace as
structural constants (a warm forkserver child confirms in ~3-4 s; the window is the
pool's existing init budget). Never fold "the child confirmed ready" into
process liveness (`proc.is_alive`) or the task idle rail: a child deadlocked
@ -2683,6 +2677,24 @@ omits setup jobs with unconfirmed termination. Service quiescence excludes
zombie-only groups, but checks every member before releasing a writer fence
(`tests/test_claudexor_custody_lifetime.py`, `tests/test_process_custody_liveness.py`).
Out-of-process extension HTTP responses execute their standard Starlette ASGI
response in the child. The runner owns staging, Popen registration and cleanup;
`extension_route_stream` owns only portable pipe frames and ASGI delivery. Preserve
ordered headers, HEAD/Range and background actions. Consumer backpressure is not
an idle failure and no total/pre-header response deadline applies. Bind cancellation
to the existing loaded bundle, before spawning, and detach on completion. A
cancellation during startup retains the worker future and process context until
it exits; move context entry/exit, spawn registration, pipe shutdown and termination
off the ASGI event loop. Static captured module sources spawn no child. Chunk size
and post-response cleanup grace come from the runtime-limits owner via config.py
and are not stream deadlines. Bound each frame before reading its payload, not
the cumulative response; retain sanitized bounded stderr and the actual exit
code after draining a child that dies abnormally.
Failed final sends remain delivery failures; background failure after a successful final
body is a separate diagnostic. Widget pull credits bound transport buffering;
large URL downloads use the existing native/browser file owner, never an automatic
HTTP-stream-to-Blob conversion.
## Platform Abstraction Rule
Platform-specific code goes through `ouroboros/platform_layer.py`: platform
@ -2933,6 +2945,23 @@ or require a class where established function owners already preserve the
boundary. Enforcement: CHECKLISTS item 17 (`gateway_parity`) and
`tests/test_gateway_parity.py`.
Named skill chat ingress uses the existing message-bus canonical writer before
enqueue. Its internal callback retains request-owned upload copies once the
canonical write is attempted, even if that write fails with an unknown outcome;
replays do not adopt new copies. The existing settled HTTP-worker wait retains
copy/admission until completion under cancellation, and a local cleanup stack
removes only unaccepted destinations. Accepted attachments remain through an
unknown write/queue outcome or disconnect. Such failure may retain an unused
copy but never claims acceptance. Cancellation support and intent writes must
address the same installation root as the operation being read. The server consumes that internal accepted source without logging a
second row. Operation reads/cancel verify the complete source against actual
task or direct-turn ownership; presentation annotations are discovery hints.
Named waits use that operation's state, while unnamed legacy waits retain their
chat callback. Tests: `test_host_service_operation_identity.py` and
`test_host_service_operations.py`. Successful child WS relay failures travel as
bounded counters through the existing process-facts channel, preserving the
producer result and best-effort `None` API (`test_extension_ws_diagnostics.py`).
## Build & CI
### Python dependency locks

View file

@ -21,14 +21,14 @@ The manifest is the SSOT of the module→domain assignment (1:1, complete over t
| D11 | Gateway, server & Web UI | 49 | 0 |
| D12 | Settings & configuration | 15 | 0 |
| D13 | Safety, guards & runtime mode | 9 | 0 |
| D14 | Skills & extensions | 53 | 0 |
| D14 | Skills & extensions | 54 | 0 |
| D15 | Memory, knowledge, consciousness & self-evolution | 17 | 0 |
| D16 | Observability, usage accounting & cost | 11 | 0 |
| D17 | Projects, workspaces & task results | 20 | 0 |
| D18 | Launcher, packaging, platform & shared substrate | 12 | 0 |
| D19 | Frozen contracts (ABI) | 10 | 0 |
| D20 | Presence | 9 | 0 |
| **total** | | **519** | **0** |
| **total** | | **520** | **0** |
## Dependency direction matrix (strict, pinned)
@ -617,6 +617,7 @@ No function body (≥ 10 normalized lines) is shared verbatim across domains. Ne
- `ouroboros/extension_process_runner.py`
- `ouroboros/extension_reconcile_queue.py`
- `ouroboros/extension_registry_state.py`
- `ouroboros/extension_route_stream.py`
- `ouroboros/extension_surface_names.py`
- `ouroboros/extension_ui_validation.py`
- `ouroboros/marketplace/__init__.py`

View file

@ -2,14 +2,14 @@
AST-derived inventory of compatibility facades, regenerated by `python scripts/regenerate_inventories.py`. Do not edit. A facade row is any runtime module whose top-level `from <population module> import ...` statements carry the `noqa: F401` re-export marker — the codebase's declared "this binding exists for its binding, not for this module's own use" convention (reference FACADE_CONSUMERS method). Leaf domains come from `ouroboros/domains.toml`; a leaf outside the facade's domain is marked ✗ (that edge also appears in the manifest's pinned direction matrix). `tests/test_generated_inventories.py` pins byte-identity, so any re-export surface change must regenerate this file.
- facade modules: **55**; marked re-export bindings: **2302**; cross-domain facade→leaf pairs: **130**
- facade modules: **55**; marked re-export bindings: **2316**; cross-domain facade→leaf pairs: **130**
| facade | domain | bindings | leaves |
|---|---|---:|---|
| `launcher.py` | D18 | 2 | `ouroboros/launcher_windows_runtime.py` (2) |
| `ouroboros/agent.py` | D01 | 31 | `ouroboros/agent_dispatch.py` (14)<br>`ouroboros/agent_startup_checks.py` (4)<br>`ouroboros/config.py` (2 ✗D12)<br>`ouroboros/subagent_dispatch_notes.py` (4 ✗D07)<br>`ouroboros/subagents.py` (7 ✗D07) |
| `ouroboros/agent_task_pipeline.py` | D01 | 26 | `ouroboros/dialogue_provenance.py` (2 ✗D15)<br>`ouroboros/post_task_synthesis.py` (11)<br>`ouroboros/synthesis_cost_text.py` (5)<br>`ouroboros/task_finalization.py` (8) |
| `ouroboros/config.py` | D12 | 92 | `ouroboros/model_slots.py` (13)<br>`ouroboros/provider_models.py` (6 ✗D02)<br>`ouroboros/review_model_routes.py` (10)<br>`ouroboros/runtime_limits.py` (33)<br>`ouroboros/settings_defaults.py` (15)<br>`ouroboros/settings_scales.py` (13)<br>`ouroboros/update_channels.py` (2) |
| `ouroboros/config.py` | D12 | 106 | `ouroboros/model_slots.py` (13)<br>`ouroboros/provider_models.py` (6 ✗D02)<br>`ouroboros/review_model_routes.py` (10)<br>`ouroboros/runtime_limits.py` (47)<br>`ouroboros/settings_defaults.py` (15)<br>`ouroboros/settings_scales.py` (13)<br>`ouroboros/update_channels.py` (2) |
| `ouroboros/context.py` | D03 | 4 | `ouroboros/context_runtime_facts.py` (4) |
| `ouroboros/delegate_custody.py` | D07 | 10 | `ouroboros/delegate_custody_reconcile.py` (9)<br>`ouroboros/delegate_evidence.py` (1) |
| `ouroboros/extension_loader.py` | D14 | 99 | `ouroboros/contracts/plugin_api.py` (7 ✗D19)<br>`ouroboros/extension_child_catalog.py` (8)<br>`ouroboros/extension_companion.py` (3)<br>`ouroboros/extension_import_staging.py` (6)<br>`ouroboros/extension_isolated_deps.py` (4)<br>`ouroboros/extension_liveness.py` (8)<br>`ouroboros/extension_plugin_api.py` (6)<br>`ouroboros/extension_registry_state.py` (20)<br>`ouroboros/extension_surface_names.py` (12)<br>`ouroboros/extension_ui_validation.py` (5)<br>`ouroboros/gateway/host_service.py` (1 ✗D11)<br>`ouroboros/provider_models.py` (1 ✗D02)<br>`ouroboros/skill_loader.py` (13)<br>`ouroboros/skill_token.py` (1)<br>`ouroboros/tools/skill_exec.py` (1)<br>`ouroboros/utils.py` (3 ✗D18) |

View file

@ -85,6 +85,20 @@ from ouroboros.review_model_routes import (
resolved_review_model_target, # noqa: F401
)
from ouroboros.runtime_limits import (
WORKER_SPAWN_GRACE_SEC, # noqa: F401
WORKER_READY_WINDOW_SEC, # noqa: F401
WORKER_READY_MAX_ATTEMPTS, # noqa: F401
EXTENSION_STREAM_CHUNK_BYTES, # noqa: F401
EXTENSION_CHILD_CLEANUP_GRACE_SEC, # noqa: F401
NESTED_SETTLEMENT_MARGIN_SEC, # noqa: F401
NETWORK_WAIT_NOTE_INTERVAL_SEC, # noqa: F401
NETWORK_WAIT_BACKOFF_START_SEC, # noqa: F401
TCP_KEEPALIVE_IDLE_SEC, # noqa: F401
TCP_KEEPALIVE_INTERVAL_SEC, # noqa: F401
TCP_KEEPALIVE_PROBE_COUNT, # noqa: F401
EXTENSION_STREAM_METADATA_BYTES, # noqa: F401
WS_RELAY_BURST, # noqa: F401
WS_RELAY_REFILL_PER_SEC, # noqa: F401
DELEGATE_WAIT_CEILING_SEC, # noqa: F401
DELEGATE_WAIT_WINDOW_MAX_SEC, # noqa: F401
MAX_ACTIVE_SUBAGENTS_HARD_CAP, # noqa: F401
@ -159,28 +173,6 @@ SettingsIntegrityError = _settings_integrity.SettingsIntegrityError
RESTART_EXIT_CODE = 42
PANIC_EXIT_CODE = 99
AGENT_SERVER_PORT = 8765
NESTED_SETTLEMENT_MARGIN_SEC = 30 # Structural ordering margin, not a cognition timeout.
# Owner-note cadence while a task waits out a provider-connection outage; the effective interval is min(this, idle_timeout/2) so the notes also keep the idle rail alive.
NETWORK_WAIT_NOTE_INTERVAL_SEC = 300
# First free-redial pause of a transport-wait episode; doubles per wait iteration up to the existing 60s transient backoff cap (Q10: an existing bound, not a new knob).
NETWORK_WAIT_BACKOFF_START_SEC = 4.0
# TCP keepalive for long-lived remote LLM sockets (idle threshold, probe interval, probe count): kernel probes
# detect a silently dropped NAT/VPN mapping instead of hanging to the read timeout; platform_layer builds the options.
TCP_KEEPALIVE_IDLE_SEC = 60
TCP_KEEPALIVE_INTERVAL_SEC = 60
TCP_KEEPALIVE_PROBE_COUNT = 5
# Worker-pool spawn bounds (structural constants, not env knobs). Grace after a full-pool spawn before the crash
# detector counts dead workers (up to ~60s to init: spawn + pip); workers.py binds it as `_SPAWN_GRACE_SEC`, the extension import-staging sweep reads it too.
WORKER_SPAWN_GRACE_SEC = 90.0
# Readiness window for ONE spawned/respawned slot: unassignable until the child's own `worker_ready` row lands; alive
# but silent past this = torn down and replaced. Sized to the spawn grace (the pool's existing init budget): a warm
# forkserver child boots in ~3-4s (G13 mock lane: 3.5-4.9s startup, 2.5-3.2s respawn), a cold 4-vCPU CI runner well under 60s (its 21-scenario mock lane runs in ~80s), and the E2E
# scenarios wait 240s per task, so a wedged child is a fast, named failure. A contract distinct from process liveness
# (`proc.is_alive`, worker_health.py) and from the task idle rail (queue_timeouts.py): a deadlocked child is alive.
WORKER_READY_WINDOW_SEC = 90.0
# Consecutive readiness failures of one slot before it is parked and reported (three strikes, like the crash-storm fence).
WORKER_READY_MAX_ATTEMPTS = 3
# --- Usage-ledger compaction policy (CPL4-C6, owner sanction 1A) -------------
# docs/v7next/DESIGN_USAGE_COMPACTION.md. Constants, not env knobs. Compact the
# monetary ledger once its byte size reaches ~0.2s-per-cold-replay scale, well

View file

@ -147,6 +147,7 @@ D20 = "Presence"
"ouroboros/extension_registry_state.py" = "D14"
"ouroboros/extension_surface_names.py" = "D14"
"ouroboros/extension_process_runner.py" = "D14"
"ouroboros/extension_route_stream.py" = "D14"
"ouroboros/extension_reconcile_queue.py" = "D14"
"ouroboros/extension_ui_validation.py" = "D14"
"ouroboros/fallback_cooldown.py" = "D02"

View file

@ -20,6 +20,7 @@ import pathlib
import secrets
import sys
import threading
import urllib.error
import urllib.request
import uuid
from typing import Any, Callable, Dict, Optional, Sequence
@ -181,6 +182,21 @@ def mint_skill_token(state_dir: pathlib.Path, skill_name: str, skill_dir: Option
_ws_broadcaster: Optional[Callable[[dict], None]] = None
_child_ws_relay_failures: Dict[str, int] = {}
def take_child_ws_relay_failures() -> Dict[str, int]:
"""Drain this per-call child's fixed-category aggregate, including its threads."""
with _lock:
failures = dict(_child_ws_relay_failures)
_child_ws_relay_failures.clear()
return failures
def _record_ws_relay_failure(reason: str) -> None:
# The per-call child owns this aggregate; no message/URL/exception text is kept.
with _lock:
_child_ws_relay_failures[reason] = _child_ws_relay_failures.get(reason, 0) + 1
def set_ws_broadcaster(broadcaster: Callable[[dict], None] | None) -> None:
@ -646,7 +662,7 @@ class PluginAPIImpl:
base_url = (os.environ.get("HOST_SERVICE_URL") or "").strip()
token = (os.environ.get("HOST_SERVICE_TOKEN") or "").strip()
if not base_url or not token:
log.debug("extension %s dropped WS message %s: no host bridge env", self._skill, short)
_record_ws_relay_failure("missing_transport")
return
body = json.dumps({"message_type": short, "data": data}).encode("utf-8")
request = urllib.request.Request(
@ -658,8 +674,17 @@ class PluginAPIImpl:
try:
with urllib.request.urlopen(request, timeout=2): # noqa: S310 - loopback Host Service
return
except urllib.error.HTTPError as exc:
reason = "rate_limited" if exc.code == 429 else (
"http_client_error" if 400 <= exc.code < 500 else
"http_server_error" if 500 <= exc.code < 600 else "http_error")
_record_ws_relay_failure(reason)
try:
exc.close()
except OSError:
pass
except Exception:
log.debug("extension %s host WS relay failed for %s", self._skill, short, exc_info=True)
_record_ws_relay_failure("transport_error")
def on_unload(self, callback: Callable[[], Any]) -> None:
_reject_extension_child_side_effect("on_unload")

View file

@ -2,11 +2,15 @@
The host process may safely catalog and dispatch extensions whose isolated
dependencies include native wheels: plugin import and handler execution happen
in a short-lived child process, so Rust/C aborts cannot take down server.py.
in a per-call child process, so Rust/C aborts cannot take down server.py.
HTTP response children live through streaming and cleanup; one-shot calls keep
their existing timeout and result-size contracts.
"""
from __future__ import annotations
from contextlib import contextmanager
import asyncio
import base64
import inspect
@ -26,8 +30,9 @@ from types import SimpleNamespace
from typing import Any, Callable, Dict, List
from starlette.requests import Request
from starlette.responses import FileResponse, Response, StreamingResponse
from starlette.responses import JSONResponse, Response
from ouroboros.config import EXTENSION_CHILD_CLEANUP_GRACE_SEC
from ouroboros.provider_models import MODEL_PROVIDER_CREDENTIAL_KEYS
from ouroboros.platform_layer import (
merge_hidden_kwargs, posix_signal_name, subprocess_new_group_kwargs,
@ -265,6 +270,8 @@ def _child_env(
granted_keys: List[str],
) -> Dict[str, str]:
env = _scrub_env(env_allowlist, skill_state_dir(drive_root, skill_name), skill_name, granted_keys=granted_keys)
env["OUROBOROS_APP_ROOT"] = str(pathlib.Path(drive_root).parent)
env["OUROBOROS_SETTINGS_PATH"] = str(pathlib.Path(drive_root) / "settings.json")
env["OUROBOROS_DATA_DIR"] = str(drive_root)
env["OUROBOROS_REPO_DIR"] = str(repo_dir)
env["OUROBOROS_EXTENSION_PROCESS_CHILD"] = "1"
@ -298,7 +305,7 @@ def _child_env(
return env
def _drain(pipe: Any, cap: int, out: bytearray, overflow: Dict[str, bool], label: str) -> None:
def _drain(pipe: Any, cap: int, out: bytearray, overflow: Dict[str, bool], label: str, *, discard_excess: bool = False) -> None:
try:
while True:
chunk = pipe.read(4096)
@ -307,11 +314,14 @@ def _drain(pipe: Any, cap: int, out: bytearray, overflow: Dict[str, bool], label
remaining = cap - len(out)
if remaining <= 0:
overflow[label] = True
if discard_excess:
continue
return
out.extend(chunk[:remaining])
if len(chunk) > remaining:
overflow[label] = True
return
if not discard_excess:
return
except (OSError, ValueError):
return
@ -392,7 +402,7 @@ def _mark_child_spawned(exc: BaseException) -> BaseException:
def _publish_child_facts(
proc: Any, started_ts: float, *, timed_out: bool = False,
killed_by_host: bool = False,
killed_by_host: bool = False, ws_relay_failures=None, skill_name: str = "",
) -> None:
"""Publish the extension child's typed process facts to the call's channel.
@ -406,26 +416,26 @@ def _publish_child_facts(
plus the kill facts, which is the whole truth on Windows too.
"""
try:
publish_process_facts(
facts = publish_process_facts(
returncode=proc.poll(),
started_ts=started_ts,
timed_out=timed_out,
killed_by_host=killed_by_host,
ws_relay_failures=ws_relay_failures,
)
if facts.get("ws_relay_failures"):
log.warning("extension child WS relay failures for %s (best effort): %s", skill_name, facts["ws_relay_failures"])
except Exception: # never replace an in-flight child failure
log.debug("extension child process facts could not be published", exc_info=True)
def _run_child(
payload: Dict[str, Any],
*,
skill_dir: pathlib.Path,
drive_root: pathlib.Path,
repo_dir: pathlib.Path,
env: Dict[str, str],
timeout_sec: int,
on_spawn: Callable[[], None] | None = None,
) -> Dict[str, Any]:
@contextmanager
def _child_process(
payload: Dict[str, Any], *, skill_dir: pathlib.Path, drive_root: pathlib.Path,
repo_dir: pathlib.Path, env: Dict[str, str],
on_spawn: Callable[[], None] | None = None, stream: bool = False,
):
"""One staging, spawn, registration and cleanup owner for every child mode."""
calls_dir = skill_state_dir(drive_root, str(payload.get("skill_name") or "")) / "extension_calls"
_ensure_private_dir(calls_dir)
input_path = calls_dir / f"{uuid.uuid4().hex}.json"
@ -445,12 +455,19 @@ def _run_child(
kwargs: Dict[str, Any] = {
"cwd": str(repo_dir),
"env": env,
"stdin": subprocess.DEVNULL,
"stdin": subprocess.PIPE if stream else subprocess.DEVNULL,
"stdout": subprocess.PIPE,
"stderr": subprocess.PIPE,
}
kwargs.update(subprocess_new_group_kwargs())
proc = subprocess.Popen(cmd, **merge_hidden_kwargs(kwargs)) # noqa: S603 - argv is host-constructed
try:
proc = subprocess.Popen(cmd, **merge_hidden_kwargs(kwargs)) # noqa: S603 - argv is host-constructed
except BaseException:
try:
input_path.unlink(missing_ok=True)
except OSError:
pass
raise
child_started_ts = time.monotonic()
# ABI-9 spawn boundary: from here on the child EXISTS. Every exception
# leaving this function — process registration, the on_spawn durable
@ -460,11 +477,62 @@ def _run_child(
try:
with _subprocess_lock:
_active_subprocesses.add(proc)
if stream:
from ouroboros.process_custody import record_process
record_process(drive_root, pid=proc.pid, cmd=cmd,
purpose=f"extension_route:{payload.get('skill_name')}", scope="session")
if on_spawn is not None:
# Popen has already dispatched the child. A failed durable
# disclosure must stop it rather than leave an untracked external
# execution: the finally block below kills and reaps the child.
on_spawn()
yield SimpleNamespace(proc=proc, result_path=result_path, started_ts=child_started_ts)
except BaseException as exc:
# Everything in this block runs AFTER Popen: the typed marker lets the
# dispatcher stamp physical_dispatch on real post-spawn failures only.
raise _mark_child_spawned(exc)
finally:
try:
try:
if proc.poll() is None:
_kill_process_group(proc)
proc.wait(timeout=EXTENSION_CHILD_CLEANUP_GRACE_SEC)
except Exception:
pass
with _subprocess_lock:
_active_subprocesses.discard(proc)
try:
input_path.unlink(missing_ok=True)
except OSError:
pass
try:
result_path.unlink(missing_ok=True)
except OSError:
pass
shutil.rmtree(import_root_base, ignore_errors=True)
for pipe in (proc.stdin, proc.stdout, proc.stderr):
try:
if pipe:
pipe.close()
except OSError:
pass
except BaseException as cleanup_exc:
# A cleanup failure must never REPLACE a marked in-flight
# exception with an unmarked one (nor itself escape unmarked on
# the success path): the replacing exception carries the marker
# too — the child really did run.
raise _mark_child_spawned(cleanup_exc)
def _run_child(
payload: Dict[str, Any], *, skill_dir: pathlib.Path, drive_root: pathlib.Path,
repo_dir: pathlib.Path, env: Dict[str, str], timeout_sec: int,
on_spawn: Callable[[], None] | None = None,
) -> Dict[str, Any]:
with _child_process(payload, skill_dir=skill_dir, drive_root=drive_root,
repo_dir=repo_dir, env=env, on_spawn=on_spawn) as child:
proc, result_path = child.proc, child.result_path
child_started_ts = child.started_ts
stdout = bytearray()
stderr = bytearray()
overflow = {"stdout": False, "stderr": False}
@ -488,8 +556,8 @@ def _run_child(
failure_kind="timeout",
)
time.sleep(0.05)
out_thread.join(timeout=2)
err_thread.join(timeout=2)
out_thread.join(timeout=EXTENSION_CHILD_CLEANUP_GRACE_SEC)
err_thread.join(timeout=EXTENSION_CHILD_CLEANUP_GRACE_SEC)
_publish_child_facts(proc, child_started_ts)
if proc.returncode != 0:
stderr_text = stderr.decode("utf-8", errors="replace").strip()
@ -505,44 +573,11 @@ def _run_child(
result = json.loads(result_path.read_text(encoding="utf-8") or "{}")
except json.JSONDecodeError as exc:
raise ExtensionProcessError(f"extension child returned invalid JSON: {exc}") from exc
_publish_child_facts(proc, child_started_ts, ws_relay_failures=result.get("ws_relay_failures"),
skill_name=str(payload.get("skill_name") or ""))
if not result.get("ok", False):
raise ExtensionProcessError(str(result.get("error") or "extension child failed"))
return dict(result)
except BaseException as exc:
# Everything in this block runs AFTER Popen: the typed marker lets the
# dispatcher stamp physical_dispatch on real post-spawn failures only.
raise _mark_child_spawned(exc)
finally:
try:
try:
if proc.poll() is None:
_kill_process_group(proc)
proc.wait(timeout=2)
except Exception:
pass
with _subprocess_lock:
_active_subprocesses.discard(proc)
try:
input_path.unlink(missing_ok=True)
except OSError:
pass
try:
result_path.unlink(missing_ok=True)
except OSError:
pass
shutil.rmtree(import_root_base, ignore_errors=True)
for pipe in (proc.stdout, proc.stderr):
try:
if pipe:
pipe.close()
except OSError:
pass
except BaseException as cleanup_exc:
# A cleanup failure must never REPLACE a marked in-flight
# exception with an unmarked one (nor itself escape unmarked on
# the success path): the replacing exception carries the marker
# too — the child really did run.
raise _mark_child_spawned(cleanup_exc)
def _record_extension_dispatch(
@ -724,13 +759,16 @@ def dispatch_extension_tool_subprocess(ext_tool: Dict[str, Any], ctx: ToolContex
return str(result.get("result") or "")
def dispatch_extension_route_subprocess(spec: Dict[str, Any], request_payload: Dict[str, Any], *, drive_root: pathlib.Path, repo_dir: pathlib.Path) -> Dict[str, Any]:
def dispatch_extension_route_subprocess(spec: Dict[str, Any], request_payload: Dict[str, Any], *, drive_root: pathlib.Path, repo_dir: pathlib.Path) -> Response:
skills_repo_path = pathlib.Path(str(spec.get("skills_repo_path") or repo_dir))
skill = _skill_for_dispatch(str(spec.get("skill") or ""), pathlib.Path(drive_root), skills_repo_path)
env = _base_env_for_skill(skill, pathlib.Path(drive_root), pathlib.Path(repo_dir))
model_capable = _extension_has_model_credentials(skill, pathlib.Path(drive_root))
dispatch_id = f"extension:route:{uuid.uuid4().hex}"
return _run_child(
from functools import partial
from ouroboros.extension_route_stream import RouteStreamResponse
child_factory = partial(_child_process,
{
"mode": "route",
"skill_name": skill.name,
@ -744,7 +782,7 @@ def dispatch_extension_route_subprocess(spec: Dict[str, Any], request_payload: D
drive_root=pathlib.Path(drive_root),
repo_dir=pathlib.Path(repo_dir),
env=env,
timeout_sec=max(1, int(spec.get("timeout_sec") or 60)),
stream=True,
on_spawn=((
lambda: _record_extension_dispatch(
dispatch_id=dispatch_id,
@ -756,6 +794,8 @@ def dispatch_extension_route_subprocess(spec: Dict[str, Any], request_payload: D
) if model_capable else None),
)
return RouteStreamResponse(spec, child_factory)
def dispatch_extension_ws_subprocess(spec: Dict[str, Any], msg: Dict[str, Any], *, drive_root: pathlib.Path, repo_dir: pathlib.Path) -> Any:
skills_repo_path = pathlib.Path(str(spec.get("skills_repo_path") or repo_dir))
@ -798,7 +838,7 @@ def _skill_for_dispatch(skill_name: str, drive_root: pathlib.Path, skills_repo_p
return skill
def _load_child_extension(skill_name: str, drive_root: pathlib.Path, repo_dir: pathlib.Path, skills_repo_path: pathlib.Path) -> None:
def _load_child_extension(skill_name: str, drive_root: pathlib.Path, repo_dir: pathlib.Path, skills_repo_path: pathlib.Path) -> Any:
from ouroboros.config import load_settings
from ouroboros.extension_loader import load_extension
from ouroboros.skill_loader import discover_skills
@ -817,6 +857,7 @@ def _load_child_extension(skill_name: str, drive_root: pathlib.Path, repo_dir: p
)
if err:
raise ExtensionProcessError(err)
return skill
def _surface_catalog() -> Dict[str, Any]:
@ -908,13 +949,15 @@ def _handler_wants_ctx(handler: Any) -> bool:
return False
async def _request_from_payload(payload: Dict[str, Any], drive_root: pathlib.Path, repo_dir: pathlib.Path) -> Request:
async def _request_from_payload(payload: Dict[str, Any], drive_root: pathlib.Path, repo_dir: pathlib.Path, disconnect=None) -> Request:
body = base64.b64decode(str(payload.get("body_b64") or ""))
sent = False
async def receive() -> Dict[str, Any]:
nonlocal sent
if sent:
if disconnect is not None:
return await disconnect()
return {"type": "http.request", "body": b"", "more_body": False}
sent = True
return {"type": "http.request", "body": body, "more_body": False}
@ -938,103 +981,25 @@ async def _request_from_payload(payload: Dict[str, Any], drive_root: pathlib.Pat
return Request(scope, receive)
async def _response_to_payload(response: Response, scope: Dict[str, Any]) -> Dict[str, Any]:
started: Dict[str, Any] = {
"status_code": int(getattr(response, "status_code", 200) or 200),
"headers": dict(getattr(response, "headers", {}) or {}),
}
body = bytearray()
received = False
async def receive() -> Dict[str, Any]:
nonlocal received
if received:
return {"type": "http.disconnect"}
received = True
return {"type": "http.request", "body": b"", "more_body": False}
async def send(message: Dict[str, Any]) -> None:
msg_type = str(message.get("type") or "")
if msg_type == "http.response.start":
started["status_code"] = int(message.get("status") or started["status_code"])
headers = {}
for raw_key, raw_value in message.get("headers") or []:
key = bytes(raw_key).decode("latin-1", errors="ignore")
value = bytes(raw_value).decode("latin-1", errors="ignore")
if key.lower() != "content-length":
headers[key] = value
started["headers"] = headers
elif msg_type == "http.response.body":
chunk = bytes(message.get("body") or b"")
if len(body) + len(chunk) > _RESULT_CAP:
raise ExtensionProcessError("extension child response body exceeded safety cap")
body.extend(chunk)
await response(scope, receive, send)
headers = dict(started.get("headers") or {})
headers.pop("content-length", None)
return {
"kind": "response",
"status_code": int(started.get("status_code") or 200),
"headers": headers,
"media_type": getattr(response, "media_type", None),
"body_b64": base64.b64encode(bytes(body)).decode("ascii"),
}
async def _streaming_response_to_payload(response: StreamingResponse) -> Dict[str, Any]:
body = bytearray()
async for chunk in response.body_iterator:
if isinstance(chunk, str):
chunk = chunk.encode(getattr(response, "charset", "utf-8") or "utf-8")
chunk = bytes(chunk or b"")
if len(body) + len(chunk) > _RESULT_CAP:
raise ExtensionProcessError("extension child response body exceeded safety cap")
body.extend(chunk)
headers = dict(response.headers)
headers.pop("content-length", None)
return {
"kind": "response",
"status_code": int(response.status_code),
"headers": headers,
"media_type": getattr(response, "media_type", None),
"body_b64": base64.b64encode(bytes(body)).decode("ascii"),
}
def _file_response_to_payload(response: FileResponse) -> Dict[str, Any]:
path = pathlib.Path(response.path)
if path.stat().st_size > _RESULT_CAP:
raise ExtensionProcessError("extension child response body exceeded safety cap")
headers = dict(response.headers)
headers.pop("content-length", None)
return {
"kind": "response",
"status_code": int(response.status_code),
"headers": headers,
"media_type": getattr(response, "media_type", None),
"body_b64": base64.b64encode(path.read_bytes()).decode("ascii"),
}
async def _call_route(surface: str, request_payload: Dict[str, Any], drive_root: pathlib.Path, repo_dir: pathlib.Path) -> Dict[str, Any]:
async def _call_route(surface: str, request_payload: Dict[str, Any], drive_root: pathlib.Path,
repo_dir: pathlib.Path, skill: Any, channel: Any) -> None:
from ouroboros.extension_loader import list_routes
from ouroboros.extension_isolated_deps import async_isolated_site_dirs_scope
from ouroboros.skill_dependencies import auto_install_specs_for_skill
channel.bind()
spec = list_routes().get(surface)
if not spec or not callable(spec.get("handler")):
raise ExtensionProcessError(f"extension route {surface!r} is not registered")
request = await _request_from_payload(request_payload, drive_root, repo_dir)
result = spec["handler"](request)
result = await _run_maybe_async(result)
if isinstance(result, FileResponse):
return _file_response_to_payload(result)
if isinstance(result, StreamingResponse):
return await _streaming_response_to_payload(result)
if isinstance(result, Response):
return await _response_to_payload(result, request.scope)
if isinstance(result, (dict, list)):
return {"kind": "json", "status_code": 200, "data": _json_safe(result)}
return {"kind": "text", "status_code": 200, "text": str(result)}
request = await _request_from_payload(request_payload, drive_root, repo_dir, channel.receive)
result = await _run_maybe_async(spec["handler"](request))
if not isinstance(result, Response):
result = JSONResponse(_json_safe(result)) if isinstance(result, (dict, list)) else Response(str(result))
# Handler scope ended before its lazy body/background runs. The child owns
# this skill's response scope too, so delayed dependency imports still work.
async with async_isolated_site_dirs_scope(skill.skill_dir,
enabled=bool(auto_install_specs_for_skill(drive_root, skill))):
await result(request.scope, request.receive, channel.send)
async def _call_ws(surface: str, msg: Dict[str, Any]) -> Any:
@ -1047,36 +1012,62 @@ async def _call_ws(surface: str, msg: Dict[str, Any]) -> Any:
def _child_main(input_path: str) -> None:
from ouroboros.extension_plugin_api import take_child_ws_relay_failures
payload = json.loads(pathlib.Path(input_path).read_text(encoding="utf-8"))
drive_root = pathlib.Path(payload["drive_root"])
repo_dir = pathlib.Path(payload["repo_dir"])
skills_repo_path = pathlib.Path(payload.get("skills_repo_path") or repo_dir)
skill_name = str(payload["skill_name"])
take_child_ws_relay_failures()
mode = str(payload.get("mode") or "")
channel = None
if mode == "route":
from ouroboros.extension_route_stream import ChildResponseChannel
channel = ChildResponseChannel()
_bootstrap_quiet_child_crash_reporting()
try:
_load_child_extension(skill_name, drive_root, repo_dir, skills_repo_path)
mode = str(payload.get("mode") or "")
skill = _load_child_extension(skill_name, drive_root, repo_dir, skills_repo_path)
if mode == "catalog":
result = _surface_catalog()
elif mode == "tool":
result = {"result": _json_safe(asyncio.run(_call_tool(str(payload.get("surface") or ""), dict(payload.get("args") or {}), drive_root, repo_dir, dict(payload.get("ctx") or {}))))}
elif mode == "route":
result = {"route": asyncio.run(_call_route(str(payload.get("surface") or ""), dict(payload.get("request") or {}), drive_root, repo_dir))}
asyncio.run(_call_route(str(payload.get("surface") or ""), dict(payload.get("request") or {}), drive_root, repo_dir, skill, channel))
elif mode == "ws":
result = {"result": _json_safe(asyncio.run(_call_ws(str(payload.get("surface") or ""), dict(payload.get("message") or {}))))}
else:
raise ExtensionProcessError(f"unknown extension child mode {mode!r}")
_write_child_result(payload, {"ok": True, **result})
if channel is None:
outcome = {"ok": True, **result}
except BaseException as exc:
_write_child_result(payload, {"ok": False, "error": sanitize_tool_result_for_log(f"{type(exc).__name__}: {exc}")})
raise SystemExit(0)
if channel is not None:
try:
channel.error(exc)
except OSError:
pass
else:
outcome = {"ok": False, "error": sanitize_tool_result_for_log(f"{type(exc).__name__}: {exc}")}
finally:
try:
from ouroboros.extension_loader import unload_extension
unload_extension(skill_name)
except Exception:
pass
except Exception as exc:
if channel is not None:
try:
channel.error(exc)
except OSError:
pass
failures = take_child_ws_relay_failures()
if channel is not None:
try:
channel.finish(failures)
except OSError:
pass
else:
if failures:
outcome["ws_relay_failures"] = failures
_write_child_result(payload, outcome)
if __name__ == "__main__":

View file

@ -14,6 +14,7 @@ import hashlib
import json
import pathlib
import threading
from contextlib import contextmanager
from dataclasses import dataclass, field
from types import ModuleType
from typing import Any, Callable, Dict, List, Optional, Sequence
@ -302,3 +303,24 @@ def get_tool_stamped(name: str) -> Optional[Dict[str, Any]]:
def _record_companion_name(bundle: _ExtensionRegistrations, name: str) -> None:
if name not in bundle.companion_names:
bundle.companion_names.append(name)
@contextmanager
def extension_work_scope(spec: Dict[str, Any], work: Any):
"""Attach cancellable work to the exact publication already owning dispatch."""
from ouroboros.contracts.plugin_api import ExtensionRegistrationError
name = str(spec.get("skill") or "")
generation = str(spec.get("extension_generation") or "")
with _lock:
bundle = _extensions.get(name)
if bundle is None or name in _unloading or bundle.generation_digest != generation:
raise ExtensionRegistrationError("extension generation changed before response dispatch")
bundle.supervised_futures.append(work)
try:
yield
finally:
with _lock:
# Captured bundle identity: completing old work cannot detach a new
# generation's resource after unload/reload replaced the registry row.
bundle.supervised_futures[:] = [item for item in bundle.supervised_futures if item is not work]

View file

@ -0,0 +1,374 @@
"""ASGI response frames over the extension child's existing portable stdio pipes.
The runner owns process staging/custody; this module owns response delivery and
backpressure. No response lifetime timer or durable stream registry is involved.
"""
from __future__ import annotations
import asyncio
import concurrent.futures
import contextlib
import json
import logging
import os
import struct
import sys
import threading
from starlette.responses import JSONResponse, Response
from ouroboros.utils import sanitize_tool_result_for_log
from ouroboros.config import (
EXTENSION_STREAM_CHUNK_BYTES, EXTENSION_STREAM_METADATA_BYTES,
EXTENSION_CHILD_CLEANUP_GRACE_SEC,
)
log = logging.getLogger(__name__)
def _write_frame(stream, kind: bytes, payload: bytes = b"") -> None:
frame = kind + payload
for value in (struct.pack("!I", len(frame)), frame):
remaining = memoryview(value)
while remaining:
written = stream.write(remaining)
if not written:
raise OSError("extension response channel stopped accepting bytes")
remaining = remaining[written:]
stream.flush()
def _read_exact(stream, size: int) -> bytes:
parts = bytearray()
while len(parts) < size:
chunk = stream.read(size - len(parts))
if not chunk:
raise EOFError("extension response channel closed")
parts.extend(chunk)
return bytes(parts)
def _read_frame(stream):
size = struct.unpack("!I", _read_exact(stream, 4))[0]
if size < 1:
raise ValueError("empty extension response frame")
kind = _read_exact(stream, 1)
payload_limit = EXTENSION_STREAM_CHUNK_BYTES + 1 if kind == b"B" else EXTENSION_STREAM_METADATA_BYTES
if size - 1 > payload_limit:
raise ValueError("extension response frame exceeds its channel bound")
return kind, _read_exact(stream, size - 1)
class ChildResponseChannel:
"""Child ASGI send/receive: reserve stdout before loading plugin code."""
def __init__(self):
import sys
sys.stdout.flush()
self.output = os.fdopen(os.dup(sys.stdout.fileno()), "wb", buffering=0)
os.dup2(sys.stderr.fileno(), sys.stdout.fileno())
self.input = os.fdopen(os.dup(sys.stdin.fileno()), "rb", buffering=0)
self.body_complete = False
self.started = False
self.disconnected = None
self.control_thread = None
def bind(self):
loop = asyncio.get_running_loop()
self.disconnected = asyncio.Event()
def control():
try:
_read_frame(self.input)
except (EOFError, OSError, ValueError):
pass
with contextlib.suppress(RuntimeError):
loop.call_soon_threadsafe(self.disconnected.set)
self.control_thread = threading.Thread(target=control, daemon=True, name="extension-route-control")
self.control_thread.start()
async def receive(self):
await self.disconnected.wait()
return {"type": "http.disconnect"}
async def send(self, message):
if self.disconnected.is_set():
raise OSError("extension response client disconnected")
kind = message["type"]
if kind == "http.response.start":
payload = json.dumps({
"status": message["status"],
"headers": [[bytes(k).decode("latin-1"), bytes(v).decode("latin-1")]
for k, v in message.get("headers", [])],
}).encode()
await asyncio.to_thread(_write_frame, self.output, b"S", payload)
self.started = True
elif kind == "http.response.body":
body = memoryview(message.get("body") or b"")
more = bool(message.get("more_body", False))
for start in range(0, max(1, len(body)), EXTENSION_STREAM_CHUNK_BYTES):
end = min(start + EXTENSION_STREAM_CHUNK_BYTES, len(body))
payload = bytes([int(more or end < len(body))]) + body[start:end].tobytes()
await asyncio.to_thread(_write_frame, self.output, b"B", payload)
self.body_complete = not more
else:
raise ValueError(f"unsupported extension response message {kind!r}")
def error(self, exc):
message = sanitize_tool_result_for_log(f"{type(exc).__name__}: {exc}")
_write_frame(self.output, b"C" if self.body_complete else b"E", message.encode("utf-8"))
def finish(self, ws_relay_failures=None):
payload = json.dumps({"ws_relay_failures": ws_relay_failures}).encode() if ws_relay_failures else b""
_write_frame(self.output, b"X", payload)
self.output.close()
# The control thread is daemon-owned and ends with this per-call child.
# Closing stdin here could block on its pending read's file-object lock.
class RouteStreamResponse(Response):
"""A response whose handle belongs to one currently published skill bundle."""
def __init__(self, spec, child_factory):
super().__init__(status_code=200)
self.spec = spec
self.child_factory = child_factory
self.cancelled = False
self.task = self.loop = None
self.body_complete = False
self.started = False
self.client_disconnected = False
self._wire_expected = None
self._wire_sent = 0
self._finishing_send = False
self._head = False
self.ws_relay_failures = None
def cancel(self):
if self.cancelled:
return
self.cancelled = True
if self.loop is not None and self.task is not None:
with contextlib.suppress(RuntimeError):
self.loop.call_soon_threadsafe(self.task.cancel)
def _wire_complete(self):
return self.started and self._wire_expected is not None and self._wire_sent >= self._wire_expected
async def _wait_worker(self, callback, *args):
"""Keep custody of blocking work until it settles, even after cancellation."""
future = asyncio.create_task(asyncio.to_thread(callback, *args))
cancelled = False
while not future.done():
try:
await asyncio.shield(future)
except asyncio.CancelledError:
cancelled = True
except Exception:
if not cancelled:
raise
break
if cancelled:
try:
future.result()
except Exception as exc:
log.warning("extension response work failed while cancellation settled: %s", exc)
raise asyncio.CancelledError
return future.result()
@contextlib.asynccontextmanager
async def _child_scope(self):
"""The existing process context stays owned while startup runs off-loop."""
from ouroboros.extension_registry_state import extension_work_scope
stack = contextlib.ExitStack()
def start():
stack.enter_context(extension_work_scope(self.spec, self))
if self.cancelled:
raise asyncio.CancelledError
return stack.enter_context(self.child_factory())
try:
yield await self._wait_worker(start)
finally:
# Startup is settled before _wait_worker propagates cancellation;
# this stack therefore still owns any process that actually spawned.
await self._wait_worker(stack.__exit__, *sys.exc_info())
async def _send_to_client(self, send, message):
try:
await send(message)
except OSError:
# HTTP clients may close after Content-Length bytes (or HEAD
# headers), before FileResponse's final empty ASGI body. Preserve
# the child's background work when the wire body is already done.
if self._wire_complete() and (self._head or not message.get("body")):
return
self.client_disconnected = True
raise
async def _pump(self, frames, send):
from ouroboros.extension_process_runner import ExtensionProcessError
while True:
kind, payload = await frames.get()
if kind == b"S":
fields = json.loads(payload)
self.status_code = int(fields["status"])
self.raw_headers = [(k.encode("latin-1"), v.encode("latin-1"))
for k, v in fields["headers"]]
lengths = [v for k, v in self.raw_headers if k.lower() == b"content-length"]
try:
self._wire_expected = int(lengths[0]) if lengths and all(value.strip().isdigit() for value in lengths) else None
if self._wire_expected is not None and any(int(value) != self._wire_expected for value in lengths):
self._wire_expected = None
except ValueError:
self._wire_expected = None
if self._head or self.status_code in {204, 205, 304}:
self._wire_expected = 0
self._finishing_send = self._wire_expected == 0
await self._send_to_client(send, {"type": "http.response.start", "status": self.status_code,
"headers": self.raw_headers})
self.started = True
self._finishing_send = False
elif kind == b"B":
if not self.started or not payload:
raise ExtensionProcessError("extension body preceded response headers")
final = payload[0] == 0
self._finishing_send = final or (
self._wire_expected is not None and self._wire_sent + len(payload) - 1 >= self._wire_expected)
try:
await self._send_to_client(send, {"type": "http.response.body", "body": payload[1:],
"more_body": not final})
except BaseException:
# A failed final send is not a completed delivery. The flag
# only suppresses the server's normal post-response disconnect.
self.body_complete = False
raise
self._wire_sent += len(payload) - 1
self.body_complete = final
self._finishing_send = False
elif kind in {b"E", b"C"}:
detail = payload.decode("utf-8", errors="replace")
if self.body_complete:
log.warning("extension response cleanup failed after delivery: %s", detail)
if kind == b"E":
return
else:
raise ExtensionProcessError(detail)
elif kind == b"X":
if not self.body_complete:
raise ExtensionProcessError("extension response ended before its final body")
self.ws_relay_failures = json.loads(payload).get("ws_relay_failures") if payload else None
return
else:
raise ExtensionProcessError("invalid extension response frame")
async def __call__(self, scope, receive, send):
from ouroboros.extension_process_runner import (
_drain, _STDERR_CAP, _publish_child_facts, _format_child_returncode,
)
self.loop, self.task = asyncio.get_running_loop(), asyncio.current_task()
self._head = str(scope.get("method") or "GET").upper() == "HEAD"
if self.cancelled:
raise asyncio.CancelledError
try:
async with self._child_scope() as child:
proc = child.proc
frames = asyncio.Queue(maxsize=1)
stop_reader = threading.Event()
submitted = []
reader_lock = threading.Lock()
stderr, overflow = bytearray(), {"stderr": False}
def reader():
while not stop_reader.is_set():
try:
frame = _read_frame(proc.stdout)
except (EOFError, OSError, ValueError) as exc:
frame = (b"E", str(exc).encode())
with reader_lock:
if stop_reader.is_set():
return
future = asyncio.run_coroutine_threadsafe(frames.put(frame), self.loop)
submitted[:] = [future]
try:
future.result()
except concurrent.futures.CancelledError:
return
if frame[0] in {b"X", b"E"}:
return
read_thread = threading.Thread(target=reader, daemon=True, name="extension-route-response")
err_thread = threading.Thread(target=_drain,
args=(proc.stderr, _STDERR_CAP, stderr, overflow, "stderr"),
kwargs={"discard_excess": True}, daemon=True, name="extension-route-stderr")
read_thread.start()
err_thread.start()
async def disconnect():
while True:
if (await receive())["type"] == "http.disconnect":
return
await asyncio.sleep(0)
pump_task = asyncio.create_task(self._pump(frames, send))
disconnect_task = asyncio.create_task(disconnect())
try:
done, _ = await asyncio.wait({pump_task, disconnect_task}, return_when=asyncio.FIRST_COMPLETED)
if disconnect_task in done and not (
self.body_complete or self._wire_complete() or self._finishing_send
):
self.client_disconnected = True
return
await pump_task
finally:
with reader_lock:
stop_reader.set()
for future in submitted:
future.cancel()
for task in (pump_task, disconnect_task):
task.cancel()
await asyncio.gather(pump_task, disconnect_task, return_exceptions=True)
# A cancellation grants the child's existing ASGI cleanup a
# short exit opportunity; this is not a response timer.
killed_by_host = False
def cleanup():
nonlocal killed_by_host
with contextlib.suppress(OSError, ValueError, BrokenPipeError):
_write_frame(proc.stdin, b"D")
try:
proc.wait(timeout=EXTENSION_CHILD_CLEANUP_GRACE_SEC)
except Exception:
from ouroboros.tools.shell import _kill_process_group
_kill_process_group(proc)
killed_by_host = True
with contextlib.suppress(Exception):
proc.wait(timeout=EXTENSION_CHILD_CLEANUP_GRACE_SEC)
read_thread.join(EXTENSION_CHILD_CLEANUP_GRACE_SEC)
err_thread.join(EXTENSION_CHILD_CLEANUP_GRACE_SEC)
try:
await self._wait_worker(cleanup)
finally:
# Keep the existing thread-local process-fact publisher.
_publish_child_facts(proc, child.started_ts, killed_by_host=killed_by_host,
ws_relay_failures=self.ws_relay_failures,
skill_name=str(self.spec.get("skill") or ""))
if proc.returncode not in (None, 0) and not killed_by_host:
detail = sanitize_tool_result_for_log(stderr.decode("utf-8", errors="replace").strip())[-2000:]
log.warning("extension route child exited abnormally: %s; %s",
_format_child_returncode(proc.returncode), detail)
if overflow["stderr"]:
log.warning("extension route diagnostic output was truncated")
except asyncio.CancelledError:
raise
except Exception as exc:
if self.client_disconnected:
return
if self.body_complete:
log.warning("extension response cleanup failed after delivery: %s", exc)
elif self.started:
raise
else:
self.status_code = 502
await JSONResponse({"error": f"{type(exc).__name__}: {exc}"}, status_code=502)(scope, receive, send)

View file

@ -110,6 +110,7 @@ class ChatOutbound(TypedDict):
markdown: NotRequired[bool]
is_progress: NotRequired[bool]
task_id: NotRequired[str]
origin_message_ref: NotRequired[Dict[str, Any]]
# X3: a repair receipt whose managed task id does not exist yet (the router
# mints it at promotion). Typed truth instead of an invented id.
task_id_pending: NotRequired[bool]

View file

@ -635,22 +635,7 @@ async def api_extension_dispatch(request: Request) -> Response:
drive_root=drive_root,
repo_dir=_request_repo_dir(request),
)
route_result = dict(child_result.get("route") or {})
kind = str(route_result.get("kind") or "")
status_code = int(route_result.get("status_code") or 200)
if kind == "response":
headers = dict(route_result.get("headers") or {})
headers.pop("content-length", None)
body_bytes = base64.b64decode(str(route_result.get("body_b64") or ""))
return Response(
body_bytes,
status_code=status_code,
headers=headers,
media_type=route_result.get("media_type") or None,
)
if kind == "json":
return JSONResponse(route_result.get("data"), status_code=status_code)
return Response(str(route_result.get("text") or ""), status_code=status_code)
return child_result
except Exception as exc:
log.exception("extension child dispatch failure: %s", mount)
return json_error(f"{type(exc).__name__}: {exc}", 502)

View file

@ -6,6 +6,7 @@ import asyncio
import hmac
import json
import logging
import math
import os
import pathlib
import threading
@ -22,6 +23,7 @@ from starlette.websockets import WebSocket, WebSocketDisconnect
from ouroboros.contracts.chat_id_policy import A2A_CHAT_ID_MAX, A2A_CHAT_ID_MIN, is_a2a_chat_id
from ouroboros.event_bus import get_global_event_bus
from ouroboros.config import WS_RELAY_BURST, WS_RELAY_REFILL_PER_SEC
from ouroboros.gateway._helpers import run_sync_to_completion
from ouroboros.gateway.files import store_chat_upload
from ouroboros.skill_loader import (
@ -30,7 +32,7 @@ from ouroboros.skill_loader import (
load_enabled,
review_status_allows_execution,
)
from ouroboros.utils import atomic_write_json, read_json_dict, utc_now_iso
from ouroboros.utils import append_jsonl, atomic_write_json, read_json_dict, utc_now_iso
log = logging.getLogger(__name__)
_json_error = lambda message, status=500: JSONResponse({"ok": False, "error": message}, status_code=status)
@ -39,26 +41,53 @@ DEFAULT_HOST_SERVICE_HOST = "127.0.0.1"
DEFAULT_HOST_SERVICE_PORT = 8767
AUTH_TOKEN_FILENAME = "auth_token.json"
# The out-of-process WS progress relay (``POST /ui/ws-message``) is the one
# token-bucket lane: a 60-message burst reserve that refills one message per
# second (owner decision 2026-09-06: a burst must not silence a widget for the
# rest of a minute). Every other Host Service lane keeps the sliding window.
class HostServiceAuthError(Exception):
"""Raised when a skill token cannot be authenticated."""
class _RateLimiter:
def __init__(self, limit: int = 60, window_sec: float = 60.0):
"""The one Host Service admission limiter, two policies under one lock.
``allow(key)`` is the sliding window (``limit`` hits per ``window_sec``) the
chat/presence/decision/tools lanes keep unchanged. ``allow_burst(key)`` is a
token bucket for the WS relay lane only: it also AGGREGATES its refusals per
key so a burst is reported once (first refusal, then one summary when the
lane admits again or the bucket goes idle) instead of one log line per
dropped message; ``on_burst_end(key, dropped, duration_sec)`` is the sink.
"""
def __init__(
self,
limit: int = 60,
window_sec: float = 60.0,
*,
on_burst_end: Optional[Callable[[str, int, float], None]] = None,
):
self.limit = limit
self.window_sec = window_sec
self._hits: Dict[str, Deque[float]] = defaultdict(deque)
# Token buckets: key -> [tokens, last_refill, capacity, refill_per_sec].
self._buckets: Dict[str, list] = {}
# Refusals since the last admit: key -> {"dropped", "since", "last"}.
self._refused: Dict[str, Dict[str, float]] = {}
self._on_burst_end = on_burst_end
self._lock = threading.Lock()
self._last_sweep = time.monotonic()
def _sweep(self, now: float) -> None:
def _sweep(self, now: float) -> list:
# Drop keys idle past the window so _hits does not grow unbounded as
# distinct skill keys ({skill}:{endpoint}) churn over the process
# lifetime. Must pop each key's stale timestamps FIRST, then delete the
# ones left empty (an idle key still holds stale, un-popped entries).
# Collect-then-delete avoids mutating the dict during iteration.
# Caller holds self._lock.
# Caller holds self._lock. Returns the refusal bursts whose bucket went
# idle without a later admit, for the caller to report off the lock.
stale = []
for key, hits in self._hits.items():
while hits and now - hits[0] > self.window_sec:
@ -67,21 +96,87 @@ class _RateLimiter:
stale.append(key)
for key in stale:
del self._hits[key]
ended = []
for key, bucket in list(self._buckets.items()):
tokens, last, capacity, rate = bucket
if min(capacity, tokens + (now - last) * rate) >= capacity:
del self._buckets[key]
burst = self._refused.pop(key, None)
if burst:
ended.append((key, burst))
return ended
def _report(self, ended: list) -> None:
for key, burst in ended:
if self._on_burst_end is None:
continue
try:
self._on_burst_end(key, int(burst["dropped"]), float(burst["last"] - burst["since"]))
except Exception:
log.debug("Rate-limiter burst sink failed for %s", key, exc_info=True)
def allow(self, key: str) -> bool:
now = time.monotonic()
ended: list = []
with self._lock:
# Amortized cleanup: at most once per window, under the existing lock.
if now - self._last_sweep > self.window_sec:
self._sweep(now)
ended = self._sweep(now)
self._last_sweep = now
hits = self._hits[key]
while hits and now - hits[0] > self.window_sec:
hits.popleft()
if len(hits) >= self.limit:
return False
hits.append(now)
return True
admitted = False
else:
hits.append(now)
admitted = True
self._report(ended)
return admitted
def allow_burst(
self,
key: str,
*,
capacity: int = WS_RELAY_BURST,
refill_per_sec: float = WS_RELAY_REFILL_PER_SEC,
) -> Dict[str, Any]:
"""Token-bucket admission for one key.
Returns ``{"allowed", "retry_after_sec", "dropped_in_burst"}``: on a
refusal ``retry_after_sec`` is the wait until one token exists and
``dropped_in_burst`` counts this refusal and every earlier one since the
last admit (``1`` marks the first refusal of a burst).
"""
now = time.monotonic()
ended: list = []
with self._lock:
if now - self._last_sweep > self.window_sec:
ended = self._sweep(now)
self._last_sweep = now
bucket = self._buckets.get(key)
if bucket is None:
bucket = self._buckets[key] = [float(capacity), now, float(capacity), float(refill_per_sec)]
tokens = min(bucket[2], bucket[0] + (now - bucket[1]) * bucket[3])
bucket[1] = now
if tokens >= 1.0:
bucket[0] = tokens - 1.0
burst = self._refused.pop(key, None)
if burst:
ended.append((key, burst))
verdict = {"allowed": True, "retry_after_sec": 0.0, "dropped_in_burst": 0}
else:
bucket[0] = tokens
burst = self._refused.setdefault(key, {"dropped": 0, "since": now, "last": now})
burst["dropped"] += 1
burst["last"] = now
verdict = {
"allowed": False,
"retry_after_sec": (1.0 - tokens) / bucket[3],
"dropped_in_burst": int(burst["dropped"]),
}
self._report(ended)
return verdict
class HostServiceContext:
@ -101,11 +196,32 @@ class HostServiceContext:
self.tool_schemas_getter = tool_schemas_getter or self._default_tool_schemas
self.ws_broadcaster_getter = ws_broadcaster_getter or self._default_ws_broadcaster
self.presence_runner = presence_runner or self._default_presence_runner
self.rate_limiter = _RateLimiter()
self.rate_limiter = _RateLimiter(on_burst_end=self._ws_relay_burst_ended)
self._inflight: Dict[str, int] = defaultdict(int)
self._inflight_lock = threading.Lock()
self._counter_lock = threading.Lock()
def _ws_relay_burst_ended(self, key: str, dropped: int, duration_sec: float) -> None:
"""Report one aggregated WS relay refusal burst: a warning plus one
durable ``host_service_ws_relay_dropped`` row (the ``broadcast_partial_failure``
precedent in ``gateway/ws.py``), never one line per dropped message."""
skill = key.rsplit(":", 1)[0]
log.warning(
"Host Service WS relay for skill %r dropped %d message(s) over %.1fs "
"(burst reserve %d, refill %.0f/s)",
skill, dropped, duration_sec, WS_RELAY_BURST, WS_RELAY_REFILL_PER_SEC,
)
try:
append_jsonl(self.data_dir / "logs" / "events.jsonl", {
"ts": utc_now_iso(),
"type": "host_service_ws_relay_dropped",
"skill": skill,
"dropped": int(dropped),
"duration_sec": round(float(duration_sec), 3),
})
except Exception:
log.debug("Failed to record host_service_ws_relay_dropped event", exc_info=True)
def _default_bridge(self) -> Any:
from supervisor.message_bus import try_get_bridge
@ -291,6 +407,16 @@ async def _api_allocate_internal(request: Request) -> JSONResponse:
async def _api_chat_inject(request: Request) -> JSONResponse:
"""Inject one owner-channel message; correlate it when the caller names it.
``client_message_id`` is the inbound message identity (#667): the host keeps
it on the canonical inbound row it already writes, answers with the
correlated ``operation_ref`` on 202/200/504, and a repeated delivery of the
SAME message (same id, same text, same skill) REJOINS the accepted operation
instead of enqueueing a second one; a different message under a reused id
is refused (409), never mistaken for a replay. Without an id the historical
envelope is byte-identical.
"""
ctx: HostServiceContext = request.app.state.host_service_context
try:
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
@ -324,16 +450,44 @@ async def _api_chat_inject(request: Request) -> JSONResponse:
"(allocate one via /chat/allocate-internal)", 400,
)
timeout = max(1, min(int(payload.get("timeout_sec") or 1800), 1800)) if wait_for_response else 1800
try:
uploads = await run_sync_to_completion(
_inject_attachment_uploads, ctx, skill_name, payload.get("attachments"), pending_uploads,
)
except ValueError as exc:
return _json_error(str(exc), 400)
correlated = {"operation_ref": operation_ref(chat_id, client_message_id)} if client_message_id else {}
rejoined = False
if client_message_id:
rows = await asyncio.to_thread(_chat_rows, ctx, chat_id)
inbound = _inbound_row(rows, client_message_id)
if inbound is not None:
from ouroboros.project_dialogue import _text_sha256
from ouroboros.task_status import SETTLED_STATUSES
if str(inbound.get("source") or "") != f"skill:{skill_name}":
return _json_error("client_message_id is already bound to another source", 409)
logged = text.strip() or image_caption.strip() or (
"(image attached)" if str(payload.get("image_base64") or "").strip()
else "(file attached)" if payload.get("attachments") else ""
)
if _text_sha256(inbound.get("text")) != _text_sha256(logged):
return _json_error("client_message_id was already used for a different message", 409)
state = _operation_state(ctx, rows, inbound)
if state["status"] in SETTLED_STATUSES:
return JSONResponse({
"ok": True, "response": str(state.get("text") or ""),
"status": state["status"], "rejoined": True, **correlated,
})
if not wait_for_response:
return JSONResponse({"ok": True, "status": "accepted", "rejoined": True, **correlated}, status_code=202)
rejoined = True
uploads: list[dict[str, str]] = []
if not rejoined:
try:
uploads = await run_sync_to_completion(
_inject_attachment_uploads, ctx, skill_name, payload.get("attachments"), pending_uploads,
)
except ValueError as exc:
return _json_error(str(exc), 400)
bridge = ctx.bridge_getter()
response_event: asyncio.Event = asyncio.Event()
response_holder: dict[str, str] = {}
if wait_for_response:
if wait_for_response and not client_message_id:
loop = asyncio.get_running_loop()
def on_response(response_text: str) -> None:
@ -341,35 +495,60 @@ async def _api_chat_inject(request: Request) -> JSONResponse:
loop.call_soon_threadsafe(response_event.set)
subscription_id = bridge.subscribe_response(chat_id, on_response)
bridge.enqueue_local_message(
text,
chat_id=chat_id,
user_id=int(payload.get("user_id") or 0),
source=f"skill:{skill_name}",
sender_label=str(payload.get("sender_label") or skill_name),
image_base64=str(payload.get("image_base64") or ""),
image_mime=str(payload.get("image_mime") or ""),
image_caption=image_caption,
transport=payload.get("transport") if isinstance(payload.get("transport"), dict) else {},
**({"task_metadata": {"chat_attachment_uploads": uploads}} if uploads else {}),
**({"client_message_id": client_message_id} if client_message_id else {}),
)
# The accepted message now owns its copies, including on disconnect.
pending_uploads.pop_all()
if not rejoined:
message = dict(
chat_id=chat_id,
user_id=int(payload.get("user_id") or 0),
source=f"skill:{skill_name}",
sender_label=str(payload.get("sender_label") or skill_name),
image_base64=str(payload.get("image_base64") or ""),
image_mime=str(payload.get("image_mime") or ""),
image_caption=image_caption,
transport=payload.get("transport") if isinstance(payload.get("transport"), dict) else {},
**({"task_metadata": {"chat_attachment_uploads": uploads}} if uploads else {}),
**({"client_message_id": client_message_id} if client_message_id else {}),
)
if client_message_id:
from supervisor.message_bus import accept_local_message
try:
_, rejoined = await run_sync_to_completion(
accept_local_message, bridge, ctx.data_dir, text,
retain_inputs=pending_uploads.pop_all, **message,
)
except ValueError as exc:
return _json_error(str(exc), 409)
else:
bridge.enqueue_local_message(text, **message)
pending_uploads.pop_all()
if not wait_for_response:
return JSONResponse({"ok": True, "status": "queued"}, status_code=202)
if rejoined:
return JSONResponse({"ok": True, "status": "accepted", "rejoined": True, **correlated}, status_code=202)
return JSONResponse({"ok": True, "status": "queued", **correlated}, status_code=202)
deadline = time.monotonic() + timeout
while not response_event.is_set():
if client_message_id:
from ouroboros.task_status import SETTLED_STATUSES
state = await asyncio.to_thread(_owned_operation_state, ctx, skill_name, chat_id, client_message_id)
if state and state["status"] in {*SETTLED_STATUSES, "lost"}:
return JSONResponse({"ok": True, "response": str(state.get("text") or ""),
"status": state["status"], "rejoined": rejoined, **correlated})
remaining = deadline - time.monotonic()
if remaining <= 0:
return _json_error("timed out waiting for response", 504)
# The work keeps running; the ref lets the caller recover its
# late answer through /chat/operations (#667).
return JSONResponse(
{"ok": False, "error": "timed out waiting for response", **correlated},
status_code=504,
)
try:
if await request.is_disconnected():
return _json_error("client disconnected", 499)
return JSONResponse({"ok": False, "error": "client disconnected", **correlated}, status_code=499)
await asyncio.wait_for(response_event.wait(), timeout=min(1.0, remaining))
except asyncio.TimeoutError:
continue
return JSONResponse({"ok": True, "response": response_holder.get("text", "")})
return JSONResponse({"ok": True, "response": response_holder.get("text", ""), **correlated})
except json.JSONDecodeError:
return _json_error("invalid json", 400)
except Exception as exc:
@ -664,8 +843,30 @@ async def _api_ws_message(request: Request) -> JSONResponse:
return _json_error(f"skill {skill_name!r} is not installed", 403)
if "ws_handler" not in {str(p).strip() for p in (loaded.manifest.permissions or [])}:
return _json_error(f"skill {skill_name!r} lacks ws_handler permission", 403)
if not ctx.rate_limiter.allow(f"{skill_name}:ws"):
return _json_error("rate limit exceeded", 429)
verdict = ctx.rate_limiter.allow_burst(f"{skill_name}:ws")
if not verdict["allowed"]:
# Visible at the host, aggregated per burst: the first refusal logs once,
# the rest ride the counter until the bucket admits again (then the
# context's sink reports the dropped total). The child keeps its
# best-effort ``None``; ``Retry-After`` says when one token exists.
if verdict["dropped_in_burst"] == 1:
log.warning(
"Host Service WS relay for skill %r refused: burst reserve of %d "
"messages is empty (refills %.0f/s); further refusals in this burst "
"are aggregated",
skill_name, WS_RELAY_BURST, WS_RELAY_REFILL_PER_SEC,
)
retry_after = float(verdict["retry_after_sec"])
return JSONResponse(
{
"ok": False,
"error": "rate limit exceeded",
"retry_after_sec": round(retry_after, 3),
"dropped_in_burst": int(verdict["dropped_in_burst"]),
},
status_code=429,
headers={"Retry-After": str(max(1, math.ceil(retry_after)))},
)
try:
payload = await request.json()
except Exception:
@ -687,6 +888,302 @@ async def _api_ws_message(request: Request) -> JSONResponse:
return JSONResponse({"ok": True, "type": full}, status_code=202)
# --- A2A operation correlation (#667) ---------------------------------------
#
# An accepted inbound message is identified by the (chat_id, client_message_id)
# pair the caller supplied; ``operation_ref`` is that pair spelled
# ``"<chat_id>:<client_message_id>"``, so a caller can address the operation
# before its own inject wait has returned. The host keeps NO new record: the
# canonical inbound source, actual task origin and in-flight turn registry
# are the authorities. Annotation/outbound ids only help discover candidates.
def operation_ref(chat_id: int, client_message_id: str) -> str:
return f"{int(chat_id)}:{client_message_id}"
def _parse_operation_ref(value: Any) -> tuple[int, str]:
head, sep, tail = str(value or "").partition(":")
if not sep or not tail.strip():
raise ValueError("operation_ref must be '<chat_id>:<client_message_id>'")
return int(head), tail.strip()[:128]
def _chat_rows(ctx: HostServiceContext, chat_id: int) -> list:
"""Read retained canonical sources; a display tail cannot prove absence."""
from ouroboros.utils import iter_jsonl_objects, jsonl_chain_handles
with jsonl_chain_handles(ctx.data_dir / "logs" / "chat.jsonl", strict=True) as handles:
return [row for path, handle in handles for row in iter_jsonl_objects(path, _handle=handle)
if row.get("chat_id") == chat_id]
def _inbound_row(rows: list, client_message_id: str) -> Optional[Dict[str, Any]]:
return next(
(
row for row in rows
if str(row.get("direction") or "") == "in"
and str(row.get("client_message_id") or "") == client_message_id
),
None,
)
def _operation_state(ctx: HostServiceContext, rows: list, inbound: Dict[str, Any]) -> Dict[str, Any]:
"""Join the existing records into one typed view of an accepted message.
Authority is the task/turn's complete ingress origin matched to the skill's
canonical source, then that task's effective retry-aware status; otherwise
``pending`` — or ``lost``
when the host session that accepted the message is gone and nothing else
answers. ``cancel_supported`` is true only for work THIS message started
that the cancellation owner can address: a promoted task or a live direct
turn; an ephemeral decision turn (pre-promotion) and a message steered into
a pre-existing task are disclosed, not cancelled.
"""
from ouroboros.project_dialogue import latest_chat_annotations, entry_matches_source_ref, owner_message_ref_is_valid
from ouroboros.task_results import load_task_result
from ouroboros.task_status import SETTLED_STATUSES, load_effective_task_result
from supervisor.active_activity import get_direct_activity_registry
from supervisor import queue as task_queue
chat_id = int(inbound.get("chat_id") or 0)
client_message_id = str(inbound.get("client_message_id") or "")
state: Dict[str, Any] = {
"operation_ref": operation_ref(chat_id, client_message_id),
"chat_id": chat_id,
"client_message_id": client_message_id,
"accepted_at": str(inbound.get("ts") or ""),
"status": "pending",
"cancel_supported": False,
}
def owns(record: dict) -> bool:
ref = record.get("origin_message_ref")
return bool(owner_message_ref_is_valid(ref) and entry_matches_source_ref(inbound, [ref]))
# An annotation is only a discovery hint. The actual task's complete
# ingress origin, already scoped to the authenticated skill's row, is the
# authority. Live queue payloads cover tasks not yet persisted by a worker.
queued = {}
cancel_owner_matches = (task_queue.DRIVE_ROOT is not None
and pathlib.Path(task_queue.DRIVE_ROOT).resolve() == ctx.data_dir.resolve())
if cancel_owner_matches:
with task_queue._queue_lock:
for task in [*task_queue.PENDING, *(meta.get("task") or {} for meta in task_queue.RUNNING.values())]:
if owns(task):
queued[str(task.get("id") or "")] = dict(task)
receipt = latest_chat_annotations(ctx.data_dir).get(client_message_id) or {}
targets = dict.fromkeys([*queued, str(receipt.get("target") or ""),
*(str(row.get("task_id") or "") for row in reversed(rows))])
for target in targets:
if not target:
continue
try:
raw = queued.get(target) or load_task_result(ctx.data_dir, target, strict=True) or {}
except ValueError:
continue # An unreadable/disallowed discovery hint carries no authority.
if not owns(raw):
continue
stored = load_effective_task_result(ctx.data_dir, target, materialize_artifacts=False) or {}
status = str(stored.get("status") or "")
state.update({
"task_id": target,
"phase": "managed_task",
"status": status or "pending",
})
if status in SETTLED_STATUSES:
state["text"] = str(stored.get("result") or "")
# A delivered result can precede post-task work or live descendants.
# Use the same ownership facts as the existing cancellation ingress.
state["cancel_supported"] = cancel_owner_matches and (
task_queue.task_has_live_ownership(target) or task_queue.task_subtree_is_live(target)
)
else:
state["cancel_supported"] = cancel_owner_matches
if not cancel_owner_matches:
state["reason"] = "cancel_owner_unavailable"
if stored.get("cancel_state"):
state["cancel_state"] = str(stored.get("cancel_state"))
return state
for entry in get_direct_activity_registry().snapshot(chat_id):
activity = get_direct_activity_registry().get(str(entry.get("activity_id") or ""))
if activity is not None and owns({"origin_message_ref": activity.origin_message_ref}):
kind = str(entry.get("kind") or "direct_chat")
state.update({
"status": "running",
"phase": kind,
"task_id": str(entry.get("activity_id") or ""),
"cancel_supported": kind == "direct_chat" and cancel_owner_matches,
})
if kind == "direct_chat" and not cancel_owner_matches:
state["reason"] = "cancel_owner_unavailable"
return state
for row in reversed(rows):
terminal = str(row.get("task_terminal_status") or "")
if row.get("direction") == "out" and terminal in SETTLED_STATUSES and owns(row):
state.update({"status": terminal, "text": str(row.get("text") or "")})
return state
accepted_session = str(inbound.get("session_id") or "")
live_session = str((read_json_dict(ctx.data_dir / "state" / "state.json") or {}).get("session_id") or "")
if accepted_session and live_session and accepted_session != live_session:
state.update({"status": "lost", "reason": "host_restarted_before_answer"})
return state
def _owned_operation_state(
ctx: HostServiceContext, skill_name: str, chat_id: int, client_message_id: str,
) -> Optional[Dict[str, Any]]:
"""The operation view, or None unless THIS skill injected the message."""
rows = _chat_rows(ctx, chat_id)
inbound = _inbound_row(rows, client_message_id)
if inbound is None or str(inbound.get("source") or "") != f"skill:{skill_name}":
return None
return _operation_state(ctx, rows, inbound)
def _cancel_owned_operation(
ctx: HostServiceContext, skill_name: str, chat_id: int, client_message_id: str, reason: str,
) -> tuple[int, Dict[str, Any]]:
from ouroboros.task_status import SETTLED_STATUSES
state = _owned_operation_state(ctx, skill_name, chat_id, client_message_id)
if state is None:
return 404, {"ok": False, "error": "operation not found"}
base = {key: state[key] for key in ("operation_ref", "task_id", "phase", "status") if key in state}
if state["status"] in SETTLED_STATUSES and not state.get("cancel_supported"):
return 200, {"ok": True, "outcome": "already_terminal", **base}
if not state.get("cancel_supported") or not state.get("task_id"):
reason_code = state.get("reason") or {
"ephemeral_decision": "decision_turn_in_flight",
"lost": "host_restarted_before_answer",
}.get(str(state.get("phase") or state["status"]), "not_started")
return 409, {"ok": False, "outcome": "cancel_unsupported", "reason": reason_code, **base}
return _cancel_task_through_owner(skill_name, str(state["task_id"]), reason, base, ctx.data_dir)
def _cancel_task_through_owner(
skill_name: str, task_id: str, reason: str, base: Dict[str, Any], drive_root: pathlib.Path,
) -> tuple[int, Dict[str, Any]]:
"""The existing cancel ingress shape: durable intent first (fail-closed),
then the same cascade custody path the browser Stop uses; the typed outcome
is read back from the effective result, never assumed."""
from ouroboros.cancel_intents import (
CancelIntentProjectionCorrupt,
SCOPE_CASCADE,
STOP_POLICY_IMMEDIATE,
request_cancel,
)
from ouroboros.gateway.tasks import _run_cascade_cancel
from ouroboros.task_results import STATUS_CANCELLED
from ouroboros.task_status import SETTLED_STATUSES, load_effective_task_result
from supervisor.queue import DRIVE_ROOT, task_has_live_ownership, task_subtree_is_live
if DRIVE_ROOT is None or pathlib.Path(DRIVE_ROOT).resolve() != drive_root.resolve():
return 409, {"ok": False, "outcome": "cancel_unsupported", "reason": "cancel_owner_unavailable", **base}
def _status() -> str:
stored = load_effective_task_result(pathlib.Path(DRIVE_ROOT), task_id, materialize_artifacts=False) or {}
return str(stored.get("status") or "")
if not task_has_live_ownership(task_id) and not task_subtree_is_live(task_id):
status = _status()
if status in SETTLED_STATUSES:
return 200, {"ok": True, "outcome": "already_terminal", **base, "status": status}
return 503, {"ok": False, "outcome": "unresolved", "reason": "cancellation_did_not_settle", **base, "status": status}
try:
request_cancel(
DRIVE_ROOT, task_id, reason=reason, source=f"skill:{skill_name}",
scope=SCOPE_CASCADE, allow_settled_target=True,
requested_stop_policy=STOP_POLICY_IMMEDIATE,
)
except CancelIntentProjectionCorrupt:
log.error("Host Service cancel refused for %s: intent projection corrupt", task_id)
return 503, {"ok": False, "outcome": "refused", "reason": "cancel_intent_projection_corrupt", **base}
except Exception:
log.warning("Host Service cancel-intent write failed for %s", task_id, exc_info=True)
return 503, {"ok": False, "outcome": "refused", "reason": "cancel_intent_write_failed", **base}
settled = _run_cascade_cancel(task_id)
status = _status()
if not settled or status not in SETTLED_STATUSES:
return 503, {"ok": False, "outcome": "unresolved", "reason": "cancellation_did_not_settle", **base, "status": status}
return 200, {
"ok": True,
"outcome": "cancelled" if status == STATUS_CANCELLED else "already_terminal",
**base,
"status": status,
}
async def _api_chat_operation(request: Request) -> JSONResponse:
"""Read ONE accepted message this skill injected (#667).
Scoped by source provenance: the canonical inbound row must carry this
skill's ``source``; anything else is not found. No task id is accepted from
the caller, so this stays a callback boundary, not a general task API.
"""
ctx: HostServiceContext = request.app.state.host_service_context
try:
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
ctx.require_permission(skill_name, token_payload, "inject_chat")
except HostServiceAuthError as exc:
return _json_error(str(exc), 403)
if not ctx.rate_limiter.allow(f"{skill_name}:operations"):
return _json_error("rate limit exceeded", 429)
try:
chat_id, client_message_id = _parse_operation_ref(request.path_params.get("operation_ref"))
except ValueError as exc:
return _json_error(str(exc), 400)
try:
state = await asyncio.to_thread(_owned_operation_state, ctx, skill_name, chat_id, client_message_id)
except Exception as exc:
log.debug("Host service operation lookup failed", exc_info=True)
return _json_error(str(exc), 500)
if state is None:
return _json_error("operation not found", 404)
return JSONResponse({"ok": True, **state})
async def _api_chat_cancel(request: Request) -> JSONResponse:
"""Cancel the host work ONE accepted message of this skill started (#667).
The cancellation owner does the work — durable intent first
(``cancel_intents.request_cancel``), then custody through the cascade path
the browser Stop uses — and the answer is its typed outcome: ``cancelled``,
``already_terminal``, ``unresolved`` (custody did not settle; the work is
still live) or ``cancel_unsupported`` (nothing this request started is
addressable yet: still queued, an ephemeral decision turn in flight, or a
message the decision lane delivered into a pre-existing task). Never a
``cancelled`` that did not happen.
"""
ctx: HostServiceContext = request.app.state.host_service_context
try:
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
ctx.require_permission(skill_name, token_payload, "inject_chat")
except HostServiceAuthError as exc:
return _json_error(str(exc), 403)
if not ctx.rate_limiter.allow(f"{skill_name}:cancel"):
return _json_error("rate limit exceeded", 429)
try:
body = await request.json()
except Exception:
return _json_error("invalid json", 400)
if not isinstance(body, dict):
return _json_error("request body must be a JSON object", 400)
try:
chat_id, client_message_id = _parse_operation_ref(body.get("operation_ref"))
except ValueError as exc:
return _json_error(str(exc), 400)
reason = " ".join(str(body.get("reason") or "").split())[:500]
try:
status, payload = await asyncio.to_thread(
_cancel_owned_operation, ctx, skill_name, chat_id, client_message_id, reason,
)
except Exception as exc:
log.warning("Host service operation cancel failed", exc_info=True)
return _json_error(str(exc), 503)
return JSONResponse(payload, status_code=status)
async def _ws_events(websocket: WebSocket) -> None:
ctx: HostServiceContext = websocket.app.state.host_service_context
try:
@ -747,6 +1244,8 @@ def create_host_service_app(
Route("/tools/schemas", _api_tool_schemas, methods=["GET"]),
Route("/chat/allocate-internal", _api_allocate_internal, methods=["POST"]),
Route("/chat/inject", _api_chat_inject, methods=["POST"]),
Route("/chat/operations/{operation_ref:path}", _api_chat_operation, methods=["GET"]),
Route("/chat/cancel", _api_chat_cancel, methods=["POST"]),
Route("/chat/decision", _api_chat_decision, methods=["POST"]),
Route("/presence/turn", _api_presence_turn, methods=["POST"]),
Route("/presence/work/{work_ref}", _api_presence_work, methods=["GET"]),

View file

@ -44,6 +44,7 @@ from ouroboros.utils import (
emit_cognitive_operation_event,
emit_main_llm_call_state_event,
emit_log_event,
has_log_sink,
sanitize_tool_result_for_log,
truncate_review_artifact,
utc_now_iso,
@ -972,11 +973,11 @@ def _record_llm_call_error(
"llm_call_id": ctx.llm_call_id, "round": ctx.round_idx, "attempt": ctx.attempt + 1,
"model": ctx.model,
}
# ONE durable row per failure (#355): the events.jsonl append below is
# forwarded live by the worker's events tail, so a second live-only
# `llm_round_error` sibling was the same failure delivered twice to the
# Logs tab. Background Consciousness keeps its own live `llm_round_error`.
append_jsonl(ctx.drive_logs / "events.jsonl", {
# ONE error row (#355): a successful append's registered sink owns live
# delivery. Without that path, send the SAME evidence through the queue,
# preserving its identity for live/backfill dedupe. No llm_round_error
# sibling here; Background Consciousness keeps its own separate producer.
error_event = {
"ts": utc_now_iso(), "type": "llm_api_error", **identity, "error": safe_error,
"error_kind": classification.kind, "retry_same_request": will_retry,
"status_code": classification.status_code, "provider_code": classification.provider_code,
@ -984,7 +985,9 @@ def _record_llm_call_error(
**custody_fields,
**(ctx.context_fit_event_fields or {}),
"request_ref": ctx.request_ref.get("manifest_ref") if ctx.request_ref else None,
})
}
if not append_jsonl(ctx.drive_logs / "events.jsonl", error_event) or not has_log_sink():
emit_log_event(ctx.event_queue, error_event, log_label="LLM call error")
ctx.accumulated_usage.update(_last_llm_error=_short_error_text(safe_error),
_last_llm_error_kind=classification.kind, _last_llm_retry_same_request=will_retry)
if classification.retry_after_sec is not None:

View file

@ -1133,7 +1133,12 @@ class MCPManager:
if not isinstance(result, ToolResult):
raise TypeError("MCP transport returned a non-ToolResult outcome")
except asyncio.TimeoutError:
text = f"⚠️ MCP_TOOL_TIMEOUT: server {cfg.id!r} did not respond in {timeout}s"
text = (
f"⚠️ MCP_TOOL_TIMEOUT: server {cfg.id!r} did not respond in {timeout}s. "
"The remote outcome is unknown: side effects may already have happened, "
"and remote cancellation is not confirmed. Use the server-specific status/read "
"tool, if available, to reconcile the operation before retrying."
)
return ToolResult(status="timeout", code="MCP_TIMEOUT", text=text)
except BaseException as exc: # noqa: BLE001 - any failure is reported
body = f"⚠️ MCP_TOOL_ERROR: {type(exc).__name__}: {_redact_error_text(exc, cfg)}"

View file

@ -1,9 +1,9 @@
"""Ouroboros — the numeric runtime knobs and their clamps.
Worker count, task liveness windows, per-call ceilings, reviewer and acceptance
budgets, subagent caps and delegation windows. Every one of them is an
environment-or-default scalar clamped into a documented band, so a typo falls
back to the shipped value instead of disabling a rail.
budgets, subagent caps and delegation windows, plus fixed process and transport
bounds. Environment-or-default getters clamp into documented bands, so a typo
falls back to the shipped value instead of disabling a rail.
"""
from __future__ import annotations
@ -18,6 +18,39 @@ from ouroboros.settings_defaults import (
)
EXTENSION_STREAM_CHUNK_BYTES = 64 * 1024
# Exit/pipe-drain grace after a response ends; never a response lifetime timer.
EXTENSION_CHILD_CLEANUP_GRACE_SEC = 2
NESTED_SETTLEMENT_MARGIN_SEC = 30 # Structural ordering margin, not a cognition timeout.
# Owner-note cadence while a task waits out a provider-connection outage; the effective interval is min(this, idle_timeout/2) so the notes also keep the idle rail alive.
NETWORK_WAIT_NOTE_INTERVAL_SEC = 300
# First free-redial pause of a transport-wait episode; doubles per wait iteration up to the existing 60s transient backoff cap (Q10: an existing bound, not a new knob).
NETWORK_WAIT_BACKOFF_START_SEC = 4.0
# TCP keepalive for long-lived remote LLM sockets (idle threshold, probe interval, probe count): kernel probes
# detect a silently dropped NAT/VPN mapping instead of hanging to the read timeout; platform_layer builds the options.
TCP_KEEPALIVE_IDLE_SEC = 60
TCP_KEEPALIVE_INTERVAL_SEC = 60
TCP_KEEPALIVE_PROBE_COUNT = 5
# One response frame may carry metadata or a body chunk; never a total response cap.
EXTENSION_STREAM_METADATA_BYTES = 512 * 1024
# Only the out-of-process WS relay is limited: the existing reserve refills gradually.
WS_RELAY_BURST = 60
WS_RELAY_REFILL_PER_SEC = 1.0
# Worker-pool spawn bounds (structural constants, not env knobs). Grace after a full-pool spawn before the crash
# detector counts dead workers (up to ~60s to init: spawn + pip); workers.py binds it as `_SPAWN_GRACE_SEC`, the extension import-staging sweep reads it too.
WORKER_SPAWN_GRACE_SEC = 90.0
# Readiness window for ONE spawned/respawned slot: unassignable until the child's own `worker_ready` row lands; alive
# but silent past this = torn down and replaced. Sized to the spawn grace (the pool's existing init budget): a warm
# forkserver child boots in ~3-4s (G13 mock lane: 3.5-4.9s startup, 2.5-3.2s respawn), a cold 4-vCPU CI runner well under 60s (its 21-scenario mock lane runs in ~80s), and the E2E
# scenarios wait 240s per task, so a wedged child is a fast, named failure. A contract distinct from process liveness
# (`proc.is_alive`, worker_health.py) and from the task idle rail (queue_timeouts.py): a deadlocked child is alive.
WORKER_READY_WINDOW_SEC = 90.0
# Consecutive readiness failures of one slot before it is parked and reported (three strikes, like the crash-storm fence).
WORKER_READY_MAX_ATTEMPTS = 3
def _clamped_number_setting(key: str, *, low, high=float("inf"), cast=float):
"""Env-or-default numeric setting clamped to [low, high]; a typo falls back to the
shipped default. SSOT for the clamped scalar getters below — the seven of them were

View file

@ -122,10 +122,12 @@ BAND_PATHS = {
"ouroboros/context.py": "Entered the band from the 1501-1600 zone (1590 lines) by the v7 D03 extraction of the runtime-section fact builders into ouroboros/context_runtime_facts.py; shrink-only residue of the split, not new growth.",
"ouroboros/deep_self_review.py": "Deep self-review moved from one packed call to the three reviewer-row deliveries inside ONE surface module: the packed Atlas assembler and the retrieving runner (route-aware availability, provenance header, mandatory-read coverage, typed failures) share the memory whitelist, the prompt constants and the failure typing, so a split would separate the surface from its own pack contract; shrink next touch.",
"ouroboros/delegate_custody.py": "D07 DEL1 split brought the custody monolith DOWN from the 1600 hard cap into the band (1600->1305); reconcile family extracted to delegate_custody_reconcile.py, shrink-only direction",
"ouroboros/extension_plugin_api.py": "The PluginAPI transport owner retains per-child best-effort WS failure aggregation beside the relay that observes failures; the existing process-result channel carries that aggregate without a separate diagnostic service.",
"ouroboros/extension_process_runner.py": None,
"ouroboros/gateway/control.py": "Entered the band from 966 lines: the update-flow redesign added the shared stash-first prologue (_stash_local_work_fenced/_unwind_stashed_update) and the review-wave affordability floor to the update apply orchestration (update-flow-redesign sprint, Q9/Q10 owner decisions).",
"ouroboros/gateway/extensions.py": "Extensions HTTP surface re-entered the band when the module endpoint moved to the in-memory reviewed bundle (widgets lifecycle 1a); shrink next touch.",
"ouroboros/gateway/history.py": "Shrank into the band by extracting the ledger-derived cost-breakdown endpoint family (compat buckets/groups + /api/cost-breakdown) to gateway/cost_breakdown.py; the chat-history window and its lineage/quota machinery stay here.",
"ouroboros/gateway/host_service.py": "The one loopback callback boundary for reviewed skills: token auth, the chat/decision/presence/WS-relay routes and, with #667, the operation read/cancel that joins existing chat, routing, turn and task records; one trust boundary, one module.",
"ouroboros/gateway/settings.py": "Retiring persistent auto-Low removed the former giant debt; the remaining owner and reviewer settings endpoints stay centralized while tracked in the shrinking band.",
"ouroboros/loop_acceptance_review.py": "F6 upstream sync: the A-material acceptance family (paid identity, free replay, identical-refusal terminal, dialogue history) folded into the campaign review leaf per the sync principle (upstream leaf acceptance_dialogue.py retired)",
"ouroboros/loop_delivery.py": "F6 upstream sync: the delivery-protocol upstream deltas (hold-control literals, trailing-object/fence-aware protocol parsers) folded into the campaign delivery leaf (upstream leaf delivery_protocol.py retired)",

View file

@ -90,6 +90,22 @@ def provider_terminal_body(text: str, notice: str) -> str:
return (text + "\n\n" if text else "") + "[Host status]\n" + notice
def host_operation_reply_kwargs(source_ref: Any, terminal_status: str = "") -> Dict[str, Any]:
"""Correlate a host-accepted reply without inventing a durable task result.
Callers supply the host-captured source, never a transport's arbitrary
metadata. An empty status marks an acknowledgement, not completion.
"""
from ouroboros.project_dialogue import owner_message_ref_is_valid
if not owner_message_ref_is_valid(source_ref):
return {}
meta = {"origin_message_ref": dict(source_ref)}
if terminal_status:
meta["task_terminal_status"] = terminal_status
return {"progress_meta": meta}
def stamp_root_final_phase(
send_event: Dict[str, Any], task: Dict[str, Any], *, post_task_open: bool, terminal_status: str,
) -> None:
@ -117,6 +133,9 @@ def prepare_terminal_send_event(
*, ephemeral: bool, presence: bool,
) -> Dict[str, Any]:
"""Preserve raw host salvage, then build the one live/replay projection."""
if not presence and task.get("_is_direct_chat") and (task.get("metadata") or {}).get("_host_operation"):
correlation = host_operation_reply_kwargs(task.get("origin_message_ref"))
send_event.setdefault("progress_meta", {}).update(correlation.get("progress_meta", {}))
origin = str(usage.get("terminal_origin") or "")
notice = str(usage.get("terminal_provider_notice") or "")
if ephemeral and not presence:

View file

@ -70,6 +70,7 @@ PROCESS_FACT_KEYS = (
"timed_out",
"killed_by_host",
"pre_exec_failure",
"ws_relay_failures",
)
@ -92,7 +93,7 @@ def active_resolved_runtime(ctx) -> str:
def publish_process_facts(
*, returncode=None, started_ts: float, resolved_runtime: str = "",
timed_out: bool = False, killed_by_host: bool = False,
pre_exec_failure: str = "",
pre_exec_failure: str = "", ws_relay_failures=None,
) -> Dict[str, object]:
"""Publish this thread's typed process facts for the in-flight process tool.
@ -126,6 +127,18 @@ def publish_process_facts(
facts["signal"] = name
if resolved_runtime:
facts["resolved_runtime"] = resolved_runtime
# Child-reported delivery diagnostics cannot author measured exit/kill facts.
# Only fixed categories and positive integer counts cross this boundary;
# message text, URLs, tokens and arbitrary child metadata never do.
if isinstance(ws_relay_failures, dict):
failures = {
key: ws_relay_failures[key]
for key in ("missing_transport", "rate_limited", "http_client_error",
"http_server_error", "http_error", "transport_error")
if type(ws_relay_failures.get(key)) is int and ws_relay_failures[key] > 0
}
if failures:
facts["ws_relay_failures"] = failures
_process_facts_tls.facts = facts
# Returned so a producer that ALSO discloses these facts elsewhere (the
# verify_and_record receipt) copies the published ones instead of deriving

View file

@ -385,7 +385,7 @@ TOOL_CODE_SPECS: Mapping[str, ToolCodeSpec] = MappingProxyType(
"timeout",
"timeout",
"error",
"inspect MCP health before retrying",
"reconcile the remote outcome before retrying",
),
"EXTENSION_TIMEOUT": _code_spec(
"timeout",

View file

@ -28,6 +28,12 @@ def set_log_sink(fn: Optional[Callable[[Dict[str, Any]], None]]) -> None:
global _log_sink
_log_sink = fn
def has_log_sink() -> bool:
"""Whether this process registered a sink, not a delivery acknowledgement."""
return _log_sink is not None
def utc_now_iso() -> str:
return _dt.datetime.now(tz=_dt.timezone.utc).isoformat()

100
server.py
View file

@ -32,6 +32,7 @@ from ouroboros.server_auth import (
from ouroboros.server_entrypoint import bound_service_socket, find_free_port, parse_server_args, write_port_file
from ouroboros.server_web import NoCacheStaticFiles, make_index_page, resolve_web_dir
from ouroboros.usage_accounting import ensure_legacy_imported
from ouroboros.task_finalization import host_operation_reply_kwargs
from ouroboros.gateway import collect_routes
from ouroboros.gateway import settings as _gateway_settings
from ouroboros.gateway.ws import (
@ -264,14 +265,12 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
user_id = coerce_chat_identity((msg.get("from") or {}).get("id"), chat_id or 1)
text = str(msg.get("text") or "")
source = str(msg.get("source") or "web")
sender_label = str(msg.get("sender_label") or "")
sender_session_id = str(msg.get("sender_session_id") or "")
client_message_id = str(msg.get("client_message_id") or "")
transport = msg.get("transport") if isinstance(msg.get("transport"), dict) else {}
image_base64 = str(msg.get("image_base64") or "")
image_mime = str(msg.get("image_mime") or "image/jpeg")
image_caption = str(msg.get("image_caption") or "")
suppress_chat_log = bool(msg.get("suppress_chat_log"))
task_constraint = msg.get("task_constraint") if isinstance(msg.get("task_constraint"), dict) else None
task_metadata = msg.get("task_metadata") if isinstance(msg.get("task_metadata"), dict) else None
image_data = (image_base64, image_mime, image_caption) if image_base64 else None
@ -310,55 +309,21 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
if owner_id is None and external_identity_present:
owner_id = user_id
from supervisor.message_bus import log_chat
from supervisor.message_bus import record_inbound_message
# Origin identity is captured HERE, where the host writes the canonical
# row (BIBLE P2: identity by value, never re-derived from content
# downstream). Only a row that is actually logged mints a ref — a
# suppressed message must not reference a non-existent canonical row.
origin_message_ref: Optional[Dict[str, Any]] = None
if not suppress_chat_log:
log_chat(
"in",
chat_id,
user_id,
log_text,
ts=now_iso,
source=source,
sender_label=sender_label,
sender_session_id=sender_session_id,
client_message_id=client_message_id,
transport=transport,
client_surface=(
task_metadata.get("client_surface")
if isinstance(task_metadata, dict) and isinstance(task_metadata.get("client_surface"), dict)
else None
),
)
from ouroboros.project_dialogue import build_owner_message_ref
origin_message_ref = build_owner_message_ref(
chat_id=chat_id,
client_message_id=client_message_id,
ts=now_iso,
text=log_text,
)
if source != "web":
bridge.broadcast({
"type": "photo" if image_base64 else "chat",
"role": "user",
"content": text,
"caption": image_caption,
"image_base64": image_base64,
"mime": image_mime,
"ts": now_iso,
"source": source,
"sender_label": sender_label,
"sender_session_id": sender_session_id,
"client_message_id": client_message_id,
"transport": transport,
"chat_id": chat_id,
})
# The same writer mints ordinary ingress and validates preaccepted
# skill deliveries; the latter already have their one canonical row.
origin_message_ref = record_inbound_message(
bridge, msg, chat_id=chat_id, user_id=user_id,
client_message_id=client_message_id, text=log_text, ts=now_iso,
)
if task_metadata or msg.get("accepted_source_ref"):
task_metadata = {k: v for k, v in (task_metadata or {}).items() if k != "_host_operation"}
if msg.get("accepted_source_ref"):
task_metadata["_host_operation"] = True
reply_source = origin_message_ref if msg.get("accepted_source_ref") else None
def reply(body: str, status: str = "completed") -> None:
ctx.send_with_budget(chat_id, body, **host_operation_reply_kwargs(reply_source, status))
def _stamp_owner_activity(live: dict) -> None:
if live.get("owner_id") is None and external_identity_present:
live["owner_id"] = user_id
@ -372,7 +337,7 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
if is_external_transport and is_slash_command:
if not external_identity_present:
ctx.send_with_budget(chat_id, "⚠️ Command ignored: this transport did not provide owner identity.")
reply("⚠️ Command ignored: this transport did not provide owner identity.", "failed")
continue
owner_ext_id = st.get("owner_external_id")
owner_ext_chat_id = st.get("owner_external_chat_id")
@ -384,7 +349,7 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
live["owner_external_bound_at"] = now_iso
ctx.update_state(_bind_external_owner)
ctx.send_with_budget(chat_id, "✅ Owner chat registered. Send the command again to execute it.")
reply("✅ Owner chat registered. Send the command again to execute it.")
continue
try:
owner_ext_id_int = int(owner_ext_id or 0)
@ -393,21 +358,21 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
owner_ext_id_int = 0
owner_ext_chat_id_int = 0
if owner_ext_id_int != user_id or owner_ext_chat_id_int != chat_id:
ctx.send_with_budget(chat_id, "⚠️ Command ignored: this transport is not the bound owner chat.")
reply("⚠️ Command ignored: this transport is not the bound owner chat.", "failed")
continue
if lowered.startswith("/panic"):
ctx.send_with_budget(chat_id, "🛑 PANIC: killing everything. App will close.")
reply("🛑 PANIC: killing everything. App will close.", "")
_execute_panic_stop(ctx.consciousness, ctx.kill_workers)
elif lowered.startswith("/restart"):
ctx.send_with_budget(chat_id, "♻️ Restarting.")
reply("♻️ Restarting.", "")
ok, restart_msg = _safe_restart_serialized(
ctx.safe_restart,
reason="owner_restart",
unsynced_policy="rescue_and_reset",
)
if not ok:
ctx.send_with_budget(chat_id, f"⚠️ Restart cancelled: {restart_msg}")
reply(f"⚠️ Restart cancelled: {restart_msg}", "failed")
continue
state_dir = DATA_DIR / "state"
owner_restart_flag = state_dir / "owner_restart_no_resume.flag"
@ -421,7 +386,7 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
owner_restart_flag.unlink(missing_ok=True)
stable_skip_flag.unlink(missing_ok=True)
log.warning("Failed to write owner restart no-resume flag", exc_info=True)
ctx.send_with_budget(chat_id, "⚠️ Restart cancelled: could not write restart state.")
reply("⚠️ Restart cancelled: could not write restart state.", "failed")
continue
try:
ctx.kill_workers(
@ -435,12 +400,12 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
stable_skip_flag.unlink(missing_ok=True)
log.warning("Restart cancelled because worker shutdown failed", exc_info=True)
try:
ctx.send_with_budget(chat_id, "⚠️ Restart cancelled: failed to stop workers.")
reply("⚠️ Restart cancelled: failed to stop workers.", "failed")
except Exception:
pass
continue
try:
ctx.send_with_budget(chat_id, "Stopping active task. New settings apply to the next message.")
reply("Stopping active task. New settings apply to the next message.", "")
except Exception:
log.warning("Failed to send owner restart stop notice; continuing restart", exc_info=True)
_request_restart_exit(owner=True)
@ -461,7 +426,7 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
block = evolution_block_reason()
if block:
ctx.send_with_budget(chat_id, block)
reply(block, "failed")
continue
# GR4-6: clear the durable owner-stop flag BEFORE the campaign is
# minted — the old order (campaign first, flag cleared in the later
@ -485,7 +450,7 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
# an unconditional True would invent a stop that never happened.
_evo_update_state(lambda live, _v=_prior_owner_stop: live.__setitem__(
"evolution_owner_stopped", _v))
ctx.send_with_budget(chat_id, "⚠️ Evolution stayed OFF: campaign state could not be created.")
reply("⚠️ Evolution stayed OFF: campaign state could not be created.", "failed")
continue
st2 = ctx.load_state()
st2["evolution_mode_enabled"] = bool(turn_on)
@ -500,10 +465,7 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
# autostop, which would disable the owner's campaign after one cycle.
st2["post_task_autostop"] = False
ctx.save_state(st2)
ctx.send_with_budget(
chat_id,
f"🧬 Evolution campaign: {'ON' if turn_on else _owner_evolution_stop(ctx, chat_id)}",
)
reply(f"🧬 Evolution campaign: {'ON' if turn_on else _owner_evolution_stop(ctx, chat_id)}")
elif lowered.startswith("/bg"):
parts = lowered.split()
action = parts[1] if len(parts) > 1 else "status"
@ -512,21 +474,21 @@ def _process_bridge_updates(bridge, offset: int, ctx: Any) -> int:
_bg_s = ctx.load_state()
_bg_s["bg_consciousness_enabled"] = True
ctx.save_state(_bg_s)
ctx.send_with_budget(chat_id, f"🧠 {result}")
reply(f"🧠 {result}")
elif action in ("stop", "off", "0"):
result = ctx.consciousness.stop()
_bg_s = ctx.load_state()
_bg_s["bg_consciousness_enabled"] = False
ctx.save_state(_bg_s)
ctx.send_with_budget(chat_id, f"🧠 {result}")
reply(f"🧠 {result}")
else:
bg_status = "running" if ctx.consciousness.is_running else "stopped"
ctx.send_with_budget(chat_id, f"🧠 Background consciousness: {bg_status}")
reply(f"🧠 Background consciousness: {bg_status}")
elif lowered.startswith("/status"):
from supervisor.state import status_text
status = status_text(ctx.WORKERS, ctx.PENDING, ctx.RUNNING)
ctx.send_with_budget(chat_id, status)
reply(status)
else:
_route_owner_message(
bridge,

View file

@ -26,6 +26,7 @@ class DirectActivityEntry:
kind: str = "direct_chat" # "direct_chat" | "ephemeral_decision"
phase: str = "thinking"
started_at: float = field(default_factory=time.time)
origin_message_ref: Dict[str, Any] = field(default_factory=dict)
def to_dict(self) -> Dict[str, Any]:
return {
@ -55,6 +56,7 @@ class DirectActivityRegistry:
project_id: str = "",
kind: str = "direct_chat",
phase: str = "thinking",
origin_message_ref: Optional[Dict[str, Any]] = None,
) -> DirectActivityEntry:
aid = str(activity_id or "").strip()
if not aid:
@ -67,6 +69,7 @@ class DirectActivityRegistry:
kind=str(kind or "direct_chat"),
phase=str(phase or "thinking"),
started_at=time.time(),
origin_message_ref=dict(origin_message_ref or {}),
)
with self._lock:
self._activities[aid] = entry
@ -116,6 +119,7 @@ def track_direct_activity(
project_id: str = "",
kind: str = "direct_chat",
phase: str = "thinking",
origin_message_ref: Optional[Dict[str, Any]] = None,
) -> Iterator[DirectActivityEntry]:
registry = get_direct_activity_registry()
entry = registry.register(
@ -125,6 +129,7 @@ def track_direct_activity(
project_id=project_id,
kind=kind,
phase=phase,
origin_message_ref=origin_message_ref,
)
try:
yield entry

View file

@ -31,6 +31,111 @@ DATA_DIR = None # pathlib.Path
TOTAL_BUDGET_LIMIT: float = 0.0
BUDGET_REPORT_EVERY_MESSAGES: int = 10
_BRIDGE: Optional["LocalChatBridge"] = None
_INGRESS_LOCK = threading.Lock()
def accepted_chat_message(drive_root, chat_id: int, client_message_id: str) -> Optional[dict]:
"""Read a named canonical source across its retained generation chain."""
from pathlib import Path
from ouroboros.utils import iter_jsonl_objects, jsonl_chain_handles
with jsonl_chain_handles(Path(drive_root) / "logs" / "chat.jsonl", strict=True) as handles:
for path, handle in reversed(handles):
for row in iter_jsonl_objects(path, _handle=handle):
if (row.get("direction") == "in" and row.get("chat_id") == chat_id
and row.get("client_message_id") == client_message_id):
return row
return None
def accept_local_message(bridge, drive_root, text: str, *, retain_inputs=None, **message) -> tuple[dict, bool]:
"""Accept a named skill delivery once, then hand its exact source to the queue.
The canonical row is the acceptance record, written by this existing owner
BEFORE enqueue. A crash after acceptance never authorizes another enqueue;
operation reads disclose a lost host session instead. The single host's
ingress lock covers check/write/enqueue, including simultaneous HTTP calls.
"""
from ouroboros.project_dialogue import _text_sha256, build_owner_message_ref
chat_id = int(message["chat_id"])
message_id = str(message["client_message_id"])
source = str(message["source"])
logged = text.strip() or str(message.get("image_caption") or "").strip() or (
"(image attached)" if message.get("image_base64")
else "(file attached)" if (message.get("task_metadata") or {}).get("chat_attachment_uploads") else ""
)
if not logged:
raise ValueError("message is empty")
with _INGRESS_LOCK:
previous = accepted_chat_message(drive_root, chat_id, message_id)
if previous is not None:
if previous.get("source") != source:
raise ValueError("client_message_id is already bound to another source")
if _text_sha256(previous.get("text")) != _text_sha256(logged):
raise ValueError("client_message_id was already used for a different message")
return previous, True
ts = utc_now_iso()
try:
row = log_chat(
"in", chat_id, int(message.get("user_id") or 0), logged, ts=ts,
source=source, client_message_id=message_id,
sender_label=str(message.get("sender_label") or ""),
transport=message.get("transport"), drive_root=drive_root, require_write=True,
)
finally:
# Once this write is attempted, failure can leave canonical bytes.
# Transfer input custody without claiming acceptance or queue success;
# a replay/pre-write refusal never adopts this request's fresh copies.
if retain_inputs is not None:
retain_inputs()
ref = build_owner_message_ref(chat_id=chat_id, client_message_id=message_id, ts=ts, text=logged)
bridge.enqueue_local_message(text, **message, accepted_source_ref=ref)
return row, False
def record_inbound_message(bridge, message: dict, *, chat_id: int, user_id: int,
client_message_id: str, text: str, ts: str) -> Optional[dict]:
"""Keep one canonical ingress writer for dequeued and preaccepted messages."""
from ouroboros.project_dialogue import (
_text_sha256, build_owner_message_ref, entry_matches_source_ref, owner_message_ref_is_valid,
)
source = str(message.get("source") or "web")
ref = message.get("accepted_source_ref")
if ref:
row = accepted_chat_message(DATA_DIR, chat_id, client_message_id) if owner_message_ref_is_valid(ref) else None
if (not row or row.get("source") != source or row.get("chat_id") != chat_id
or row.get("client_message_id") != client_message_id or not entry_matches_source_ref(row, [ref])
or ref["text_sha256"] != _text_sha256(text)):
raise ValueError("accepted source does not match the queued message")
ref = dict(ref)
ts = ref["ts"]
elif message.get("suppress_chat_log"):
return None
else:
metadata = message.get("task_metadata") or {}
log_chat(
"in", chat_id, user_id, text, ts=ts, source=source,
sender_label=str(message.get("sender_label") or ""),
sender_session_id=str(message.get("sender_session_id") or ""),
client_message_id=client_message_id, transport=message.get("transport"),
client_surface=(metadata.get("client_surface") if isinstance(metadata, dict)
and isinstance(metadata.get("client_surface"), dict) else None),
)
ref = build_owner_message_ref(chat_id=chat_id, client_message_id=client_message_id, ts=ts, text=text)
if source != "web":
bridge.broadcast({
"type": "photo" if message.get("image_base64") else "chat", "role": "user",
"content": str(message.get("text") or ""), "caption": str(message.get("image_caption") or ""),
"image_base64": str(message.get("image_base64") or ""),
"mime": str(message.get("image_mime") or "image/jpeg"), "ts": ts, "source": source,
"sender_label": str(message.get("sender_label") or ""),
"sender_session_id": str(message.get("sender_session_id") or ""),
"client_message_id": client_message_id, "transport": message.get("transport") or {},
"chat_id": chat_id,
})
return ref
def _chat_media_download_url(task_id: str, data: bytes, mime: str) -> Tuple[str, str]:
@ -225,6 +330,7 @@ class LocalChatBridge:
"suppress_chat_log",
"task_constraint",
"task_metadata",
"accepted_source_ref",
):
value = msg.get(key)
if value not in (None, "", 0):
@ -332,6 +438,7 @@ class LocalChatBridge:
suppress_chat_log: bool = False,
task_constraint: Optional[Dict[str, Any]] = None,
task_metadata: Optional[Dict[str, Any]] = None,
accepted_source_ref: Optional[Dict[str, Any]] = None,
) -> None:
clean_text = str(text or "").strip()
caption_text = str(image_caption or "").strip()
@ -359,6 +466,7 @@ class LocalChatBridge:
"suppress_chat_log": bool(suppress_chat_log),
"task_constraint": dict(task_constraint or {}),
"task_metadata": dict(task_metadata or {}),
"accepted_source_ref": dict(accepted_source_ref or {}),
})
def send_message(
@ -1122,11 +1230,19 @@ def log_chat(
size_bytes: Optional[int] = None,
client_surface: Optional[Dict[str, Any]] = None,
message_meta: Optional[Dict[str, Any]] = None,
) -> None:
if DATA_DIR:
drive_root=None,
require_write: bool = False,
) -> Optional[dict]:
root = drive_root if drive_root is not None else DATA_DIR
if root:
from pathlib import Path
from ouroboros.utils import read_json_dict
root = Path(root)
record = {
"ts": ts or utc_now_iso(),
"session_id": load_state().get("session_id"),
"session_id": ((read_json_dict(root / "state" / "state.json") or {})
if drive_root is not None else load_state()).get("session_id"),
"direction": direction,
"chat_id": chat_id,
"user_id": user_id,
@ -1169,6 +1285,8 @@ def log_chat(
if key in meta:
record[key] = meta[key]
record.update(carry_cost_meta(meta))
if isinstance(meta.get("origin_message_ref"), dict):
record["origin_message_ref"] = dict(meta["origin_message_ref"])
if filename:
record["filename"] = filename
if mime:
@ -1190,7 +1308,13 @@ def log_chat(
record["quiz"] = dict(quiz)
if size_bytes is not None:
record["size_bytes"] = int(size_bytes)
append_jsonl(DATA_DIR / "logs" / "chat.jsonl", record)
written = append_jsonl(root / "logs" / "chat.jsonl", record, require_lock=require_write)
if require_write:
if not written:
raise RuntimeError("canonical message acceptance could not be persisted")
return record
elif require_write:
raise RuntimeError("canonical message acceptance requires a data root")
def send_with_budget(chat_id: int, text: str, log_text: Optional[str] = None,

View file

@ -163,14 +163,15 @@ def _handle_chat_direct_locked(
task_metadata: Optional[dict] = None,
) -> None:
from supervisor.state import budget_remaining, load_state
failure_meta = _host_operation_failure(task_metadata)
try:
remaining = budget_remaining(load_state(), strict=True)
except Exception:
_pool().send_with_budget(chat_id, "⚠️ Cost accounting is unavailable. Task was not dispatched; retry after ledger recovery.")
_pool().send_with_budget(chat_id, "⚠️ Cost accounting is unavailable. Task was not dispatched; retry after ledger recovery.", **failure_meta)
return
if remaining <= 0:
try:
_pool().send_with_budget(chat_id, "🚫 Budget exhausted. Task rejected. Please increase TOTAL_BUDGET in settings.")
_pool().send_with_budget(chat_id, "🚫 Budget exhausted. Task rejected. Please increase TOTAL_BUDGET in settings.", **failure_meta)
except Exception:
pass
return
@ -181,6 +182,15 @@ def _handle_chat_direct_locked(
)
def _host_operation_failure(metadata: Optional[dict]) -> dict:
"""Optional terminal correlation for the host's preaccepted skill messages."""
from ouroboros.task_finalization import host_operation_reply_kwargs
if (metadata or {}).get("_host_operation"):
return host_operation_reply_kwargs((metadata or {}).get("origin_message_ref"), "failed")
return {}
def _broadcast_task_named(msg: dict) -> None:
"""Bridge broadcast callback for the proactive namer (kept tiny + fail-soft)."""
try:
@ -285,6 +295,7 @@ def _run_chat_task(
_pool().send_with_budget(
chat_id,
f"⚠️ Task not started: every attachment was rejected.\n{rendered}",
**_host_operation_failure(task_metadata),
)
return
from ouroboros.artifacts import attachment_manifest_projection
@ -348,6 +359,7 @@ def _run_chat_task(
project_id=pid,
kind=kind,
phase="thinking",
origin_message_ref=task.get("origin_message_ref"),
):
# Announce the authoritative start immediately (owner decision 2A):
# the client's `Sending...` retires on this frame, not on a socket
@ -417,11 +429,15 @@ def _run_chat_task(
)
except Exception:
log.debug("Failed-turn typing announce failed", exc_info=True)
failure_meta = _host_operation_failure(task_metadata)
progress_meta = {"task_terminal_status": "failed"}
if failure_meta:
progress_meta["origin_message_ref"] = failure_meta["progress_meta"]["origin_message_ref"]
_pool().send_with_budget(
chat_id,
err_msg,
task_id=failed_task_id,
progress_meta={"task_terminal_status": "failed"},
progress_meta=progress_meta,
)
except Exception:
log.debug("Suppressed exception", exc_info=True)
@ -441,14 +457,15 @@ def handle_chat_ephemeral(
mode / effort, not a cheaper lane). Ephemeral turns are serialized among
themselves and are barred from long-term memory/reflection/evolution writes."""
from supervisor.state import budget_remaining, load_state
failure_meta = _host_operation_failure(task_metadata)
try:
remaining = budget_remaining(load_state(), strict=True)
except Exception:
_pool().send_with_budget(chat_id, "⚠️ Cost accounting is unavailable. Task was not dispatched; retry after ledger recovery.")
_pool().send_with_budget(chat_id, "⚠️ Cost accounting is unavailable. Task was not dispatched; retry after ledger recovery.", **failure_meta)
return
if remaining <= 0:
try:
_pool().send_with_budget(chat_id, "🚫 Budget exhausted. Task rejected. Please increase TOTAL_BUDGET in settings.")
_pool().send_with_budget(chat_id, "🚫 Budget exhausted. Task rejected. Please increase TOTAL_BUDGET in settings.", **failure_meta)
except Exception:
pass
return

View file

@ -146,6 +146,7 @@ def test_cancelled_inject_retains_copy_and_inflight_until_worker_settles(tmp_pat
asyncio.run(run())
def test_inject_refuses_files_outside_the_skill_state_and_bad_shapes(tmp_path):
bridge = FakeBridge()
client = _client(tmp_path, bridge)

View file

@ -428,19 +428,33 @@ def test_steering_and_project_mailbox_writers_pass_client_surface():
# exercised via write_owner_message round-trip; these pins catch a dropped
# kwarg at the two forwarding call sites).
steering = (REPO / "supervisor" / "steering.py").read_text(encoding="utf-8")
# v7 split: the project-mailbox write moved to the owner-routing leaf while the
# log_chat forwarding stayed in server.py, so the pin reads BOTH owners as one
# surface — the point is that neither call site loses the kwarg.
server_src = ((REPO / "server.py").read_text(encoding="utf-8")
+ (REPO / "ouroboros" / "server_owner_routing.py").read_text(encoding="utf-8"))
# The project mailbox belongs to owner routing; the ordinary inbound log
# now belongs to message_bus.record_inbound_message. Pin each actual owner
# separately so an unrelated extra forwarding site cannot hide a dropped one.
writers = {
"project mailbox": REPO / "ouroboros" / "server_owner_routing.py",
"inbound chat journal": REPO / "supervisor" / "message_bus.py",
}
assert "client_surface=" in steering, "steer mailbox write dropped client_surface"
# BOTH server call sites (project-mailbox write AND log_chat forwarding)
# must carry the kwarg — a single-substring pin went false-green when one
# of the two was dropped (final code review MAJOR).
assert server_src.count("client_surface=(") >= 2, (
"a server client_surface forwarding call site was dropped "
f"(found {server_src.count('client_surface=(')}, expected >= 2)"
for owner, path in writers.items():
assert "client_surface=(" in path.read_text(encoding="utf-8"), f"{owner} dropped client_surface"
def test_inbound_message_owner_persists_the_original_client_surface(tmp_path, monkeypatch):
from supervisor import message_bus
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
monkeypatch.setattr(message_bus, "load_state", lambda: {"session_id": "surface-fixture"})
surface = {"pywebview": False, "ua": "Phone/1.0"}
message = {"source": "web", "task_metadata": {"client_surface": surface}}
ref = message_bus.record_inbound_message(
SimpleNamespace(), message, chat_id=1, user_id=1,
client_message_id="surface-message", text="owner input", ts="2026-09-07T00:00:00Z",
)
row = json.loads((tmp_path / "logs" / "chat.jsonl").read_text(encoding="utf-8"))
assert row["client_surface"] == surface
assert row["client_message_id"] == ref["client_message_id"] == "surface-message"
assert row["text"] == "owner input" and row["direction"] == "in"
def test_presentation_env_is_stripped_by_benchmark_server_runner():

View file

@ -21,6 +21,20 @@ PACKAGE = REPO / "ouroboros"
_LEAVES = (settings_defaults, settings_scales, model_slots, review_model_routes, runtime_limits)
_MOVED_OWNERS = {
"WORKER_SPAWN_GRACE_SEC": runtime_limits,
"WORKER_READY_WINDOW_SEC": runtime_limits,
"WORKER_READY_MAX_ATTEMPTS": runtime_limits,
"EXTENSION_STREAM_CHUNK_BYTES": runtime_limits,
"EXTENSION_CHILD_CLEANUP_GRACE_SEC": runtime_limits,
"NESTED_SETTLEMENT_MARGIN_SEC": runtime_limits,
"NETWORK_WAIT_NOTE_INTERVAL_SEC": runtime_limits,
"NETWORK_WAIT_BACKOFF_START_SEC": runtime_limits,
"TCP_KEEPALIVE_IDLE_SEC": runtime_limits,
"TCP_KEEPALIVE_INTERVAL_SEC": runtime_limits,
"TCP_KEEPALIVE_PROBE_COUNT": runtime_limits,
"EXTENSION_STREAM_METADATA_BYTES": runtime_limits,
"WS_RELAY_BURST": runtime_limits,
"WS_RELAY_REFILL_PER_SEC": runtime_limits,
"ENDPOINT_AUTHORED_SETTINGS": settings_defaults,
# v6.104.0 upstream: the OpenRouter shipped-model defaults arrive in the
# vocabulary leaf the v7 split created for exactly this class of fact.

View file

@ -185,6 +185,7 @@ def test_run_child_marks_only_post_popen_exceptions(tmp_path, monkeypatch):
class _FakeProc:
pid = 4242
returncode = 0
stdin = None
stdout = io.BytesIO(b"")
stderr = io.BytesIO(b"")
@ -214,6 +215,7 @@ def _fake_proc_cls():
returncode = 0
def __init__(self):
self.stdin = None
self.stdout = io.BytesIO(b"")
self.stderr = io.BytesIO(b"")

View file

@ -178,12 +178,21 @@ def test_the_public_loader_surface_is_unchanged():
def test_extension_extraction_size_bounds_have_meaningful_headroom():
from ouroboros.review import BAND_MODULE_MAX_LINES, TARGET_MODULE_LINES
from ouroboros.size_ratchet_manifest import BAND_PATHS
counts = {
module.__name__: len(
pathlib.Path(module.__file__).read_text(encoding="utf-8").splitlines()
)
for module in (extension_loader, *_LEAVES)
}
assert counts["ouroboros.extension_loader"] <= 1000
assert all(count <= 1000 for count in counts.values())
assert 600 <= counts["ouroboros.extension_plugin_api"] <= 1000
assert counts["ouroboros.extension_loader"] <= TARGET_MODULE_LINES
for module, count in counts.items():
path = module.replace(".", "/") + ".py"
limit = TARGET_MODULE_LINES
if path in BAND_PATHS:
assert BAND_PATHS[path], f"new extension band entry needs its rationale: {path}"
limit = BAND_MODULE_MAX_LINES
assert count <= limit, f"{path}: {count} lines exceeds its {limit}-line bound"
assert counts["ouroboros.extension_plugin_api"] >= 600

View file

@ -3,6 +3,7 @@ from __future__ import annotations
import asyncio
import base64
import json
import pathlib
import sys
import time
@ -31,6 +32,7 @@ def _payload(result) -> str:
return head if sep else text
from tests._shared import clean_extension_runtime_state
from tests.test_extension_route_streaming import collect_response
from tests.test_extension_loader import (
_add_fake_native_dep,
_isolated_site_packages_dir,
@ -318,8 +320,9 @@ def test_native_risk_extension_route_dispatches_out_of_process(tmp_path):
repo_dir=pathlib.Path(__file__).resolve().parents[1],
)
assert result["route"]["kind"] == "json"
assert result["route"]["data"] == {
events = asyncio.run(collect_response(result, method="POST"))
assert events[0]["status"] == 200
assert json.loads(b"".join(event.get("body", b"") for event in events)) == {
"value": "isolated-native-risk",
"name": "anton",
"skill": "native_route",
@ -328,7 +331,7 @@ def test_native_risk_extension_route_dispatches_out_of_process(tmp_path):
assert pathlib.Path(spec["skills_repo_path"]) == repo_root
def test_native_risk_extension_streaming_route_is_materialized_out_of_process(tmp_path):
def test_native_risk_extension_streaming_route_streams_out_of_process(tmp_path):
from ouroboros.extension_process_runner import dispatch_extension_route_subprocess
plugin = (
@ -366,8 +369,9 @@ def test_native_risk_extension_streaming_route_is_materialized_out_of_process(tm
repo_dir=pathlib.Path(__file__).resolve().parents[1],
)
assert result["route"]["kind"] == "response"
assert base64.b64decode(result["route"]["body_b64"]) == b"chunk-a-chunk-b"
events = asyncio.run(collect_response(result))
assert events[0]["status"] == 200
assert b"".join(event.get("body", b"") for event in events) == b"chunk-a-chunk-b"
def test_native_risk_extension_gateway_route_child_failure_returns_502(tmp_path, monkeypatch):
@ -421,8 +425,9 @@ def test_native_risk_extension_gateway_route_child_failure_returns_502(tmp_path,
response = asyncio.run(api_extension_dispatch(request))
assert response.status_code == 502
assert b"route-child-boom" in response.body
events = asyncio.run(collect_response(response))
assert events[0]["status"] == 502
assert b"route-child-boom" in b"".join(event.get("body", b"") for event in events)
def test_native_risk_extension_gateway_route_rejects_oversized_body_before_child(tmp_path, monkeypatch):

View file

@ -0,0 +1,597 @@
"""Real dependency-bearing child responses retain ASGI semantics and lifetime."""
import asyncio
import json
import pathlib
import sys
import pytest
from ouroboros import extension_loader
from ouroboros.extension_process_runner import dispatch_extension_route_subprocess
from tests._shared import clean_extension_runtime_state
from tests.test_extension_loader import _prepare_extension, _add_fake_native_dep, _mark_isolated_deps_installed
from tests.test_widget_stream_download_ui import widget_server as widget_server
@pytest.fixture(autouse=True)
def clean_runtime(monkeypatch):
monkeypatch.setenv("OUROBOROS_RUNTIME_MODE", "advanced")
clean_extension_runtime_state()
yield
clean_extension_runtime_state()
def prepare_route(tmp_path, plugin, *, name="stream_skill"):
loaded, skills_root, drive_root = _prepare_extension(tmp_path, name, plugin,
permissions=["route"], extra_frontmatter="dependencies:\n - dummy_pkg\n")
_add_fake_native_dep(loaded)
_mark_isolated_deps_installed(drive_root, loaded)
assert extension_loader.load_extension(loaded, lambda: {}, drive_root=drive_root) is None
spec = extension_loader.list_routes()[f"/api/extensions/{name}/stream"]
return loaded, spec, drive_root
def route_response(spec, drive_root, *, method="GET", headers=()):
return dispatch_extension_route_subprocess(spec, {
"method": method, "path": spec["path"], "headers": list(headers), "body_b64": "",
}, drive_root=drive_root, repo_dir=pathlib.Path(__file__).resolve().parents[1])
async def collect_response(response, *, method="GET", observer=None):
events = []
async def receive():
await asyncio.Event().wait()
async def send(message):
events.append(message)
if observer is not None:
await observer(message)
await response({"type": "http", "method": method}, receive, send)
return events
def test_response_streams_late_dependency_import_and_keeps_duplicate_headers(tmp_path):
plugin = '''from starlette.responses import StreamingResponse
async def chunks():
import dummy_pkg
yield dummy_pkg.VALUE.encode()
yield b'-tail'
def stream(request):
r = StreamingResponse(chunks(), media_type='text/plain')
r.raw_headers.extend([(b'x-item', b'one'), (b'x-item', b'two')])
return r
def register(api):
api.register_route('stream', stream)
'''
_, spec, root = prepare_route(tmp_path, plugin)
response = route_response(spec, root)
events = asyncio.run(collect_response(response))
assert events[0]["status"] == 200, events
assert events[0]["headers"][-2:] == [(b"x-item", b"one"), (b"x-item", b"two")]
assert b"".join(e.get("body", b"") for e in events) == b"isolated-native-risk-tail"
assert events[-1]["more_body"] is False
assert list((root / 'state/skills/stream_skill/extension_calls').glob('*.result.json')) == []
def test_child_failure_before_headers_is_502(tmp_path):
_, spec, root = prepare_route(tmp_path, '''def stream(request):
raise RuntimeError('before-headers')
def register(api):
api.register_route('stream', stream)
''')
events = asyncio.run(collect_response(route_response(spec, root)))
assert events[0]["status"] == 502
assert "before-headers" in json.loads(events[-1]["body"])["error"]
@pytest.mark.parametrize(('method', 'range_header', 'status', 'expected'), [
('GET', '', 200, None), ('HEAD', '', 200, b''),
('GET', 'bytes=2-8', 206, b'2345678'), ('HEAD', 'bytes=2-8', 206, b''),
('GET', 'bytes=99999999-', 416, b''),
])
def test_file_response_preserves_large_bytes_head_and_range(tmp_path, method, range_header, status, expected):
plugin = '''from starlette.responses import FileResponse
def stream(request):
return FileResponse(request.app.state.drive_root / 'large.bin', filename='export.bin')
def register(api):
api.register_route('stream', stream)
'''
_, spec, root = prepare_route(tmp_path, plugin)
content = b'0123456789' * (128 * 1024)
(root / 'large.bin').write_bytes(content)
headers = [('range', range_header)] if range_header else []
events = asyncio.run(collect_response(route_response(spec, root, method=method, headers=headers), method=method))
assert events[0]['status'] == status, events
body = b''.join(e.get('body', b'') for e in events)
assert body == (content if expected is None else expected)
response_headers = dict(events[0]['headers'])
if status == 200:
assert response_headers[b'content-length'] == str(len(content)).encode()
assert b'export.bin' in response_headers[b'content-disposition']
elif status == 206:
assert response_headers[b'content-range'] == f'bytes 2-8/{len(content)}'.encode()
assert response_headers[b'content-length'] == b'7'
def test_first_bytes_precede_producer_completion(tmp_path):
plugin = '''import asyncio
from starlette.responses import StreamingResponse
def stream(request):
async def chunks():
yield b'first'
while not (request.app.state.drive_root / 'release').exists():
await asyncio.sleep(.01)
yield b'last'
return StreamingResponse(chunks())
def register(api):
api.register_route('stream', stream)
'''
_, spec, root = prepare_route(tmp_path, plugin)
from ouroboros.tools.shell import _active_subprocesses
async def observe(message):
if message.get('body') == b'first':
assert any(proc.poll() is None for proc in _active_subprocesses)
assert not (root / 'release').exists()
(root / 'release').touch()
events = asyncio.run(collect_response(route_response(spec, root), observer=observe))
assert b''.join(e.get('body', b'') for e in events) == b'firstlast'
assert not _active_subprocesses
@pytest.mark.parametrize('failure', [False, True])
def test_background_work_survives_normal_body_completion(tmp_path, failure, caplog):
plugin = '''import asyncio
from starlette.responses import Response
from starlette.background import BackgroundTask
def stream(request):
async def after():
import dummy_pkg
await asyncio.sleep(.05)
(request.app.state.drive_root / 'background.done').write_text(dummy_pkg.VALUE)
FAIL
return Response(b'delivered', background=BackgroundTask(after))
def register(api):
api.register_route('stream', stream)
'''.replace('FAIL', "raise RuntimeError('cleanup-failed')" if failure else 'pass')
_, spec, root = prepare_route(tmp_path, plugin)
response = route_response(spec, root)
async def run():
finished = asyncio.Event()
events = []
async def receive():
await finished.wait()
return {'type': 'http.disconnect'}
async def send(message):
events.append(message)
if message['type'] == 'http.response.body' and not message.get('more_body'):
finished.set()
await response({'type': 'http', 'method': 'GET'}, receive, send)
return events
events = asyncio.run(run())
assert events[0]['status'] == 200
assert events[-1]['body'] == b'delivered'
assert (root / 'background.done').read_text() == 'isolated-native-risk'
if failure:
assert 'cleanup-failed' in caplog.text
def test_midbody_error_aborts_instead_of_success(tmp_path):
from ouroboros.extension_process_runner import ExtensionProcessError
_, spec, root = prepare_route(tmp_path, '''from starlette.responses import StreamingResponse
async def chunks():
yield b'partial'
raise RuntimeError('midbody-failed')
def stream(request):
return StreamingResponse(chunks())
def register(api):
api.register_route('stream', stream)
''')
response = route_response(spec, root)
with pytest.raises(ExtensionProcessError, match='midbody-failed'):
asyncio.run(collect_response(response))
assert response.body_complete is False
def test_child_uses_selected_roots(tmp_path):
_, spec, root = prepare_route(tmp_path, '''import os, sys, pathlib, ouroboros
def stream(request):
return {'finder': any('__editable___ouroboros_' in getattr(f, '__module__', '') for f in sys.meta_path),
'module': ouroboros.__file__, 'python': sys.executable,
'roots': {k: os.environ[k] for k in ('OUROBOROS_APP_ROOT','OUROBOROS_REPO_DIR','OUROBOROS_DATA_DIR','OUROBOROS_SETTINGS_PATH')}}
def register(api):
api.register_route('stream', stream)
''')
events = asyncio.run(collect_response(route_response(spec, root)))
body = json.loads(b''.join(e.get('body', b'') for e in events))
assert body['python'] == sys.executable
print('CHILD_BINDING', json.dumps(body, sort_keys=True))
assert pathlib.Path(body['module']).resolve().is_relative_to(pathlib.Path(__file__).resolve().parents[1])
assert body['roots']['OUROBOROS_DATA_DIR'] == str(root)
assert body['roots']['OUROBOROS_SETTINGS_PATH'] == str(root / 'settings.json')
assert body['roots']['OUROBOROS_APP_ROOT'] == str(root.parent)
def test_disconnect_before_headers_reaps_the_child(tmp_path):
_, spec, root = prepare_route(tmp_path, '''import asyncio
async def stream(request):
await asyncio.sleep(90)
def register(api):
api.register_route('stream', stream)
''')
from ouroboros.tools.shell import _active_subprocesses
response = route_response(spec, root)
async def run():
async def receive():
while not _active_subprocesses:
await asyncio.sleep(.01)
return {'type': 'http.disconnect'}
async def send(_message):
pytest.fail('disconnected request should not publish headers')
await response({'type': 'http', 'method': 'GET'}, receive, send)
assert response.client_disconnected is True
asyncio.run(run())
assert not _active_subprocesses
def test_unload_cancels_only_the_observed_skill_instance(tmp_path):
plugin = '''import asyncio
from starlette.responses import StreamingResponse
def stream(request):
async def chunks():
yield b'first'
await asyncio.sleep(90)
return StreamingResponse(chunks())
def register(api):
api.register_route('stream', stream)
'''
one_root = tmp_path / 'one'
two_root = tmp_path / 'two'
one_root.mkdir()
two_root.mkdir()
_, one, root_one = prepare_route(one_root, plugin, name='one')
_, two, root_two = prepare_route(two_root, plugin, name='two')
first = route_response(one, root_one)
second = route_response(two, root_two)
async def run():
first_ready, second_ready = asyncio.Event(), asyncio.Event()
async def observe_one(message):
if message.get('body') == b'first':
first_ready.set()
async def observe_two(message):
if message.get('body') == b'first':
second_ready.set()
tasks = [asyncio.create_task(collect_response(first, observer=observe_one)),
asyncio.create_task(collect_response(second, observer=observe_two))]
try:
await asyncio.wait_for(asyncio.gather(first_ready.wait(), second_ready.wait()), 20)
assert extension_loader.unload_extension('one', expected_generation='not-this-generation') is False
assert not first.cancelled and not second.cancelled
assert extension_loader.unload_extension('one', expected_generation=one['extension_generation']) is True
with pytest.raises(asyncio.CancelledError):
await tasks[0]
assert not second.cancelled and not tasks[1].done()
finally:
second.cancel()
await asyncio.gather(*tasks, return_exceptions=True)
asyncio.run(run())
from ouroboros.tools.shell import _active_subprocesses
assert not _active_subprocesses
@pytest.mark.integration
def test_slow_reader_outlives_the_retired_sixty_second_limit(tmp_path):
"""A real child blocked by delivery backpressure remains alive beyond 60s."""
plugin = '''from starlette.responses import FileResponse
def stream(request):
return FileResponse(request.app.state.drive_root / 'large.bin')
def register(api):
api.register_route('stream', stream)
'''
_, spec, root = prepare_route(tmp_path, plugin)
(root / 'large.bin').write_bytes(b'x' * (2 * 1024 * 1024))
from ouroboros.tools.shell import _active_subprocesses
paused = False
async def slow(message):
nonlocal paused
if message.get('body') and not paused:
paused = True
await asyncio.sleep(65)
assert any(proc.poll() is None for proc in _active_subprocesses)
events = asyncio.run(collect_response(route_response(spec, root), observer=slow))
assert paused
assert sum(len(e.get('body', b'')) for e in events) == 2 * 1024 * 1024
assert not _active_subprocesses
def test_failure_sending_final_body_is_not_cleanup_success(tmp_path):
_, spec, root = prepare_route(tmp_path, "def stream(request):\n return 'final'\ndef register(api):\n api.register_route('stream', stream)\n")
response = route_response(spec, root)
async def fail_send(message):
if message['type'] == 'http.response.body':
raise OSError('client disconnected during final send')
asyncio.run(collect_response(response, observer=fail_send))
assert response.client_disconnected is True
assert response.body_complete is False
@pytest.mark.parametrize('method', ['GET', 'HEAD'])
def test_content_length_close_preserves_file_background(tmp_path, method):
_, spec, root = prepare_route(tmp_path, """from starlette.responses import FileResponse
from starlette.background import BackgroundTask
def stream(request):
def after():
(request.app.state.drive_root / 'file-background').write_text('finished')
return FileResponse(request.app.state.drive_root / 'large.bin', background=BackgroundTask(after))
def register(api):
api.register_route('stream', stream)
""")
size = 64 * 1024
(root / 'large.bin').write_bytes(b'x' * size)
response = route_response(spec, root, method=method)
async def run():
disconnected, sent = asyncio.Event(), 0
async def receive():
await disconnected.wait()
return {'type':'http.disconnect'}
async def send(message):
nonlocal sent
if message['type'] == 'http.response.start' and method == 'HEAD':
disconnected.set()
await asyncio.sleep(.01)
elif message['type'] == 'http.response.body':
if disconnected.is_set():
raise OSError('client closed after promised body')
sent += len(message.get('body', b''))
if sent == size:
assert message['more_body'] is True, 'exact FileResponse block ends with a later empty body'
disconnected.set()
await asyncio.sleep(.01)
await response({'type':'http','method':method}, receive, send)
assert sent == (0 if method == 'HEAD' else size)
asyncio.run(run())
assert response.body_complete is True
assert (root / 'file-background').read_text() == 'finished'
def test_disconnect_during_body_runs_child_generator_cleanup(tmp_path):
_, spec, root = prepare_route(tmp_path, """import asyncio
from starlette.responses import StreamingResponse
def stream(request):
async def chunks():
try:
yield b'first'
await asyncio.sleep(90)
finally:
(request.app.state.drive_root / 'cancel-cleanup').write_text('done')
return StreamingResponse(chunks())
def register(api):
api.register_route('stream', stream)
""")
response = route_response(spec, root)
async def run():
first = asyncio.Event()
async def receive():
await first.wait()
return {'type':'http.disconnect'}
async def send(message):
if message.get('body') == b'first':
first.set()
await response({'type':'http','method':'GET'}, receive, send)
asyncio.run(run())
assert response.client_disconnected and not response.body_complete
assert (root / 'cancel-cleanup').read_text() == 'done'
from ouroboros.tools.shell import _active_subprocesses
assert not _active_subprocesses
def test_frame_writer_handles_partial_pipe_writes():
import io
from ouroboros.extension_route_stream import _write_frame, _read_frame
class PartialPipe(io.BytesIO):
def write(self, data):
return super().write(data[:3])
pipe = PartialPipe()
_write_frame(pipe, b'B', b'bytes across several writes')
pipe.seek(0)
assert _read_frame(pipe) == (b'B', b'bytes across several writes')
def test_reload_does_not_let_old_response_cancel_new_generation(tmp_path):
loaded, spec, root = prepare_route(tmp_path, '''import asyncio
from starlette.responses import StreamingResponse
def stream(request):
async def chunks():
yield b'first'
await asyncio.sleep(90)
return StreamingResponse(chunks())
def register(api):
api.register_route('stream', stream)
''')
old = route_response(spec, root)
async def run():
ready = asyncio.Event()
async def observe(message):
if message.get('body') == b'first':
ready.set()
old_task = asyncio.create_task(collect_response(old, observer=observe))
await asyncio.wait_for(ready.wait(), 20)
extension_loader.unload_extension(loaded.name, expected_generation=spec['extension_generation'])
with pytest.raises(asyncio.CancelledError):
await old_task
assert await asyncio.to_thread(extension_loader.load_extension, loaded, lambda: {}, drive_root=root) is None
newer = extension_loader.list_routes()[spec['path']]
assert newer['extension_generation'] != spec['extension_generation']
current = route_response(newer, root)
ready.clear()
current_task = asyncio.create_task(collect_response(current, observer=observe))
try:
await asyncio.wait_for(ready.wait(), 20)
old.cancel()
assert extension_loader.unload_extension(loaded.name, expected_generation=spec['extension_generation']) is False
assert not current.cancelled and not current_task.done()
finally:
current.cancel()
await asyncio.gather(current_task, return_exceptions=True)
asyncio.run(run())
@pytest.mark.serial
def test_real_http_second_route_and_module_respond_during_child_startup(widget_server, monkeypatch):
import threading
import httpx
from ouroboros import process_custody
from ouroboros.tools.shell import _active_subprocesses
from ouroboros.extension_route_stream import RouteStreamResponse
settled = threading.Event()
completed = 0
original_response = RouteStreamResponse.__call__
async def observe_completion(self, *args):
nonlocal completed
try:
return await original_response(self, *args)
finally:
completed += 1
if completed == 2:
settled.set()
monkeypatch.setattr(RouteStreamResponse, "__call__", observe_completion)
entered, release = threading.Event(), threading.Event()
calls = []
original = process_custody.record_process
def held_record(*args, **kwargs):
calls.append(kwargs["pid"])
if len(calls) == 1:
entered.set()
assert release.wait(15), "test did not release its startup barrier"
return original(*args, **kwargs)
monkeypatch.setattr(process_custody, "record_process", held_record)
async def run():
async with httpx.AsyncClient(base_url=widget_server["url"], timeout=10) as client:
first = asyncio.create_task(client.get("/api/extensions/export_widget/stream"))
try:
assert await asyncio.to_thread(entered.wait, 10)
module = await client.get("/api/extensions/export_widget/module/widget.js")
assert module.status_code == 200 and b"OuroborosWidget" in module.content
assert len(calls) == 1, "captured static module sources must not spawn children"
second = await client.get("/api/extensions/export_widget/export")
assert second.status_code == 200
assert second.content == widget_server["expected"]["widget-large.bin"]
assert len(calls) == 2 and not first.done()
finally:
release.set()
(widget_server["root"] / "release-stream").touch()
result = await first
assert result.status_code == 200 and result.content == b"firstlast"
asyncio.run(run())
# Receiving the final HTTP byte precedes the permitted child/background cleanup.
assert settled.wait(10), "both actual response handles must finish cleanup"
assert not _active_subprocesses
@pytest.mark.serial
@pytest.mark.parametrize("startup_fails", [False, True])
def test_real_http_startup_cancel_retains_spawned_child_until_cleanup(widget_server, monkeypatch, startup_fails):
import threading
import httpx
from ouroboros import process_custody, extension_registry_state
from ouroboros.tools.shell import _active_subprocesses
entered, release = threading.Event(), threading.Event()
children = []
original = process_custody.record_process
def held_record(*args, **kwargs):
children.extend(proc for proc in _active_subprocesses if proc.pid == kwargs["pid"])
entered.set()
assert release.wait(15), "test did not release its startup barrier"
result = original(*args, **kwargs)
if startup_fails:
raise RuntimeError("controlled failure after real process registration")
return result
monkeypatch.setattr(process_custody, "record_process", held_record)
async def run():
async with httpx.AsyncClient(base_url=widget_server["url"], timeout=10) as client:
request = asyncio.create_task(client.get("/api/extensions/export_widget/stream"))
try:
assert await asyncio.to_thread(entered.wait, 10)
assert len(children) == 1 and children[0].poll() is None
with extension_registry_state._lock:
bundle = extension_registry_state._extensions["export_widget"]
work, = bundle.supervised_futures
work.cancel()
work.loop.call_soon_threadsafe(work.task.cancel) # repeated cancellation
module = await client.get("/api/extensions/export_widget/module/widget.js")
assert module.status_code == 200
assert not work.task.done(), "startup worker still owns an unreturned child"
finally:
release.set()
try:
response = await request
assert response.status_code == 500 # cancelled before ASGI headers
except httpx.RemoteProtocolError:
pass # uvicorn may close a cancelled unstarted response
for _ in range(100):
if work.task.done():
break
await asyncio.sleep(.01)
assert work.task.done()
assert all(proc.poll() is not None for proc in children)
assert not _active_subprocesses
assert work not in bundle.supervised_futures
calls_dir = widget_server["root"] / "state/skills/export_widget/extension_calls"
assert list(calls_dir.iterdir()) == []
asyncio.run(run())
@pytest.mark.parametrize("kind,limit", [(b"B", "body"), (b"S", "metadata")])
def test_stream_frame_allocation_is_bounded_without_limiting_response_length(kind, limit):
import io
import struct
from ouroboros.config import EXTENSION_STREAM_CHUNK_BYTES, EXTENSION_STREAM_METADATA_BYTES
from ouroboros.extension_route_stream import _read_frame
bound = EXTENSION_STREAM_CHUNK_BYTES + 1 if limit == "body" else EXTENSION_STREAM_METADATA_BYTES
class Observed(io.BytesIO):
def __init__(self, value):
super().__init__(value)
self.read_sizes = []
def read(self, size=-1):
self.read_sizes.append(size)
return super().read(size)
over = Observed(struct.pack("!I", bound + 2) + kind)
with pytest.raises(ValueError, match="channel bound"):
_read_frame(over)
assert over.read_sizes == [4, 1], "oversized payload must be refused before allocation"
huge = Observed(struct.pack("!I", 0xffffffff) + kind)
with pytest.raises(ValueError, match="channel bound"):
_read_frame(huge)
assert huge.read_sizes == [4, 1]
# Multiple legal frames remain readable; there is no cumulative body cap.
frame = struct.pack("!I", bound + 1) + kind + b"x" * bound
valid = Observed(frame * 3)
for _ in range(3):
assert _read_frame(valid) == (kind, b"x" * bound)
def test_native_route_crash_retains_drained_stderr_and_exit_code(tmp_path, caplog):
import logging
from ouroboros.tools.shell import _active_subprocesses
_, spec, drive = prepare_route(tmp_path, """import os, sys
def stream(request):
sys.stderr.write('CONTROLLED_NATIVE_CRASH\\n')
sys.stderr.flush()
os._exit(7)
def register(api):
api.register_route('stream', stream)
""")
with caplog.at_level(logging.WARNING, logger="ouroboros.extension_route_stream"):
messages = asyncio.run(collect_response(route_response(spec, drive)))
assert messages[0]["status"] == 502
assert "CONTROLLED_NATIVE_CRASH" in caplog.text and "returncode=7" in caplog.text
assert not _active_subprocesses

View file

@ -0,0 +1,199 @@
"""Successful extension children disclose best-effort relay loss without failing work."""
import concurrent.futures
import contextlib
import http.server
import json
import os
from pathlib import Path
import socket
import threading
import time
import pytest
from tests._extension_loader_shared import _prepare_extension
from tests._shared import clean_extension_runtime_state
@pytest.fixture(autouse=True)
def relay_state(monkeypatch):
from ouroboros.extension_plugin_api import take_child_ws_relay_failures
from ouroboros.tools.process_facts import consume_last_process_facts
monkeypatch.setenv("OUROBOROS_RUNTIME_MODE", "advanced")
clean_extension_runtime_state()
take_child_ws_relay_failures()
consume_last_process_facts()
yield
take_child_ws_relay_failures()
consume_last_process_facts()
clean_extension_runtime_state()
@contextlib.contextmanager
def relay_endpoint(status):
"""Own every endpoint; never use the installation's real Host Service."""
calls = []
if status == "missing_env":
yield {}, calls
return
if status == "connection_refused":
with socket.socket() as sock:
sock.bind(("127.0.0.1", 0))
yield {"HOST_SERVICE_URL": f"http://127.0.0.1:{sock.getsockname()[1]}",
"HOST_SERVICE_TOKEN": "ws-test-fixture"}, calls
return
class Handler(http.server.BaseHTTPRequestHandler):
def do_POST(self):
self.rfile.read(int(self.headers.get("content-length", "0")))
calls.append(self.path)
body = json.dumps({"ok": status == 202}).encode()
self.send_response(status)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, *_args):
pass
server = http.server.HTTPServer(("127.0.0.1", 0), Handler)
thread = threading.Thread(target=server.serve_forever)
thread.start()
try:
yield {"HOST_SERVICE_URL": f"http://127.0.0.1:{server.server_address[1]}",
"HOST_SERVICE_TOKEN": "ws-test-fixture"}, calls
finally:
server.shutdown()
server.server_close()
thread.join(5)
assert not thread.is_alive()
def child_env(tmp_path, drive, repo):
env = dict(os.environ)
env.pop("HOST_SERVICE_TOKEN", None)
env.pop("HOST_SERVICE_URL", None)
env.update({"OUROBOROS_EXTENSION_PROCESS_CHILD": "1",
"OUROBOROS_APP_ROOT": str(tmp_path), "OUROBOROS_REPO_DIR": str(repo),
"OUROBOROS_DATA_DIR": str(drive),
"OUROBOROS_SETTINGS_PATH": str(drive / "settings.json")})
return env
def run_relay_child(tmp_path, status, *, mode="tool", concurrent=False):
from ouroboros import extension_process_runner as runner
from ouroboros.extension_surface_names import extension_surface_name
plugin = '''def register(api):
def send(index):
assert api.send_ws_message('progress', {'index': index}) is None
def run(ctx):
SEND
return 'successful producer result'
api.register_tool('run', run, description='relay fixture', schema={})
api.register_ws_handler('run', run)
api.register_route('run', lambda request: {'value': run(request)})
CATALOG
'''.replace('SEND', "import concurrent.futures\n with concurrent.futures.ThreadPoolExecutor(max_workers=4) as pool:\n list(pool.map(send, range(1000)))" if concurrent else "for index in range(5):\n send(index)")
plugin = plugin.replace('CATALOG', 'run(None)' if mode == 'catalog' else 'pass')
skill, skills_repo, drive = _prepare_extension(
tmp_path, "ws_diagnostic", plugin, permissions=["tool", "ws_handler", "route"])
repo = Path(__file__).resolve().parents[1]
env = child_env(tmp_path, drive, repo)
with relay_endpoint(status) as (bridge, requests):
env.update(bridge)
surface = (f"/api/extensions/{skill.name}/run" if mode == 'route'
else extension_surface_name(skill.name, "run"))
result = runner._run_child(
{"mode": mode, "skill_name": skill.name, "surface": surface,
"args": {}, "message": {}, "request": {"method": "GET"},
"drive_root": str(drive), "repo_dir": str(repo),
"skills_repo_path": str(skills_repo)},
skill_dir=skill.skill_dir, drive_root=drive, repo_dir=repo,
env=env, timeout_sec=30)
return result, requests
@pytest.mark.parametrize(("status", "reason"), [
(202, None), (429, "rate_limited"), (418, "http_client_error"),
(500, "http_server_error"), (302, "http_error"),
("connection_refused", "transport_error"), ("missing_env", "missing_transport"),
])
def test_successful_child_publishes_relay_aggregate(tmp_path, status, reason, caplog):
from ouroboros.tools.process_facts import consume_last_process_facts
result, requests = run_relay_child(tmp_path, status)
assert result["ok"] is True
assert result["result"] == "successful producer result"
facts = consume_last_process_facts()
assert facts["exit_code"] == 0
reports = [row for row in caplog.records if "WS relay failures" in row.getMessage()]
if reason:
assert result["ws_relay_failures"] == facts["ws_relay_failures"] == {reason: 5}
assert len(reports) == 1
assert "ws_diagnostic" in reports[0].getMessage()
assert "ws-test-fixture" not in reports[0].getMessage()
assert "http://" not in reports[0].getMessage()
else:
assert "ws_relay_failures" not in result and "ws_relay_failures" not in facts
assert reports == [] # Accepted does not become a browser-delivery claim.
assert len(requests) == (5 if isinstance(status, int) else 0)
@pytest.mark.parametrize("mode", ["catalog", "ws"])
def test_one_shot_modes_share_diagnostic_delivery(tmp_path, mode, caplog):
from ouroboros.tools.process_facts import consume_last_process_facts
result, _ = run_relay_child(tmp_path, "missing_env", mode=mode)
assert result["ok"] is True
assert result["ws_relay_failures"] == {"missing_transport": 5}
assert consume_last_process_facts()["ws_relay_failures"] == {"missing_transport": 5}
assert len([row for row in caplog.records if "WS relay failures" in row.getMessage()]) == 1
def test_concurrent_child_sends_have_one_bounded_aggregate(tmp_path, caplog):
result, _ = run_relay_child(tmp_path, "missing_env", concurrent=True)
assert result["result"] == "successful producer result"
assert result["ws_relay_failures"] == {"missing_transport": 1000}
assert len(json.dumps(result["ws_relay_failures"])) < 100
assert len([row for row in caplog.records if "WS relay failures" in row.getMessage()]) == 1
def test_reported_diagnostics_cannot_replace_measured_process_facts():
from ouroboros.tools.process_facts import consume_last_process_facts, publish_process_facts
publish_process_facts(returncode=0, started_ts=time.monotonic(), ws_relay_failures={
"exit_code": -9, "url": "private", "rate_limited": 3,
"http_error": "private", "missing_transport": True,
"http_client_error": -1, "http_server_error": 0,
})
facts = consume_last_process_facts()
assert facts["exit_code"] == 0 and "signal" not in facts
assert facts["ws_relay_failures"] == {"rate_limited": 3}
publish_process_facts(returncode=0, started_ts=time.monotonic())
assert "ws_relay_failures" not in consume_last_process_facts()
def test_collector_drains_without_cross_call_leakage():
from ouroboros.extension_plugin_api import _record_ws_relay_failure, take_child_ws_relay_failures
with concurrent.futures.ThreadPoolExecutor(max_workers=4) as pool:
list(pool.map(lambda _: _record_ws_relay_failure("transport_error"), range(2000)))
assert take_child_ws_relay_failures() == {"transport_error": 2000}
assert take_child_ws_relay_failures() == {}
def test_real_child_diagnostic_reaches_successful_tool_trace(tmp_path):
from tests.test_process_signal_observability import _run_single
def execute(_name, _args):
result, _ = run_relay_child(tmp_path, "missing_env")
return result["result"]
result = _run_single(tmp_path, execute, tool="ext_ws_diagnostic_run")
assert result["result"] == "successful producer result"
assert result["result_meta"]["exit_code"] == 0
assert result["result_meta"]["ws_relay_failures"] == {"missing_transport": 5}
rows = [json.loads(line) for line in (tmp_path / "logs" / "tools.jsonl").read_text().splitlines()]
assert rows[-1]["ws_relay_failures"] == {"missing_transport": 5}

View file

@ -0,0 +1,97 @@
"""The route completion frame carries the same relay aggregate after body delivery."""
import asyncio
import json
from pathlib import Path
import pytest
from ouroboros import extension_process_runner as runner
from ouroboros import extension_loader
from ouroboros.extension_route_stream import RouteStreamResponse
from ouroboros.tools.process_facts import consume_last_process_facts
from tests._extension_loader_shared import _prepare_extension
from tests.test_extension_ws_diagnostics import child_env, relay_endpoint, relay_state # noqa: F401
@pytest.mark.parametrize(("status", "reason"), [
(202, None), (429, "rate_limited"), (500, "http_server_error"),
("connection_refused", "transport_error"), ("missing_env", "missing_transport"),
])
def test_real_route_completion_discloses_background_relay_failures(tmp_path, status, reason, caplog):
# Sending after the HTTP body proves that the final X carries the diagnostic;
# an earlier result snapshot would miss every failure in this fixture.
plugin = '''from starlette.background import BackgroundTask
from starlette.responses import Response
def register(api):
def stream(request):
def after():
for index in range(5):
assert api.send_ws_message('progress', {'index': index}) is None
return Response(b'already delivered', background=BackgroundTask(after))
api.register_route('stream', stream)
'''
skill, skills_repo, drive = _prepare_extension(tmp_path, "ws_stream_diagnostic", plugin,
permissions=["route", "ws_handler"])
repo = Path(__file__).resolve().parents[1]
env = child_env(tmp_path, drive, repo)
assert extension_loader.load_extension(skill, lambda: {}, drive_root=drive,
repo_path=str(skills_repo)) is None
spec = extension_loader.list_routes()[f"/api/extensions/{skill.name}/stream"]
with relay_endpoint(status) as (bridge, requests):
env.update(bridge)
def child_factory():
return runner._child_process(
{"mode": "route", "skill_name": skill.name, "surface": spec["path"],
"request": {"method": "GET", "path": spec["path"]},
"drive_root": str(drive), "repo_dir": str(repo), "skills_repo_path": str(skills_repo)},
skill_dir=skill.skill_dir, drive_root=drive, repo_dir=repo, env=env, stream=True)
response = RouteStreamResponse(spec, child_factory)
events = []
async def run():
delivered = asyncio.Event()
async def receive():
await delivered.wait()
return {"type": "http.disconnect"}
async def send(message):
events.append(message)
if message["type"] == "http.response.body" and not message.get("more_body"):
delivered.set()
await response({"type": "http", "method": "GET"}, receive, send)
asyncio.run(run())
assert events[0]["status"] == 200, events
assert b"".join(row.get("body", b"") for row in events) == b"already delivered"
assert response.body_complete is True
facts = consume_last_process_facts()
assert facts["exit_code"] == 0
if reason:
assert facts["ws_relay_failures"] == response.ws_relay_failures == {reason: 5}
assert len([row for row in caplog.records if "WS relay failures" in row.getMessage()]) == 1
else:
assert response.ws_relay_failures is None and "ws_relay_failures" not in facts
assert not [row for row in caplog.records if "WS relay failures" in row.getMessage()]
assert len(requests) == (5 if isinstance(status, int) else 0)
assert not list((drive / "state" / "skills" / skill.name / "extension_calls").glob("*.result.json"))
assert not runner._active_subprocesses
def test_legacy_empty_completion_frame_remains_compatible():
response = RouteStreamResponse({}, None)
frames = asyncio.Queue()
frames.put_nowait((b"S", json.dumps({"status": 200, "headers": []}).encode()))
frames.put_nowait((b"B", b"\x00complete"))
frames.put_nowait((b"X", b""))
events = []
async def send(message):
events.append(message)
asyncio.run(response._pump(frames, send))
assert events[-1]["body"] == b"complete"
assert response.ws_relay_failures is None

View file

@ -277,6 +277,13 @@ def test_extension_dispatch_surfaces_disclose_once(kind, tmp_path, monkeypatch):
return {"result": "ok"}
monkeypatch.setattr(extension_runner, "_run_child", fake_run)
if kind == "route":
from contextlib import contextmanager
@contextmanager
def fake_child(payload, **kwargs):
fake_run(payload, **kwargs)
yield None
monkeypatch.setattr(extension_runner, "_child_process", fake_child)
expected_task = "extension:alpha"
expected_root = "extension:alpha"
expected_parent = ""
@ -304,12 +311,15 @@ def test_extension_dispatch_surfaces_disclose_once(kind, tmp_path, monkeypatch):
expected_task, expected_root, expected_parent = "child-task", "root-task", "parent-task"
expected_source = "extension_tool:alpha:echo"
elif kind == "route":
extension_runner.dispatch_extension_route_subprocess(
response = extension_runner.dispatch_extension_route_subprocess(
{"skill": "alpha", "path": "/hello", "skills_repo_path": str(tmp_path)},
{},
drive_root=drive_root,
repo_dir=repo_dir,
)
assert _external_rows(drive_root) == [], "preparing a response is not a physical dispatch"
with response.child_factory():
pass
ledger_root = drive_root
expected_source = "extension_route:alpha:/hello"
else:
@ -623,6 +633,7 @@ def test_extension_child_spawn_failure_records_nothing(tmp_path, monkeypatch):
def test_extension_child_timeout_keeps_one_post_spawn_disclosure(tmp_path, monkeypatch):
class HangingProcess:
def __init__(self):
self.stdin = None
self.stdout = io.BytesIO()
self.stderr = io.BytesIO()
self.returncode = None

View file

@ -0,0 +1,155 @@
"""Named ingress keeps terminal reply identity and the normal cancellation owner."""
from types import SimpleNamespace
import pytest
from ouroboros import agent_task_pipeline as pipeline
from ouroboros.task_results import load_task_result, write_task_result
from ouroboros.utils import iter_jsonl_objects
from supervisor import message_bus, state
from supervisor.events import _handle_send_message
from tests.test_cancel_cascade_v664 import _fake_worker, _install_worker, _isolate_queue
from tests.test_host_service_operations import CHAT, MSG, _client, _headers, _inbound, _origin_ref, _receipt
@pytest.mark.parametrize("ephemeral", [True, False])
@pytest.mark.parametrize("host_operation", [True, False])
def test_actual_final_producer_preserves_named_reply_and_rejoin(tmp_path, monkeypatch, ephemeral, host_operation):
bridge = message_bus.LocalChatBridge()
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
monkeypatch.setattr(message_bus, "_BRIDGE", bridge)
monkeypatch.setattr(message_bus, "load_state", lambda: {})
monkeypatch.setattr(state, "reconstruct_task_cost", lambda *a, **k: {})
monkeypatch.setattr(pipeline, "_run_post_task_processing_async", lambda *a, **k: None)
client = _client(tmp_path, bridge)
body = {"chat_id": CHAT, "client_message_id": MSG, "text": "answer once"}
assert client.post("/chat/inject", headers=_headers(), json=body).status_code == 202
ref = bridge.get_updates(0, timeout=0)[0]["message"]["accepted_source_ref"]
task_id = f"inline-{ephemeral}-{host_operation}"
task = {"id": task_id, "type": "task", "chat_id": CHAT, "text": "answer once",
"_is_direct_chat": True, "_ephemeral_turn": ephemeral,
"origin_message_ref": ref, "metadata": {"_host_operation": host_operation}}
from ouroboros.agent import OuroborosAgent
env = SimpleNamespace(drive_root=tmp_path, repo_dir=tmp_path)
agent = OuroborosAgent.__new__(OuroborosAgent)
agent.env = env
agent._persist_running_record(task)
events = []
pipeline.emit_task_results(
env, None, None, events, task,
"The exact requested answer.", {"execution_status": "ok", "reason_code": "final_message"},
{"tool_calls": [], "reasoning_notes": []}, 0.0, tmp_path / "logs",
)
final = next(row for row in events if row["type"] == "send_message")
assert final["progress_meta"].get("origin_message_ref") == (ref if host_operation else None)
if host_operation:
assert final["progress_meta"]["origin_message_ref"] is not ref
if ephemeral:
assert not (tmp_path / "task_results" / f"{task_id}.json").exists()
assert final["progress_meta"]["task_terminal_status"] == "completed"
else:
assert final["progress_meta"]["task_phase"] == "finalizing"
_handle_send_message(final, SimpleNamespace(
DRIVE_ROOT=tmp_path, send_with_budget=message_bus.send_with_budget, append_jsonl=lambda *a, **k: None,
))
result = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
if host_operation or not ephemeral: # Managed records already carry their exact origin.
assert result["status"] == "completed" and result["text"] == "The exact requested answer."
rejoined = client.post("/chat/inject", headers=_headers(), json={**body, "wait_for_response": True})
assert rejoined.status_code == 200 and rejoined.json()["rejoined"]
assert rejoined.json()["response"] == result["text"]
else:
assert result["status"] == "pending" and "text" not in result
assert bridge._inbox.empty()
@pytest.mark.parametrize("case,chat,user,owner,expected", [
("missing_identity", -42, 0, {}, "failed"),
("registration", 42, 42, {}, "completed"),
("status", 42, 42, {"owner_external_id": 42, "owner_external_chat_id": 42}, "completed"),
("wrong_owner", 43, 43, {"owner_external_id": 42, "owner_external_chat_id": 42}, "failed"),
])
def test_server_inline_replies_settle_their_exact_ingress(tmp_path, monkeypatch, case, chat, user, owner, expected):
import server
bridge = message_bus.LocalChatBridge()
live_state = {"owner_id": 1, **owner}
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
monkeypatch.setattr(message_bus, "_BRIDGE", bridge)
monkeypatch.setattr(message_bus, "load_state", lambda: live_state)
monkeypatch.setattr(state, "status_text", lambda *a: "actual status reply")
ctx = SimpleNamespace(load_state=lambda: dict(live_state), update_state=lambda fn: fn(live_state),
send_with_budget=message_bus.send_with_budget, WORKERS={}, PENDING=[], RUNNING={})
client = _client(tmp_path, bridge)
body = {"chat_id": chat, "user_id": user, "client_message_id": case, "text": "/status"}
assert client.post("/chat/inject", headers=_headers(), json=body).status_code == 202
server._process_bridge_updates(bridge, 0, ctx)
rows = list(iter_jsonl_objects(tmp_path / "logs/chat.jsonl"))
assert len(rows) == 2 and rows[-1]["task_terminal_status"] == expected
assert rows[-1]["origin_message_ref"]["client_message_id"] == case
result = client.get(f"/chat/operations/{chat}:{case}", headers=_headers()).json()
assert result["status"] == expected and result["text"] == rows[-1]["text"]
rejoined = client.post("/chat/inject", headers=_headers(), json={**body, "wait_for_response": chat < 0})
assert rejoined.status_code == 200 and rejoined.json()["rejoined"]
assert bridge._inbox.empty() and not list((tmp_path / "task_results").glob("*.json"))
@pytest.mark.parametrize("kind", ["root_worker", "child_worker", "settled"])
def test_completed_result_uses_existing_live_cancellation_owner(tmp_path, monkeypatch, kind):
from supervisor import workers
task_queue, _ = _isolate_queue(monkeypatch, tmp_path, [])
monkeypatch.setattr(task_queue, "_emit_cancel_task_done", lambda *a, **k: None)
client = _client(tmp_path)
_inbound(tmp_path, "work")
ref = _origin_ref(tmp_path)
task = {"id": "root", "root_task_id": "root", "chat_id": CHAT, "origin_message_ref": ref}
_receipt(tmp_path, "promote_chat_to_task", "scheduled", "root")
write_task_result(tmp_path, "root", "completed", result="Final answer", origin_message_ref=ref)
if kind != "settled":
running = dict(task) if kind == "root_worker" else {
"id": "child", "root_task_id": "root", "parent_task_id": "root", "chat_id": CHAT,
}
task_queue.RUNNING[running["id"]] = {"task": running}
worker, physical = _fake_worker(running["id"])
_install_worker(monkeypatch, worker)
monkeypatch.setattr("ouroboros.platform_layer.kill_pid_tree", lambda *a, **k: physical.__setitem__("alive", False))
monkeypatch.setattr(workers, "get_event_q", __import__("queue").Queue)
view = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
assert view["status"] == "completed" and view["cancel_supported"] is (kind != "settled")
response = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": f"{CHAT}:{MSG}"})
assert response.status_code == 200 and response.json()["outcome"] == "already_terminal"
assert load_task_result(tmp_path, "root")["status"] == "completed"
assert load_task_result(tmp_path, "root")["result"] == "Final answer"
if kind != "settled":
assert not physical["alive"] and (tmp_path / "state/cancel_intents.json").exists()
assert not task_queue.RUNNING
else:
assert not (tmp_path / "state/cancel_intents.json").exists()
def test_restart_acknowledgement_does_not_claim_the_operation_finished(tmp_path, monkeypatch):
import server
bridge = message_bus.LocalChatBridge()
live_state = {"owner_id": 1, "owner_external_id": 42, "owner_external_chat_id": 42}
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
monkeypatch.setattr(message_bus, "_BRIDGE", bridge)
monkeypatch.setattr(message_bus, "load_state", lambda: live_state)
monkeypatch.setattr(server, "_safe_restart_serialized", lambda *a, **k: (False, "controlled refusal"))
ctx = SimpleNamespace(load_state=lambda: dict(live_state), update_state=lambda fn: fn(live_state),
send_with_budget=message_bus.send_with_budget, safe_restart=object())
client = _client(tmp_path, bridge)
body = {"chat_id": 42, "user_id": 42, "client_message_id": "restart-refused", "text": "/restart"}
assert client.post("/chat/inject", headers=_headers(), json=body).status_code == 202
server._process_bridge_updates(bridge, 0, ctx)
inbound, acknowledgement, refused = list(iter_jsonl_objects(tmp_path / "logs/chat.jsonl"))
assert inbound["direction"] == "in"
assert acknowledgement["text"] == "♻️ Restarting."
assert acknowledgement["origin_message_ref"] == refused["origin_message_ref"]
assert "task_terminal_status" not in acknowledgement
assert refused["task_terminal_status"] == "failed" and "controlled refusal" in refused["text"]
assert not acknowledgement.get("is_progress") and not refused.get("is_progress")
view = client.get("/chat/operations/42:restart-refused", headers=_headers()).json()
assert view["status"] == "failed" and view["text"] == refused["text"]

View file

@ -0,0 +1,340 @@
"""Real ingress and task-authority controls for repeated skill deliveries."""
from concurrent.futures import ThreadPoolExecutor
import pytest
from ouroboros.project_dialogue import build_owner_message_ref
from ouroboros.task_results import write_task_result
from ouroboros.utils import iter_jsonl_objects
from supervisor import message_bus
from tests.test_host_service_operations import (
CHAT, MSG, _client, _headers, _chat_row, _inbound, _isolate_queue, _origin_ref, _receipt,
)
def test_atomic_acceptance_before_dequeue_and_exact_source_consumption(tmp_path, monkeypatch):
bridge = message_bus.LocalChatBridge()
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
client = _client(tmp_path, bridge)
body = {"chat_id": CHAT, "client_message_id": MSG, "text": "one message",
"accepted_source_ref": {"chat_id": 999}, "task_metadata": {"origin_message_ref": {"chat_id": 999}}}
with ThreadPoolExecutor(max_workers=4) as pool:
replies = list(pool.map(lambda _: client.post("/chat/inject", headers=_headers(), json=body), range(4)))
assert all(reply.status_code == 202 for reply in replies)
assert bridge._inbox.qsize() == 1
assert sum(reply.json().get("rejoined", False) for reply in replies) == 3
rows = list(iter_jsonl_objects(tmp_path / "logs/chat.jsonl"))
assert len(rows) == 1
update = bridge.get_updates(0, timeout=0)[0]["message"]
ref = message_bus.record_inbound_message(
bridge, update, chat_id=CHAT, user_id=0, client_message_id=MSG, text="one message", ts="later",
)
assert ref == build_owner_message_ref(chat_id=CHAT, client_message_id=MSG, ts=rows[0]["ts"], text="one message")
assert len(list(iter_jsonl_objects(tmp_path / "logs/chat.jsonl"))) == 1
assert not update.get("task_metadata", {}).get("origin_message_ref")
replay = client.post("/chat/inject", headers=_headers(), json=body)
assert replay.json()["rejoined"] is True and bridge._inbox.empty()
distinct = client.post("/chat/inject", headers=_headers(), json={**body, "client_message_id": "message-two"})
assert distinct.status_code == 202 and bridge._inbox.qsize() == 1
def test_failed_acceptance_never_queues_and_restart_never_requeues(tmp_path, monkeypatch):
from ouroboros.utils import atomic_write_json
bridge = message_bus.LocalChatBridge()
client = _client(tmp_path, bridge)
atomic_write_json(tmp_path / "state/state.json", {"session_id": "before"})
body = {"chat_id": CHAT, "client_message_id": MSG, "text": "one message"}
with monkeypatch.context() as scoped:
scoped.setattr(message_bus, "append_jsonl", lambda *a, **kw: False)
assert client.post("/chat/inject", headers=_headers(), json=body).status_code == 500
assert bridge._inbox.empty()
assert client.post("/chat/inject", headers=_headers(), json=body).status_code == 202
atomic_write_json(tmp_path / "state/state.json", {"session_id": "after"})
restarted = message_bus.LocalChatBridge()
restarted_client = _client(tmp_path, restarted)
assert restarted_client.post("/chat/inject", headers=_headers(), json=body).json()["rejoined"]
state = restarted_client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
assert state["status"] == "lost" and restarted._inbox.empty()
@pytest.mark.parametrize("mismatch", ["chat_id", "ts", "text_sha256", "source"])
def test_annotation_collision_never_authorizes_foreign_task(tmp_path, monkeypatch, mismatch):
client = _client(tmp_path)
_inbound(tmp_path, "own source")
ref = _origin_ref(tmp_path)
if mismatch == "source":
_chat_row(tmp_path, "in", "foreign source", chat_id=1, client_message_id=MSG, source="web")
ref = build_owner_message_ref(chat_id=1, client_message_id=MSG, ts="foreign", text="foreign source")
else:
ref[mismatch] = 1 if mismatch == "chat_id" else "0" * 64 if mismatch == "text_sha256" else "foreign"
_, pending = _isolate_queue(monkeypatch, tmp_path, [{"id": "foreign", "chat_id": 1, "origin_message_ref": ref}])
write_task_result(tmp_path, "foreign", "scheduled", origin_message_ref=ref, result="foreign answer")
_receipt(tmp_path, "promote_chat_to_task", "scheduled", "foreign")
state = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
assert state["status"] == "pending" and "task_id" not in state and "text" not in state
result = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": f"{CHAT}:{MSG}"})
assert result.status_code == 409 and len(pending) == 1
assert not (tmp_path / "state/cancel_intents.json").exists()
def test_interleaved_same_chat_answers_are_matched_by_task_origin(tmp_path):
client = _client(tmp_path)
_inbound(tmp_path, "request A")
ref_a = _origin_ref(tmp_path)
_chat_row(tmp_path, "in", "request B", client_message_id="B")
write_task_result(tmp_path, "task-a", "completed", origin_message_ref=ref_a, result="answer A")
_chat_row(tmp_path, "out", "answer A", task_id="task-a")
a = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
b = client.get(f"/chat/operations/{CHAT}:B", headers=_headers()).json()
assert a["status"] == "completed" and a["text"] == "answer A"
assert b["status"] == "pending" and "text" not in b
@pytest.mark.parametrize("lane", ["_handle_chat_direct_locked", "handle_chat_ephemeral"])
def test_pre_task_budget_refusal_remains_an_exact_failed_operation(tmp_path, monkeypatch, lane):
from supervisor import state, worker_chat_lane, workers
bridge = message_bus.LocalChatBridge()
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
monkeypatch.setattr(message_bus, "_BRIDGE", bridge)
monkeypatch.setattr(message_bus, "load_state", lambda: {})
monkeypatch.setattr(state, "budget_remaining", lambda *_a, **_kw: 0)
monkeypatch.setattr(workers, "send_with_budget", message_bus.send_with_budget)
client = _client(tmp_path, bridge)
assert client.post("/chat/inject", headers=_headers(), json={
"chat_id": CHAT, "client_message_id": MSG, "text": "do work",
}).status_code == 202
ref = bridge.get_updates(0, timeout=0)[0]["message"]["accepted_source_ref"]
getattr(worker_chat_lane, lane)(CHAT, "do work", task_metadata={"origin_message_ref": ref, "_host_operation": True})
result = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
assert result["status"] == "failed" and "Budget exhausted" in result["text"]
@pytest.mark.parametrize("chat_id", [1, 123456])
def test_ordinary_main_and_transport_refusals_keep_their_original_envelope(monkeypatch, chat_id):
from supervisor import state, worker_chat_lane, workers
sent = []
monkeypatch.setattr(state, "budget_remaining", lambda *_a, **_kw: 0)
monkeypatch.setattr(workers, "send_with_budget", lambda *a, **kw: sent.append((a, kw)))
worker_chat_lane._handle_chat_direct_locked(chat_id, "hello", task_metadata={
"origin_message_ref": build_owner_message_ref(chat_id=chat_id, client_message_id="ordinary", ts="now", text="hello"),
})
assert sent == [((chat_id, "🚫 Budget exhausted. Task rejected. Please increase TOTAL_BUDGET in settings."), {})]
@pytest.mark.parametrize("ephemeral", [False, True])
@pytest.mark.parametrize("host_operation", [False, True])
def test_chat_crash_preserves_only_the_host_accepted_operation(tmp_path, monkeypatch, ephemeral, host_operation):
from queue import SimpleQueue
from ouroboros import project_naming
from supervisor import worker_chat_lane, workers
class CrashingAgent:
def handle_task(self, task):
raise RuntimeError("test chat execution failed")
bridge = message_bus.LocalChatBridge()
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
monkeypatch.setattr(message_bus, "_BRIDGE", bridge)
monkeypatch.setattr(message_bus, "load_state", lambda: {})
monkeypatch.setattr(workers, "DRIVE_ROOT", tmp_path)
monkeypatch.setattr(workers, "get_event_q", SimpleQueue)
monkeypatch.setattr(workers, "send_with_budget", message_bus.send_with_budget)
monkeypatch.setattr(project_naming, "spawn_proactive_namer", lambda *a, **kw: None)
client = _client(tmp_path, bridge)
assert client.post("/chat/inject", headers=_headers(), json={
"chat_id": CHAT, "client_message_id": MSG, "text": "do work",
}).status_code == 202
ref = bridge.get_updates(0, timeout=0)[0]["message"]["accepted_source_ref"]
worker_chat_lane._run_chat_task(
CrashingAgent(), CHAT, "do work", ephemeral=ephemeral,
task_metadata={"origin_message_ref": ref, "_host_operation": host_operation},
)
rows = list(iter_jsonl_objects(tmp_path / "logs/chat.jsonl"))
final = rows[-1]
assert "test chat execution failed" in final["text"]
assert final["task_terminal_status"] == "failed"
assert final.get("origin_message_ref") == (ref if host_operation else None)
result = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
if host_operation:
assert result["status"] == "failed" and "test chat execution failed" in result["text"]
else:
assert result["status"] == "pending" and "text" not in result
def test_real_server_consumes_one_source_and_authors_only_its_operation_marker(tmp_path, monkeypatch):
from types import SimpleNamespace
import server
bridge = message_bus.LocalChatBridge()
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
monkeypatch.setattr(message_bus, "load_state", lambda: {})
routed = []
monkeypatch.setattr(server, "_route_owner_message", lambda bridge, ctx, row: routed.append(row))
ctx = SimpleNamespace(load_state=lambda: {"owner_id": 1}, update_state=lambda mutator: mutator({"owner_id": 1}))
client = _client(tmp_path, bridge)
client.post("/chat/inject", headers=_headers(), json={"chat_id": CHAT, "client_message_id": MSG, "text": "hello"})
ref = _origin_ref(tmp_path)
server._process_bridge_updates(bridge, 0, ctx)
assert routed[-1]["origin_message_ref"] == ref
assert routed[-1]["task_metadata"]["_host_operation"] is True
assert len(list(iter_jsonl_objects(tmp_path / "logs/chat.jsonl"))) == 1
bridge.enqueue_local_message("ordinary", task_metadata={"_host_operation": True})
server._process_bridge_updates(bridge, 1, ctx)
assert "_host_operation" not in routed[-1]["task_metadata"]
def test_rotation_during_source_read_cannot_turn_replay_into_new_work(tmp_path, monkeypatch):
import sys
from ouroboros import utils
from supervisor.state import rotate_chat_log_if_needed
bridge = message_bus.LocalChatBridge()
client = _client(tmp_path, bridge)
body = {"chat_id": CHAT, "client_message_id": MSG, "text": "hello"}
accepted = client.post("/chat/inject", headers=_headers(), json=body)
assert accepted.status_code == 202
source = _origin_ref(tmp_path)
original = utils.jsonl_archive_segments
outcomes = []
def rotate_after_live_open(path, **kwargs):
if not outcomes:
try:
rotate_chat_log_if_needed(tmp_path, max_bytes=1)
except PermissionError:
# Windows can defer the independent rotator while the reader
# owns the live handle. Its failure is not a reader failure.
assert sys.platform == "win32"
assert (tmp_path / "logs/chat.jsonl").stat().st_size > 0
outcomes.append("open_handle")
else:
assert original(path)
outcomes.append("rotated")
return original(path, **kwargs)
monkeypatch.setattr(utils, "jsonl_archive_segments", rotate_after_live_open)
replay = client.post("/chat/inject", headers=_headers(), json=body)
assert replay.status_code == 202 and replay.json()["rejoined"]
assert bridge._inbox.qsize() == 1 and _origin_ref(tmp_path) == source
assert outcomes in (["rotated"], ["open_handle"])
if sys.platform != "win32":
assert outcomes == ["rotated"]
# With all read handles closed, both platforms must really rotate and
# recover the same source from its archive without another enqueue.
rotate_chat_log_if_needed(tmp_path, max_bytes=1)
assert original(tmp_path / "logs/chat.jsonl")
archived_replay = client.post("/chat/inject", headers=_headers(), json=body)
assert archived_replay.status_code == 202 and archived_replay.json()["rejoined"]
assert bridge._inbox.qsize() == 1 and _origin_ref(tmp_path) == source
def test_racing_named_upload_rejoin_keeps_only_the_accepted_copy(tmp_path, monkeypatch):
import pathlib
import threading
from ouroboros.gateway import host_service
bridge = message_bus.LocalChatBridge()
client = _client(tmp_path, bridge)
source = tmp_path / "state/skills/a2a/input.pdf"
source.write_bytes(b"complete attachment")
barrier = threading.Barrier(2)
original = host_service.store_chat_upload
def copy(*args, **kwargs):
result = original(*args, **kwargs)
barrier.wait(timeout=5)
return result
monkeypatch.setattr(host_service, "store_chat_upload", copy)
body = {"chat_id": CHAT, "client_message_id": MSG, "text": "one message",
"attachments": [{"path": str(source)}]}
with ThreadPoolExecutor(max_workers=2) as pool:
responses = list(pool.map(lambda _: client.post("/chat/inject", headers=_headers(), json=body), range(2)))
assert all(response.status_code == 202 for response in responses)
assert sum(response.json().get("rejoined", False) for response in responses) == 1
updates = bridge.get_updates(0, timeout=0)
assert len(updates) == 1 and bridge._inbox.empty()
stored = pathlib.Path(updates[0]["message"]["task_metadata"]["chat_attachment_uploads"][0]["path"])
assert list((tmp_path / "uploads").iterdir()) == [stored]
assert stored.read_bytes() == source.read_bytes()
@pytest.mark.parametrize("failure", ["write_unknown_empty", "write_unknown_landed", "queue_unknown"])
def test_named_upload_custody_survives_unknown_write_or_queue_outcome(tmp_path, monkeypatch, failure):
bridge = message_bus.LocalChatBridge()
client = _client(tmp_path, bridge)
source = tmp_path / "state/skills/a2a/input.pdf"
source.write_bytes(b"complete attachment")
original = message_bus.log_chat
def fail(*args, **kwargs):
if failure == "write_unknown_landed":
original(*args, **kwargs)
raise OSError("controlled admission failure")
if failure.startswith("write_unknown"):
monkeypatch.setattr(message_bus, "log_chat", fail)
else:
monkeypatch.setattr(bridge, "enqueue_local_message", fail)
response = client.post("/chat/inject", headers=_headers(), json={
"chat_id": CHAT, "client_message_id": MSG, "text": "one message",
"attachments": [{"path": str(source)}],
})
assert response.status_code == 500 and bridge._inbox.empty()
copies = list((tmp_path / "uploads").iterdir())
assert len(copies) == 1
assert copies[0].read_bytes() == source.read_bytes()
row = message_bus.accepted_chat_message(tmp_path, CHAT, MSG)
assert bool(row) is (failure != "write_unknown_empty")
assert source.read_bytes() == b"complete attachment"
@pytest.mark.parametrize("cancel_mode", ["asyncio", "anyio"])
def test_cancelled_named_acceptance_settles_before_upload_cleanup(tmp_path, monkeypatch, cancel_mode):
import asyncio
import threading
import anyio
from types import SimpleNamespace
from ouroboros.gateway import host_service
bridge = message_bus.LocalChatBridge()
client = _client(tmp_path, bridge)
source = tmp_path / "state/skills/a2a/input.pdf"
source.write_bytes(b"complete attachment")
entered, release = threading.Event(), threading.Event()
original = message_bus.log_chat
def held(*args, **kwargs):
row = original(*args, **kwargs)
entered.set()
assert release.wait(5)
return row
monkeypatch.setattr(message_bus, "log_chat", held)
async def body():
return {"chat_id": CHAT, "client_message_id": MSG, "text": "one message",
"attachments": [{"path": str(source)}]}
request = SimpleNamespace(app=client.app, headers={k.lower(): v for k, v in _headers().items()}, json=body)
ctx = client.app.state.host_service_context
async def run():
scope = anyio.CancelScope()
async def call():
with scope:
return await host_service._api_chat_inject(request)
task = asyncio.create_task(call())
try:
assert await asyncio.to_thread(entered.wait, 5)
scope.cancel() if cancel_mode == "anyio" else task.cancel()
await asyncio.sleep(0)
assert not task.done() and ctx._inflight["a2a"] == 1
finally:
release.set()
if cancel_mode == "anyio":
await task
assert scope.cancelled_caught
else:
with pytest.raises(asyncio.CancelledError):
await task
assert ctx._inflight["a2a"] == 0
asyncio.run(run())
message = bridge.get_updates(0, timeout=0)[0]["message"]
from pathlib import Path
stored = Path(message["task_metadata"]["chat_attachment_uploads"][0]["path"])
assert stored.read_bytes() == source.read_bytes()

View file

@ -0,0 +1,492 @@
"""Host Service A2A operation correlation (#667) and the WS relay burst reserve.
An injected message that carries a ``client_message_id`` becomes an addressable
operation: ``/chat/inject`` answers with its ``operation_ref``, a repeated SAME
message rejoins instead of enqueueing again, ``/chat/operations/{ref}`` joins the
existing records (inbound chat row, routing receipt, live turn registry, task
result, outbound answer row) into one typed view scoped to the injecting skill,
and ``/chat/cancel`` runs the existing cancellation owner and reports only what
it proved. The WS relay lane is a 60-message token bucket refilling one per
second; refusals are aggregated per burst.
"""
from __future__ import annotations
import json
import pathlib
import pytest
from starlette.testclient import TestClient
from ouroboros.gateway.host_service import (
WS_RELAY_BURST,
create_host_service_app,
operation_ref,
)
from ouroboros.utils import append_jsonl, utc_now_iso
from tests.test_host_service_api import FakeBridge, _seed_token
SKILL = "a2a"
TOKEN = "a2a-token"
CHAT = -42
MSG = "a2a:task-1:msg-1"
def _client(tmp_path: pathlib.Path, bridge=None, **kwargs) -> TestClient:
_seed_token(tmp_path, skill=SKILL, token=TOKEN, permissions=["inject_chat"])
bridge = bridge or FakeBridge()
return TestClient(create_host_service_app(tmp_path, bridge_getter=lambda: bridge, **kwargs))
def _chat_row(tmp_path: pathlib.Path, direction: str, text: str, *, chat_id: int = CHAT,
client_message_id: str = "", source: str = f"skill:{SKILL}", task_id: str = "",
session_id: str = "s1", record_type: str = "") -> None:
"""The canonical row ``supervisor.message_bus.log_chat`` writes (same keys)."""
row = {
"ts": utc_now_iso(), "session_id": session_id, "direction": direction, "chat_id": chat_id,
"user_id": 0, "text": text, "format": "", "source": source, "sender_label": "",
"sender_session_id": "", "client_message_id": client_message_id, "transport": {},
"task_id": task_id,
}
if record_type:
row["type"] = record_type
append_jsonl(tmp_path / "logs" / "chat.jsonl", row)
def _inbound(tmp_path: pathlib.Path, text: str = "hello", **kwargs) -> None:
_chat_row(tmp_path, "in", text, client_message_id=MSG, **kwargs)
def _headers() -> dict:
return {"X-Skill-Token": TOKEN}
def _origin_ref(tmp_path):
from ouroboros.project_dialogue import build_owner_message_ref
from supervisor.message_bus import accepted_chat_message
row = accepted_chat_message(tmp_path, CHAT, MSG)
return build_owner_message_ref(chat_id=CHAT, client_message_id=MSG, ts=row["ts"], text=row["text"])
def _answer(tmp_path, text, task_id="40cc86d9"):
from ouroboros.task_results import write_task_result
write_task_result(tmp_path, task_id, "completed", result=text, origin_message_ref=_origin_ref(tmp_path))
_chat_row(tmp_path, "out", text, task_id=task_id)
def _receipt(tmp_path: pathlib.Path, action: str, status: str, target: str) -> None:
from ouroboros.project_dialogue import append_chat_annotation
assert append_chat_annotation(tmp_path, MSG, action=action, target=target, status=status)
# --- inject: correlation and rejoin -----------------------------------------
def test_inject_returns_operation_ref_only_when_the_caller_names_the_message(tmp_path):
bridge = FakeBridge()
client = _client(tmp_path, bridge)
plain = client.post("/chat/inject", headers=_headers(), json={"text": "hi", "chat_id": CHAT})
assert plain.status_code == 202 and plain.json() == {"ok": True, "status": "queued"}
named = client.post("/chat/inject", headers=_headers(),
json={"text": "hi", "chat_id": CHAT, "client_message_id": MSG})
assert named.status_code == 202
assert named.json() == {"ok": True, "status": "queued", "operation_ref": operation_ref(CHAT, MSG)}
assert bridge.messages[1]["client_message_id"] == MSG
_answer(tmp_path, "reply from host")
waited = client.post("/chat/inject", headers=_headers(), json={
"text": "hi", "chat_id": CHAT, "client_message_id": MSG, "wait_for_response": True, "timeout_sec": 5,
})
assert waited.status_code == 200
assert waited.json()["operation_ref"] == operation_ref(CHAT, MSG)
def test_inject_wait_expiry_carries_the_operation_ref(tmp_path):
class SilentBridge(FakeBridge):
def enqueue_local_message(self, text, **kwargs):
self.messages.append({"text": text, **kwargs})
client = _client(tmp_path, SilentBridge())
response = client.post("/chat/inject", headers=_headers(), json={
"text": "slow", "chat_id": CHAT, "client_message_id": MSG, "wait_for_response": True, "timeout_sec": 1,
})
assert response.status_code == 504
assert response.json() == {
"ok": False, "error": "timed out waiting for response", "operation_ref": operation_ref(CHAT, MSG),
}
def test_same_message_rejoins_and_a_different_message_is_refused(tmp_path):
bridge = FakeBridge()
client = _client(tmp_path, bridge)
_inbound(tmp_path, "hello")
rejoin = client.post("/chat/inject", headers=_headers(),
json={"text": " hello ", "chat_id": CHAT, "client_message_id": MSG})
assert rejoin.status_code == 202
assert rejoin.json() == {
"ok": True, "status": "accepted", "rejoined": True, "operation_ref": operation_ref(CHAT, MSG),
}
assert bridge.messages == [], "a rejoin must not enqueue a second message"
reused = client.post("/chat/inject", headers=_headers(),
json={"text": "something else", "chat_id": CHAT, "client_message_id": MSG})
assert reused.status_code == 409
assert "different message" in reused.json()["error"]
assert bridge.messages == []
def test_rejoin_of_an_answered_message_returns_the_late_answer(tmp_path):
bridge = FakeBridge()
client = _client(tmp_path, bridge)
_inbound(tmp_path, "hello")
_answer(tmp_path, "late but real answer")
response = client.post("/chat/inject", headers=_headers(), json={
"text": "hello", "chat_id": CHAT, "client_message_id": MSG, "wait_for_response": True, "timeout_sec": 5,
})
assert response.status_code == 200
body = response.json()
assert body["response"] == "late but real answer" and body["status"] == "completed" and body["rejoined"] is True
assert bridge.messages == []
def test_rejoin_refuses_a_message_id_bound_to_another_source(tmp_path):
client = _client(tmp_path, FakeBridge())
_inbound(tmp_path, "hello", source="skill:telegram")
response = client.post("/chat/inject", headers=_headers(),
json={"text": "hello", "chat_id": CHAT, "client_message_id": MSG})
assert response.status_code == 409
assert "another source" in response.json()["error"]
# --- operations read ----------------------------------------------------------
def test_operation_read_is_scoped_to_the_injecting_skill(tmp_path):
client = _client(tmp_path)
ref = operation_ref(CHAT, MSG)
assert client.get(f"/chat/operations/{ref}", headers=_headers()).status_code == 404
_inbound(tmp_path, "hello", source="skill:telegram")
assert client.get(f"/chat/operations/{ref}", headers=_headers()).status_code == 404
assert client.get("/chat/operations/not-a-ref", headers=_headers()).status_code == 400
assert client.get(f"/chat/operations/{ref}").status_code == 403
def test_operation_read_reports_pending_then_the_durable_answer(tmp_path):
client = _client(tmp_path)
ref = operation_ref(CHAT, MSG)
_chat_row(tmp_path, "out", "an older answer in the same chat") # before the message: never its answer
_inbound(tmp_path, "hello")
pending = client.get(f"/chat/operations/{ref}", headers=_headers()).json()
assert pending["status"] == "pending" and pending["cancel_supported"] is False
assert pending["chat_id"] == CHAT and pending["client_message_id"] == MSG
_chat_row(tmp_path, "out", "", record_type="routing_options") # host-authored, not an answer
_answer(tmp_path, "the answer")
done = client.get(f"/chat/operations/{ref}", headers=_headers()).json()
assert done["status"] == "completed" and done["text"] == "the answer" and done["task_id"] == "40cc86d9"
def test_operation_read_reports_a_live_direct_or_ephemeral_turn(tmp_path, monkeypatch):
_isolate_queue(monkeypatch, tmp_path, [])
from supervisor.active_activity import get_direct_activity_registry
client = _client(tmp_path)
_inbound(tmp_path, "hello")
registry = get_direct_activity_registry()
registry.clear()
try:
registry.register("40cc86d9", CHAT, client_message_id=MSG, kind="ephemeral_decision", origin_message_ref=_origin_ref(tmp_path))
state = client.get(f"/chat/operations/{operation_ref(CHAT, MSG)}", headers=_headers()).json()
assert state["status"] == "running" and state["phase"] == "ephemeral_decision"
assert state["task_id"] == "40cc86d9" and state["cancel_supported"] is False
registry.clear()
registry.register("40cc86d9", CHAT, client_message_id=MSG, kind="direct_chat", origin_message_ref=_origin_ref(tmp_path))
state = client.get(f"/chat/operations/{operation_ref(CHAT, MSG)}", headers=_headers()).json()
assert state["phase"] == "direct_chat" and state["cancel_supported"] is True
finally:
registry.clear()
def test_operation_read_follows_the_routing_receipt_to_the_promoted_task(tmp_path, monkeypatch):
_isolate_queue(monkeypatch, tmp_path, [])
from ouroboros.task_results import STATUS_COMPLETED, STATUS_SCHEDULED, write_task_result
client = _client(tmp_path)
_inbound(tmp_path, "do the thing")
_receipt(tmp_path, "promote_chat_to_task", "scheduled", "task-abc")
write_task_result(tmp_path, "task-abc", STATUS_SCHEDULED, description="do the thing", origin_message_ref=_origin_ref(tmp_path))
state = client.get(f"/chat/operations/{operation_ref(CHAT, MSG)}", headers=_headers()).json()
assert state["task_id"] == "task-abc" and state["phase"] == "managed_task"
assert state["status"] == STATUS_SCHEDULED and state["cancel_supported"] is True
write_task_result(tmp_path, "task-abc", STATUS_COMPLETED, result="final answer")
state = client.get(f"/chat/operations/{operation_ref(CHAT, MSG)}", headers=_headers()).json()
assert state["status"] == STATUS_COMPLETED and state["text"] == "final answer"
assert state["cancel_supported"] is False
def test_operation_read_reports_lost_after_a_host_restart(tmp_path):
from ouroboros.utils import atomic_write_json
client = _client(tmp_path)
_inbound(tmp_path, "hello", session_id="old-session")
atomic_write_json(tmp_path / "state" / "state.json", {"session_id": "new-session"})
state = client.get(f"/chat/operations/{operation_ref(CHAT, MSG)}", headers=_headers()).json()
assert state["status"] == "lost" and state["reason"] == "host_restarted_before_answer"
# --- cancel -------------------------------------------------------------------
def _isolate_queue(monkeypatch, tmp_path, tasks):
from supervisor import queue, workers
pending = [dict(task) for task in tasks]
monkeypatch.setattr(queue, "DRIVE_ROOT", tmp_path)
monkeypatch.setattr(queue, "PENDING", pending)
monkeypatch.setattr(queue, "RUNNING", {})
monkeypatch.setattr(workers, "WORKERS", {}, raising=False)
monkeypatch.setattr(workers, "_chat_agent", None, raising=False)
monkeypatch.setattr(queue, "persist_queue_snapshot", lambda reason="": None)
return queue, pending
def test_cancel_before_any_addressable_work_is_disclosed_not_faked(tmp_path, monkeypatch):
from supervisor.active_activity import get_direct_activity_registry
_isolate_queue(monkeypatch, tmp_path, [])
client = _client(tmp_path)
ref = operation_ref(CHAT, MSG)
assert client.post("/chat/cancel", headers=_headers(), json={"operation_ref": ref}).status_code == 404
_inbound(tmp_path, "hello")
queued = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": ref})
assert queued.status_code == 409
assert queued.json()["outcome"] == "cancel_unsupported" and queued.json()["reason"] == "not_started"
registry = get_direct_activity_registry()
registry.clear()
try:
registry.register("40cc86d9", CHAT, client_message_id=MSG, kind="ephemeral_decision", origin_message_ref=_origin_ref(tmp_path))
deciding = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": ref})
finally:
registry.clear()
assert deciding.status_code == 409 and deciding.json()["reason"] == "decision_turn_in_flight"
assert client.post("/chat/cancel", headers=_headers(), json={"operation_ref": "junk"}).status_code == 400
assert client.post("/chat/cancel", headers=_headers(), json=["not", "an", "object"]).status_code == 400
def test_cancel_of_a_message_steered_into_a_foreign_task_is_unsupported(tmp_path, monkeypatch):
from ouroboros.task_results import STATUS_RUNNING, write_task_result
_isolate_queue(monkeypatch, tmp_path, [{"id": "foreign", "chat_id": 1, "root_task_id": "foreign"}])
client = _client(tmp_path)
_inbound(tmp_path, "please also do X")
_receipt(tmp_path, "steer_task", "delivered", "foreign")
write_task_result(tmp_path, "foreign", STATUS_RUNNING, description="owner work")
response = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": operation_ref(CHAT, MSG)})
assert response.status_code == 409
assert response.json()["outcome"] == "cancel_unsupported"
assert response.json()["reason"] == "not_started"
assert "task_id" not in response.json(), "a presentation hint must not expose unrelated task state"
def test_cancel_runs_the_real_cancellation_owner_on_the_promoted_task(tmp_path, monkeypatch):
"""A real queued task (PENDING, no worker) bound to the message by the
promotion receipt is cancelled through the durable intent + custody path;
the answer is read back from the durable result, and a second cancel says
``already_terminal``."""
from ouroboros.task_results import STATUS_CANCELLED, STATUS_SCHEDULED, load_task_result, write_task_result
queue, pending = _isolate_queue(monkeypatch, tmp_path, [
{"id": "task-abc", "chat_id": CHAT, "root_task_id": "task-abc"},
])
client = _client(tmp_path)
_inbound(tmp_path, "do the thing")
_receipt(tmp_path, "promote_chat_to_task", "scheduled", "task-abc")
write_task_result(tmp_path, "task-abc", STATUS_SCHEDULED, description="do the thing", origin_message_ref=_origin_ref(tmp_path))
pending[0]["origin_message_ref"] = _origin_ref(tmp_path)
response = client.post("/chat/cancel", headers=_headers(),
json={"operation_ref": operation_ref(CHAT, MSG), "reason": "peer gave up"})
assert response.status_code == 200, response.json()
assert response.json()["outcome"] == "cancelled" and response.json()["status"] == STATUS_CANCELLED
assert response.json()["task_id"] == "task-abc"
assert pending == []
assert load_task_result(tmp_path, "task-abc")["status"] == STATUS_CANCELLED
intents = json.loads((tmp_path / "state" / "cancel_intents.json").read_text(encoding="utf-8"))
assert "task-abc" not in (intents.get("intents") or {}), "the intent settled"
again = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": operation_ref(CHAT, MSG)})
assert again.status_code == 200 and again.json()["outcome"] == "already_terminal"
state = client.get(f"/chat/operations/{operation_ref(CHAT, MSG)}", headers=_headers()).json()
assert state["status"] == STATUS_CANCELLED
def test_cancel_reports_unresolved_when_custody_does_not_settle(tmp_path, monkeypatch):
from ouroboros.gateway import tasks as gateway_tasks
from ouroboros.task_results import STATUS_RUNNING, load_task_result, write_task_result
from supervisor import queue
_isolate_queue(monkeypatch, tmp_path, [])
monkeypatch.setattr(queue, "task_has_live_ownership", lambda task_id: True)
monkeypatch.setattr(queue, "task_subtree_is_live", lambda task_id, **_kw: True)
monkeypatch.setattr(gateway_tasks, "_run_cascade_cancel", lambda task_id: False)
client = _client(tmp_path)
_inbound(tmp_path, "do the thing")
_receipt(tmp_path, "promote_chat_to_task", "scheduled", "task-abc")
write_task_result(tmp_path, "task-abc", STATUS_RUNNING, description="do the thing", origin_message_ref=_origin_ref(tmp_path))
response = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": operation_ref(CHAT, MSG)})
assert response.status_code == 503
assert response.json()["outcome"] == "unresolved" and response.json()["ok"] is False
assert load_task_result(tmp_path, "task-abc")["status"] == STATUS_RUNNING
intents = json.loads((tmp_path / "state" / "cancel_intents.json").read_text(encoding="utf-8"))
assert "task-abc" in (intents.get("intents") or {}), "the durable intent stays open for the watchdog"
def test_cancel_refuses_when_the_durable_intent_cannot_be_recorded(tmp_path, monkeypatch):
from ouroboros import cancel_intents
from ouroboros.task_results import STATUS_RUNNING, write_task_result
from supervisor import queue
_isolate_queue(monkeypatch, tmp_path, [])
monkeypatch.setattr(queue, "task_has_live_ownership", lambda task_id: True)
def _boom(*_args, **_kwargs):
raise OSError("disk full")
monkeypatch.setattr(cancel_intents, "request_cancel", _boom)
client = _client(tmp_path)
_inbound(tmp_path, "do the thing")
_receipt(tmp_path, "promote_chat_to_task", "scheduled", "task-abc")
write_task_result(tmp_path, "task-abc", STATUS_RUNNING, description="do the thing", origin_message_ref=_origin_ref(tmp_path))
response = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": operation_ref(CHAT, MSG)})
assert response.status_code == 503
assert response.json()["outcome"] == "refused" and response.json()["reason"] == "cancel_intent_write_failed"
# --- WS relay burst reserve --------------------------------------------------
def test_ws_relay_burst_reserve_refuses_with_diagnostics_then_refills(tmp_path, monkeypatch):
from types import SimpleNamespace
from ouroboros.gateway import host_service
clock = [0.0]
# Only this Host module's limiter sees the controlled clock; ASGI keeps time.
monkeypatch.setattr(host_service, "time", SimpleNamespace(monotonic=lambda: clock[0]))
_seed_token(tmp_path, skill="wsskill", token="tok", manifest_permissions=["ws_handler"])
sent: list[dict] = []
app = create_host_service_app(tmp_path, ws_broadcaster_getter=lambda: sent.append)
client = TestClient(app)
payload = {"message_type": "progress", "data": {"pct": 1}}
for _ in range(WS_RELAY_BURST):
assert client.post("/ui/ws-message", headers={"X-Skill-Token": "tok"}, json=payload).status_code == 202
assert len(sent) == WS_RELAY_BURST
first = client.post("/ui/ws-message", headers={"X-Skill-Token": "tok"}, json=payload)
second = client.post("/ui/ws-message", headers={"X-Skill-Token": "tok"}, json=payload)
assert first.status_code == 429 and second.status_code == 429
assert first.json()["dropped_in_burst"] == 1 and second.json()["dropped_in_burst"] == 2
assert 0 < first.json()["retry_after_sec"] <= 1.0
assert first.headers["Retry-After"] == "1"
assert len(sent) == WS_RELAY_BURST, "a refused relay reaches no browser client"
# One second later exactly one token exists again; the burst summary is
# recorded ONCE, durably, with the aggregate dropped count.
clock[0] += 1.0
assert client.post("/ui/ws-message", headers={"X-Skill-Token": "tok"}, json=payload).status_code == 202
assert client.post("/ui/ws-message", headers={"X-Skill-Token": "tok"}, json=payload).status_code == 429
rows = [json.loads(line) for line in (tmp_path / "logs" / "events.jsonl").read_text(encoding="utf-8").splitlines()]
dropped = [row for row in rows if row.get("type") == "host_service_ws_relay_dropped"]
assert [(row["skill"], row["dropped"]) for row in dropped] == [("wsskill", 2)]
def test_in_process_send_ws_message_is_not_bounded_by_the_host_bucket(tmp_path):
"""The bucket replaces only the Host Service WS lane; the in-process
broadcast path (``PluginAPIImpl.send_ws_message``) keeps no bound."""
from ouroboros import extension_plugin_api as plugin_api
from ouroboros.extension_loader import extension_surface_name
received: list[dict] = []
previous = plugin_api._ws_broadcaster
plugin_api._ws_broadcaster = received.append
try:
api = plugin_api.PluginAPIImpl.__new__(plugin_api.PluginAPIImpl)
api._skill = "burst_skill"
api._permissions = {"ws_handler"}
api._runtime_closing = False
api._runtime_closed = False
api._api_lock = __import__("threading").RLock()
for index in range(WS_RELAY_BURST * 3):
api.send_ws_message("progress", {"n": index})
finally:
plugin_api._ws_broadcaster = previous
assert len(received) == WS_RELAY_BURST * 3
assert received[-1]["type"] == extension_surface_name("burst_skill", "progress")
@pytest.mark.parametrize("value", ["", ":", "abc", "12:"])
def test_operation_ref_parsing_rejects_malformed_refs(value):
from ouroboros.gateway.host_service import _parse_operation_ref
with pytest.raises(ValueError):
_parse_operation_ref(value)
@pytest.mark.parametrize("status", ["running", "scheduled", "completed", "cancelled"])
def test_unowned_cancel_reports_only_the_observed_terminal_state(tmp_path, monkeypatch, status):
from ouroboros import cancel_intents
from ouroboros.task_results import write_task_result, load_task_result
from ouroboros.task_status import SETTLED_STATUSES
_isolate_queue(monkeypatch, tmp_path, [])
client = _client(tmp_path)
_inbound(tmp_path, "accepted work")
_receipt(tmp_path, "promote_chat_to_task", "scheduled", "unowned")
write_task_result(tmp_path, "unowned", status, origin_message_ref=_origin_ref(tmp_path))
def no_new_cancel(*args, **kwargs):
pytest.fail("absent physical ownership must not mint a second cancellation attempt")
monkeypatch.setattr(cancel_intents, "request_cancel", no_new_cancel)
response = client.post("/chat/cancel", headers=_headers(),
json={"operation_ref": operation_ref(CHAT, MSG)})
body = response.json()
assert body["status"] == load_task_result(tmp_path, "unowned")["status"] == status
if status in SETTLED_STATUSES:
assert response.status_code == 200 and body["ok"] is True
assert body["outcome"] == "already_terminal"
else:
assert response.status_code == 503 and body["ok"] is False
assert body["outcome"] == "unresolved"
assert body["reason"] == "cancellation_did_not_settle"
def test_cancel_never_targets_a_different_queue_root(tmp_path, monkeypatch):
from ouroboros.task_results import write_task_result
from ouroboros.gateway import tasks
host, other = tmp_path / "host", tmp_path / "other"
client = _client(host)
_inbound(host, "own request")
_receipt(host, "promote_chat_to_task", "scheduled", "same-id")
write_task_result(host, "same-id", "scheduled", origin_message_ref=_origin_ref(host))
write_task_result(other, "same-id", "scheduled", description="unrelated")
_isolate_queue(monkeypatch, other, [{"id": "same-id", "chat_id": 1}])
monkeypatch.setattr(tasks, "_run_cascade_cancel", lambda *_a: pytest.fail("foreign cancellation"))
view = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
assert view["task_id"] == "same-id" and view["cancel_supported"] is False
response = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": f"{CHAT}:{MSG}"})
assert response.status_code == 409
assert response.json()["reason"] == "cancel_owner_unavailable"
assert not (host / "state/cancel_intents.json").exists()
assert not (other / "state/cancel_intents.json").exists()
def test_direct_operation_with_a_different_cancel_owner_is_explicit(tmp_path, monkeypatch):
from supervisor.active_activity import get_direct_activity_registry
host, other = tmp_path / "host", tmp_path / "other"
client = _client(host)
_inbound(host, "own request")
_isolate_queue(monkeypatch, other, [])
registry = get_direct_activity_registry()
registry.clear()
try:
registry.register("direct", CHAT, client_message_id=MSG, kind="direct_chat", origin_message_ref=_origin_ref(host))
state = client.get(f"/chat/operations/{CHAT}:{MSG}", headers=_headers()).json()
assert state["phase"] == "direct_chat" and state["cancel_supported"] is False
assert state["reason"] == "cancel_owner_unavailable"
response = client.post("/chat/cancel", headers=_headers(), json={"operation_ref": f"{CHAT}:{MSG}"})
assert response.status_code == 409 and response.json()["reason"] == "cancel_owner_unavailable"
finally:
registry.clear()

View file

@ -1,16 +1,20 @@
"""Regression tests for the host-service rate limiter (#27/#30).
"""Regression tests for the host-service rate limiter (#27/#30) and its token bucket.
The `_RateLimiter._hits` defaultdict previously grew unbounded: keys (one per
`{skill}:{endpoint}`) were never deleted, only their stale timestamps popped. A
periodic sweep now frees keys that have gone idle past the window, without
changing any rate-limit decision.
The WS relay lane (`allow_burst`) is a token bucket — a 60-message burst reserve
refilling one message per second — whose refusals are aggregated per burst and
reported once through the `on_burst_end` sink.
"""
from __future__ import annotations
import time
from ouroboros.gateway.host_service import _RateLimiter
from ouroboros.gateway.host_service import WS_RELAY_BURST, WS_RELAY_REFILL_PER_SEC, _RateLimiter
def test_sweep_frees_idle_keys():
@ -57,3 +61,56 @@ def test_swept_key_is_recreated_cleanly():
assert rl.allow("k") is True
assert rl.allow("k") is True
assert rl.allow("k") is False
def test_burst_reserve_admits_capacity_then_refills_one_per_second():
ended: list[tuple[str, int, float]] = []
rl = _RateLimiter(on_burst_end=lambda key, dropped, duration: ended.append((key, dropped, duration)))
assert WS_RELAY_BURST == 60 and WS_RELAY_REFILL_PER_SEC == 1.0
for _ in range(WS_RELAY_BURST):
assert rl.allow_burst("s:ws")["allowed"] is True
refused = rl.allow_burst("s:ws")
assert refused["allowed"] is False and refused["dropped_in_burst"] == 1
assert 0 < refused["retry_after_sec"] <= 1.0
assert rl.allow_burst("s:ws")["dropped_in_burst"] == 2
# Half a second later there is still no whole token.
rl._buckets["s:ws"][1] -= 0.5
assert rl.allow_burst("s:ws")["allowed"] is False
# A full second after the last refill exactly one message passes, which
# closes the burst and reports its aggregate ONCE.
rl._buckets["s:ws"][1] -= 0.5
assert rl.allow_burst("s:ws")["allowed"] is True
assert len(ended) == 1 and ended[0][0] == "s:ws" and ended[0][1] == 3
fresh = rl.allow_burst("s:ws")
assert fresh["allowed"] is False
assert fresh["dropped_in_burst"] == 1, "a new burst starts counting from one"
# The sliding-window lanes are untouched by the bucket.
assert rl.allow("s:inject") is True
assert "s:ws" not in rl._hits
def test_burst_reserve_keys_are_independent_and_idle_buckets_are_swept():
ended: list[tuple[str, int, float]] = []
rl = _RateLimiter(window_sec=0.01, on_burst_end=lambda *args: ended.append(args))
for _ in range(WS_RELAY_BURST):
rl.allow_burst("a:ws")
assert rl.allow_burst("a:ws")["allowed"] is False
assert rl.allow_burst("b:ws")["allowed"] is True, "another skill keeps its own reserve"
# The skill stops sending after the refusal: the sweep (any lane's admission,
# once per window) drops the refilled bucket and still reports the burst.
rl._buckets["a:ws"][1] -= WS_RELAY_BURST / WS_RELAY_REFILL_PER_SEC
rl._last_sweep = time.monotonic() - 100
assert rl.allow("c:inject") is True
assert "a:ws" not in rl._buckets
assert [(key, dropped) for key, dropped, _duration in ended] == [("a:ws", 1)]
def test_burst_sink_failure_never_breaks_admission():
def _boom(*_args):
raise RuntimeError("sink down")
rl = _RateLimiter(on_burst_end=_boom)
for _ in range(WS_RELAY_BURST + 1):
rl.allow_burst("s:ws")
rl._buckets["s:ws"][1] -= 1.0
assert rl.allow_burst("s:ws")["allowed"] is True

View file

@ -553,44 +553,74 @@ def test_logs_js_backfills_all_streams_and_dedupes_without_dropping_preconnect()
assert "loadStart" not in src
def test_llm_call_failure_reaches_the_live_log_exactly_once(tmp_path, production_sink):
"""#355: one LLM failure was two Logs rows — the durable `llm_api_error`
append (forwarded live by the events tail) plus a live-only
`llm_round_error` sibling from the same producer. The producer now writes
the durable row only; through the production sink that is ONE frame."""
import inspect
import json
import queue as queue_mod
@pytest.mark.parametrize("write_succeeds", [True, False])
@pytest.mark.parametrize("sink_installed", [True, False])
def test_llm_call_failure_reaches_the_live_log_exactly_once(
tmp_path, production_sink, monkeypatch, write_succeeds, sink_installed,
):
"""#355: one failure remains visible with or without its durable write/sink.
A successful append with the production sink sends the only live frame.
A failed append or absent sink falls back to the existing live queue, with
the same timestamp and complete evidence, without inventing persistence.
"""
import queue as queue_mod
from ouroboros import loop_llm_call, utils
from ouroboros.loop_llm_call import _LlmErrorContext, _record_llm_call_error
from supervisor.events import _handle_log_event
src = inspect.getsource(_record_llm_call_error)
assert '"llm_round_error"' not in src
assert '"type": "llm_api_error"' in src
if not sink_installed:
utils.set_log_sink(None)
monkeypatch.setattr(loop_llm_call, "utc_now_iso", lambda: "2026-09-07T13:05:00Z")
production_sink.running["m1"] = {"task": {"id": "m1", "chat_id": 42}}
live = queue_mod.Queue()
events_file = tmp_path / "logs/events.jsonl"
attempted_rows = []
append = loop_llm_call.append_jsonl
def append_observed(path, row):
attempted_rows.append(dict(row))
return append(path, row)
monkeypatch.setattr(loop_llm_call, "append_jsonl", append_observed)
if not write_succeeds:
write = utils._write_fd_fully
def refuse_event_write(fd, data, path):
if pathlib.Path(path) == events_file:
raise OSError("controlled journal write refusal")
return write(fd, data, path)
monkeypatch.setattr(utils, "_write_fd_fully", refuse_event_write)
ctx = _LlmErrorContext(
task_id="m1", task_type="task", execution_id="exec-1", round_id="round-1",
llm_call_id="call-1", round_idx=1, attempt=0, model="provider/model",
request_ref=None, drive_logs=tmp_path / "logs", event_queue=live,
accumulated_usage={}, context_fit_event_fields={},
request_ref={"manifest_ref": {"path": "request/ref"}},
drive_logs=tmp_path / "logs", event_queue=live,
accumulated_usage={}, context_fit_event_fields={"selected_context_mode": "test-context"},
)
(tmp_path / "logs").mkdir()
class _ProviderError(RuntimeError):
class ProviderError(RuntimeError):
status_code = 503
_record_llm_call_error(_ProviderError("upstream unavailable"), ctx)
rows = [json.loads(line) for line in (tmp_path / "logs" / "events.jsonl").read_text().splitlines()]
assert [row["type"] for row in rows if row["type"].startswith("llm_")] == ["llm_api_error"]
# The producer's live queue carries no second copy of the same failure.
live_items = []
_record_llm_call_error(ProviderError("upstream unavailable"), ctx)
[error_row] = attempted_rows
assert error_row["type"] == "llm_api_error"
queued = []
while not live.empty():
live_items.append(live.get_nowait())
assert not any(
str((item.get("data") or item).get("type") or "") == "llm_round_error"
for item in live_items if isinstance(item, dict)
), live_items
# Through the production sink the durable append is the one live frame.
assert [f["type"] for f in production_sink.frames if str(f.get("type", "")).startswith("llm_")] == ["llm_api_error"]
queued.append(live.get_nowait())
assert len(queued) == (0 if write_succeeds and sink_installed else 1)
if queued:
assert queued[0] == {"type": "log_event", "data": error_row}
_handle_log_event(queued[0], _ctx(production_sink))
[frame] = production_sink.frames
assert frame == {**error_row, "chat_id": 42}
assert frame["ts"] == "2026-09-07T13:05:00Z"
assert frame["execution_id"] == "exec-1" and frame["llm_call_id"] == "call-1"
assert frame["status_code"] == 503 and frame["request_ref"] == {"path": "request/ref"}
assert frame["selected_context_mode"] == "test-context"
rows = [json.loads(line) for line in events_file.read_text(encoding="utf-8").splitlines()]
assert rows == ([error_row] if write_succeeds else [])

View file

@ -0,0 +1,159 @@
"""A transport timeout leaves remote effects unknown until a server-specific read."""
from __future__ import annotations
import asyncio
import json
import sys
from contextlib import asynccontextmanager
from types import SimpleNamespace
import pytest
from ouroboros import mcp_client
from ouroboros.tools.registry import ToolRegistry
from ouroboros.tools.tool_result import TOOL_CODE_SPECS
@pytest.fixture
def registry(tmp_path, monkeypatch):
mcp_client.reset_manager_for_tests()
monkeypatch.setattr("ouroboros.safety.check_safety", lambda *a, **kw: (True, ""))
yield ToolRegistry(repo_dir=tmp_path, drive_root=tmp_path)
mcp_client.reset_manager_for_tests()
def _configure(server, *, timeout=1):
manager = mcp_client.get_manager()
manager.reconfigure({"MCP_ENABLED": True, "MCP_TOOL_TIMEOUT_SEC": timeout,
"MCP_SERVERS": [{"id": "controlled", "enabled": True, **server}]})
return manager
def _assert_unknown(text):
assert "MCP_TOOL_TIMEOUT" in text
assert "remote outcome is unknown" in text
assert "side effects may already have happened" in text
assert "remote cancellation is not confirmed" in text
assert "server-specific status/read" in text
assert "before retrying" in text
@pytest.mark.parametrize("timeout_phase", ["initialize", "after_effect"])
@pytest.mark.parametrize("projection", ["typed", "text"])
def test_timeout_disclosure_survives_registry_and_status_read(
tmp_path, monkeypatch, registry, timeout_phase, projection,
):
"""Use real client wait/cancel and registry dispatch; only the peer is controlled."""
calls = []
effects = []
closed_sessions = []
first_session = True
@asynccontextmanager
async def transport(_cfg):
yield None, None
class Session:
def __init__(self, *_streams):
nonlocal first_session
self.initial_attempt = first_session
first_session = False
async def __aenter__(self):
return self
async def __aexit__(self, *_exc):
closed_sessions.append(self.initial_attempt)
async def initialize(self):
if self.initial_attempt and timeout_phase == "initialize":
await asyncio.Event().wait()
async def call_tool(self, name, arguments):
calls.append((name, arguments))
if name == "apply":
effects.append(arguments["operation_id"])
await asyncio.Event().wait() # The response never arrives.
return SimpleNamespace(content=[SimpleNamespace(text=json.dumps(effects))], isError=False)
monkeypatch.setattr(mcp_client, "_MCP_SDK_AVAILABLE", True)
monkeypatch.setattr(mcp_client, "_transport_factory", transport)
monkeypatch.setattr(mcp_client, "ClientSession", Session)
manager = _configure({"url": "https://controlled.invalid/mcp"})
async def list_tools(_cfg, _timeout):
return [{"name": name, "input_schema": {}} for name in ("apply", "status")]
manager._async_list_tools = list_tools
assert manager.refresh_server("controlled")["ok"]
args = {"operation_id": "op-667"}
if projection == "typed":
result = registry.execute_result("mcp_controlled__apply", args)
assert (result.status, result.code) == ("timeout", "MCP_TIMEOUT")
text = result.text
else:
text = registry.execute("mcp_controlled__apply", args)
_assert_unknown(text)
assert "reconcile" in TOOL_CODE_SPECS["MCP_TIMEOUT"].recovery
expected = ["op-667"] if timeout_phase == "after_effect" else []
assert effects == expected
assert len(calls) == len(expected) # No implicit replay after timeout.
assert closed_sessions == [True] # Local close is not a remote cancel receipt.
status = registry.execute_result("mcp_controlled__status", args)
assert (status.status, status.code) == ("ok", "OK")
assert status.text.endswith(json.dumps(expected))
assert calls[-1] == ("status", args)
assert effects == expected
assert closed_sessions == [True, False]
def test_real_stdio_effect_before_timeout_is_reconciled_by_status_read(tmp_path, registry):
"""A real SDK peer persists the effect but omits its reply, then serves status."""
pytest.importorskip("mcp")
effects = tmp_path / "effects.jsonl"
calls = tmp_path / "calls.jsonl"
script = tmp_path / "controlled_mcp_timeout.py"
script.write_text('''import json, pathlib, sys
effects, calls = map(pathlib.Path, sys.argv[1:])
for line in sys.stdin:
request = json.loads(line)
if "id" not in request:
continue
method = request["method"]
if method == "initialize":
result = {"protocolVersion": request["params"]["protocolVersion"],
"capabilities": {"tools": {}}, "serverInfo": {"name": "controlled", "version": "1"}}
elif method == "tools/list":
result = {"tools": [{"name": name, "inputSchema": {"type": "object"}}
for name in ("apply", "status")]}
elif method == "tools/call":
params = request["params"]
with calls.open("a", encoding="utf-8") as out:
out.write(json.dumps(params) + "\\n")
if params["name"] == "apply":
with effects.open("a", encoding="utf-8") as out:
out.write(json.dumps(params["arguments"]) + "\\n")
continue # Effect committed; deliberately no response.
rows = [json.loads(line) for line in effects.read_text().splitlines()] if effects.exists() else []
result = {"content": [{"type": "text", "text": json.dumps(rows)}], "isError": False}
else:
result = {}
print(json.dumps({"jsonrpc": "2.0", "id": request["id"], "result": result}), flush=True)
''', encoding="utf-8")
manager = _configure({"transport": "stdio", "command": sys.executable,
"args": [str(script), str(effects), str(calls)]}, timeout=3)
assert manager.refresh_server("controlled")["ok"]
args = {"operation_id": "op-667-real"}
timeout = registry.execute_result("mcp_controlled__apply", args)
assert (timeout.status, timeout.code) == ("timeout", "MCP_TIMEOUT")
_assert_unknown(timeout.text)
assert [json.loads(line) for line in effects.read_text().splitlines()] == [args]
assert [json.loads(line)["name"] for line in calls.read_text().splitlines()] == ["apply"]
status = registry.execute_result("mcp_controlled__status", args)
assert (status.status, status.code) == ("ok", "OK")
assert status.text.endswith(json.dumps([args]))
assert [json.loads(line)["name"] for line in calls.read_text().splitlines()] == ["apply", "status"]
assert [json.loads(line) for line in effects.read_text().splitlines()] == [args]

View file

@ -532,7 +532,7 @@ def scan_data_paths(root: pathlib.Path = REPO) -> frozenset[str]:
# root-task projection with its gaps ledger (``state/skill_review_root_tasks*``)
# and the per-project retirement locks (``state/delegate_project_retirements/``)
# — while the retired acceptance api-fallback record left the population.
EXPECTED_SCAN_PATHS = 286 # Combined skill dependencies, Go caches and artifact capture paths.
EXPECTED_SCAN_PATHS = 286 # Combined Artifact, Skills and Host path owners.
# Scanned paths that must always be present — guards the scanner itself
# against a silent regression that would shrink coverage while keeping counts

View file

@ -57,7 +57,7 @@ _POPEN_ALLOWLIST = {
# so /panic's tracked-subprocess sweep can never observe it alive but
# untracked (isolated_deps._run template).
"ouroboros/claudexor_daemon.py",
"ouroboros/extension_process_runner.py", # waited extension child
"ouroboros/extension_process_runner.py", # waited calls and session-owned response streams
"ouroboros/workspace_executor.py", # custody write-through added at spawn
"ouroboros/local_model.py", # custody record added at spawn
"ouroboros/extension_companion.py", # custody write-through added at spawn

View file

@ -797,6 +797,7 @@ def _extension_proc(returncode):
pid = 9001
def __init__(self):
self.stdin = None # Popen with DEVNULL has no parent-side input pipe.
self.stdout = io.BytesIO(b"")
self.stderr = io.BytesIO(b"")
self.returncode = returncode
@ -833,6 +834,7 @@ def test_extension_child_timeout_publishes_host_kill(tmp_path, monkeypatch):
pid = 9002
def __init__(self):
self.stdin = None # Popen with DEVNULL has no parent-side input pipe.
self.stdout = io.BytesIO(b"")
self.stderr = io.BytesIO(b"")
self.returncode = None

View file

@ -340,7 +340,6 @@ def test_restart_watchdog_waits_for_uvicorn_exit():
def test_owner_restart_copy_is_explicit_about_stopped_task():
source = _read("server.py")
assert 'ctx.send_with_budget(chat_id, "♻️ Restarting.")' in source
assert "Stopping active task. New settings apply to the next message." in source
assert "owner_restart_no_resume.flag" in source
assert "owner_restart_no_resume" in source

View file

@ -536,3 +536,39 @@ def test_skill_exec_extension_message_reports_typed_liveness(tmp_path, monkeypat
assert "SKILL_EXEC_EXTENSION" in out
assert "live_loaded=False" in out
assert "has already been called" not in out
@pytest.mark.parametrize('registration', [True, False])
def test_documented_module_widget_recipe_creates_the_registered_tab(tmp_path, registration):
"""Execute the documentation example; manifest metadata alone stays metadata."""
import re
from tests._extension_loader_shared import _write_ext_skill
from ouroboros import extension_loader
from ouroboros.skill_loader import find_skill, save_enabled, save_review_state, SkillReviewState
from tests._shared import clean_extension_runtime_state
text = (pathlib.Path(__file__).resolve().parents[1] / 'docs/CREATING_SKILLS.md').read_text(encoding='utf-8')
section = text.split('### `kind: "module"` widgets', 1)[1].split('#### The in-frame bridge', 1)[0]
declaration = re.search(r'```yaml\n(.*?)```', section, re.S).group(1)
plugin = re.search(r'```python\n(.*?)```', section, re.S).group(1)
root, skills = tmp_path / 'drive', tmp_path / 'skills'
skill_dir = _write_ext_skill(skills, 'recipe', permissions=['widget'],
plugin_body=plugin if registration else 'def register(api):\n pass\n',
extra_frontmatter=declaration)
widget = "document.getElementById('root').textContent = 'Recipe widget ✓';"
(skill_dir / 'widget.js').write_text(widget, encoding='utf-8')
assert (skill_dir / 'widget.js').read_bytes() == widget.encode('utf-8')
loaded = find_skill(root, 'recipe', repo_path=str(skills))
save_enabled(root, 'recipe', True)
save_review_state(root, 'recipe', SkillReviewState(status='pass', content_hash=loaded.content_hash))
loaded = find_skill(root, 'recipe', repo_path=str(skills))
clean_extension_runtime_state()
try:
assert extension_loader.load_extension(loaded, lambda: {}, drive_root=root, repo_path=str(skills)) is None
rows = extension_loader.live_widget_projection('recipe')
assert len(rows) == (1 if registration else 0)
if registration:
assert rows[0]['tab']['render']['entry'] == 'widget.js'
assert rows[0]['tab']['title'] == 'Editor'
finally:
clean_extension_runtime_state()

View file

@ -0,0 +1,297 @@
"""Optional real-browser/native-window widget exports through installed skill routes."""
from __future__ import annotations
import ast
import base64
import hashlib
import json
import logging
import os
from pathlib import Path
import shutil
import socket
import tempfile
import threading
import time
from types import SimpleNamespace
import urllib.parse
import urllib.request
import pytest
pytestmark = [pytest.mark.ui_browser, pytest.mark.serial]
REPO = Path(__file__).resolve().parents[1]
_WIDGET = r'''
const root = document.getElementById('root');
root.innerHTML = `<style>body{font:16px sans-serif;background:#18212d;color:#edf4ff;padding:20px}button{display:block;margin:14px 0;padding:12px 20px;min-width:280px}#status{padding-top:12px}</style>
<h2>Export from an isolated skill</h2>
<button id="blob">Export Blob</button><button id="data">Export data URL</button>
<button id="explicit">Save with widget API</button><button id="route">Download large file</button><button id="stream">Read stream</button><p id="status">Ready</p>`;
const anchor = (href, filename) => {
const a = document.createElement('a'); a.href=href; a.download=filename;
document.body.append(a); a.click(); a.remove();
};
document.getElementById('blob').onclick = () => {
const url=URL.createObjectURL(new Blob(['blob export'], {type:'text/plain'}));
anchor(url,'widget-blob.txt'); URL.revokeObjectURL(url);
};
document.getElementById('data').onclick = () => anchor('data:text/plain;base64,ZGF0YSBleHBvcnQ=','widget-data.txt');
document.getElementById('explicit').onclick = async () => {
const result=await OuroborosWidget.download('widget-explicit.txt',new Blob(['explicit export']));
document.getElementById('status').textContent=result.native?'Saved by the desktop app':'Browser download started';
};
document.getElementById('route').onclick = () => anchor('/api/extensions/export_widget/export','widget-large.bin');
document.getElementById('stream').onclick = async () => {
const reader=(await OuroborosWidget.fetch('/api/extensions/export_widget/stream')).body.getReader();
let text=new TextDecoder().decode((await reader.read()).value);
window.parent.postMessage({type:'fixture-first'},'*');
await new Promise(resolve=>{const next=e=>{if(e.data?.fixtureContinue){window.removeEventListener('message',next);resolve()}};window.addEventListener('message',next)});
for(;;){const chunk=await reader.read();if(chunk.done)break;text+=new TextDecoder().decode(chunk.value)}
document.getElementById('status').textContent='Stream: '+text;
window.parent.postMessage({type:'fixture-stream-done',text},'*');
};
requestAnimationFrame(() => window.parent.postMessage({type:'fixture-layout', buttons:
Object.fromEntries(['blob','data','explicit','route','stream'].map(id=>{const r=document.getElementById(id).getBoundingClientRect();return [id,{x:r.x+r.width/2,y:r.y+r.height/2}]}))},'*'));
'''
_PLUGIN = '''import asyncio
from starlette.responses import FileResponse, StreamingResponse
def export(request):
return FileResponse(request.app.state.drive_root / 'large.bin', filename='widget-large.bin')
def stream(request):
async def chunks():
yield b'first'
while not (request.app.state.drive_root / 'release-stream').exists():
await asyncio.sleep(.01)
yield b'last'
return StreamingResponse(chunks())
def register(api):
api.register_route('export', export)
api.register_route('stream', stream)
api.register_ui_tab('main', 'Exports', render={'kind':'module','entry':'widget.js'})
'''
_HTML = '''<!doctype html><html><head><meta charset="utf-8"><link rel="stylesheet" href="/web/style.css"></head>
<body style="padding:30px"><h1>Skill export check</h1><section data-widget-key="export_widget:main"><div class="widgets-card-status"></div><div id="mount"></div></section>
<script type="module">
import {mountModuleWidget} from '/web/modules/widget_module.js';
window.addEventListener('message',e=>{
if(e.data?.type==='fixture-layout')window.fixtureButtons=e.data.buttons;
if(e.data?.type==='fixture-first')window.fixtureFirst=true;
if(e.data?.type==='fixture-stream-done')window.fixtureStreamText=e.data.text;
});
const nativeFetch=window.fetch.bind(window);
window.fixtureReads=0;
window.fetch=async (...args)=>{
const response=await nativeFetch(...args);
if(String(args[0]).endsWith('/stream')){
const get=response.body.getReader.bind(response.body);
response.body.getReader=(...options)=>{const reader=get(...options),read=reader.read.bind(reader);reader.read=(...values)=>{window.fixtureReads++;return read(...values)};return reader};
}
return response;
};
window.disposeWidget=await mountModuleWidget(document.getElementById('mount'),{skill:'export_widget',ws_prefix:''},{entry:'widget.js',height:520});
window.fixtureReady=true;
</script></body></html>'''
@pytest.fixture
def widget_server(tmp_path, monkeypatch):
from starlette.applications import Starlette
from starlette.responses import HTMLResponse
from starlette.routing import Mount, Route
from starlette.staticfiles import StaticFiles
import uvicorn
from ouroboros import extension_loader
from ouroboros.gateway.extensions import api_extension_module, api_extension_dispatch
from ouroboros.skill_loader import find_skill, save_enabled, save_review_state, SkillReviewState
from tests._extension_loader_shared import _write_ext_skill, _add_fake_native_dep, _mark_isolated_deps_installed
from tests._shared import clean_extension_runtime_state
root, skills = tmp_path / 'drive', tmp_path / 'skills'
root.mkdir()
monkeypatch.setenv('OUROBOROS_RUNTIME_MODE', 'advanced')
monkeypatch.setattr('ouroboros.config.get_skills_repo_path', lambda: str(skills))
skill_dir = _write_ext_skill(skills, 'export_widget', plugin_body=_PLUGIN, permissions=['route','widget'],
extra_frontmatter='dependencies:\n - dummy_pkg\n')
(skill_dir / 'widget.js').write_text(_WIDGET)
loaded = find_skill(root, 'export_widget', repo_path=str(skills))
save_enabled(root, loaded.name, True)
save_review_state(root, loaded.name, SkillReviewState(status='pass', content_hash=loaded.content_hash))
loaded = find_skill(root, loaded.name, repo_path=str(skills))
_add_fake_native_dep(loaded)
_mark_isolated_deps_installed(root, loaded)
clean_extension_runtime_state()
assert extension_loader.load_extension(loaded, lambda: {}, drive_root=root, repo_path=str(skills)) is None
(root / 'large.bin').write_bytes(b'large export\n' * (128 * 1024))
expected = {'widget-blob.txt': b'blob export', 'widget-data.txt': b'data export',
'widget-explicit.txt': b'explicit export', 'widget-large.bin': (root / 'large.bin').read_bytes()}
app = Starlette(routes=[Route('/', lambda _r: HTMLResponse(_HTML)),
Route('/api/extensions/{skill}/module/{entry:path}', api_extension_module),
Route('/api/extensions/{skill}/{rest:path}', api_extension_dispatch, methods=['GET','HEAD']),
Mount('/web', app=StaticFiles(directory=REPO / 'web'))])
app.state.drive_root, app.state.repo_dir = root, REPO
sock = socket.socket()
sock.bind(('127.0.0.1', 0))
server = uvicorn.Server(uvicorn.Config(app, log_level='warning'))
thread = threading.Thread(target=server.run, kwargs={'sockets':[sock]}, daemon=True)
thread.start()
deadline = time.monotonic() + 10
while not server.started and thread.is_alive() and time.monotonic() < deadline:
time.sleep(.01)
assert server.started
try:
yield {'url':f'http://127.0.0.1:{sock.getsockname()[1]}', 'port':sock.getsockname()[1], 'root':root, 'expected':expected}
finally:
server.should_exit = True
thread.join(10)
sock.close()
clean_extension_runtime_state()
assert not thread.is_alive()
def native_file_api(port, read_sizes):
"""Reuse the sibling lane's AST harness for this checkout's real launcher owner."""
source = REPO / 'launcher.py'
names = {'_resolve_bridge_file_url','_unique_bridge_target','_fetch_bridge_url_to','MainApi'}
selected = [node for node in ast.walk(ast.parse(source.read_text()))
if isinstance(node,(ast.FunctionDef,ast.ClassDef)) and node.name in names]
assert {node.name for node in selected} == names
def open_recorded(*args, **kwargs):
response = urllib.request.urlopen(*args, **kwargs)
original = response.read
def read(size=-1):
assert 0 < size <= 1024 * 1024
read_sizes.append(size)
return original(size)
response.read = read
return response
namespace = {'actual_port':port,'pathlib':__import__('pathlib'),'shutil':shutil,'tempfile':tempfile,
'base64':base64,'log':logging.getLogger(__name__),
'urllib':SimpleNamespace(parse=urllib.parse,request=SimpleNamespace(urlopen=open_recorded))}
exec(compile(ast.Module(body=selected,type_ignores=[]),str(source),'exec'),namespace)
return namespace['MainApi'](), hashlib.sha256(source.read_bytes()).hexdigest()
def _evidence(tmp_path):
root = Path(os.environ.get('OUROBOROS_UI_EVIDENCE_DIR', str(tmp_path / 'evidence')))
root.mkdir(parents=True, exist_ok=True)
return root
def test_browser_widget_exports(widget_server, tmp_path):
from playwright.sync_api import sync_playwright
evidence = _evidence(tmp_path)
with sync_playwright() as playwright:
browser = playwright.chromium.launch(executable_path=os.environ.get('OUROBOROS_TEST_CHROME') or None)
try:
page = browser.new_page(viewport={'width':1100,'height':850}, accept_downloads=True)
page.goto(widget_server['url'])
frame = page.frame_locator('iframe')
for control, name in zip(['blob','data','explicit','route'], widget_server['expected']):
with page.expect_download() as pending:
frame.locator('#'+control).click()
path = evidence / ('browser-' + name)
pending.value.save_as(path)
assert path.read_bytes() == widget_server['expected'][name]
frame.locator('#stream').click()
page.wait_for_function('window.fixtureFirst === true')
assert page.evaluate('window.fixtureReads') == 1
(widget_server['root'] / 'release-stream').touch()
page.evaluate('document.querySelector("iframe").contentWindow.postMessage({fixtureContinue:true},"*")')
page.wait_for_function('window.fixtureStreamText === "firstlast"')
page.screenshot(path=str(evidence / 'widget-browser.png'))
finally:
browser.close()
def _native_controller(window, actions, expected, root, evidence, failures):
"""Drive real Qt mouse events into the opaque iframe and verify saved files."""
from qtpy.QtCore import QPoint, Qt
from qtpy.QtTest import QTest
def gui(function):
done, result = threading.Event(), []
def execute():
try:
result.append(function())
except BaseException as exc:
result.append(exc)
finally:
done.set()
actions.call.emit(execute)
assert done.wait(10), 'GUI action did not complete'
if isinstance(result[0], BaseException):
raise result[0]
return result[0]
try:
deadline = time.monotonic()+30
while time.monotonic()<deadline:
if window.evaluate_js('Boolean(window.fixtureButtons && window.pywebview?.api?.save_bytes_to_downloads)'):
break
time.sleep(.05)
else:
raise AssertionError('native widget and bridge did not become ready')
coords = window.evaluate_js('({buttons:window.fixtureButtons,frame:(()=>{const r=document.querySelector("iframe").getBoundingClientRect();return {x:r.x,y:r.y}})()})')
for control, name in zip(['blob','data','explicit','route'], expected):
point = coords['buttons'][control]
x, y = round(coords['frame']['x']+point['x']+1), round(coords['frame']['y']+point['y']+1)
gui(lambda: QTest.mouseClick(window.native.webview.focusProxy() or window.native.webview,
Qt.MouseButton.LeftButton, pos=QPoint(x,y)))
path = Path.home() / 'Downloads' / name
deadline = time.monotonic()+20
while time.monotonic()<deadline:
if path.exists() and path.read_bytes()==expected[name]:
break
time.sleep(.05)
else:
raise AssertionError(f'native export did not save {name}')
shutil.copy2(path,evidence / ('native-'+name))
point = coords['buttons']['stream']
x, y = round(coords['frame']['x']+point['x']+1), round(coords['frame']['y']+point['y']+1)
gui(lambda: QTest.mouseClick(window.native.webview.focusProxy() or window.native.webview,
Qt.MouseButton.LeftButton, pos=QPoint(x,y)))
deadline = time.monotonic()+20
while not window.evaluate_js('window.fixtureFirst === true') and time.monotonic()<deadline:
time.sleep(.02)
assert window.evaluate_js('window.fixtureFirst === true')
assert window.evaluate_js('window.fixtureReads') == 1
(root / 'release-stream').touch()
window.evaluate_js('document.querySelector("iframe").contentWindow.postMessage({fixtureContinue:true},"*")')
deadline = time.monotonic()+20
while window.evaluate_js('window.fixtureStreamText') != 'firstlast' and time.monotonic()<deadline:
time.sleep(.02)
assert window.evaluate_js('window.fixtureStreamText') == 'firstlast'
from qtpy.QtWidgets import QApplication
gui(lambda: QTest.qWait(100))
gui(lambda: QApplication.primaryScreen().grabWindow(int(window.native.winId())).save(str(evidence / 'widget-native-screen.png')))
except BaseException as exc:
failures.append(exc)
finally:
window.destroy()
def test_native_widget_exports(widget_server, tmp_path, monkeypatch):
runtime = tmp_path / 'qt-runtime'
runtime.mkdir(mode=0o700)
monkeypatch.setenv('XDG_RUNTIME_DIR', str(runtime))
webview = pytest.importorskip('webview')
pytest.importorskip('qtpy')
from qtpy.QtCore import QObject, Signal, Slot
if os.environ.get('PYWEBVIEW_GUI') != 'qt':
pytest.skip('native Qt probe is explicitly selected by its isolated launcher')
class Actions(QObject):
call = Signal(object)
@Slot(object)
def execute(self, function):
function()
actions = Actions()
actions.call.connect(actions.execute)
evidence = _evidence(tmp_path)
read_sizes, failures = [], []
api, source_hash = native_file_api(widget_server['port'], read_sizes)
window = webview.create_window('Widget export test', widget_server['url'], js_api=api, width=1100,height=850)
webview.start(_native_controller, (window,actions,widget_server['expected'],widget_server['root'],evidence,failures), gui='qt')
assert not failures, failures
assert read_sizes and all(size > 0 for size in read_sizes)
(evidence / 'native-receipt.json').write_text(json.dumps({'launcher_sha256':source_hash,
'reads':read_sizes,'files':{name:{'size':len(data),'sha256':hashlib.sha256(data).hexdigest()}
for name,data in widget_server['expected'].items()}},indent=2)+'\n')

View file

@ -231,7 +231,7 @@ def test_widgets_keep_iframe_sandbox_locked_down():
assert "const csp = moduleFrameCsp(tab.skill);" in module
assert "connect-src" not in source
assert "'unsafe-eval'" not in source
assert "window.OuroborosWidget = { fetch: request, onEvent };" in source
assert "window.OuroborosWidget = { fetch: request, onEvent, download };" in source
assert "module widget fetch outside extension route prefix" in source

View file

@ -237,6 +237,8 @@
* @property {boolean=} markdown
* @property {boolean=} is_progress
* @property {string=} task_id
* @property {Object=} origin_message_ref
* Host-captured inbound identity for a correlated operation's terminal reply.
* @property {boolean=} ephemeral_decision
* @property {string=} task_phase
* "finalizing" on a root's early final answer: post-task synthesis still

View file

@ -11,6 +11,7 @@ export function bridgeChunkBuffer(view) {
// Child side of the one bridge grammar (nonce-bound, parent ⇄ frame):
// child → parent ouro-widget-fetch {id, url, init} · ouro-widget-fetch-abort {id}
// ouro-widget-fetch-pull {id} · ouro-widget-download {id, name, source}
// ouro-widget-events {op: subscribe | unsubscribe} · ouro-widget-disposed
// ouro-widget-error {kind: error | rejection | csp, message, source, line}
// parent → child ouro-widget-fetch-chunk {id, phase: headers | data | end | error, …}
@ -19,16 +20,18 @@ export function bridgeChunkBuffer(view) {
// ReadableStream fed by `data` frames (binary by default), so text/json/blob
// and incremental body reads all work. No default timeout — `init.timeoutMs`
// is the author's opt-in bound; `init.signal` aborts through the parent.
export function moduleBridgeScript(nonce) {
export function moduleBridgeScript(nonce, routeBase = '') {
return `
(() => {
const nonce = ${JSON.stringify(nonce)};
const routeBase = ${JSON.stringify(routeBase)};
let seq = 0;
let disposing = false;
let disposed = false;
// id → in-flight bridged fetch: settles its Response on the headers
// frame, then feeds, ends or errors that Response's body stream.
const pending = new Map();
const downloads = new Map();
const cleanup = new Set();
const eventListeners = new Set();
const post = (message) => window.parent.postMessage({ ...message, nonce }, '*');
@ -49,10 +52,16 @@ export function moduleBridgeScript(nonce) {
const hooks = Array.from(cleanup);
cleanup.clear();
await Promise.allSettled(hooks.map((fn) => Promise.resolve().then(fn)));
window.document?.removeEventListener('click', clickDownload);
blobUrls.clear();
if (createUrl) urlApi.createObjectURL = createUrl;
if (revokeUrl) urlApi.revokeObjectURL = revokeUrl;
post({ type: 'ouro-widget-disposed' });
disposed = true;
pending.forEach((item) => item.fail(new Error('widget disposed')));
pending.clear();
downloads.forEach(({ reject }) => reject(new Error('widget disposed')));
downloads.clear();
eventListeners.clear();
window.removeEventListener('message', onMessage);
window.removeEventListener('error', onError);
@ -76,6 +85,14 @@ export function moduleBridgeScript(nonce) {
});
return;
}
if (msg.type === 'ouro-widget-download-result') {
const item = downloads.get(msg.id);
if (!item) return;
downloads.delete(msg.id);
if (msg.result?.ok) item.resolve(msg.result);
else item.reject(new Error(msg.result?.error || 'widget download failed'));
return;
}
if (msg.type !== 'ouro-widget-fetch-chunk') return;
pending.get(msg.id)?.frame(msg);
};
@ -93,9 +110,12 @@ export function moduleBridgeScript(nonce) {
const method = String(init.method || 'GET').toUpperCase();
let settled = false;
let body = null;
let pulled = null;
const finish = () => {
pending.delete(id);
signal?.removeEventListener('abort', onAbort);
pulled?.();
pulled = null;
};
const fail = (error) => {
finish();
@ -122,8 +142,14 @@ export function moduleBridgeScript(nonce) {
const nullBody = method === 'HEAD' || [204, 205, 304].includes(Number(msg.status));
const stream = nullBody ? null : new ReadableStream({
start(controller) { body = controller; },
pull() {
return new Promise((done) => {
pulled = done;
post({ type: 'ouro-widget-fetch-pull', id });
});
},
cancel,
});
}, { highWaterMark: 0 });
try {
resolve(new Response(stream, {
status: Number(msg.status) || 200,
@ -140,6 +166,8 @@ export function moduleBridgeScript(nonce) {
}
if (msg.phase === 'data') {
try { body?.enqueue(new Uint8Array(msg.chunk)); } catch {}
pulled?.();
pulled = null;
return;
}
if (msg.phase === 'end') {
@ -203,7 +231,40 @@ export function moduleBridgeScript(nonce) {
window.addEventListener('message', onMessage);
window.__ouroWidgetOnDispose = onDispose;
window.fetch = request;
window.OuroborosWidget = { fetch: request, onEvent };
const download = (name, source) => new Promise((resolve, reject) => {
if (disposed) { reject(new Error('widget disposed')); return; }
const id = ++seq;
downloads.set(id, { resolve, reject });
try { post({ type: 'ouro-widget-download', id, name: String(name || 'download'), source: typeof source === 'string' ? (blobUrls.get(source) || source) : source }); }
catch (error) { downloads.delete(id); reject(error); }
});
// Remember the Blob behind a frame-owned URL: the opaque frame's
// URL cannot be fetched by the host, and its CSP forbids script IO.
const blobUrls = new Map();
const urlApi = window.URL;
const createUrl = urlApi?.createObjectURL?.bind(urlApi);
const revokeUrl = urlApi?.revokeObjectURL?.bind(urlApi);
if (createUrl) urlApi.createObjectURL = (source) => {
const url = createUrl(source);
if (source instanceof Blob) blobUrls.set(url, source);
return url;
};
if (revokeUrl) urlApi.revokeObjectURL = (url) => {
blobUrls.delete(String(url));
return revokeUrl(url);
};
const clickDownload = (event) => {
if (event.defaultPrevented || event.button > 0) return;
const anchor = event.target?.closest?.('a[download]');
if (!anchor) return;
const href = String(anchor.href || '');
const source = blobUrls.get(href) || href;
if (!blobUrls.has(href) && !href.startsWith('data:') && !(routeBase && href.startsWith(routeBase))) return;
event.preventDefault();
download(anchor.download, source).catch((error) => fault('error', error.message, '', 0));
};
window.document?.addEventListener('click', clickDownload);
window.OuroborosWidget = { fetch: request, onEvent, download };
})();
`;
}

View file

@ -10,6 +10,7 @@ import { escapeHtmlAttr as escapeHtml } from './utils.js';
import { bridgeChunkBuffer, moduleBridgeScript, moduleResizeScript } from './widget_frame.js';
import { boundedNumber, WIDGET_DISPOSE_ACK_TIMEOUT_MS, WIDGET_REQUEST_TIMEOUT_MS } from './widget_job.js';
import { setWidgetCardFault } from './widget_card.js';
import { downloadViaHostBridge, downloadBlobViaHostBridge } from './ui_helpers.js';
export const WIDGET_FRAME_DEFAULT_HEIGHT = 320;
export const WIDGET_FRAME_MAX_HEIGHT = 8192;
@ -136,7 +137,7 @@ export async function mountModuleWidget(mount, tab, render, mountSignal = null,
.replace(/<!--/g, '<\\!--');
const autoHeight = render.height === undefined || render.height === null;
const maxHeight = frameMaxHeight(render);
const bridge = moduleBridgeScript(nonce);
const bridge = moduleBridgeScript(nonce, `${window.location.origin}${expectedPrefix}`);
const resizeBridge = autoHeight
? moduleResizeScript(
nonce, WIDGET_FRAME_DEFAULT_HEIGHT, maxHeight, WIDGET_FRAME_BORDER_RESERVE,
@ -176,6 +177,22 @@ export async function mountModuleWidget(mount, tab, render, mountSignal = null,
const id = msg.id;
const init = msg.init || {};
const controller = new AbortController();
let credit = false;
let wake = null;
controller.pull = () => { credit = true; wake?.(); };
const waitForPull = async () => {
if (!credit) await new Promise((resolve, reject) => {
const abort = () => { wake = null; reject(new DOMException('Aborted', 'AbortError')); };
wake = () => {
wake = null;
controller.signal.removeEventListener('abort', abort);
resolve();
};
if (controller.signal.aborted) abort();
else controller.signal.addEventListener('abort', abort, { once: true });
});
credit = false;
};
pendingRequests.set(id, controller);
let timedOut = false;
const timeoutMs = Number(init.timeoutMs);
@ -208,6 +225,7 @@ export async function mountModuleWidget(mount, tab, render, mountSignal = null,
const reader = r.body?.getReader();
while (reader) {
if (!iframe.isConnected) controller.abort();
await waitForPull();
const { done, value } = await reader.read();
if (done) break;
const chunk = bridgeChunkBuffer(value);
@ -221,6 +239,25 @@ export async function mountModuleWidget(mount, tab, render, mountSignal = null,
pendingRequests.delete(id);
}
};
const relayDownload = async (msg) => {
let result;
try {
const source = msg.source;
const name = String(msg.name || 'download');
if (source instanceof Blob || (typeof source === 'string' && source.startsWith('data:'))) {
result = await downloadBlobViaHostBridge(source, name);
} else {
const parsed = new URL(String(source || ''), window.location.origin);
if (parsed.origin !== window.location.origin || !parsed.pathname.startsWith(expectedPrefix)) {
throw new Error('module widget download outside extension route prefix');
}
result = await downloadViaHostBridge(parsed.pathname + parsed.search, name, { streaming: true });
}
} catch (error) {
result = { ok: false, error: error?.message || String(error) };
}
post({ type: 'ouro-widget-download-result', id: msg.id, result });
};
const onMessage = (event) => {
if (disposed || !iframe || event.source !== iframe.contentWindow) return;
const msg = event.data || {};
@ -255,6 +292,14 @@ export async function mountModuleWidget(mount, tab, render, mountSignal = null,
else if (msg.op === 'unsubscribe') messageHandlers?.delete(onWsMessage);
return;
}
if (msg.type === 'ouro-widget-download') {
relayDownload(msg);
return;
}
if (msg.type === 'ouro-widget-fetch-pull') {
pendingRequests.get(msg.id)?.pull();
return;
}
if (msg.type === 'ouro-widget-fetch-abort') {
pendingRequests.get(msg.id)?.abort();
return;

View file

@ -241,3 +241,47 @@ test('the fault channel caps at 10 posts per frame and is removed on dispose', a
assert.equal(listeners.has(type), false, type);
}
});
test('body pull grants one next chunk only when the consumer reads', async () => {
const { window, posted, chunk, flush } = bridgeHarness();
const pending = window.fetch('/api/extensions/s/stream');
chunk(1, 'headers', { status: 200, headers: [] });
const response = await pending;
const pulls = () => posted.filter((message) => message.type === 'ouro-widget-fetch-pull');
await flush();
assert.equal(pulls().length, 0);
const reader = response.body.getReader();
const first = reader.read();
await flush();
assert.equal(pulls().length, 1);
chunk(1, 'data', { chunk: bytes(1, 2) });
assert.deepEqual(Array.from((await first).value), [1, 2]);
await flush();
assert.equal(pulls().length, 1);
const last = reader.read();
await flush();
assert.equal(pulls().length, 2);
chunk(1, 'end');
assert.equal((await last).done, true);
});
test('download uses the nonce bridge and settles from the actual host outcome', async () => {
const { window, posted, deliver, listeners, flush } = bridgeHarness();
const blob = new Blob(['report'], { type: 'text/plain' });
let settled = false;
const pending = window.OuroborosWidget.download('report.txt', blob).then((result) => { settled = true; return result; });
assert.equal(posted[0].type, 'ouro-widget-download');
assert.equal(posted[0].source, blob);
assert.equal(posted[0].name, 'report.txt');
listeners.get('message')({ source: window.parent, data: {
nonce: 'foreign', type: 'ouro-widget-download-result', id: 1, result: { ok: true },
} });
await flush();
assert.equal(settled, false);
deliver({ type: 'ouro-widget-download-result', id: 1, result: { ok: true, native: true, filename: 'report.txt' } });
assert.deepEqual(await pending, { ok: true, native: true, filename: 'report.txt' });
const failure = window.OuroborosWidget.download('bad.txt', blob);
deliver({ type: 'ouro-widget-download-result', id: 2, result: { ok: false, error: 'disk full' } });
await assert.rejects(failure, /disk full/);
});