diff --git a/docs/architecture/06-agent-core.md b/docs/architecture/06-agent-core.md index 57bc9d1e6..07a427ec5 100644 --- a/docs/architecture/06-agent-core.md +++ b/docs/architecture/06-agent-core.md @@ -126,7 +126,7 @@ Retry budgets are failure-class specific: empty/incomplete responses and transie Transport failures are classified by physical-attempt custody; each class has its own owner and rails. A REMOTE pre-dispatch failure is `transport_unavailable` (`loop_llm_call.classify_llm_exception`: released custody, $0, on a non-local provider). It takes one physical attempt per call; pacing belongs to the round-level wait episode (`ouroboros/loop_transport.py`), which tries the fallback chain once only when `USE_LOCAL_FALLBACK` makes it local (remote candidates never dial over a proven dead egress), waits (durable `network_wait` events, progress that keeps the idle rail alive, owner-interruptible backoff), then redials the SAME round for free. A managed task waits as long as its existing rails allow — owner deadline minus the dispatch-admission reserve, budget, Stop, absolute ceiling; never a new setting, and `OUROBOROS_TRANSIENT_RETRY_MAX` bounds only transient PROVIDER failures — and its exhaustion is its terminal result. The interactive class (every turn stamped direct-chat: owner chat, Presence turns, consciousness wake-ups) waits the same way but has no queue rails and ordinarily no owner deadline, so its episode is bounded by the task idle timeout (`OUROBOROS_TASK_IDLE_TIMEOUT_SEC`) measured from each episode's entry: it limits idle WAITING only and never cancels in-flight work, the shorter of it and an explicit deadline binds, the durable `ended` detail names the expired rail (`interactive_wait_window_exhausted`, or the deadline's own), and exhaustion emits an owner note alongside the provider-outage terminal. Closing `network_wait` rows exist only for cooperative exits — an external kill, Panic, crash or direct-turn hard stop leaves none — so correlate a confirmed `task_done` or durable terminal result; silence never proves completion. -The mid-flight class is `provider_outcome_unknown`: a dispatched request without a terminal provider fact keeps its unresolved monetary bound and is never resent. A NEW logical request afterwards needs a unique host-attested input; who gets one is split by caller class. Ordinary managed tasks and native API children use the upstream-observation continuation (§6 Review delivery): the same wait admits a marked new physical attempt once connectivity is observed. A configured-session supervisor with exactly one live delegated leaf first latches a durable unknown-provider hold (`ouroboros/delegate_hold.py`, closing any transport episode with `network_wait ended / hold_latched`) and parks in the ordinary `supervised_wait`, resuming with a NEW round only on a meaningful leaf wake whose receipt is that input (control wakes and an ineligible hold grant nothing; budget admission stays fail-closed against the unresolved bound), and uses the managed continuation when no hold applies. Interactive primary rounds (direct-chat and Presence turns) get the bounded transport-death repeat below: no-effect by construction (no response, so nothing ran or was sent; §12), not an event retry. Every other surface — forced-final, fallback-chain candidates, review actors, safety, probes, web search, consolidation/summary/reflection, external-harness delegated runs — keeps no-resend (budget 0, or it never enters `call_llm_with_retry`), and an unknown outcome that is not a typed transport death is never resent. +The mid-flight class is `provider_outcome_unknown`: a dispatched request with no terminal provider fact keeps its unresolved monetary bound and is never resent. A NEW logical request afterwards needs a unique host-attested input, granted per caller class. Ordinary managed tasks and native API children use the upstream-observation continuation (§6 Review delivery), whose wait admits a marked new physical attempt on observed connectivity. A configured-session supervisor with exactly one live delegated leaf latches a durable unknown-provider hold (`ouroboros/delegate_hold.py`, closing any transport episode as `network_wait ended / hold_latched`) and parks in the ordinary `supervised_wait`, resuming in a NEW round only on a meaningful leaf wake whose receipt is that input (control wakes and an ineligible hold grant nothing; budget admission stays fail-closed on the unresolved bound); without a hold it uses the managed continuation. Interactive primary rounds get the bounded transport-death repeat below: no-effect by construction (no Host-executed tool or correspondent action can stem from the failed attempt; priced read-only provider-owned retrieval may rerun; §12), not an event retry. Every other surface — forced-final, fallback-chain candidates, review actors, safety, probes, web search, consolidation/summary/reflection, external-harness delegated runs — keeps no-resend (budget 0, or never enters `call_llm_with_retry`); an unknown outcome not typed as a transport death is never resent. The transport-death repeat rail: a typed death (`transport_custody.is_retryable_transport_death` — httpx `ReadError`/`WriteError`/`RemoteProtocolError` through the explicit `__cause__` chain, or a requests wrapper carrying `ProtocolError`/`RemoteDisconnected`; never a timeout, a provider status/body error, a pre-dispatch failure, a local provider or a loopback route) lets the PRIMARY main-loop round dispatch alone (`_dispatch_round_model`, `transport_death_retries`) repeat the SAME logical request at most twice per round. Each repeat is a NEW physical attempt with its own ledger row, and the earlier rows stay `unresolved` at their upper bound (a fully dead round reserves up to three: honest accounting over a cheaper rail). A round record (`execution_id:round:round_idx`) counts the repeats and names the class the terminal reports; `llm_non_retryable_same_request` marks only the exhaustion. A current typed `finalize_now` control or the deadline can refuse a not-yet-sent repeat (`finalize_control_pending` names the control case); the round then exits through the unknown no-resend terminal. A round holding a repeat record sends nothing but those typed-death repeats: a repeat failing with any other class ends the round on the same terminal — no transient burst, empty-response retry, compaction retry, forced-final dial, fallback chain or local-only pass — because the earlier request may still be live; only a repeat that never left the host (released custody) stays the free wait episode's to redial, and when that window closes the terminal is still `provider_outcome_unknown_no_resend`. diff --git a/docs/architecture/12-host-service-companions-and-chat-ids.md b/docs/architecture/12-host-service-companions-and-chat-ids.md index 2e6d5e146..fe706f713 100644 --- a/docs/architecture/12-host-service-companions-and-chat-ids.md +++ b/docs/architecture/12-host-service-companions-and-chat-ids.md @@ -6,23 +6,23 @@ Chat uploads have one storage owner, `gateway.files`, for Host-confined paths an Telegram document mirroring resolves the captured `file_ref` and streams an owned file handle through its multipart client (legacy inline bytes still accepted). Its outgoing 50 MiB boundary is separate from the integration's inbound 10 MiB download limit. Oversized files stay saved in the application: a ready, already-running owner-authenticated Mini App can be opened through its entry button; otherwise the message says it cannot mirror the file and directs the owner to the app. Delivery never starts a tunnel or publishes a new public or token-bearing artifact URL, and the notice never claims the bytes were uploaded to Telegram. -The Host Service is a loopback, authenticated callback boundary for reviewed skills (`ouroboros/gateway/host_service.py`, `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}`). Every request authenticates an opaque `x-skill-token` bound to the skill's content hash, executable review, enablement and grants — secrets never enter the token, a payload edit stales it, and the client wrapper refuses stringification (`skill_token.py`). The frozen route family is exactly: `/identity`, `/tools/schemas`, `/chat/allocate-internal`, `/chat/inject`, `/chat/operations/{operation_ref}`, `/chat/cancel`, `/chat/decision`, `/presence/turn`, `/presence/work/{work_ref}`, `/presence/delivery`, `/ui/ws-message`, and WS `/events`; permissions still decide which route works for a given skill. Review of a transport skill evaluates identity binding, attribution, polling bounds, panic cleanup, token confinement and exfiltration — an owner-bound reviewed transport may be a first-class control surface, not a screen-only integration. External slash commands bind a separate positive-identity external owner slot (`supervisor/state.py`), so an unidentified transport can never bind commands and the local web owner can never lock out a real remote owner. `wait_for_response` remains restricted to A2A-allocated chats; named messages wait for their exact operation, and only the legacy unnamed path uses chat-wide response subscriptions. +The Host Service is a loopback, authenticated callback boundary for reviewed skills (`ouroboros/gateway/host_service.py`, `127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}`). Every request authenticates an opaque `x-skill-token` bound to the skill's content hash, executable review, enablement and grants — secrets never enter the token, a payload edit stales it, and the client wrapper refuses stringification (`skill_token.py`). The frozen route family is exactly: `/identity`, `/tools/schemas`, `/chat/allocate-internal`, `/chat/inject`, `/chat/operations/{operation_ref}`, `/chat/cancel`, `/chat/decision`, `/presence/turn`, `/presence/work/{work_ref}`, `/presence/delivery`, `/ui/ws-message`, and WS `/events`; permissions still decide which route works per skill. Review of a transport skill evaluates identity binding, attribution, polling bounds, panic cleanup, token confinement and exfiltration — an owner-bound reviewed transport may be a first-class control surface, not a screen-only integration. External slash commands bind a separate positive-identity external owner slot (`supervisor/state.py`), so an unidentified transport can never bind commands and the local web owner can never lock out a real remote owner. `wait_for_response` remains restricted to A2A-allocated chats; only the legacy unnamed path uses chat-wide response subscriptions. Admission is one limiter with two policies (`host_service._RateLimiter`): the WS relay lane (`/ui/ws-message`) is a token bucket — a 60-message burst reserve per skill refilling one message per second, so a burst does not silence a widget for the rest of a minute — while every other lane keeps its 60-per-60-seconds sliding window. A refused relay is visible at the host and aggregated per burst: one warning when the reserve empties, then one warning plus one durable `host_service_ws_relay_dropped` row in `logs/events.jsonl` with the dropped count when the lane admits again or the idle bucket is swept; the 429 carries `retry_after_sec` and `dropped_in_burst`, the child's `send_ws_message` stays best-effort `None`, and the in-process broadcast path has no bound. Successful extension children also return a bounded `ws_relay_failures` count map through their result envelope, reported as one host warning without message bodies, URLs or credentials; a child that dies before its final envelope may lose that aggregate, and the measured death/timeout facts stay authoritative. -Operation correlation: a named injected message has `operation_ref=:` on 202, 200, 504 and disconnect responses. `supervisor.message_bus.accept_local_message` serializes check, canonical inbound-row acceptance and enqueue, and the `log_chat` write must succeed before work is queued. Repeated same-id, same-text, same-skill delivery rejoins even before supervisor dequeue; changed content or source is refused with 409. The queue stays in-memory: a crash after acceptance can lose delivery and reads honestly as `lost`, never authorizing a second enqueue. Routing annotations and outbound task ids are discovery hints only — task reads and cancellation require the actual queue/task record's complete `origin_message_ref` to match the authenticated skill's canonical source (`DirectActivityRegistry` carries the same origin), and named response waits poll that exact operation and its retry-aware effective task result, because chat ordering alone never proves a reply. Cancel enters the durable intent and cascade-custody owner (§5) only for work with that origin and the same installation root; a different or unavailable owner root and unaddressable or foreign work are `cancel_unsupported` before any intent is written, and unresolved custody never becomes a false `cancelled`. +Operation correlation: a named injected message has `operation_ref=:` on 202, 200, 504 and disconnect responses. `supervisor.message_bus.accept_local_message` serializes check, canonical inbound-row acceptance and enqueue, and the `log_chat` write must succeed before work is queued. Repeated same-id, same-text, same-skill delivery rejoins even before supervisor dequeue; changed content or source is refused with 409. The queue stays in-memory: a crash after acceptance can lose delivery and reads as `lost`, never authorizing a second enqueue. Routing annotations and outbound task ids are discovery hints only — task reads and cancellation require the queue/task record's complete `origin_message_ref` to match the authenticated skill's canonical source (`DirectActivityRegistry` carries the same origin), and named response waits poll that exact operation and its retry-aware effective task result, because chat ordering alone never proves a reply. Cancel enters the durable intent and cascade-custody owner (§5) only for work with that origin and the same installation root; a different or unavailable owner root and unaddressable or foreign work are `cancel_unsupported` before any intent is written, and unresolved custody never becomes a false `cancelled`. -Presence (`presence_runner.py`): `POST /presence/turn` admits exact provider/account/conversation/thread facts, skill-confined files and a reviewed binding. Auth/grant discovery, pre-turn admission, staging and replay reads share a two-permit async admission off the ASGI loop, leaving executor capacity for owner state and controls during slow probes. Conversation and active-turn gates hold cross-process leases: queued turns wait as coroutines, executing turns own ContextVars-preserving threads, and an HTTP disconnect does not cancel work or return capacity early. The Host joins a live retry by turn ID only when the canonical conversation, actor and text match (reporting version and staged-file paths are not event identity); a mismatch is typed `presence_event_identity_conflict` (409, rejected), never another room's answer. A legacy result whose source cannot be reconstructed refuses replay. +Presence (`presence_runner.py`): `POST /presence/turn` admits exact provider/account/conversation/thread facts, skill-confined files and a reviewed binding. Auth/grant discovery, pre-turn admission, staging and replay reads share a two-permit async admission off the ASGI loop, so slow probes leave executor capacity for owner state and controls. Conversation and active-turn gates hold cross-process leases: queued turns wait as coroutines, executing turns own ContextVars-preserving threads, and an HTTP disconnect does not cancel work or return capacity early. The Host joins a live retry by turn ID only when the canonical conversation, actor and text match (reporting version and staged-file paths are not event identity); a mismatch is typed `presence_event_identity_conflict` (409, rejected), never another room's answer. A legacy result whose source cannot be reconstructed refuses replay. -Before `handle_task` can call any model or tool, Presence writes and reads back a source-bound RUNNING row in the task-result store. A failed/ambiguous start write is `presence_start_unwritable` (409, retry); a persisted RUNNING/INTERRUPTED row or reconciled placeholder with no terminal is `presence_attempt_outcome_unknown` (409, retry), not permission to regenerate. A cancelled turn stays blocked. A later same-source terminal can outrank an old quarantined row; an unreadable or identity-mismatched row cannot. Provider monetary abandonment does not prove no external effect. Owner-only Main recovery notices serialize same-task replay readers through the Presence gate's process/file locks without a turn slot, and need a chat write on a fresh JSONL boundary before the notified stamp; a crash between the two can repeat a notice. When primary or fallback routes refuse quota, even if a later fallback errors or overflows, `resource_refusal_no_resend` is a retryable `presence_resources_unavailable` (409) without terminal speech: the event stays with its transport, the gate is released and Main gets one deduplicated notice. An admitted child keeps its `work_ref` on this refusal for `/presence/work` polling; that handoff precludes the no-effect retry proof. A failed row is reusable only with terminal `presence_retry_proof`: first-round temporary refusal, *all* route/rotation operations engine-confirmed `not_started` and ledger-released, no response/tool/handoff, and a future reset. After reset the conversation gate writes `presence_retry_next`, a locked link to a new physical task ID derived from the predecessor's, retaining the old result; Host joins by original event ID, the live/orphan gate and `turn_ref` use the physical ID. A crash after linking resumes the unused successor; another refusal requires its own proof and reset, not transport-cadence retries. Old/unproven/undated rows stay empty 409 and owner-only; the input is not logged twice. No arbitrary external API exactly-once claim. The §6 bounded same-round transport-death repeat is a no-effect repeat by construction (tools and correspondent sends follow a landed response), not the forbidden event retry; the 409 and owner notice apply once its repeats are spent. The durable row is the one authority the Host reads back after `handle_task` (`_terminal_refusal`): a terminal under the unknown-outcome fence (the pipeline stamps `metadata.presence_unknown_outcome` from the loop's no-resend predicate; the forced rail words it `provider_unavailable` and may salvage a draft as `model_final`) and a RUNNING row, or one host-reconciled from RUNNING, without its own terminal are both `presence_attempt_outcome_unknown`, never completed/silent; other infrastructure terminals keep the deferred projection (host text never speaks, an admitted child stays pollable); the envelope supplies only the admitted `work_ref`. +Before `handle_task` calls any model or tool, Presence writes and reads back a source-bound RUNNING row in the task-result store. A failed/ambiguous start write is `presence_start_unwritable`; a persisted RUNNING/INTERRUPTED row or reconciled placeholder with no terminal is `presence_attempt_outcome_unknown` (both 409, retry), not permission to regenerate. A cancelled turn stays blocked. A later same-source terminal can outrank an old quarantined row; an unreadable or identity-mismatched row cannot. Provider monetary abandonment does not prove no external effect. Owner-only Main recovery notices serialize same-task replay readers through the Presence gate's process/file locks without a turn slot and need a chat write on a fresh JSONL boundary before the notified stamp; a crash between the two can repeat a notice. When primary or fallback routes refuse quota, even if a later fallback errors or overflows, `resource_refusal_no_resend` is a retryable `presence_resources_unavailable` (409) without terminal speech: the event stays with its transport, the gate is released and Main gets one deduplicated notice. An admitted child keeps its `work_ref` on this refusal for `/presence/work` polling; that handoff precludes the no-effect retry proof. A failed row is reusable only with terminal `presence_retry_proof`: first-round temporary refusal, *all* route/rotation operations engine-confirmed `not_started` and ledger-released, no response/tool/handoff, and a future reset. After reset the conversation gate writes `presence_retry_next`, a locked link to a new physical task ID derived from the predecessor's, retaining the old result; Host joins by original event ID, the live/orphan gate and `turn_ref` use the physical ID. A crash after linking resumes the unused successor; another refusal requires its own proof and reset, not transport-cadence retries. Old/unproven/undated rows stay empty 409 and owner-only; the input is not logged twice. No arbitrary external API exactly-once claim. The §6 same-round transport-death repeat is no-effect by construction (no Host-executed tool or correspondent action can stem from the failed attempt; priced, read-only, non-mutating provider-owned retrieval may rerun), not an event retry; the 409 and owner notice apply when the round ends unresolved without a further permitted repeat. The durable row is the one authority the Host reads back after `handle_task` (`_terminal_refusal`): a terminal under the unknown-outcome fence (the pipeline stamps `metadata.presence_unknown_outcome` from the loop's no-resend predicate; the forced rail words it `provider_unavailable` and may salvage a draft as `model_final`) and a RUNNING row, or one host-reconciled from RUNNING, without its own terminal are both `presence_attempt_outcome_unknown`, never completed/silent; other infrastructure terminals keep the deferred projection (host text never speaks, an admitted child stays pollable); the envelope supplies only the admitted `work_ref`. -All entry/receipt paths derive `presence_bindings.conversation_key` (empty thread = `0`). The Presence-local live set protects in-process work from orphan reaping but is not owner-addressable; update drains do not wait for it. The rebuildable last-turn pointer stores completed speech and real deferred `work_ref`; `/presence/work/{work_ref}` polls only correlated work, not the general task API. Turn, receipt and inject budgets are independent. `message`, `silent`, `tool_delivered`, `deferred` are distinct outcomes; deferred requires admitted work, orphaned children replay silent. Promotion preserves ceiling, cost and return context, refuses unusable folders and cannot widen Project/workspace/source. `presence_cancel_work` requires the current binding. Agents share one autobiography. +All entry/receipt paths derive `presence_bindings.conversation_key` (empty thread = `0`). The Presence-local live set protects in-process work from orphan reaping but is not owner-addressable; update drains never wait for it. The rebuildable last-turn pointer stores completed speech and real deferred `work_ref`; `/presence/work/{work_ref}` polls only correlated work, not the general task API. Turn, receipt and inject budgets are independent. `message`, `silent`, `tool_delivered`, `deferred` are distinct outcomes; deferred requires admitted work, orphaned children replay silent. Promotion preserves ceiling, cost and return context, refuses unusable folders and cannot widen Project/workspace/source. `presence_cancel_work` requires the current binding. Agents share one autobiography. A binding's own work is an independent root carrying its nonempty binding id, from any of its conversations (`presence_related_work`); other bindings, owner roots, inline turns, children and rows without that provenance are never attributed. The context lists its first page (canonical root, read gaps stated); new ceilings add `recent_tasks`/`get_task_result` host-bound to `presence_scope=own_binding` plus `steer_task`, and steer, `presence_cancel_work` and a selected cancel or forward reach only that work (cancel/forward also the caller's own tree); a selected global reader stays global. A delegated descendant inherits the ceiling plus only `metadata.presence_binding_authority`, never the speaker's `metadata.presence` (no forced reply): its readers, steer and forward reach that work and its own tree, root included; its promote and that root's follow-ups carry the same carrier (`presence_root_carrier`): related work that never speaks. Steer, like read/cancel, trusts a canonical row over a live one; an unstamped steer is fenced by the sender's live row. Owner chat and consciousness at Act or above may `initiate_presence` on an enabled binding. -The ceiling includes knowledge, scratchpad, identity and chat history, not correspondent tool or owner-command authority. Unselected baseline `chat_history` is unbound; older ceilings keep their digest without gaining baseline tools. `presence_context.py` supplies route facts and frames incoming speech as observation, distinct from initiation or inherited work. Source text stays separate from host attachment declarations; optional provider identity/mention/root facts do not determine the addressee. The model chooses useful, concise participation or silence; early speech follows that choice through selected send tools, never Working forwarding. `presence_finish` enters common completion checks after the tool/control/budget tail; an omitted message/deferred body keeps an answer round. Revised/failed work discards stale completion text, and the next round is told so with the task's confirmed sends. Ordinary authored best-effort replies remain speech. A forced final is the internal record: only a valid `presence_finish` declared inside that one call speaks (a declared partial even on failure, never a `tool_delivered` note); anything else says nothing new and records `presence_declaration`. Host-authored text stays owner-side; an ordinary final's own wording is not screened. A tool-send finish note stays a separate `finish_note` in the previous-turn pointer even when owed work changes its outcome to deferred; it is never prior speech. Host-only terminals retain scheduled child custody as deferred with an empty body. Live/cache/work readers share authorship and empty-body rules; unknown-origin legacy text speaks only for completed rows. Synthesis preserves the recorded adapter body, outcome/origin/work reference and internal result, not today's replay policy; preparation is no delivery receipt. +The ceiling includes knowledge, scratchpad, identity and chat history, not correspondent tool or owner-command authority. Unselected baseline `chat_history` is unbound; older ceilings keep their digest without gaining baseline tools. `presence_context.py` supplies route facts and frames incoming speech as observation, distinct from initiation or inherited work. Source text stays separate from host attachment declarations; optional provider identity/mention/root facts do not determine the addressee. The model chooses useful, concise participation or silence; early speech follows that choice through selected send tools, never Working forwarding. `presence_finish` enters common completion checks after the tool/control/budget tail; an omitted message/deferred body keeps an answer round. Revised/failed work discards stale completion text; the next round is told so with the task's confirmed sends. Ordinary authored best-effort replies remain speech. A forced final is the internal record: only a valid `presence_finish` declared inside that one call speaks (a declared partial even on failure, never a `tool_delivered` note); anything else says nothing new and records `presence_declaration`. Host-authored text stays owner-side; an ordinary final's own wording is not screened. A tool-send finish note stays a separate `finish_note` in the previous-turn pointer even when owed work changes its outcome to deferred; it is never prior speech. Host-only terminals retain scheduled child custody as deferred with an empty body. Live/cache/work readers share authorship and empty-body rules; unknown-origin legacy text speaks only for completed rows. Synthesis preserves the recorded adapter body, outcome/origin/work reference and internal result, not today's replay policy; preparation is no delivery receipt. -Receipt reporting is negotiated through `/identity` (`presence_delivery_version: 1`) and optional `delivery_reporting_version: 1` on a turn; mode survives cached/deferred results. Mode1 writes outgoing history on provider receipts; mode0 retains an authored, delivery-unconfirmed row. `presence_delivery.py` accepts authenticated `POST /presence/delivery` observations through the chat writer. Parts retain target, text/format and provider facts; SMTP `accepted` means provider acceptance only. Speech is typed `presence_delivery`, never a task-finalizing untyped row; failed/uncertain attempts are System facts, queued remains in tool/outbox receipts. Memory, history and consolidation retain destination/state/details. A Host-context projection rebuilds once from retained chat generations and updates after required writes; the receipt writer starts each record on a JSONL boundary even after a torn tail. Readable identical reports deduplicate and readable conflicting reports refuse. A malformed retained row makes the rebuilt projection `history_coverage=gapped`: later reports may be recorded, but `duplicate=false` means only that no matching *readable* receipt was indexed, not that an earlier hidden effect never happened. The projection names `indexed` otherwise (the last rebuild plus its own writes, not an attestation that every concurrent chat writer has been re-scanned); a physically unreadable archive still refuses instead of pretending to be an empty history. Transport outboxes own provider receipts and separate report ACK/backoff: a slow or failed report neither resends nor blocks provider delivery. No second store or scheduler; old queues are not imported. Wire fields and procedure: CREATING_SKILLS, “Reporting actual Presence delivery”. +Receipt reporting is negotiated through `/identity` (`presence_delivery_version: 1`) and optional `delivery_reporting_version: 1` on a turn; mode survives cached/deferred results. Mode1 writes outgoing history on provider receipts; mode0 retains an authored, delivery-unconfirmed row. `presence_delivery.py` accepts authenticated `POST /presence/delivery` observations through the chat writer. Parts retain target, text/format and provider facts; SMTP `accepted` means provider acceptance only. Speech is typed `presence_delivery`, never a task-finalizing untyped row; failed/uncertain attempts are System facts, queued remains in tool/outbox receipts. Memory, history and consolidation retain destination/state/details. A Host-context projection rebuilds once from retained chat generations and updates after required writes; the receipt writer starts each record on a JSONL boundary even after a torn tail. Readable identical reports deduplicate and readable conflicting reports refuse. A malformed retained row makes the rebuilt projection `history_coverage=gapped`: later reports may be recorded, but `duplicate=false` means only that no matching *readable* receipt was indexed, not that an earlier hidden effect never happened. The projection names `indexed` otherwise (the last rebuild plus its own writes, not an attestation that every concurrent chat writer has been re-scanned); a physically unreadable archive still refuses rather than pose as an empty history. Transport outboxes own provider receipts and separate report ACK/backoff: a slow or failed report neither resends nor blocks provider delivery. No second store or scheduler; old queues are not imported. Wire fields and procedure: CREATING_SKILLS, “Reporting actual Presence delivery”. Companion processes are host-supervised: reviewed manifest-declared descriptors enter durable custody, reconcile after lifecycle changes and restart, and stop on disable/unload/panic. `state/extension_generation.json` carries the opposite direction — the server's published live set, which a running task worker adopts at a task's start or at a dispatch miss, so an enable after boot is not invisible until the pool respawns. Worker-side changes write durable reconcile requests (`state/extension_reconcile/`) rather than spawning server-owned children, and every reconcile state names the process that answered and whether that marker request was written. Health observations are process-qualified: aggregate and Skills UI health use the server observation as authority and expose the worker observation only as a qualifier, so a failed handoff cannot advance `last_known_good`; restart-budget exhaustion persists a terminal reason, cleared only by a later successful start. A companion's cwd is the reviewed payload directory, so a payload edit stales review before reload instead of silently mutating a live process. The live projection is `state/extension_companions.json`. diff --git a/tests/test_presence_refusals.py b/tests/test_presence_refusals.py index b7c32f677..f8f1695b5 100644 --- a/tests/test_presence_refusals.py +++ b/tests/test_presence_refusals.py @@ -668,11 +668,16 @@ def test_inline_turn_repeats_a_dispatched_transport_death_as_a_no_effect_attempt An inline Presence turn is a direct-chat task, so its PRIMARY round keeps the paid repeat rail (``loop_llm_call._TRANSPORT_DEATH_RETRIES``): a DISPATCHED request whose socket died with a typed transport death is sent once more in the same round as a NEW physical attempt - with its own ledger row. That repeat has no prior effect BY CONSTRUCTION: the host executes - tools and the transport speaks only after a model response has landed, and none landed. So - the owner rule (the same event is retried only with positive proof of no prior effect and a - new physical identity) is met, and the unknown-outcome refusal (empty 409 plus owner notice) - applies only once the round's repeats are exhausted. Real Host endpoint, runner, loop, round + with its own ledger row. That repeat is no-effect BY CONSTRUCTION in this narrow sense: no + Host-executed tool and no correspondent action can stem from the failed attempt, because the + host executes tools and the transport speaks only after a model response has landed, and + none landed. Provider-owned retrieval (e.g. the server-side web search a provider runs + before the response lands) may execute again on the repeat: read-only, non-mutating, priced. + So the owner rule (the same event is retried only with positive proof of no prior effect and + a new physical identity) is met, and the unknown-outcome refusal (empty 409 plus owner + notice) applies when the round ends unresolved without a further permitted repeat: the + deadline or a typed finalize control can refuse even the first repeat, and a repeat failing + with any other class ends the round at once. Real Host endpoint, runner, loop, round dispatcher, terminal pipeline and attempt ledger: the socket is scripted, the backoff sleep is recorded instead of slept, and the fallback chain is a tripwire. """ diff --git a/tests/test_provider_terminal_notice.py b/tests/test_provider_terminal_notice.py index 58156ed94..277f81922 100644 --- a/tests/test_provider_terminal_notice.py +++ b/tests/test_provider_terminal_notice.py @@ -137,7 +137,9 @@ def test_actual_presence_and_cached_read_keep_authored_speech_and_owner_notice_s assert body["error"] == "presence_attempt_outcome_unknown: source_event_id" assert (body["turn_ref"], body["work_ref"]) == (task_id, "next-task") assert all(RAW not in str(value) and notice not in str(value) for value in body.values()) - assert notices == [task_id] * 2 # the owner alone hears the notice, once per refusal + # each refusal consults the owner-notice writer (mocked here; production dedups on + # presence_recovery_owner_notified, so the owner hears it once) + assert notices == [task_id] * 2 view = presence_result_from_stored(stored, task_id) assert (view.outcome, view.text, view.work_ref) == ("silent", "", "next-task")