Commit graph

8 commits

Author SHA1 Message Date
Ouroboros
6784af2e3c delegation: truthful supervising-wait facts across runs, sleeps and wakes
The child-delivery cursor is task-scoped: _load_state now carries the
committed coordination_cursor across a run-id change, and the recovery
handoff snapshots it and the restore re-installs it, so an acknowledged
child terminal or beacon is never re-announced while an unacknowledged
one (whose cursor never committed) still re-emits exactly once.

Each supervised_wait call opens a sleep_entry snapshot (labelled after a
worker-loss adoption or an interrupted call, never carrying time over).
The single wake-publication point stamps whole-call sleep facts on every
wake, terminal included, before pending_wake is stored, so a replay is
exact; quiet-status wakes drop the last tick's waited_sec/quiet_for_sec
and generic keep-watching/cancel note but keep the PAUSED note (now the
shared delegate_progress.paused_note). cache_horizon_note is computed
once per wake from the last recorded model response (_last_llm_call_meta
ts, a lower bound on cache age); the two dead per-tick call sites in
_delegate_wait are removed.

Every wake also carries leaf_live_input, the route row's declared
liveInput read once at entry through the loop transport (bounded by one
beat; unknown when unread, none when the engine has no such field), and
the supervision state records dated observation facts
(last_answered_observation_at, observation_failure) instead of relying
on the file's write time. The spill envelope keeps sleep,
cache_horizon_note and leaf_live_input.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-26 08:10:00 +03:00
Ouroboros
e33fd69115 Make the external-executor verbs answer with a native typed result
`delegate_shared._fail` rendered `{"status":"refused", ...}` as a plain string,
and the registry's legacy text adapter classified it as OK: it only understands
a top-level `ok:false` or a first-line `⚠️ IDENTIFIER` marker. So a refused
`delegate_wait`/`delegate_cancel` — daemon unreachable, run not owned, a
containment fault, a cancel the daemon refused — was recorded as a SUCCESSFUL
tool call on the outcome axis, in the acceptance packet and in the supervising
task's own reasoning.

Owner decision Q8A: the fix is a native structured result INSIDE the family,
not a repo-wide ABI migration. `_fail` now returns a `ToolResult` whose text is
the same JSON the callers emitted, plus two ADDITIVE envelope keys — `ok:false`
and `host_code` — written beside the domain payload. The domain `reason` is
never renamed into `ToolResult.code`, and the domain extras
(`definitely_unrun`, `pending_invocation_id`, `run_id`, `reset_at`, the custody
facts) keep their places. One exact, closed table maps a reason to its class:
substrate refusals (the daemon, the engine, custody or the run said no) are
`TOOL_REPORTED_FAILURE`, recorded and never degrading; malformed or
self-contradictory calls are `TOOL_ARG_ERROR`, which degrades and feeds
reflection. Neither is a timeout or the generic tool error, and an unclassified
reason defaults to the substrate class — the safe direction.

The second literal refusal author, `supervised_wait`'s checkpoint argument
check, folds into `_fail`. `_delegate_cancel`'s own outcomes join it: `failed`
and `containment_fault_run_may_still_be_live` report a run that may still be
live and mutating, so they publish as failures, while `confirmed` and
`requested` stay successful observations of the control surface.

Publication happens once, at the four REGISTERED entries, after every
decoration and immediately before the string is returned: `subagent_runtime`
mutates the start payload after `_delegate_start`, and an earlier publish fails
the registry's equality gate silently. `exact_start` now reads the native
payload, adds the actor identity and the work-order source, and returns
`_replace_tool_result`; its `json.loads → TypeError → return result` bypass,
which dropped that decoration without saying so, is deleted.
`_mark_actor_physical_start` still reads the DOMAIN `started` /
`started_uncustodied` status, never the host class.

The consumers migrate in the same change as the type they consume: the
configured-session bootstrap, the recovery handoff, the unknown-provider hold
and the pending-wake replay all read the producer's own payload rather than a
stringified result — without this, every leaf wake would have failed its
acknowledgement and taken the no-resend terminal. The wake envelope carries
`ok`/`host_code` through both fitted-spill shapes, so a refusal too large to
inline cannot read as a successful wait. `wait_once` deliberately keeps its
`str` tick contract, and `integrate_delegated_patch` keeps its own string ABI
at the one helper the two families share.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-15 18:01:50 +03:00
Ouroboros
27e91b44dd P2-fix1f: Close a daemon-outage episode only on real contact
The recovery gate accepted any tick that was not the outage itself, so a
refusal raised BEFORE any byte left the host closed the episode: a missing
descriptor, an unreadable token or an engine below the version floor all
arrive as status refused, and the owner was told the daemon was reachable
again at the moment it became less reachable. With the constant key of the
previous shape that false recovery also burned the recovery toast for the
rest of the session.

The two facts that already shared a shape now share a name: a tick that
returned no daemon answer, whether an observation hole or a pre-flight
refusal, drops the loop's cached transport AND cannot close an outage. Only a
payload the daemon itself produced is contact. One frozenset, one predicate,
no new state.

Tests: an outage followed by refused/daemon_not_discovered leaves the episode
open and emits no recovery line, while an answered read still closes it.
2026-09-12 17:08:03 +03:00
Ouroboros
2193b7f02a P2-fix1e: Separate a slow daemon from an unreachable one in the observation hole
S1(b) relayed the gateway's typed code as the quiet observation reason, but
_request raises EVERY httpx error with the single code daemon_unreachable, and
tools/delegate is the only producer of observation_pending. The supervision
loop's unreachable discriminator was therefore true for every hole, so a plain
read timeout against a daemon that is alive and answering raised the
owner-facing outage incident. A finder probe drove the real stack against a
real loopback daemon that answered the first GET after the bound and the
second immediately: the run settled from that daemon while the owner was told
it was unreachable and then reachable again. The base emitted nothing.

This is S1's own letter, classification by transport class, applied where the
observation is produced. The gateway maps the httpx type to the observation
reason: ReadTimeout is our own bound expiring and says nothing about the
daemon (observation_read_timeout, quiet), while a socket that could not be
opened or broke mid-exchange carried no answer (daemon_unreachable, the half
worth an owner line). Both stay quiet renewals on the same beat.

The reason rides BESIDE code, not instead of it. Overloading code would
silently change three unrelated readers that key on daemon_unreachable: the
model-control retry loop in llm_claudexor (a read timeout would stop being a
control outage and become an unknown-outcome error), the daemon liveness
probe, which lists httpx.ReadTimeout as a transport-unreachable cause, and the
definite-unrun refusal set. One of those files belongs to another phase.
observation_reason is set only when the failure is already an observation
retry, so a received 4xx carries none.

Tests: the parametrized gateway case now pins the reason per httpx class
including ReadTimeout, a received 401 carries no observation reason at all,
the observing wait relays both reasons, and the supervision case that stays
silent is now driven by the value production really mints.
Docs: ARCHITECTURE and DEVELOPMENT name the two halves of the class instead of
claiming one reason for all of them.
2026-09-12 17:08:03 +03:00
Ouroboros
5ed3ba6762 P2-fix1d: Stamp the daemon-outage owner line per episode, not per task
S1(c) promised ONE owner line per unreachable-daemon EPISODE plus one
recovery line, with the client's toast_once providing the dedup. _owner_line
built a constant key and a constant text per task, so episodes 2..N were
dropped by both surfaces: chat_media keeps a module-level set of shown
toast_once keys, and the chat timeline dedupes a live card on
phase|headline|body. In the very wave this step exists to close, 15:05-16:25,
that is loss, not the noise the risk note predicted.

The open episode now carries its own identity, following the
loop_transport.incident precedent that stamps an episode's millisecond entry
for exactly this pair: the boolean latch becomes the episode's UTC start plus
its millisecond stamp, the stamp discriminates both toast keys, and each line
names the episode it belongs to, so the text-keyed timeline dedup sees a
second episode too. The outage and its recovery share one stamp, so a reader
can tell which outage a recovery closes. No durable state is written and the
beat is untouched.

Residual: two episodes inside the same second render the same text and can
still collapse on the timeline; their toast keys never collide, so the fact
reaches the owner either way.

Tests: two outage and recovery episodes in one wait now produce four distinct
toast keys (two before), each pair sharing its own stamp, and the beat stays
one tick per unreachable observation.
2026-09-12 17:08:03 +03:00
Anton
5164e8e0e8 P2-S2: Hold one handshaken gateway per supervision loop
Every supervision tick built a fresh ClaudexorGateway and re-handshaked
(tools/delegate.py), two HTTP requests plus TCP setup per nanny per 3 s,
against a daemon the previous step just classified as unreachable.

_delegate_wait gains keyword-only gateway=None: a supplied transport
replaces the per-call ClaudexorGateway(), is handshaken only when it has
not been yet (the gateway's engine_version is the handshake receipt, set
solely by a successful handshake), and is never closed by the borrower.
supervised_wait owns one transport for the loop when it built the
observing wait itself: constructed lazily without a handshake (so a
refusal still classifies through the wait's own typed path), lent to
every tick, dropped and closed after any tick that returned no
observation (observation_pending or refused) so the next tick re-reads
the descriptor and re-handshakes, and closed in try/finally around the
loop. A caller-supplied wait_once keeps today's four-positional call
shape and receives no keyword: the borrowed transport is meaningful only
to the real observing wait, and 21 test fakes across five modules stay
untouched. Same contract and transport; only who holds the object moves,
per call to per loop.

Tests: the pinned public sequence becomes one handshake per loop plus one
GET per tick; a fake gateway proves N quiet ticks = one handshake and a
failed tick drops and closes the cached gateway before a fresh one is
handshaken.
2026-09-12 17:07:10 +03:00
Anton
1b431085ab P2-S1: Classify a dead daemon socket as a quiet observation renewal
Widen the observation-only retry predicate in gateways/claudexor._request
from httpx.ReadTimeout alone to the read-only-retryable transport classes
(ConnectError, ConnectTimeout, PoolTimeout, ReadError, WriteError,
RemoteProtocolError, ReadTimeout), classified by exception type plus the
received status; a received 4xx/5xx still wins and _problem is untouched.
The observing delegate_wait relays the gateway's own typed code as the
quiet reason (exc.code, so daemon_unreachable) instead of the hardcoded
observation_read_timeout literal, and the supervision loop treats such a
tick as a quiet renewal on the unchanged 3 s beat: no backoff, no durable
counter, no daemon_outage latch (dispositions R4/R20/R29; the model can do
nothing with a dead socket, and every refused tick was a wasted round, 359
of them on 2026-09-10).

The owner hears about the outage exactly once per episode and once on
recovery, through the existing ctx.emit_progress_fn with the typed
task_incident/toast_once pair the client already dedupes (A11: the
incident= keyword is already accepted by agent._emit_progress; this is a
new call site in delegate_supervision, not a new channel). The
delegate_wait schema description now discloses the class and the existing
~60 s reaction delay of finalize_now/hurry while an observation read
hangs, a documentation debt of the current contract, not new behaviour.

Tests: retryable classes are typed observation holes while 401/403/503
stay refusals; the relayed reason is the typed code; N unreachable ticks
give one owner line and the first answered read one recovery line; another
typed reason gives nothing; deadline and cancellation still cut a long
unobserved stretch; the unknown-provider hold rides out a dead socket
instead of taking the refused exit.
2026-09-12 17:07:10 +03:00
Anton
13c5dcd223 Preserve transport work and reconcile exact review outcomes
Recover complete reviewer results from their original operation CAS, preserve terminal custody separately from model narrative, and consume remote streams within physical-attempt accounting. Apply caller deadlines before every recovery send, publish known tool refusals accurately, and retain raw reviewer PASS separately from completion eligibility.

Managed work waits through unknown provider outcomes and continues with a marked new attempt after route-bound upstream observation, retaining old monetary custody. Metadata HEAD observations reuse the connection bound; ordinary no-deadline clients retain their existing timeout defaults.

Preserves the complete interrupted implementation and source-only version policy. Focused consumer, golden, lifecycle, domain, inventory, lint and size checks pass. The initial full preflight failed on stale consumers and a missing manifest compatibility field; the final full preflight is still required on this committed candidate.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-11 03:45:18 +03:00