P2-fix1e: Separate a slow daemon from an unreachable one in the observation hole

S1(b) relayed the gateway's typed code as the quiet observation reason, but
_request raises EVERY httpx error with the single code daemon_unreachable, and
tools/delegate is the only producer of observation_pending. The supervision
loop's unreachable discriminator was therefore true for every hole, so a plain
read timeout against a daemon that is alive and answering raised the
owner-facing outage incident. A finder probe drove the real stack against a
real loopback daemon that answered the first GET after the bound and the
second immediately: the run settled from that daemon while the owner was told
it was unreachable and then reachable again. The base emitted nothing.

This is S1's own letter, classification by transport class, applied where the
observation is produced. The gateway maps the httpx type to the observation
reason: ReadTimeout is our own bound expiring and says nothing about the
daemon (observation_read_timeout, quiet), while a socket that could not be
opened or broke mid-exchange carried no answer (daemon_unreachable, the half
worth an owner line). Both stay quiet renewals on the same beat.

The reason rides BESIDE code, not instead of it. Overloading code would
silently change three unrelated readers that key on daemon_unreachable: the
model-control retry loop in llm_claudexor (a read timeout would stop being a
control outage and become an unknown-outcome error), the daemon liveness
probe, which lists httpx.ReadTimeout as a transport-unreachable cause, and the
definite-unrun refusal set. One of those files belongs to another phase.
observation_reason is set only when the failure is already an observation
retry, so a received 4xx carries none.

Tests: the parametrized gateway case now pins the reason per httpx class
including ReadTimeout, a received 401 carries no observation reason at all,
the observing wait relays both reasons, and the supervision case that stays
silent is now driven by the value production really mints.
Docs: ARCHITECTURE and DEVELOPMENT name the two halves of the class instead of
claiming one reason for all of them.
This commit is contained in:
Ouroboros 2026-09-12 10:37:06 +03:00
parent 5ed3ba6762
commit 2193b7f02a
5 changed files with 91 additions and 19 deletions

View file

@ -1543,7 +1543,7 @@ An undisclosed spend contributes `0.0` to `accounted_usd` — inventing a conser
**Work orders.** The compiler sends the entire chosen assignment, preserving context and instruction roles without an arbitrary host-size cutoff. Direct starts carry the normalized host contract once in instructions and the chosen assignment separately in prompt; coordination context remains complete. Real native/HTTP limits return their actual failure with the original input and any pending invocation retained. Exact-source readers remain optional capabilities. Legacy partial starts still verify their original renderer digest, selector and source intervals, and recover by replaying the recorded request body; removing new partial-start production never certifies an incomplete old run or launches a duplicate after ambiguous dispatch.
**Supervision.** `delegate_wait` is model-visible as an event-only sleep, not a caller-sized poll: `delegate_supervision.supervised_wait` renews bounded transport windows in host code at zero LLM calls, and only a meaningful event (settlement, interaction, fault, addressed message, child signal, control, recovery judgment) becomes a coalesced durable wake, replayed across worker interruption until acknowledged; deadline, ceiling, budget, and cancellation remain outer bounds. Every receipt and wake carries one host-rendered `coordination_context` (intent note, deadline remaining, root-tree spend, active descendants, remaining paid acceptance capacity) — facts for LLM judgment, never thresholds. A requested future inspection (`checkpoint_after_sec`) wakes once and is consumed by any earlier real event — no cadence, stall classifier, or hidden polling. An observation the transport could not complete at all (a dead socket, not a received failure) is a quiet renewal carrying its typed reason (`daemon_unreachable`), never a refusal that spends a model round on something the model cannot act on; the beat stays three seconds with no backoff, durable counter or outage latch, and deadline, ceiling, budget and cancellation still cut a long unobserved stretch. The outage reaches the owner exactly once per episode, with one recovery line, on the supervising task's own progress surface. The loop holds ONE handshaken gateway for its whole wait and re-establishes it after any tick that returned no observation. On a wake the nanny holds its full tool surface and the parent's captured model/effort. Nanny economics are structurally quiet: the burn baseline resets only on real acts of delegation, and coordination verbs never buy metered silence (`nanny_pacing.py`).
**Supervision.** `delegate_wait` is model-visible as an event-only sleep, not a caller-sized poll: `delegate_supervision.supervised_wait` renews bounded transport windows in host code at zero LLM calls, and only a meaningful event (settlement, interaction, fault, addressed message, child signal, control, recovery judgment) becomes a coalesced durable wake, replayed across worker interruption until acknowledged; deadline, ceiling, budget, and cancellation remain outer bounds. Every receipt and wake carries one host-rendered `coordination_context` (intent note, deadline remaining, root-tree spend, active descendants, remaining paid acceptance capacity) — facts for LLM judgment, never thresholds. A requested future inspection (`checkpoint_after_sec`) wakes once and is consumed by any earlier real event — no cadence, stall classifier, or hidden polling. An observation the transport could not complete at all (a socket failure, not a received one) is a quiet renewal carrying its typed reason, never a refusal that spends a model round on something the model cannot act on; the reason separates our own read bound expiring against a live daemon (`observation_read_timeout`, quiet and nothing more) from a socket that carried no answer (`daemon_unreachable`, the only half worth an owner line); the beat stays three seconds with no backoff, durable counter or outage latch, and deadline, ceiling, budget and cancellation still cut a long unobserved stretch. The outage reaches the owner exactly once per episode, with one recovery line, on the supervising task's own progress surface. The loop holds ONE handshaken gateway for its whole wait and re-establishes it after any tick that returned no observation. On a wake the nanny holds its full tool surface and the parent's captured model/effort. Nanny economics are structurally quiet: the burn baseline resets only on real acts of delegation, and coordination verbs never buy metered silence (`nanny_pacing.py`).
**A run's question is the nanny's to answer.** Supervision wakes immediately with typed `status="waiting_on_user"` on a NEW `pendingInteractions` entry instead of burning the engine's answer timeout in dead polling; answer keys are echoed verbatim into `delegate_answer` (custody-gated like cancel, relaying the engine's typed outcomes including `subscription_window_exhausted`), and delivered interaction ids are acknowledged only after transcript injection, so a question neither re-triggers a round nor disappears across recovery. A question above the nanny's authority escalates to the nearest live ancestor. The codex lane has no mid-run channel: a terminal with `outcome_facts.reason=input_required` is answered by a plain new `delegate_start(subagent_id=..., prompt=...)` — never the engine's rerun verb, which would start a run outside this task's custody trail.

View file

@ -2105,7 +2105,7 @@ owner, owed terminal delivery, cascade postconditions — lives in ARCHITECTURE
- Stream consumption completes inside physical accounting. Preserve indexed tools, native signatures, complete final framing and cumulative usage snapshots. An EOF/error/cancellation retains private wire evidence and cannot produce a usable partial answer. Only a structural parameter rejection uses the existing wire recovery; never infer a retry from missing stream text or ping cadence. Compatible async tool calls now use the same normalizer/validation path; local, GigaChat and Claudexor retain their separate wire contracts.
- Late reviewer reuse resolves the exact operation's complete producer receipt from existing CAS, with original task/root/attempt, slot/route, subject, contract, roster/epoch and delegated invocation where present. The current surface remains the sole wave writer and reducer. No source file existence, preview or matching prompt prose alone grants authority; missing/partial/error/mismatched custody never buys another same-operation dispatch.
- Managed unknown-outcome recovery uses the existing network-wait owner, with non-generating upstream observations and an explicit new-attempt notice after connectivity returns. Keep old outcome/cost unknown and apply current budget/Stop/deadline before dispatch. Subscription catalogs prove reachability only with generic `provenance="provider_http"` plus `observedAt` after wait entry and exact source/model/effective account; legacy/static catalogs remain unknown. A control-channel outage first rejoins the same accepted operation. Non-generating HEAD uses the existing connection allowance for every socket phase, narrowed by the owner remainder, rather than inheriting a cognitive read window without its lease. No scheduler, provider/model table, paid readiness probe or automatic manual-restart recovery is introduced.
- `delegate_wait` supervision's three-second observation beat is separate from its HTTP read allowance. A typed read-only-retryable transport failure (read timeout, connect error or timeout, pool timeout, read/write error, protocol error) is a quiet observation hole carrying its typed reason and the actual elapsed time; the beat does not slow, no durable counter or outage latch is kept, and the outage is disclosed to the owner once per episode with one recovery line. Received auth/protocol failures and owner controls remain meaningful. After terminal cleanup, use the current custody host notice alongside the original answer/narrative. Genuine builtin refusals publish typed non-success at their producer; successful warnings and existing review/Git warning buckets keep their semantics. Acceptance JSON validity and completion cleanliness remain separate decisions.
- `delegate_wait` supervision's three-second observation beat is separate from its HTTP read allowance. A typed read-only-retryable transport failure (read timeout, connect error or timeout, pool timeout, read/write error, protocol error) is a quiet observation hole carrying its typed reason and the actual elapsed time; the beat does not slow and no durable counter or outage latch is kept. The reason is per class, because our own read bound expiring against a live daemon is not the same fact as a socket that carried no answer: only the second is disclosed to the owner, once per episode with one recovery line, each stamped with that episode. Received auth/protocol failures and owner controls remain meaningful. After terminal cleanup, use the current custody host notice alongside the original answer/narrative. Genuine builtin refusals publish typed non-success at their producer; successful warnings and existing review/Git warning buckets keep their semantics. Acceptance JSON validity and completion cleanliness remain separate decisions.
Focused regressions: `test_review_late_cas_recovery.py`, `test_delivery_control_lineage.py`, `test_terminal_custody_notice.py`, `test_delegate_observation_transport.py`, `test_delegate_hold.py`, `test_configured_session_wake_rail.py`, `test_health_invariants_ownership.py`, `test_transport_b_stream_deadlines.py`, `test_transport_unknown_continuation.py`, `test_builtin_refusal_results.py` and `test_v671_acceptance_convergence.py`. Use the ordinary isolated preflight runner; full provider/renderer smoke remains separate from local fake-provider evidence.

View file

@ -68,6 +68,21 @@ _OBSERVATION_RETRYABLE_ERRORS = (
httpx.ReadTimeout, httpx.ConnectError, httpx.ConnectTimeout, httpx.PoolTimeout,
httpx.ReadError, httpx.WriteError, httpx.RemoteProtocolError,
)
# ...and WHICH hole it was, for the observer that has to tell the owner apart a
# daemon that is merely slow from one that is not there. Our own read bound
# expiring says nothing about the daemon; a socket that could not be opened or
# that broke mid-exchange says it did not answer. Classified by exception TYPE,
# never by prose, and carried beside ``code`` rather than replacing it: the
# model-control retry loop, the daemon liveness probe and the definite-unrun set
# all key on ``daemon_unreachable`` for reasons that have nothing to do with an
# observation.
_OBSERVATION_READ_TIMEOUT = "observation_read_timeout"
def _observation_reason(exc: BaseException) -> str:
"""The typed reason for a read-only observation hole."""
return (_OBSERVATION_READ_TIMEOUT if isinstance(exc, httpx.ReadTimeout)
else "daemon_unreachable")
_ATTEMPTS_REL = "attempts"
_ATTEMPT_RECORD = "attempt.yaml"
@ -86,7 +101,8 @@ class ClaudexorUnavailable(RuntimeError):
"""
def __init__(self, code: str, message: str, *, status_code: int = 0,
required_actions: tuple[str, ...] = (), observation_timeout: bool = False) -> None:
required_actions: tuple[str, ...] = (), observation_timeout: bool = False,
observation_reason: str = "") -> None:
super().__init__(message)
self.code = str(code or "claudexor_unavailable")
self.status_code = int(status_code or 0)
@ -94,6 +110,11 @@ class ClaudexorUnavailable(RuntimeError):
# Read-only observers may retry this exact HTTP read without claiming
# anything about the worker. A received HTTP refusal still wins.
self.observation_timeout = bool(observation_timeout)
# Which hole it was, for the observer only (``_observation_reason``):
# a read bound that expired against a live daemon is not the same fact
# as a socket that never carried an answer, and only the second is
# worth an owner-facing outage line.
self.observation_reason = str(observation_reason or "")
# Cross-repo contract (B1): the engine's window-exhausted RunFailure codes. A
@ -341,12 +362,14 @@ class ClaudexorGateway:
headers=headers or None, **bound) as response:
response.read()
except httpx.HTTPError as exc:
retryable = (isinstance(exc, _OBSERVATION_RETRYABLE_ERRORS)
and (response is None or response.status_code < 400))
raise ClaudexorUnavailable(
"daemon_unreachable",
f"Claudexor daemon unreachable: {type(exc).__name__}: {exc}",
status_code=response.status_code if response is not None else 0,
observation_timeout=isinstance(exc, _OBSERVATION_RETRYABLE_ERRORS)
and (response is None or response.status_code < 400),
observation_timeout=retryable,
observation_reason=_observation_reason(exc) if retryable else "",
) from exc
if response.status_code >= 400:
raise self._problem(response)

View file

@ -857,9 +857,13 @@ def _delegate_wait(ctx: ToolContext, run_id: str, wait_sec: Optional[int] = None
def read_failure(exc: ClaudexorUnavailable) -> str:
if observation_only and exc.observation_timeout:
# The gateway's per-class reason, not the generic transport code: a
# read bound that expired against a live daemon is a quiet hole and
# nothing more, while a socket that carried no answer is the outage
# the owner is told about once per episode.
return json.dumps({
"status": "observation_pending", "run_id": rid,
"reason": exc.code, "detail": str(exc),
"reason": exc.observation_reason or exc.code, "detail": str(exc),
"waited_sec": time.monotonic() - started,
})
return _fail("delegate_wait", exc.code, str(exc), run_id=rid)

View file

@ -131,17 +131,26 @@ def test_existing_control_returns_without_starting_another_observation(tmp_path,
assert calls == []
@pytest.mark.parametrize("failure", [
httpx.ConnectError("connection refused"), httpx.ConnectTimeout("connect timed out"),
httpx.PoolTimeout("pool exhausted"), httpx.ReadError("connection reset"),
httpx.WriteError("broken pipe"), httpx.RemoteProtocolError("server disconnected"),
@pytest.mark.parametrize("failure,reason", [
(httpx.ReadTimeout("read timed out"), "observation_read_timeout"),
(httpx.ConnectError("connection refused"), "daemon_unreachable"),
(httpx.ConnectTimeout("connect timed out"), "daemon_unreachable"),
(httpx.PoolTimeout("pool exhausted"), "daemon_unreachable"),
(httpx.ReadError("connection reset"), "daemon_unreachable"),
(httpx.WriteError("broken pipe"), "daemon_unreachable"),
(httpx.RemoteProtocolError("server disconnected"), "daemon_unreachable"),
])
def test_read_only_retryable_transport_failures_are_typed_observation_holes(failure):
def test_read_only_retryable_transport_failures_are_typed_observation_holes(failure, reason):
"""A socket that delivered no daemon answer is the same unresolved read as a read
timeout: typed ``daemon_unreachable`` with ``observation_timeout`` set, so the
supervising wait renews quietly instead of waking the model on every 3 s beat
(I1: 359 refusals on 2026-09-10, two of them ReadError). Classified by the
exception TYPE, never by prose; a received status still wins (test above)."""
timeout: ``observation_timeout`` set, so the supervising wait renews quietly
instead of waking the model on every 3 s beat (I1: 359 refusals on 2026-09-10,
two of them ReadError). Classified by the exception TYPE, never by prose; a
received status still wins (test above).
The OBSERVATION reason separates the two halves of that class: our own read
bound expiring says nothing about the daemon, while a socket that could not be
opened or that broke mid-exchange did not carry an answer. The transport
``code`` stays ``daemon_unreachable`` for every other reader of it."""
def _raise(_request):
raise failure
@ -154,17 +163,36 @@ def test_read_only_retryable_transport_failures_are_typed_observation_holes(fail
gateway.get_run("run-existing")
assert caught.value.code == "daemon_unreachable"
assert caught.value.observation_timeout is True
assert caught.value.observation_reason == reason
assert caught.value.status_code == 0
assert caught.value.__cause__ is failure
finally:
gateway.close()
def test_a_received_refusal_carries_no_observation_reason():
"""A 4xx the daemon actually sent is not an observation hole at all."""
def _refuse(_request):
return httpx.Response(401, json={"code": "http_401", "message": "unauthorized"})
gateway = gateway_module.ClaudexorGateway(gateway_module.DaemonEndpoint("127.0.0.1", 1, "fixture"))
gateway._client.close()
gateway._client = httpx.Client(base_url="http://127.0.0.1:1", transport=httpx.MockTransport(_refuse))
try:
with pytest.raises(gateway_module.ClaudexorUnavailable) as caught:
gateway.get_run("run-existing")
assert caught.value.observation_timeout is False
assert caught.value.observation_reason == ""
finally:
gateway.close()
def test_observation_read_failure_carries_the_gateway_typed_code(tmp_path, monkeypatch):
"""The observing wait relays the transport's own typed code as the quiet reason
"""The observing wait relays the transport's own per-class observation reason
(no hardcoded ``observation_read_timeout``), so the supervision loop can tell an
unreachable daemon apart from any other typed reason; a received refusal keeps
its refusal shape."""
unreachable daemon apart from a daemon that was merely slow; a received refusal
keeps its refusal shape."""
ctx = _delegating_ctx(tmp_path, acting=False)
entry = delegate._RunCustody(task_id=ctx.task_id, route_id="fixture", model="fixture",
project_id="fixture", project_owned=False, access="readonly")
@ -180,10 +208,19 @@ def test_observation_read_failure_carries_the_gateway_typed_code(tmp_path, monke
monkeypatch.setattr(gateway_module, "ClaudexorGateway", lambda: _Dead())
refusals.append(gateway_module.ClaudexorUnavailable(
"daemon_unreachable", "ConnectError: [Errno 61]", observation_timeout=True))
"daemon_unreachable", "ConnectError: [Errno 61]", observation_timeout=True,
observation_reason="daemon_unreachable"))
quiet = json.loads(delegate._delegate_wait(ctx, "run-dead", observation_only=True))
assert quiet["status"] == "observation_pending" and quiet["run_id"] == "run-dead"
assert quiet["reason"] == "daemon_unreachable"
# A slow but LIVE daemon is the same quiet hole with a different typed reason,
# so the supervision loop does not raise the outage line for it.
refusals.append(gateway_module.ClaudexorUnavailable(
"daemon_unreachable", "ReadTimeout", observation_timeout=True,
observation_reason="observation_read_timeout"))
slow = json.loads(delegate._delegate_wait(ctx, "run-dead", observation_only=True))
assert slow["status"] == "observation_pending"
assert slow["reason"] == "observation_read_timeout"
refusals.append(gateway_module.ClaudexorUnavailable("http_401", "unauthorized", status_code=401))
refused = json.loads(delegate._delegate_wait(ctx, "run-dead", observation_only=True))
assert refused["status"] == "refused" and refused["reason"] == "http_401"
@ -244,6 +281,14 @@ def test_unreachable_daemon_episode_is_one_owner_line_each_way(tmp_path, monkeyp
def test_other_typed_observation_reasons_say_nothing_to_the_owner(tmp_path, monkeypatch):
"""A daemon that was merely SLOW stays silent to the owner.
``observation_read_timeout`` is what the gateway mints for an httpx
ReadTimeout and what the observing wait relays (both pinned above), so this
is the production value of a live daemon that answered after our own read
bound, not a hand-fed string: the run settles from that same daemon and the
owner is never told it was unreachable.
"""
ctx = _delegating_ctx(tmp_path, acting=False)
notes = []
ctx.emit_progress_fn = lambda text, *, incident=None: notes.append((text, incident))