mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
ouroboros: checkpoint after task 410dfca20a30497b — Presence ТЗ3 — завершить сохранённый кандидат
This commit is contained in:
parent
ce9e304afe
commit
b299367135
23 changed files with 3169 additions and 420 deletions
|
|
@ -157,7 +157,7 @@ scanned data-relative path to be covered by a row here (count-anchored both ways
|
|||
| `state/skills/<name>/dependency_cache/` | `marketplace/isolated_deps.py` via existing installer subprocesses and verified resource downloads | downloaded resources keyed by sha256; package-manager cache formats | reused across payload/environment replacement; follows existing skill-state cleanup | resources are verified/downloaded again and package caches rebuilt |
|
||||
| `state/skills/<name>/go/`, `state/skills/<name>/go-cache/` | Go compiler launched by `tools/skill_exec.py:_run_go_skill`, with GOPATH/GOCACHE bound to skill_state_dir | Go-owned cache formats | reused across script runs; follows existing skill-state cleanup | compiler recreates caches; skill payload and review remain unchanged |
|
||||
| `task_results/<id>.json` | `ouroboros/task_results.py` (locked merge) | `_schema_version: 1` (ABI-2); unstamped/future/malformed → quarantine, no conversion — with one carve-out: the boot latch migration re-stamps unstamped rows still at `cancel_requested` in place (same status, one typed `task_result_cancel_latch_admitted` event) so a wedged task still reaches its `cancelled` terminal | UNBOUNDED — one file per task forever, no GC; lifecycle authority is retained deliberately | lifecycle authority lost; drive prunes degrade to age-only; strict authority reads break |
|
||||
| `task_results/quarantine/` | `ouroboros/task_result_schema.py` (same-dir rename) | quarantined bytes unchanged | NEVER GC'd (pinned); recovery is manual owner re-stamp | quarantined evidence lost |
|
||||
| `task_results/quarantine/` | `ouroboros/task_result_schema.py` (same-dir rename); `presence_runner.py` reads the retained id before a transport retry | quarantined bytes unchanged; a Presence retry with quarantined authority refuses regeneration | NEVER GC'd (pinned); recovery is manual owner re-stamp | quarantined evidence lost; a previous Presence attempt could be mistaken for a new event |
|
||||
| `task_results/artifacts/<id>/**` (+`verification_receipts.jsonl`, `.directory.*.tmp`, `.directory.*.json.tmp`), `task_results/artifact_versions/` | `ouroboros/artifacts.py`, `headless.py`, `outcome_receipt_store.py` | artifact and complete directory manifests `schema_version: 1`; scratch manifest 2 | artifact versions bounded (5 per name); artifacts live with their result; directory capture stages ZIP and manifest beside the result, removing owned temporaries on caught failures | deliverable bytes lost; results keep dangling manifests |
|
||||
| `task_drives/<id>/**` (+`tmp_scripts/`) | `ouroboros/headless.py`, `tools/tool_context.py`, `tools/shell.py` | child stamps as above | GC-retention prune at startup (terminal + age, default 7 d); the `data/tmp_scripts` fallback's hard-kill orphans are in `sweep_stale_temp_files` scope (startup-only when no script can be live) | scratch lost; canonical artifacts survive |
|
||||
| `task_trees/<root>/blackboard.jsonl` | `ouroboros/task_tree_ledger.py` | rows unversioned; snapshot digest `schema_version: 1` | GC-retention prune at startup (root terminal + age) | swarm coordination facts lost for live trees |
|
||||
|
|
|
|||
|
|
@ -248,7 +248,7 @@ Host-authored sampling defaults stay separate from explicit parameters: review r
|
|||
|
||||
#### Quota and auth waits
|
||||
|
||||
`model_wait.py` holds confirmed quota/auth waits inside the live call. Typed owner actions use the existing mailbox and the `/api/decisions` family `model_wait:<task>:<wait>`; task-result rows are projections, not restartable stack checkpoints or a second attempt ledger, and metadata polling makes no generation. On Auto the host may send the last successful same-route account as a preference, but omits it for the next request after a status-null or typed per-subject refusal on that account in the same execution; the engine remains the account chooser and, without refusal evidence, may select the same top-headroom account again — rotation is possible, not guaranteed. Pin never rotates. A route may also answer with ANOTHER model than the one requested, with no error and nothing in the account's quota: the engine states that as a typed fact on its completed result, and the transport never accepts such an answer as the round. It is retained and acknowledged like any other answer, so the horizon it would have run under stays reconstructible, but it adopts no turn state, runs no tool call, and is replaced by a NEW operation that names no account — the engine, which deprioritises that account for a while, alone chooses where every redo lands. `OUROBOROS_SERVED_MODEL_REDOS` bounds those redos; a pinned account names one account, an already admitted candidate is bound to exact bytes, and an actor with no send left on a bounded budget would otherwise read that rail's own budget error instead of this refusal, and a spent owner window buys no further send, so none of those is asked again. An outcome nobody knows is never a substitution: the transport raises it before the rail sees a result, so the existing unknown-custody fence, not this budget, governs what happens next. Each discarded generation's retention and acknowledgement state travels with its durable row, so a failed one is visible rather than assumed. A discarded generation teaches the density witness nothing, because a witness belongs to the model that produced it. Once no redo is available the typed `model_substituted` refusal reaches the caller, which is an ordinary cross-model fallback trigger for a root and a typed failure for an exact-route child; it never cools the requested model, because one substituting account says nothing about the model. The first substitution of each model in a task also becomes one timeline row of that task's card. The rule covers this transport only: an Agent-session delegation is a separate engine capability with its own model evidence. The fallback budget is unchanged (defaults: one attempt per model and a 120-second cross-model cooldown), and even a configured API fallback waits for an explicit owner switch after a quota refusal. A temporary switch changes only the waiting role; optional persistence uses the ordinary settings writer, and pending acceptance, settings saved and worker application are separate revisioned facts. Fallback Local remains a group-wide setting: a one-role persistence request may change model/account only when Local is unchanged, a different Local choice is task-only, and permanent group changes remain in Models. The current task-attempt identity selects actionable wait rows; a failed mailbox delivery retries the same accepted request without authoring a new one; terminal mailbox cleanup waits for post-task synthesis and pending accepted attachment promotion to settle.
|
||||
`model_wait.py` holds confirmed quota/auth waits inside the live call. Typed owner actions use the existing mailbox and the `/api/decisions` family `model_wait:<task>:<wait>`; task-result rows are projections, not restartable stack checkpoints or a second attempt ledger, and metadata polling makes no generation. Auto may prefer the last successful same-route account unless it got a status-null or typed per-subject refusal this execution. An HTTP vendor quota refusal of a named account is re-asked with the same payload, no preference; a pre-dispatch verdict never is; only its `poolCause` proves the pool spent; an account chosen twice stops rotation, claiming nothing. Pin never rotates. A route may also answer with ANOTHER model than the one requested, with no error and nothing in the account's quota: the engine states that as a typed fact on its completed result, and the transport never accepts such an answer as the round. It is retained and acknowledged like any other answer, so the horizon it would have run under stays reconstructible, but it adopts no turn state, runs no tool call, and is replaced by a NEW operation that names no account — the engine, which deprioritises that account for a while, alone chooses where every redo lands. `OUROBOROS_SERVED_MODEL_REDOS` bounds those redos; a pinned account names one account, an already admitted candidate is bound to exact bytes, and an actor with no send left on a bounded budget would otherwise read that rail's own budget error instead of this refusal, and a spent owner window buys no further send, so none of those is asked again. An outcome nobody knows is never a substitution: the transport raises it before the rail sees a result, so the existing unknown-custody fence, not this budget, governs what happens next. Each discarded generation's retention and acknowledgement state travels with its durable row, so a failed one is visible rather than assumed. A discarded generation teaches the density witness nothing, because a witness belongs to the model that produced it. Once no redo is available the typed `model_substituted` refusal reaches the caller, which is an ordinary cross-model fallback trigger for a root and a typed failure for an exact-route child; it never cools the requested model, because one substituting account says nothing about the model. The first substitution of each model in a task also becomes one timeline row of that task's card. The rule covers this transport only: an Agent-session delegation is a separate engine capability with its own model evidence. Fallback keeps one attempt per model and a 120-second cooldown; on quota each configured route runs before the owner wait, opened from the retained refusal (no generation first); nothing sleeps to a reset; inline Presence never waits (typed `resource_refusal`, no forced final). A temporary switch changes only the waiting role; optional persistence uses the ordinary settings writer, and pending acceptance, settings saved and worker application are separate revisioned facts. Fallback Local remains a group-wide setting: a one-role persistence request may change model/account only when Local is unchanged, a different Local choice is task-only, and permanent group changes remain in Models. The current task-attempt identity selects actionable wait rows; a failed mailbox delivery retries the same accepted request without authoring a new one; terminal mailbox cleanup waits for post-task synthesis and pending accepted attachment promotion to settle.
|
||||
|
||||
`ouroboros/gateway/task_model_wait.py` owns model-wait decision effects and the bounded settings-writer receipts; `task_decision.answer_decision` returns `(status, payload)` to both the Web and authenticated Host Service transports (they supply the installation root — no synthetic HTTP request or parallel owner registry); `supervisor/task_model_wait.py` projects worker events and quota clocks through the existing supervisor, and `ouroboros/gateway/decision_contracts.py` owns the decision-family vocabulary. Quota pauses use union duration across simultaneous waits and do not consume internal execution time, while explicit calendar deadlines remain fixed; the worker and project-write ownership stay held (a one-worker queue waits), completed tools and reviewer results remain on the same live stack, closing a browser does not stop the wait, and stopping the Ouroboros process ends this continuation guarantee. Image tools retain the tracked VLM child: `vision_process.py` carries typed errors, physical capture, cancellation and operation checkpoints across that process boundary, and a lost child result remains unknown, never free.
|
||||
|
||||
|
|
|
|||
|
|
@ -12,7 +12,7 @@ Admission is one limiter with two policies (`host_service._RateLimiter`): the WS
|
|||
|
||||
Operation correlation: a named injected message has `operation_ref=<chat_id>:<client_message_id>` on 202, 200, 504 and disconnect responses. `supervisor.message_bus.accept_local_message` serializes check, canonical inbound-row acceptance and enqueue, and the `log_chat` write must succeed before work is queued. Repeated same-id, same-text, same-skill delivery rejoins even before supervisor dequeue; changed content or source is refused with 409. The queue stays in-memory: a crash after acceptance can lose delivery and reads honestly as `lost`, never authorizing a second enqueue. Routing annotations and outbound task ids are discovery hints only — task reads and cancellation require the actual queue/task record's complete `origin_message_ref` to match the authenticated skill's canonical source (`DirectActivityRegistry` carries the same origin), and named response waits poll that exact operation and its retry-aware effective task result, because chat ordering alone never proves a reply. Cancel enters the durable intent and cascade-custody owner (§5) only for work with that origin and the same installation root; a different or unavailable owner root and unaddressable or foreign work are `cancel_unsupported` before any intent is written, and unresolved custody never becomes a false `cancelled`.
|
||||
|
||||
Presence (`presence_runner.py`): `POST /presence/turn` admits an exact provider/account/conversation/thread event and skill-state-confined files under the hash-bound `presence` permission and an owner binding (`state/presence_bindings.json`). Agents share one autobiography. Event-derived IDs deduplicate retries; cross-process locks cap concurrency and serialize conversations. Inbound, initiated and receipt paths derive one conversation key and chat id (`presence_bindings.conversation_key`, empty thread = `0`). A host-reconciled placeholder, or a still-running row of a dead attempt, is not a result: the same event id re-runs told what that attempt confirmed sending (unknown once its rows left the live chat generation), and the row shows `superseded_placeholder`. A presence-local liveness set shields an executing turn from orphan reconciliation without making it an owner-addressable direct actor, so update drains do not wait for it and after a restart the transport's retry re-runs it. Each conversation's last executed turn is a rebuildable pointer (`state/presence_turn_gate/last-<sha256>.json`; canonical sources: task results and receipts; a replay of a settled, authored turn that lost it rebuilds it) shown to the next turn. Turns, receipts and inject have separate per-skill in-flight budgets. Outcomes are message/silent/tool_delivered/deferred; deferred needs a correlated `work_ref`, read via `GET /presence/work/{work_ref}`, not the general task API; an orphan-reconciled child replays silent (nobody re-runs it). Promotion discards requested Project/workspace/source widening and copies the ceiling, cost and return context; unusable folders return `workspace_unusable` with repair detail. `presence_cancel_work` requires the current binding and conversation. Owner chat and consciousness at Act or above may `initiate_presence` on an enabled binding.
|
||||
Presence (`presence_runner.py`): `POST /presence/turn` admits provider/account/conversation/thread facts and skill-confined files under the hash-bound permission and owner binding (`state/presence_bindings.json`). Cross-process locks cap concurrency and serialize conversations. Host auth/admission/file checks run off-loop; gate waits are coroutines, then `PresenceTurnExecutions` starts a ContextVars-preserving daemon thread. Same-event retries join it; HTTP cancellation retains work, result and capacity until settlement. No model or gate wait occupies the default executor. Presence errors retain HTTP/error fields and add semantic `code`/`disposition`: auth/configuration `blocked`, invalid events/conflicting receipts `rejected`, unavailable I/O/capacity `retry`; old adapters still follow their HTTP policy. All entry/receipt paths derive keys via `presence_bindings.conversation_key` (empty thread = `0`). A dead RUNNING/INTERRUPTED attempt or reconciled placeholder is NOT safe to regenerate: an unknown provider/tool effect leaves the original event with the transport and returns `presence_attempt_outcome_unknown` (409, retry, `turn_ref`) without starting an agent; an owner-only Main question is recorded once when its message bridge is available. Administrative monetary abandonment is not no-effect proof. A cancelled turn similarly refuses regeneration rather than acknowledging the correspondent. Strict result reads and retained quarantine identity make unreadable prior authority a typed retry refusal, not a fresh event. A genuinely pre-thread failure wrote no running row and can start on retry; a canonical terminal result replays without a new generation. Exact-operation reconciliation and explicit owner resolution remain necessary when the prior effect is permanently unknown; the host does not promise eventual answering or exactly-once from an unreadable source. The local live set prevents orphan reaping but is not owner-addressable; update drains do not wait for it. The next turn reads the last-turn pointer (`state/presence_turn_gate/last-<sha256>.json`), rebuildable from task/receipt facts on replay. Turn, receipt and inject budgets are separate. Outcomes: message/silent/tool_delivered/deferred; deferred requires real `work_ref` through `/presence/work/{work_ref}`, never general task API. Orphaned children replay silent, not restarted. Promotion preserves ceiling/cost/return context, ignores requested Project/workspace/source widening, and refuses unusable folders as `workspace_unusable` with repair detail. `presence_cancel_work` requires current binding/conversation. Owner chat and consciousness Act+ may `initiate_presence` on an enabled binding. Agents share one autobiography.
|
||||
|
||||
The ceiling includes knowledge, scratchpad, identity and chat history, not correspondent tool or owner-command authority. Unselected baseline `chat_history` is unbound; older ceilings keep their digest without gaining baseline tools. `presence_context.py` supplies route facts and frames incoming speech as observation, distinct from initiation or inherited work. Source text stays separate from host attachment declarations; optional provider identity/mention/root facts do not determine the addressee. The model chooses useful, concise participation or silence; early speech follows that choice through selected send tools, never Working forwarding. `presence_finish` enters common completion checks after the tool/control/budget tail; an omitted message/deferred body keeps an answer round. Revised/failed work discards stale completion text. Authored best-effort replies remain speech; host diagnostics stay owner-side. Host-only terminals retain scheduled child custody as deferred with an empty body. Live/cache/work readers share authorship and empty-body rules; unknown-origin legacy text speaks only for completed rows. Synthesis preserves the recorded adapter body, outcome/origin/work reference and internal result, not today's replay policy; preparation is no delivery receipt.
|
||||
|
||||
|
|
|
|||
|
|
@ -3,6 +3,7 @@
|
|||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import functools
|
||||
import hmac
|
||||
import json
|
||||
import logging
|
||||
|
|
@ -204,6 +205,9 @@ class HostServiceContext:
|
|||
self._inflight_lock = threading.Lock()
|
||||
self._counter_lock = threading.Lock()
|
||||
self.presence_deliveries = PresenceDeliveryRecorder(self.data_dir)
|
||||
from ouroboros.presence_runner import PresenceTurnExecutions
|
||||
|
||||
self.presence_turns = PresenceTurnExecutions()
|
||||
|
||||
def _ws_relay_burst_ended(self, key: str, dropped: int, duration_sec: float) -> None:
|
||||
"""Report one aggregated WS relay refusal burst: a warning plus one
|
||||
|
|
@ -260,13 +264,16 @@ class HostServiceContext:
|
|||
**kwargs,
|
||||
)
|
||||
|
||||
async def admit_presence_gate(self, conversation_key: str) -> Any:
|
||||
"""Coroutine admission into the configured presence gate; the lease the turn runs under."""
|
||||
from ouroboros.presence_runner import admit_configured_gate
|
||||
|
||||
return await admit_configured_gate(self.data_dir, conversation_key)
|
||||
|
||||
@property
|
||||
def skills_state_dir(self) -> pathlib.Path:
|
||||
return self.data_dir / "state" / "skills"
|
||||
|
||||
def authenticate_token(self, raw_token: str) -> str:
|
||||
return self.authenticate_token_payload(raw_token)[0]
|
||||
|
||||
def authenticate_token_payload(self, raw_token: str) -> tuple[str, Dict[str, Any]]:
|
||||
token = str(raw_token or "").strip()
|
||||
if not token:
|
||||
|
|
@ -345,6 +352,26 @@ class HostServiceContext:
|
|||
return chat_id
|
||||
|
||||
|
||||
async def _authenticated(
|
||||
ctx: HostServiceContext, raw_token: str, permission: str = "",
|
||||
) -> tuple[str, Dict[str, Any]]:
|
||||
"""Resolve the request's skill (and one grant) off the event loop.
|
||||
|
||||
Token discovery reads every registered skill's token, review, enablement and grants
|
||||
from disk; inline, one slow walk stalled every request on the loop the owner's API
|
||||
shares. It is short read-only work, so the default executor serves it: Presence turns
|
||||
no longer wait there (``presence_runner.PresenceTurnExecutions``).
|
||||
"""
|
||||
|
||||
def resolve() -> tuple[str, Dict[str, Any]]:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(raw_token)
|
||||
if permission:
|
||||
ctx.require_permission(skill_name, token_payload, permission)
|
||||
return skill_name, token_payload
|
||||
|
||||
return await asyncio.to_thread(resolve)
|
||||
|
||||
|
||||
def _token_from_websocket(websocket: WebSocket) -> str:
|
||||
header = websocket.headers.get("x-skill-token", "")
|
||||
if header:
|
||||
|
|
@ -357,12 +384,7 @@ def _token_from_websocket(websocket: WebSocket) -> str:
|
|||
return ""
|
||||
|
||||
|
||||
async def _api_identity(request: Request) -> JSONResponse:
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
ctx.authenticate_token(request.headers.get("x-skill-token", ""))
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
def _identity_facts(ctx: HostServiceContext) -> tuple[str, str]:
|
||||
identity_path = ctx.data_dir / "memory" / "identity.md"
|
||||
name = "Ouroboros"
|
||||
description = ""
|
||||
|
|
@ -378,6 +400,16 @@ async def _api_identity(request: Request) -> JSONResponse:
|
|||
break
|
||||
except Exception:
|
||||
log.debug("Failed to read identity for host service", exc_info=True)
|
||||
return name, description
|
||||
|
||||
|
||||
async def _api_identity(request: Request) -> JSONResponse:
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
await _authenticated(ctx, request.headers.get("x-skill-token", ""))
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
name, description = await asyncio.to_thread(_identity_facts, ctx)
|
||||
return JSONResponse({"ok": True, "name": name, "description": description,
|
||||
"presence_delivery_version": DELIVERY_VERSION})
|
||||
|
||||
|
|
@ -385,28 +417,26 @@ async def _api_identity(request: Request) -> JSONResponse:
|
|||
async def _api_tool_schemas(request: Request) -> JSONResponse:
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name = ctx.authenticate_token(request.headers.get("x-skill-token", ""))
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""))
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
if not ctx.rate_limiter.allow(f"{skill_name}:tools"):
|
||||
return _json_error("rate limit exceeded", 429)
|
||||
schemas = ctx.tool_schemas_getter()
|
||||
schemas = await asyncio.to_thread(ctx.tool_schemas_getter)
|
||||
return JSONResponse({"ok": True, "tools": schemas})
|
||||
|
||||
|
||||
async def _api_allocate_internal(request: Request) -> JSONResponse:
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
try:
|
||||
ctx.require_permission(skill_name, token_payload, "inject_chat")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "inject_chat")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
try:
|
||||
payload = await request.json()
|
||||
chat_id = ctx.allocate_internal_chat_id(skill_name, str(payload.get("range_name") or "a2a"))
|
||||
chat_id = await run_sync_to_completion(
|
||||
ctx.allocate_internal_chat_id, skill_name, str(payload.get("range_name") or "a2a"),
|
||||
)
|
||||
except Exception as exc:
|
||||
return _json_error(str(exc), 400)
|
||||
return JSONResponse({"ok": True, "chat_id": chat_id})
|
||||
|
|
@ -425,11 +455,7 @@ async def _api_chat_inject(request: Request) -> JSONResponse:
|
|||
"""
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
try:
|
||||
ctx.require_permission(skill_name, token_payload, "inject_chat")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "inject_chat")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
if not ctx.rate_limiter.allow(f"{skill_name}:inject"):
|
||||
|
|
@ -646,58 +672,108 @@ async def _api_presence_delivery(request: Request) -> JSONResponse:
|
|||
"""Record exact provider receipts without sending or starting model work."""
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(
|
||||
request.headers.get("x-skill-token", "")
|
||||
)
|
||||
ctx.require_permission(skill_name, token_payload, "presence")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "presence")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
return _presence_error(str(exc), 403, "presence_auth_blocked", "blocked")
|
||||
except OSError:
|
||||
return _presence_error("presence authentication unavailable", 500, "presence_auth_unavailable", "retry")
|
||||
if not ctx.rate_limiter.allow(f"{skill_name}:presence_delivery"):
|
||||
return _json_error("rate limit exceeded", 429)
|
||||
return _presence_error("rate limit exceeded", 429, "presence_rate_limited", "retry")
|
||||
# Receipts, turns and inject have separate in-flight budgets: five long turns must
|
||||
# not starve the receipts those turns' own sends produce.
|
||||
if not ctx._enter_inflight(f"{skill_name}:delivery"):
|
||||
return _json_error("too many in-flight presence requests", 429)
|
||||
return _presence_error("too many in-flight presence requests", 429, "presence_capacity_full", "retry")
|
||||
try:
|
||||
payload = await request.json()
|
||||
result = await run_sync_to_completion(ctx.presence_deliveries.record, skill_name, payload)
|
||||
return JSONResponse(result)
|
||||
except PresenceDeliveryConflict as exc:
|
||||
return _json_error(str(exc), 409)
|
||||
return _presence_error(str(exc), 409, "presence_receipt_conflict", "rejected")
|
||||
except (ValueError, TypeError) as exc:
|
||||
return _json_error(str(exc), 400)
|
||||
return _presence_error(str(exc), 400, "presence_receipt_invalid", "rejected")
|
||||
except Exception:
|
||||
log.warning("Presence delivery history write failed for skill %s", skill_name, exc_info=True)
|
||||
return _json_error("presence delivery history write failed; retry the same receipt", 503)
|
||||
return _presence_error("presence delivery history write failed; retry the same receipt", 503,
|
||||
"presence_receipt_unwritable", "retry")
|
||||
finally:
|
||||
ctx._leave_inflight(f"{skill_name}:delivery")
|
||||
|
||||
|
||||
def _presence_error(message: str, status: int, code: str, disposition: str, **facts: Any) -> JSONResponse:
|
||||
"""Add transport recovery facts without changing the legacy HTTP/error contract.
|
||||
|
||||
The producer chooses semantics, not the status class: a stale token is blocked,
|
||||
an origin mismatch rejected, and a failed history write retryable.
|
||||
"""
|
||||
return JSONResponse({"ok": False, "error": message, "code": code,
|
||||
"disposition": disposition, **facts}, status_code=status)
|
||||
|
||||
|
||||
def _presence_exception(exc: Exception, status: int) -> JSONResponse:
|
||||
code = str(getattr(exc, "code", "presence_admission_failed"))
|
||||
disposition = "blocked"
|
||||
if code in {"chat_log_unwritable", "presence_result_missing", "presence_bindings_unreadable",
|
||||
"presence_resources_unavailable", "presence_attempt_outcome_unknown",
|
||||
"presence_result_unreadable"}:
|
||||
disposition = "retry"
|
||||
elif code in {"presence_attachment_admission_rejected", "presence_binding_wrong_transport",
|
||||
"presence_conversation_key_required", "presence_admission_conversation_mismatch"}:
|
||||
disposition = "rejected"
|
||||
facts = {"field": getattr(exc, "field", "presence")}
|
||||
if getattr(exc, "turn_ref", ""):
|
||||
facts["turn_ref"] = exc.turn_ref
|
||||
manifest = getattr(exc, "attachment_manifest", None)
|
||||
if isinstance(manifest, list):
|
||||
facts["attachment_manifest"] = [dict(row) for row in manifest if isinstance(row, dict)]
|
||||
return _presence_error(str(exc), status, code, disposition, **facts)
|
||||
|
||||
|
||||
def _admit_presence(ctx: HostServiceContext, skill_name: str, binding_id: str) -> Any:
|
||||
from ouroboros.loop import _resolve_loop_max_rounds
|
||||
from ouroboros.presence_admission import admit_presence_turn
|
||||
|
||||
return admit_presence_turn(
|
||||
drive_root=ctx.data_dir,
|
||||
authenticated_transport_skill=skill_name,
|
||||
binding_id=binding_id,
|
||||
global_max_rounds=_resolve_loop_max_rounds(),
|
||||
)
|
||||
|
||||
|
||||
async def _api_presence_turn(request: Request) -> JSONResponse:
|
||||
"""Run one non-owner event under a host-resolved reviewed profile ceiling."""
|
||||
"""Run one non-owner event under a host-resolved reviewed profile ceiling.
|
||||
|
||||
Authentication, admission and file confinement run off the event loop. The turn is
|
||||
host work this request only waits on (``presence_runner.PresenceTurnExecutions``): it
|
||||
queues on the gate as a coroutine, runs on its own thread, keeps its in-flight slot
|
||||
until it settles, and a retry of the same event joins it instead of running it twice.
|
||||
"""
|
||||
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(
|
||||
request.headers.get("x-skill-token", "")
|
||||
)
|
||||
ctx.require_permission(skill_name, token_payload, "presence")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "presence")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
return _presence_error(str(exc), 403, "presence_auth_blocked", "blocked")
|
||||
except OSError:
|
||||
return _presence_error("presence authentication unavailable", 500, "presence_auth_unavailable", "retry")
|
||||
if not ctx.rate_limiter.allow(f"{skill_name}:presence"):
|
||||
return _json_error("rate limit exceeded", 429)
|
||||
from ouroboros.presence_admission import PresenceAdmissionError, admit_presence_turn
|
||||
return _presence_error("rate limit exceeded", 429, "presence_rate_limited", "retry")
|
||||
from ouroboros.presence_admission import PresenceAdmissionError
|
||||
from ouroboros.presence_bindings import conversation_key
|
||||
from ouroboros.presence_runner import PresenceTurnError, PresenceTurnEvent
|
||||
from ouroboros.presence_runner import (
|
||||
PresenceTurnError,
|
||||
PresenceTurnEvent,
|
||||
PresenceTurnNotStarted,
|
||||
presence_turn_replay,
|
||||
presence_turn_task_id,
|
||||
)
|
||||
|
||||
if not ctx._enter_inflight(f"{skill_name}:presence"):
|
||||
return _json_error("too many in-flight presence requests", 429)
|
||||
try:
|
||||
payload = await request.json()
|
||||
if not isinstance(payload, dict) or set(payload) - {
|
||||
"binding_id", "event", "staged_files", "delivery_reporting_version",
|
||||
}:
|
||||
return _json_error("invalid presence payload", 400)
|
||||
return _presence_error("invalid presence payload", 400, "presence_payload_invalid", "rejected")
|
||||
reporting_version = delivery_reporting_version(payload.get("delivery_reporting_version", 0))
|
||||
event_payload = payload.get("event")
|
||||
expected = {
|
||||
|
|
@ -705,15 +781,9 @@ async def _api_presence_turn(request: Request) -> JSONResponse:
|
|||
"conversation_key", "actor", "conversation", "message", "text",
|
||||
}
|
||||
if not isinstance(event_payload, dict) or set(event_payload) != expected:
|
||||
return _json_error("invalid presence event", 400)
|
||||
return _presence_error("invalid presence event", 400, "presence_event_invalid", "rejected")
|
||||
|
||||
from ouroboros.loop import _resolve_loop_max_rounds
|
||||
admission = admit_presence_turn(
|
||||
drive_root=ctx.data_dir,
|
||||
authenticated_transport_skill=skill_name,
|
||||
binding_id=str(payload.get("binding_id") or ""),
|
||||
global_max_rounds=_resolve_loop_max_rounds(),
|
||||
)
|
||||
admission = await asyncio.to_thread(_admit_presence, ctx, skill_name, str(payload.get("binding_id") or ""))
|
||||
provider = str(event_payload.get("provider") or "").strip()
|
||||
account_id = str(event_payload.get("account_id") or "").strip()
|
||||
conversation_id = str(event_payload.get("conversation_id") or "").strip()
|
||||
|
|
@ -727,7 +797,8 @@ async def _api_presence_turn(request: Request) -> JSONResponse:
|
|||
)
|
||||
or (admission.origin.thread_id and thread_id != admission.origin.thread_id)
|
||||
):
|
||||
return _json_error("presence event does not match its owner-created binding", 403)
|
||||
return _presence_error("presence event does not match its owner-created binding", 403,
|
||||
"presence_origin_mismatch", "rejected")
|
||||
event = PresenceTurnEvent(
|
||||
source_event_id=str(event_payload["source_event_id"] or "").strip(),
|
||||
provider=provider,
|
||||
|
|
@ -749,13 +820,26 @@ async def _api_presence_turn(request: Request) -> JSONResponse:
|
|||
delivery_reporting_version=reporting_version,
|
||||
)
|
||||
if not event.source_event_id or not event.conversation_key or not event.actor:
|
||||
return _json_error("presence event is missing identity facts", 400)
|
||||
result = await asyncio.to_thread(
|
||||
ctx.presence_runner,
|
||||
admission=admission,
|
||||
event=event,
|
||||
staged_files=_presence_staged_files(ctx, skill_name, payload.get("staged_files")),
|
||||
)
|
||||
return _presence_error("presence event is missing identity facts", 400,
|
||||
"presence_identity_missing", "rejected")
|
||||
staged_files = await asyncio.to_thread(_presence_staged_files, ctx, skill_name, payload.get("staged_files"))
|
||||
turn_id = presence_turn_task_id(admission.binding_id, event.source_event_id)
|
||||
# A settled turn answers from its durable row without queueing behind its conversation.
|
||||
result = await asyncio.to_thread(presence_turn_replay, ctx.data_dir, turn_id, event.conversation_key)
|
||||
if result is None:
|
||||
budget = f"{skill_name}:presence"
|
||||
execution, _started = ctx.presence_turns.start_or_join(
|
||||
turn_id,
|
||||
reserve=functools.partial(ctx._enter_inflight, budget),
|
||||
release=functools.partial(ctx._leave_inflight, budget),
|
||||
admit=functools.partial(ctx.admit_presence_gate, event.conversation_key),
|
||||
run=lambda lease: ctx.presence_runner(
|
||||
admission=admission, event=event, staged_files=staged_files, admitted=lease,
|
||||
),
|
||||
)
|
||||
if execution is None:
|
||||
return _presence_error("too many in-flight presence requests", 429, "presence_capacity_full", "retry")
|
||||
result = await asyncio.wrap_future(execution.result)
|
||||
return JSONResponse({
|
||||
"ok": True,
|
||||
"status": "completed",
|
||||
|
|
@ -766,22 +850,48 @@ async def _api_presence_turn(request: Request) -> JSONResponse:
|
|||
"delivery_reporting_version": getattr(result, "delivery_reporting_version", 0),
|
||||
})
|
||||
except json.JSONDecodeError:
|
||||
return _json_error("invalid json", 400)
|
||||
return _presence_error("invalid json", 400, "presence_payload_invalid", "rejected")
|
||||
except (PresenceAdmissionError, PresenceTurnError) as exc:
|
||||
payload = {"ok": False, "error": str(exc), "code": exc.code, "field": exc.field}
|
||||
attachment_manifest = getattr(exc, "attachment_manifest", None)
|
||||
if isinstance(attachment_manifest, list):
|
||||
payload["attachment_manifest"] = [
|
||||
dict(row) for row in attachment_manifest if isinstance(row, dict)
|
||||
]
|
||||
return JSONResponse(payload, status_code=409)
|
||||
except (OSError, ValueError) as exc:
|
||||
return _json_error(str(exc), 400)
|
||||
return _presence_exception(exc, 409)
|
||||
except PresenceTurnNotStarted as exc:
|
||||
return _presence_error(str(exc), 503, "presence_turn_not_started", "retry")
|
||||
except OSError as exc:
|
||||
return _presence_error(str(exc), 400, "presence_io_unavailable", "retry")
|
||||
except ValueError as exc:
|
||||
return _presence_error(str(exc), 400, "presence_payload_invalid", "rejected")
|
||||
except Exception as exc:
|
||||
log.debug("Host service presence turn failed", exc_info=True)
|
||||
return _json_error(str(exc), 500)
|
||||
finally:
|
||||
ctx._leave_inflight(f"{skill_name}:presence")
|
||||
return _presence_error(str(exc), 500, "presence_turn_unavailable", "retry")
|
||||
|
||||
|
||||
def _presence_work_view(
|
||||
ctx: HostServiceContext, skill_name: str, work_ref: str, binding_id: str,
|
||||
) -> tuple[int, Dict[str, Any]]:
|
||||
from ouroboros.presence_bindings import load_presence_binding
|
||||
from ouroboros.task_results import load_task_result
|
||||
|
||||
load_presence_binding(ctx.data_dir, skill_name, binding_id)
|
||||
stored = load_task_result(ctx.data_dir, work_ref) or {}
|
||||
metadata = stored.get("metadata") if isinstance(stored.get("metadata"), dict) else {}
|
||||
presence = metadata.get("presence") if isinstance(metadata.get("presence"), dict) else {}
|
||||
if str(presence.get("binding_id") or "") != binding_id:
|
||||
return 404, {"ok": False, "error": "presence work reference not found",
|
||||
"code": "presence_work_not_found", "disposition": "rejected"}
|
||||
status = str(stored.get("status") or "")
|
||||
if status not in {"completed", "failed", "cancelled"}:
|
||||
return 202, {"ok": True, "status": "pending", "work_ref": work_ref,
|
||||
"delivery_reporting_version": presence.get("delivery_reporting_version", 0)}
|
||||
from ouroboros.presence_runner import presence_result_from_stored
|
||||
|
||||
result = presence_result_from_stored(stored, work_ref)
|
||||
return 200, {
|
||||
"ok": True,
|
||||
"status": status,
|
||||
"outcome": result.outcome,
|
||||
"text": result.text,
|
||||
"work_ref": work_ref,
|
||||
"delivery_reporting_version": result.delivery_reporting_version,
|
||||
}
|
||||
|
||||
|
||||
async def _api_presence_work(request: Request) -> JSONResponse:
|
||||
|
|
@ -789,45 +899,22 @@ async def _api_presence_work(request: Request) -> JSONResponse:
|
|||
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(
|
||||
request.headers.get("x-skill-token", "")
|
||||
)
|
||||
ctx.require_permission(skill_name, token_payload, "presence")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "presence")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
return _presence_error(str(exc), 403, "presence_auth_blocked", "blocked")
|
||||
except OSError:
|
||||
return _presence_error("presence authentication unavailable", 500, "presence_auth_unavailable", "retry")
|
||||
work_ref = str(request.path_params.get("work_ref") or "").strip()
|
||||
binding_id = str(request.query_params.get("binding_id") or "").strip()
|
||||
try:
|
||||
from ouroboros.presence_bindings import load_presence_binding
|
||||
from ouroboros.task_results import load_task_result
|
||||
|
||||
load_presence_binding(ctx.data_dir, skill_name, binding_id)
|
||||
stored = load_task_result(ctx.data_dir, work_ref) or {}
|
||||
metadata = stored.get("metadata") if isinstance(stored.get("metadata"), dict) else {}
|
||||
presence = metadata.get("presence") if isinstance(metadata.get("presence"), dict) else {}
|
||||
if str(presence.get("binding_id") or "") != binding_id:
|
||||
return _json_error("presence work reference not found", 404)
|
||||
status = str(stored.get("status") or "")
|
||||
if status not in {"completed", "failed", "cancelled"}:
|
||||
return JSONResponse({"ok": True, "status": "pending", "work_ref": work_ref,
|
||||
"delivery_reporting_version": presence.get("delivery_reporting_version", 0)}, status_code=202)
|
||||
from ouroboros.presence_runner import presence_result_from_stored
|
||||
|
||||
result = presence_result_from_stored(stored, work_ref)
|
||||
return JSONResponse({
|
||||
"ok": True,
|
||||
"status": status,
|
||||
"outcome": result.outcome,
|
||||
"text": result.text,
|
||||
"work_ref": work_ref,
|
||||
"delivery_reporting_version": result.delivery_reporting_version,
|
||||
})
|
||||
status, body = await asyncio.to_thread(_presence_work_view, ctx, skill_name, work_ref, binding_id)
|
||||
except Exception as exc:
|
||||
code = str(getattr(exc, "code", ""))
|
||||
if code:
|
||||
return _json_error(str(exc), 404)
|
||||
return _presence_exception(exc, 404)
|
||||
log.debug("Host service presence work lookup failed", exc_info=True)
|
||||
return _json_error("presence work lookup failed", 500)
|
||||
return _presence_error("presence work lookup failed", 500, "presence_work_unavailable", "retry")
|
||||
return JSONResponse(body, status_code=status)
|
||||
|
||||
|
||||
async def _api_chat_decision(request: Request) -> JSONResponse:
|
||||
|
|
@ -840,8 +927,7 @@ async def _api_chat_decision(request: Request) -> JSONResponse:
|
|||
"""
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
|
||||
ctx.require_permission(skill_name, token_payload, "inject_chat")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "inject_chat")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
if not ctx.rate_limiter.allow(f"{skill_name}:decision"):
|
||||
|
|
@ -871,10 +957,10 @@ async def _api_ws_message(request: Request) -> JSONResponse:
|
|||
"""
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, _payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
|
||||
skill_name, _payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""))
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
loaded = find_skill(ctx.data_dir, skill_name)
|
||||
loaded = await asyncio.to_thread(find_skill, ctx.data_dir, skill_name)
|
||||
if loaded is None:
|
||||
return _json_error(f"skill {skill_name!r} is not installed", 403)
|
||||
if "ws_handler" not in {str(p).strip() for p in (loaded.manifest.permissions or [])}:
|
||||
|
|
@ -1157,8 +1243,7 @@ async def _api_chat_operation(request: Request) -> JSONResponse:
|
|||
"""
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
|
||||
ctx.require_permission(skill_name, token_payload, "inject_chat")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "inject_chat")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
if not ctx.rate_limiter.allow(f"{skill_name}:operations"):
|
||||
|
|
@ -1190,8 +1275,7 @@ async def _api_chat_cancel(request: Request) -> JSONResponse:
|
|||
"""
|
||||
ctx: HostServiceContext = request.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(request.headers.get("x-skill-token", ""))
|
||||
ctx.require_permission(skill_name, token_payload, "inject_chat")
|
||||
skill_name, _token_payload = await _authenticated(ctx, request.headers.get("x-skill-token", ""), "inject_chat")
|
||||
except HostServiceAuthError as exc:
|
||||
return _json_error(str(exc), 403)
|
||||
if not ctx.rate_limiter.allow(f"{skill_name}:cancel"):
|
||||
|
|
@ -1220,7 +1304,7 @@ async def _api_chat_cancel(request: Request) -> JSONResponse:
|
|||
async def _ws_events(websocket: WebSocket) -> None:
|
||||
ctx: HostServiceContext = websocket.app.state.host_service_context
|
||||
try:
|
||||
skill_name, token_payload = ctx.authenticate_token_payload(_token_from_websocket(websocket))
|
||||
skill_name, token_payload = await _authenticated(ctx, _token_from_websocket(websocket))
|
||||
except HostServiceAuthError:
|
||||
await websocket.close(code=1008)
|
||||
return
|
||||
|
|
@ -1237,7 +1321,9 @@ async def _ws_events(websocket: WebSocket) -> None:
|
|||
elif message.get("type") == "subscribe":
|
||||
topic = str(message.get("topic") or "")
|
||||
try:
|
||||
ctx.require_permission(skill_name, token_payload, f"subscribe_event:{topic}")
|
||||
await asyncio.to_thread(
|
||||
ctx.require_permission, skill_name, token_payload, f"subscribe_event:{topic}",
|
||||
)
|
||||
except HostServiceAuthError as exc:
|
||||
await websocket.send_json({"type": "error", "error": str(exc)})
|
||||
continue
|
||||
|
|
|
|||
|
|
@ -61,8 +61,8 @@ from ouroboros.gateways.claudexor import (
|
|||
)
|
||||
from ouroboros.llm_attempt import _attempt_request, _candidate_before_dispatch
|
||||
from ouroboros.llm_substitution import (
|
||||
SubstitutionBudget, substitution_fact, failed_account_preference,
|
||||
remember_failed_profile, take_failed_account_preference)
|
||||
AccountRotation, SubstitutionBudget, substitution_fact, failed_account_preference,
|
||||
take_failed_account_preference)
|
||||
from ouroboros.model_slots import MODEL_ACCOUNTS_KEY, model_role_option
|
||||
from ouroboros.model_wait import ModelWaitInterrupted, current_model_wait, prepared_call_scope
|
||||
from ouroboros.observability import persist_call
|
||||
|
|
@ -816,13 +816,13 @@ def _processing_retry_payload(payload: dict, error: ClaudexorModelNotDispatched)
|
|||
|
||||
|
||||
def chat_claudexor(target: dict, messages: list, tools: list | None, **parameters: Any) -> tuple[dict, dict]:
|
||||
"""One generation, with one no-start repair per continuation/processing axis."""
|
||||
"""One generation, one no-start repair per continuation/processing axis, quota re-asks on Auto."""
|
||||
target = prepare_processing_target(target)
|
||||
payload = _request(target, messages, tools, parameters)
|
||||
retry_preparation = None
|
||||
native_repaired = processing_repaired = False
|
||||
substitution = SubstitutionBudget(ClaudexorModelError)
|
||||
for _preparation in range(3 + substitution.redos):
|
||||
substitution, rotation = SubstitutionBudget(ClaudexorModelError), AccountRotation()
|
||||
for _preparation in range(3 + substitution.redos + rotation.CEILING):
|
||||
invocation = _ModelInvocation(target, payload, parameters)
|
||||
try:
|
||||
with prepared_call_scope(retry_preparation or {}) as prepared:
|
||||
|
|
@ -839,7 +839,7 @@ def chat_claudexor(target: dict, messages: list, tools: list | None, **parameter
|
|||
continue
|
||||
adopt_turn_state((prepared or parameters).get("model_turn_state"),
|
||||
invocation.payload, result)
|
||||
return substitution.disclose(invocation.finish(result))
|
||||
return rotation.disclose(substitution.disclose(invocation.finish(result)))
|
||||
except ClaudexorModelNotDispatched as error:
|
||||
if invocation.response_ref:
|
||||
invocation.acknowledge()
|
||||
|
|
@ -852,13 +852,14 @@ def chat_claudexor(target: dict, messages: list, tools: list | None, **parameter
|
|||
updated = _reset_native(payload, error, invocation)
|
||||
native_repaired = updated is not None
|
||||
if updated is None:
|
||||
remember_failed_profile(target, parameters, error)
|
||||
rotation.refused_call(target, parameters, invocation, error) # an engine verdict is never re-asked
|
||||
raise
|
||||
payload = updated
|
||||
retry_preparation = _native_retry_preparation(target, payload, parameters, error)
|
||||
except ClaudexorModelError as error:
|
||||
remember_failed_profile(target, parameters, error)
|
||||
raise
|
||||
if not rotation.refused_call(target, parameters, invocation, error):
|
||||
raise
|
||||
retry_preparation, parameters, payload = rotation.reask(parameters, payload)
|
||||
except PhysicalAttemptPreparationFailed as error:
|
||||
cause = error.__cause__
|
||||
if isinstance(cause, ClaudexorModelError):
|
||||
|
|
@ -938,8 +939,8 @@ async def chat_claudexor_async(target: dict, messages: list, tools: list | None,
|
|||
payload = _request(target, messages, tools, parameters)
|
||||
retry_preparation = None
|
||||
native_repaired = processing_repaired = False
|
||||
substitution = SubstitutionBudget(ClaudexorModelError)
|
||||
for _preparation in range(3 + substitution.redos):
|
||||
substitution, rotation = SubstitutionBudget(ClaudexorModelError), AccountRotation()
|
||||
for _preparation in range(3 + substitution.redos + rotation.CEILING):
|
||||
invocation = _ModelInvocation(target, payload, parameters)
|
||||
try:
|
||||
with prepared_call_scope(retry_preparation or {}) as prepared:
|
||||
|
|
@ -964,7 +965,7 @@ async def chat_claudexor_async(target: dict, messages: list, tools: list | None,
|
|||
continue
|
||||
adopt_turn_state((prepared or parameters).get("model_turn_state"),
|
||||
invocation.payload, result)
|
||||
return substitution.disclose(await invocation.offload(invocation.finish, result))
|
||||
return rotation.disclose(substitution.disclose(await invocation.offload(invocation.finish, result)))
|
||||
except ClaudexorModelNotDispatched as error:
|
||||
if invocation.response_ref:
|
||||
await invocation.offload(invocation.acknowledge)
|
||||
|
|
@ -977,13 +978,14 @@ async def chat_claudexor_async(target: dict, messages: list, tools: list | None,
|
|||
updated = _reset_native(payload, error, invocation)
|
||||
native_repaired = updated is not None
|
||||
if updated is None:
|
||||
remember_failed_profile(target, parameters, error)
|
||||
rotation.refused_call(target, parameters, invocation, error) # an engine verdict is never re-asked
|
||||
raise
|
||||
payload = updated
|
||||
retry_preparation = _native_retry_preparation(target, payload, parameters, error)
|
||||
except ClaudexorModelError as error:
|
||||
remember_failed_profile(target, parameters, error)
|
||||
raise
|
||||
if not rotation.refused_call(target, parameters, invocation, error):
|
||||
raise
|
||||
retry_preparation, parameters, payload = rotation.reask(parameters, payload)
|
||||
except PhysicalAttemptPreparationFailed as error:
|
||||
cause = error.__cause__
|
||||
if isinstance(cause, (ClaudexorModelError, asyncio.CancelledError)):
|
||||
|
|
|
|||
|
|
@ -29,6 +29,21 @@ configured fallback chain owns what happens next.
|
|||
The budget is its own counter, never the transport's no-start preparation
|
||||
loop: a repair fixes a request that was never sent, while a redo spends a
|
||||
fresh generation on an answer that already arrived.
|
||||
|
||||
## Asking again after a vendor refused one account's quota
|
||||
|
||||
The same question as a redo, for a refusal instead of an answer. The serving
|
||||
engine (3.14.0) resolves ONE account per model operation before dispatch and
|
||||
invokes it once, so a quota refusal before dispatch is the engine's own verdict
|
||||
on its pool or a pin: it is never asked again, and only its ``poolCause`` says
|
||||
that every compatible account is blocked. An HTTP-level vendor refusal is a
|
||||
settled fact about the one account it names, refused before any response
|
||||
stream began, so nothing was generated; the engine turns its reset evidence
|
||||
into that account's cooldown. The SAME payload, naming no account, then lets
|
||||
the engine choose again. An account it selects twice ends the rotation without
|
||||
claiming the pool is spent, and so does every reason a redo is refused. A pin,
|
||||
an unknown outcome and a stream-level refusal (generation may have begun) are
|
||||
never asked again.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -234,3 +249,64 @@ def _refusal(budget: "SubstitutionBudget", invocation: Any, fact: dict,
|
|||
if invocation.capture is not None:
|
||||
error.physical_attempt_capture = invocation.capture
|
||||
return error
|
||||
|
||||
|
||||
class AccountRotation:
|
||||
"""One call's unpinned re-asks of the same payload after a vendor quota refusal.
|
||||
|
||||
A re-ask follows only a newly refused account, so the engine's pool bounds
|
||||
it; the ceiling is defensive and, like every other stop, claims nothing
|
||||
about the accounts that were not tried.
|
||||
"""
|
||||
|
||||
CEILING = 8
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.refused: list[str] = []
|
||||
|
||||
def refused_call(self, target: dict, parameters: dict, invocation: Any, error: Any) -> bool:
|
||||
"""Record one refusal; True only when this call may ask the engine again.
|
||||
|
||||
A refusal that reaches the caller leaves its account unpreferred for the next
|
||||
matching dispatch; a re-ask names no account, so it remembers nothing.
|
||||
"""
|
||||
from ouroboros.model_wait import model_wait_reason
|
||||
|
||||
if model_wait_reason(error) not in {"quota", "auth_quota"}:
|
||||
remember_failed_profile(target, parameters, error)
|
||||
return False
|
||||
account = str((error.route or {}).get("credentialProfileId") or "")
|
||||
repeated = account in self.refused
|
||||
if account and not repeated:
|
||||
self.refused.append(account)
|
||||
state = getattr(getattr(error, "physical_attempt_capture", None), "state", None)
|
||||
stop = ("pinned_account" if (invocation.payload.get("account") or {}).get("mode") == "pin"
|
||||
else "pool_exhausted" if ((error.problem or {}).get("context") or {}).get("poolCause") in {"quota", "mixed"}
|
||||
else "refused_before_dispatch" if state == "released"
|
||||
else "outcome_not_settled" if state != "settled"
|
||||
else "generation_not_excluded" if not error.status_code
|
||||
else "account_unobserved" if not account
|
||||
else "engine_reselected_refused_account" if repeated
|
||||
else "rotation_ceiling" if len(self.refused) > self.CEILING
|
||||
else _redo_allowed(invocation.payload))
|
||||
append_jsonl(invocation.root / "logs" / "events.jsonl", {
|
||||
"ts": utc_now_iso(), "type": "model_account_rotation", "task_id": invocation.task_id,
|
||||
"model_role": invocation.role, "operation_id": invocation.operation_id, "account": account,
|
||||
"disposition": stop or "reask", "refused_accounts": list(self.refused)})
|
||||
if stop:
|
||||
remember_failed_profile(target, parameters, error)
|
||||
error.account_rotation = {"refused_accounts": list(self.refused), "stop": stop,
|
||||
"pool_exhausted": stop == "pool_exhausted"}
|
||||
return not stop
|
||||
|
||||
@staticmethod
|
||||
def reask(parameters: dict, payload: dict) -> tuple[None, dict, dict]:
|
||||
"""No preparation to rebind, and no account preference on any later request of this call."""
|
||||
return None, {**parameters, "_no_account_preference": True}, {**payload, "account": {"mode": "auto"}}
|
||||
|
||||
def disclose(self, answer: tuple[dict, dict]) -> tuple[dict, dict]:
|
||||
"""An answer that followed a rotation names the accounts that refused first."""
|
||||
message, usage = answer
|
||||
if self.refused and isinstance(usage.get("claudexor"), dict):
|
||||
usage["claudexor"]["account_rotation"] = {"refused_accounts": list(self.refused)}
|
||||
return message, usage
|
||||
|
|
|
|||
|
|
@ -25,7 +25,7 @@ from ouroboros.deadline_utils import (
|
|||
from ouroboros.llm import LLMClient, LocalContextTooLargeError, add_usage
|
||||
from ouroboros.llm_claudexor import propagate_model_error
|
||||
from ouroboros.llm_substitution import same_route_refusal, stamp_substitutions
|
||||
from ouroboros.model_wait import propagate_model_control
|
||||
from ouroboros.model_wait import current_model_wait, model_wait_reason, propagate_model_control
|
||||
from ouroboros.openai_chat_dispatch import CUSTOM_RECEIPTS_USAGE_KEY, pop_custom_validation_receipts
|
||||
from ouroboros.llm_attempt import PROVIDER_POLICY_REFUSAL, _is_provider_policy_refusal # typed-refusal contract owner
|
||||
from ouroboros.observability import new_call_id, new_execution_id, persist_call
|
||||
|
|
@ -148,8 +148,8 @@ _COOLDOWN_ERROR_KINDS = _TRANSIENT_RETRY_KINDS | frozenset({"rate_limit"})
|
|||
# A subscription window that is spent but heals on a timer. SSOT for the name.
|
||||
# Deliberately NOT `quota_exhausted`: that class is classified PERMANENT, which is
|
||||
# correct for a billing refusal (402, no credits) and wrong for a plan window whose
|
||||
# only cure is waiting. Scheduling follows `reset_at`, not the 60s-capped exponential
|
||||
# backoff, so a six-hour window never becomes sixty one-minute retries.
|
||||
# only cure is waiting. Scheduling follows `reset_at` (never for a confirmed quota refusal, which
|
||||
# model_wait recovers, nor inline Presence), so a six-hour window never becomes sixty 1-minute retries.
|
||||
SUBSCRIPTION_WINDOW_EXHAUSTED = "subscription_window_exhausted"
|
||||
_TRANSIENT_RETRY_DEFAULT = 6
|
||||
_TRANSIENT_BACKOFF_CAP_SEC = NETWORK_WAIT_BACKOFF_MAX_SEC
|
||||
|
|
@ -305,9 +305,7 @@ def _cooldown_kind_for_empty_response(body_error: Dict[str, Any], event_type: st
|
|||
return body_kind if body_kind in _COOLDOWN_ERROR_KINDS else event_type
|
||||
|
||||
|
||||
def _retry_backoff_sec(
|
||||
accumulated_usage: Dict[str, Any], error_kind: str, attempt: int, is_transient: bool,
|
||||
) -> float:
|
||||
def _retry_backoff_sec(accumulated_usage: Dict[str, Any], error_kind: str, attempt: int, is_transient: bool) -> float:
|
||||
"""Seconds to wait before retrying the same request.
|
||||
|
||||
A KNOWN reset instant wins over guesswork: a spent subscription window is scheduled
|
||||
|
|
@ -789,7 +787,7 @@ def classify_llm_exception(exc: Exception, safe_error: str = "") -> LlmErrorClas
|
|||
str(getattr(capture, "state", "") or "") == "released"
|
||||
and str(getattr(capture, "provider", "") or "") != "local"
|
||||
and not bool(getattr(capture, "route_is_loopback", False))
|
||||
and is_pre_dispatch_transport_failure(exc)
|
||||
and is_pre_dispatch_transport_failure(exc) and not model_wait_reason(exc) # an engine resource verdict is no outage
|
||||
):
|
||||
# Typed $0 fact: the request never left this host toward a REMOTE
|
||||
# provider (released custody + typed pre-dispatch transport failure).
|
||||
|
|
@ -983,7 +981,7 @@ def _record_llm_call_error(
|
|||
**{key: error_event[key] for key in ("ts", "llm_call_id", "model", "requested_profile", "observed_route")},
|
||||
"failure_code": classification.kind, "reset_at": classification.reset_at,
|
||||
})
|
||||
ctx.accumulated_usage.update(_last_llm_error=_short_error_text(display_error),
|
||||
ctx.accumulated_usage.update(_last_llm_error=_short_error_text(display_error), _last_llm_resource_refusal=model_wait_reason(error),
|
||||
_last_llm_error_kind=classification.kind, _last_llm_retry_same_request=will_retry)
|
||||
if classification.retry_after_sec is not None:
|
||||
ctx.accumulated_usage.update(_last_llm_retry_after_sec=classification.retry_after_sec,
|
||||
|
|
@ -1030,9 +1028,10 @@ def _stop_after_llm_error(ctx: _LlmErrorContext) -> bool:
|
|||
is_transient = error_kind in _TRANSIENT_RETRY_KINDS
|
||||
# Non-transient retryables: max_retries capped by the loop ceiling (primary: no-op).
|
||||
attempt_budget = ctx.transient_budget if is_transient else min(ctx.max_retries, ctx.transient_budget)
|
||||
backoff = (
|
||||
backoff = ( # a confirmed resource refusal and inline Presence neither resend nor sleep to a reset
|
||||
_retry_backoff_sec(accumulated_usage, error_kind, ctx.attempt, is_transient)
|
||||
if ctx.attempt < attempt_budget - 1 else None
|
||||
if ctx.attempt < attempt_budget - 1 and not accumulated_usage.get("_last_llm_resource_refusal")
|
||||
and (error_kind != SUBSCRIPTION_WINDOW_EXHAUSTED or getattr(current_model_wait(), "waits_allowed", True)) else None
|
||||
)
|
||||
if backoff is not None:
|
||||
if _sleep_within_deadline(backoff, ctx.deadline_ts):
|
||||
|
|
@ -1062,6 +1061,7 @@ def provider_no_call_source(accumulated_usage: Dict[str, Any], deadline_exhauste
|
|||
cleared only by a usable response) forbids the resend whatever the sticky kind."""
|
||||
if str(accumulated_usage.get("_last_llm_error_kind") or "") == "provider_outcome_unknown" or isinstance(accumulated_usage.get(TRANSPORT_DEATHS_KEY), dict):
|
||||
return "provider_outcome_unknown_no_resend", False
|
||||
if accumulated_usage.get("resource_refusal"): return "resource_refusal_no_resend", True # typed temporary refusal: no route or owner wait recovered it
|
||||
if same_route_refusal(accumulated_usage): return "same_route_refusal_no_resend", False # a salvage call is the same request on the same route, and the accounts that refused it were just marked
|
||||
if not deadline_exhausted and bool(accumulated_usage.get(RETRY_WALL_EXHAUSTED_KEY)):
|
||||
return "retry_wall_exhausted_no_repay", True
|
||||
|
|
@ -1436,7 +1436,7 @@ def call_llm_with_retry(
|
|||
context_fit_event_fields = _context_fit_event_fields(accumulated_usage) if physical_context is not None else {}
|
||||
_take_custom_receipts(usage, msg, accumulated_usage)
|
||||
for stale in ("_last_llm_error", "_last_llm_error_kind", "_last_llm_retry_same_request",
|
||||
"_last_llm_status_code", "_last_llm_provider_code"):
|
||||
"_last_llm_status_code", "_last_llm_provider_code", "_last_llm_resource_refusal"):
|
||||
accumulated_usage.pop(stale, None)
|
||||
cost, display_model, provider, cost_estimated = _normalize_usage_cost(
|
||||
usage, model=model, use_local=use_local,
|
||||
|
|
|
|||
|
|
@ -28,6 +28,10 @@ from ouroboros.usage_accounting import PhysicalAttemptContext, PhysicalAttemptPr
|
|||
|
||||
|
||||
log = logging.getLogger("ouroboros.loop")
|
||||
# The typed, temporary non-success of a round whose resource refusal no route recovered and
|
||||
# no owner wait follows (inline Presence, or a chain stopped by its own fence): nothing more is sent.
|
||||
RESOURCE_REFUSAL_KEY = "resource_refusal"
|
||||
_CHAIN_STOP_KINDS = ("provider_outcome_unknown", "deadline_exhausted", "transport_unavailable")
|
||||
|
||||
|
||||
def _loop():
|
||||
|
|
@ -182,6 +186,11 @@ def _run_cross_model_fallback_chain(
|
|||
attempt_cap = _fcd.attempts_per_model()
|
||||
waiter = current_model_wait()
|
||||
configured_chain = parse_fallback_chain()
|
||||
# A primary quota refusal returned here, not waited: the configured routes come
|
||||
# first, and the owner question is that refusal's own wait, asked after them.
|
||||
deferred, tools._ctx._deferred_resource_refusal = getattr(tools._ctx, "_deferred_resource_refusal", None), None
|
||||
owner_question = deferred is not None and waiter is not None and waiter.waits_allowed
|
||||
candidates, tried = fallback_candidate_targets(active_model), []
|
||||
# The notice names the model that was actually just tried. `active_model`
|
||||
# stays the primary until a candidate succeeds, so a second switch would
|
||||
# otherwise read "primary -> B" beside B's predecessor's failure reason.
|
||||
|
|
@ -193,7 +202,7 @@ def _run_cross_model_fallback_chain(
|
|||
# dispatch lane is the single global USE_LOCAL_FALLBACK flag above (the
|
||||
# pre-existing chain contract): the ladder's `provider_route` stays the ""
|
||||
# sentinel rather than fabricating a per-candidate fact nothing consumes.
|
||||
for candidate in fallback_candidate_targets(active_model):
|
||||
for index, candidate in enumerate(candidates):
|
||||
fallback_model = candidate.model_id
|
||||
# The role belongs to the configured row, before active-model removal
|
||||
# and deduplication. Those filters must not shift its account binding.
|
||||
|
|
@ -261,7 +270,9 @@ def _run_cross_model_fallback_chain(
|
|||
attempt_cap=attempt_cap,
|
||||
model_role=fallback_role,
|
||||
emit_progress=emit_progress,
|
||||
defer_resource_wait=owner_question or _route_follows(candidates[index + 1:], fallback_use_local),
|
||||
)
|
||||
tried.append(fallback_model)
|
||||
msg, _cost, candidate_mode = _loop()._call_round_model(candidate_call)
|
||||
if msg is not None:
|
||||
(
|
||||
|
|
@ -296,10 +307,28 @@ def _run_cross_model_fallback_chain(
|
|||
tools._ctx.messages = messages
|
||||
tools._ctx.active_context_mode = active_context_mode
|
||||
_restore_context_fit_usage(accumulated_usage, primary_context_usage)
|
||||
if str(accumulated_usage.get("_last_llm_error_kind") or "") in ("provider_outcome_unknown", "deadline_exhausted", "transport_unavailable"):
|
||||
if str(accumulated_usage.get("_last_llm_error_kind") or "") in _CHAIN_STOP_KINDS:
|
||||
break
|
||||
_cooled(fallback_model, fallback_use_local)
|
||||
previous_model, previous_tag = fallback_model, ftag
|
||||
accumulated_usage.pop(RESOURCE_REFUSAL_KEY, None)
|
||||
if msg is None and owner_question and str(accumulated_usage.get("_last_llm_error_kind") or "") not in _CHAIN_STOP_KINDS:
|
||||
# The owner wait of the primary's retained refusal: catalog checks only, so no
|
||||
# generation precedes the answer. The primary sends again only after it resolves.
|
||||
deferred.ask_owner(waiter)
|
||||
primary_call = _loop()._RoundModelCallContext(
|
||||
llm=llm, messages=messages, tools=tools, context_fit_plan=context_fit_plan,
|
||||
active_model=active_model, tool_schemas=tool_schemas, active_effort=active_effort,
|
||||
max_retries=max_retries, drive_logs=drive_logs, task_id=task_id, round_idx=round_idx,
|
||||
event_queue=event_queue, accumulated_usage=accumulated_usage, task_type=task_type,
|
||||
active_use_local=active_use_local, active_context_mode=active_context_mode,
|
||||
drive_root=pathlib.Path(drive_logs).parent, emit_progress=emit_progress, defer_resource_wait=False)
|
||||
msg, _cost, active_context_mode = _loop()._call_round_model(primary_call)
|
||||
active_model, active_use_local = primary_call.active_model, primary_call.active_use_local
|
||||
context_fit_plan = primary_call.context_fit_plan
|
||||
elif msg is None and deferred is not None: # the refused primary buys no forced final either
|
||||
accumulated_usage[RESOURCE_REFUSAL_KEY] = deferred.terminal(
|
||||
fallbacks_tried=tried, owner_wait="not_asked" if owner_question else "not_allowed")
|
||||
return (
|
||||
msg,
|
||||
active_model,
|
||||
|
|
@ -459,6 +488,26 @@ class _RoundModelCallContext:
|
|||
# callable documented to accept incident=; the ToolContext ABI's
|
||||
# emit_progress_fn takes a single argument and must not carry the pair.
|
||||
emit_progress: Optional[Callable[..., None]] = None
|
||||
# None: a primary round, which defers a resource refusal while a configured route
|
||||
# follows or the turn may not wait. The chain sets it for each candidate.
|
||||
defer_resource_wait: Optional[bool] = None
|
||||
|
||||
|
||||
def _route_follows(candidates: List[Any], use_local: bool) -> bool:
|
||||
from ouroboros import fallback_cooldown
|
||||
|
||||
return any(not fallback_cooldown.is_cooling_down(item.model_id, use_local) for item in candidates)
|
||||
|
||||
|
||||
def _fallback_route_follows(ctx: _RoundModelCallContext) -> bool:
|
||||
"""Whether this primary round's fallback chain would still dial a configured route."""
|
||||
from ouroboros.config import fallback_candidate_targets
|
||||
|
||||
if bool(getattr(ctx.tools._ctx, "exact_model_route", False)) or isinstance(
|
||||
ctx.accumulated_usage.get(TRANSPORT_DEATHS_KEY), dict):
|
||||
return False
|
||||
use_local = runtime_setting("USE_LOCAL_FALLBACK", "").lower() in ("true", "1")
|
||||
return _route_follows(fallback_candidate_targets(ctx.active_model), use_local)
|
||||
|
||||
|
||||
def _context_fit_round_id(ctx: _RoundModelCallContext) -> str:
|
||||
|
|
@ -587,14 +636,21 @@ def _dispatch_round_model(
|
|||
"model_role": getattr(ctx, "model_role", ""),
|
||||
"task_metadata": getattr(ctx.tools._ctx, "task_metadata", {})},
|
||||
context_fit_plan=plan, overrides=waiter.overrides if waiter else None)
|
||||
binding = (waiter.register_reprepare(role, lambda kwargs: _reprepare_waiting_main(ctx, kwargs))
|
||||
if waiter is not None else contextlib.nullcontext())
|
||||
primary = getattr(ctx, "defer_resource_wait", None) is None
|
||||
if primary:
|
||||
ctx.tools._ctx._deferred_resource_refusal = None
|
||||
ctx.accumulated_usage.pop(RESOURCE_REFUSAL_KEY, None)
|
||||
previous_call = ctx.accumulated_usage.get("_last_llm_call_meta")
|
||||
from ouroboros.acceptance_settlement import expose_acceptance_feedback
|
||||
|
||||
observe_feedback = lambda sent: expose_acceptance_feedback(
|
||||
getattr(ctx.tools._ctx, "_execution_trace", {}), sent, str(ctx.task_id))
|
||||
with binding:
|
||||
with contextlib.ExitStack() as binding:
|
||||
deferral = None
|
||||
if waiter is not None:
|
||||
binding.enter_context(waiter.register_reprepare(role, lambda kwargs: _reprepare_waiting_main(ctx, kwargs)))
|
||||
if ((not waiter.waits_allowed or _fallback_route_follows(ctx)) if primary else ctx.defer_resource_wait):
|
||||
deferral = binding.enter_context(waiter.defer_resource_wait(role))
|
||||
result = _loop().call_llm_with_retry(
|
||||
ctx.llm, ctx.messages, ctx.active_model, ctx.tool_schemas,
|
||||
ctx.active_effort, ctx.max_retries, ctx.drive_logs, ctx.task_id,
|
||||
|
|
@ -616,6 +672,10 @@ def _dispatch_round_model(
|
|||
model_turn_state=getattr(ctx.tools._ctx, "model_turn_state", None),
|
||||
model_context_observer=observe_feedback,
|
||||
)
|
||||
if primary and deferral is not None and deferral.fact and result[0] is None:
|
||||
ctx.tools._ctx._deferred_resource_refusal = deferral
|
||||
if not waiter.waits_allowed: # typed at once: the terminal may come before any chain
|
||||
ctx.accumulated_usage[RESOURCE_REFUSAL_KEY] = deferral.terminal(fallbacks_tried=[], owner_wait="not_allowed")
|
||||
pending_wait_handover = getattr(ctx.tools._ctx, "_pending_model_wait_handover", None)
|
||||
if pending_wait_handover is not None:
|
||||
if result[0] is not None:
|
||||
|
|
|
|||
|
|
@ -974,10 +974,23 @@ def provider_recovery_hint(accumulated_usage: Dict[str, Any]) -> str:
|
|||
if kind == "subscription_window_exhausted":
|
||||
reset_at = str(accumulated_usage.get("_last_llm_reset_at") or "").strip()
|
||||
when = f" It resets at {reset_at}." if reset_at else ""
|
||||
refusal = accumulated_usage.get("resource_refusal")
|
||||
if not refusal:
|
||||
return (
|
||||
" The subscription window for the delegated route is spent. This is "
|
||||
f"TRANSIENT, not a billing refusal — waiting cures it.{when} Retrying is "
|
||||
"scheduled against that reset time, not the ordinary short backoff."
|
||||
)
|
||||
rotation, tried = refusal.get("account_rotation") or {}, ", ".join(refusal.get("fallbacks_tried") or [])
|
||||
# Only the engine's own pool verdict proves every account; any other stop claims nothing.
|
||||
accounts = (" The engine reports every compatible account blocked." if rotation.get("pool_exhausted") else
|
||||
f" Account rotation stopped ({rotation.get('stop')}); other accounts are unproven."
|
||||
if rotation else "")
|
||||
return (
|
||||
" The subscription window for the delegated route is spent. This is "
|
||||
f"TRANSIENT, not a billing refusal — waiting cures it.{when} Retrying is "
|
||||
"scheduled against that reset time, not the ordinary short backoff."
|
||||
" The subscription quota for this model route is spent. This is "
|
||||
f"TRANSIENT, not a billing refusal — waiting cures it.{when}{accounts}"
|
||||
f"{f' Configured fallbacks tried without an answer: {tried}.' if tried else ''} Nothing "
|
||||
"more was sent, and nothing sleeps to that reset."
|
||||
)
|
||||
if kind == "model_substituted":
|
||||
return (
|
||||
|
|
|
|||
|
|
@ -4,6 +4,12 @@ The existing task owns its worker, writer lane, mailbox and durable result.
|
|||
This module keeps only the live call's wait and role overrides. It neither
|
||||
schedules work nor records physical attempts: every resumed call still goes
|
||||
through LLMClient and the ordinary physical-attempt ledger.
|
||||
|
||||
Quota recovery order: the transport's Auto account rotation, then every
|
||||
configured route of the round, and only then the owner. A round whose later
|
||||
route exists returns its refusal instead of waiting (``ResourceDeferral``); the
|
||||
owner question is then opened from that retained refusal, never from another
|
||||
generation. An inline Presence turn never waits for quota or for the owner.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -160,6 +166,9 @@ _CALENDAR: contextvars.ContextVar[tuple[str, ...]] = contextvars.ContextVar(
|
|||
"ouroboros_model_wait_calendar", default=())
|
||||
_LOGICAL: contextvars.ContextVar[tuple[tuple[float, str], ...]] = contextvars.ContextVar(
|
||||
"ouroboros_model_wait_logical", default=())
|
||||
# One call's own scope, like its physical capture: never copied into a helper's context.
|
||||
_DEFERRED: contextvars.ContextVar[tuple["ResourceDeferral", ...]] = contextvars.ContextVar(
|
||||
"ouroboros_model_wait_deferred", default=())
|
||||
|
||||
|
||||
def copy_wait_context() -> contextvars.Context:
|
||||
|
|
@ -210,6 +219,38 @@ def dispatch_deadline_remaining_sec() -> float | None:
|
|||
return min(remaining) if remaining else None
|
||||
|
||||
|
||||
class ResourceDeferral:
|
||||
"""One round's resource refusal, returned to the round instead of waited in its call.
|
||||
|
||||
The round tries its configured routes first. The owner question is then this
|
||||
refusal's own wait, opened from the retained call: catalog checks only, so no
|
||||
generation precedes the owner's answer. A turn that may not wait keeps the fact.
|
||||
"""
|
||||
|
||||
def __init__(self, role: str):
|
||||
self.role = role
|
||||
self.fact: dict = {}
|
||||
self._retained: tuple | None = None
|
||||
|
||||
def retain(self, receiver: Any, error: Exception, values: dict) -> None:
|
||||
self._retained = (receiver, error, values)
|
||||
self.fact = {"reason": model_wait_reason(error), "role": self.role, "model": str(values.get("model") or ""),
|
||||
"reset_at": str(getattr(error, "reset_at", "") or ""),
|
||||
"account_rotation": copy.deepcopy(getattr(error, "account_rotation", None))}
|
||||
|
||||
def ask_owner(self, context: "TaskModelWait") -> dict:
|
||||
receiver, error, values = self._retained
|
||||
return context.wait(receiver, error, values)
|
||||
|
||||
def terminal(self, *, fallbacks_tried: list, owner_wait: str) -> dict:
|
||||
"""The typed temporary non-success when no route recovered it and no owner wait follows."""
|
||||
return {**self.fact, "fallbacks_tried": list(fallbacks_tried), "owner_wait": owner_wait, "temporary": True}
|
||||
|
||||
|
||||
def _deferral(role: str) -> ResourceDeferral | None:
|
||||
return next((item for item in reversed(_DEFERRED.get()) if item.role == role), None)
|
||||
|
||||
|
||||
def mutate_wait(root: Any, task_id: str, wait_id: str, transform: Callable) -> dict:
|
||||
"""Mutate only one wait projection in the existing schema-stamped task result."""
|
||||
from ouroboros.task_results import (
|
||||
|
|
@ -281,6 +322,12 @@ class TaskModelWait:
|
|||
self.seen_controls: set[str] = set()
|
||||
self.mailbox_stamp = None
|
||||
|
||||
@property
|
||||
def waits_allowed(self) -> bool:
|
||||
"""An inline Presence turn holds its conversation slot with no owner at the
|
||||
computer: rotation and fallback still run, then a typed temporary refusal."""
|
||||
return not self.task.get("_presence_turn")
|
||||
|
||||
def mutate_row(self, wait_id: str, transform: Callable) -> dict:
|
||||
if self.row_mutator is not None:
|
||||
return self.row_mutator(wait_id, transform)
|
||||
|
|
@ -377,6 +424,16 @@ class TaskModelWait:
|
|||
finally:
|
||||
_REPREPARE.reset(token)
|
||||
|
||||
@contextlib.contextmanager
|
||||
def defer_resource_wait(self, role: str) -> Iterator[ResourceDeferral]:
|
||||
"""A later route of this round, or a turn that may not wait, owns this role's refusal."""
|
||||
deferral = ResourceDeferral(role)
|
||||
token = _DEFERRED.set((*_DEFERRED.get(), deferral))
|
||||
try:
|
||||
yield deferral
|
||||
finally:
|
||||
_DEFERRED.reset(token)
|
||||
|
||||
def paused_seconds(self, slot_id: str = "", *, now: float | None = None) -> float:
|
||||
with self.lock:
|
||||
clock = self.clocks.get(slot_id)
|
||||
|
|
@ -560,6 +617,14 @@ class TaskModelWait:
|
|||
"worker_slot_held": self.worker_slot_held, "started_at": utc_now_iso()}
|
||||
if self.owner_id:
|
||||
row["model_wait_owner_id"] = self.owner_id
|
||||
if getattr(error, "account_rotation", None):
|
||||
row["account_rotation"] = copy.deepcopy(error.account_rotation)
|
||||
# A vendor refusal with no reset evidence gives the engine no cooldown for that
|
||||
# Auto account, so its catalog choosing the same account again proves nothing.
|
||||
unproven = str(route.get("credentialProfileId") or "") if (
|
||||
reason == "quota" and not account_intent and getattr(error, "status_code", 0)
|
||||
and getattr(getattr(error, "physical_attempt_capture", None), "state", None) == "settled"
|
||||
and not (problem_context.get("resetsAt") or problem_context.get("retryAfterMs"))) else ""
|
||||
with self.lock:
|
||||
self.waits[wait_id] = row
|
||||
if quota_wait:
|
||||
|
|
@ -610,6 +675,7 @@ class TaskModelWait:
|
|||
requested_model=native_model)
|
||||
if (catalog.get("source") == source
|
||||
and (not account or catalog.get("credentialProfileId") == account)
|
||||
and not (unproven and not account and catalog.get("credentialProfileId") == unproven)
|
||||
and any(item.get("id") == native_model for item in catalog.get("models", []))):
|
||||
resolution = "resource_available"
|
||||
return kwargs
|
||||
|
|
@ -656,7 +722,8 @@ def model_waitable(function: Callable | None = None, *, client_parameter: str =
|
|||
"""Catch resource refusals before helper catches; callers may decline waiting.
|
||||
|
||||
``wait_for_resources`` is call-local, leaving the shared task's overrides,
|
||||
controls and custody intact even when a refusal returns immediately.
|
||||
controls and custody intact even when a refusal returns immediately; so is a
|
||||
round's ``defer_resource_wait``, which retains the refusal for that round.
|
||||
"""
|
||||
if function is None:
|
||||
return functools.partial(model_waitable, client_parameter=client_parameter)
|
||||
|
|
@ -698,15 +765,23 @@ def model_waitable(function: Callable | None = None, *, client_parameter: str =
|
|||
poll._model_wait_original = original
|
||||
return {**values, "model_poll_control": poll}
|
||||
|
||||
def can_wait(context, error, values):
|
||||
def can_wait(context, receiver, error, values):
|
||||
"""Wait here, or raise; a deferring round first retains a quota refusal for its later routes."""
|
||||
from ouroboros.llm_claudexor import ClaudexorModelError
|
||||
|
||||
capture = getattr(error, "physical_attempt_capture", None)
|
||||
return bool(context and not context.closed and values.get("model_role")
|
||||
and values.get("wait_for_resources", True)
|
||||
and isinstance(error, ClaudexorModelError)
|
||||
and model_wait_reason(error)
|
||||
and getattr(capture, "state", None) in {"released", "settled"})
|
||||
reason = model_wait_reason(error)
|
||||
if not (context and not context.closed and values.get("model_role") and reason
|
||||
and isinstance(error, ClaudexorModelError)
|
||||
and getattr(capture, "state", None) in {"released", "settled"}):
|
||||
return False
|
||||
deferral = _deferral(values["model_role"])
|
||||
# Quota: account rotation, then every configured route, then the owner. Sign-in
|
||||
# keeps its own owner wait wherever waiting is possible.
|
||||
if deferral is not None and (reason != "auth" or not context.waits_allowed):
|
||||
deferral.retain(receiver, error, values)
|
||||
return False
|
||||
return bool(context.waits_allowed and values.get("wait_for_resources", True))
|
||||
|
||||
def merged(result, attempts):
|
||||
result[1]["ledger_attempt_ids"] = list(dict.fromkeys([*attempts, *result[1].get("ledger_attempt_ids", [])]))
|
||||
|
|
@ -751,7 +826,7 @@ def model_waitable(function: Callable | None = None, *, client_parameter: str =
|
|||
except Exception as error:
|
||||
error.model_role_route = route_projection(values)
|
||||
propagate_model_control(error)
|
||||
if not can_wait(context, error, values):
|
||||
if not can_wait(context, receiver, error, values):
|
||||
raise
|
||||
record_attempts(attempts, error)
|
||||
cancelled = threading.Event()
|
||||
|
|
@ -787,7 +862,7 @@ def model_waitable(function: Callable | None = None, *, client_parameter: str =
|
|||
except Exception as error:
|
||||
error.model_role_route = route_projection(values)
|
||||
propagate_model_control(error)
|
||||
if not can_wait(context, error, values):
|
||||
if not can_wait(context, receiver, error, values):
|
||||
raise
|
||||
record_attempts(attempts, error)
|
||||
values = context.wait(receiver, error, values)
|
||||
|
|
|
|||
|
|
@ -2,12 +2,15 @@
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import concurrent.futures
|
||||
import contextvars
|
||||
import hashlib
|
||||
import logging
|
||||
import os
|
||||
import threading
|
||||
import time
|
||||
from contextlib import contextmanager
|
||||
from contextlib import ExitStack, contextmanager
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Callable, Mapping, Sequence
|
||||
|
|
@ -17,6 +20,7 @@ from ouroboros.contracts.task_contract import attach_task_contract
|
|||
from ouroboros.presence_admission import PresenceAdmission
|
||||
from ouroboros.presence_authority import presence_ceiling_payload
|
||||
from ouroboros.task_results import (
|
||||
STATUS_CANCELLED,
|
||||
STATUS_COMPLETED,
|
||||
STATUS_FAILED,
|
||||
STATUS_INTERRUPTED,
|
||||
|
|
@ -24,6 +28,7 @@ from ouroboros.task_results import (
|
|||
is_reconciled_presence_placeholder,
|
||||
load_task_result,
|
||||
reopen_reconciled_presence_placeholder,
|
||||
task_result_path,
|
||||
)
|
||||
from ouroboros.utils import append_jsonl, atomic_write_json, iter_jsonl_objects, read_json_dict, utc_now_iso
|
||||
|
||||
|
|
@ -37,9 +42,11 @@ class PresenceTurnError(ValueError):
|
|||
field: str,
|
||||
*,
|
||||
attachment_manifest: Sequence[Mapping[str, Any]] = (),
|
||||
turn_ref: str = "",
|
||||
) -> None:
|
||||
self.code = str(code or "presence_turn_failed")
|
||||
self.field = str(field or "presence_turn")
|
||||
self.turn_ref = str(turn_ref or "")
|
||||
self.attachment_manifest = [
|
||||
dict(row) for row in attachment_manifest if isinstance(row, Mapping)
|
||||
]
|
||||
|
|
@ -154,8 +161,40 @@ def build_presence_result_event(task: dict[str, Any], text: str, ctx: Any, *, te
|
|||
}
|
||||
|
||||
|
||||
_GATE_POLL_SEC = 0.05
|
||||
|
||||
|
||||
def _try_exclusive(fd: int) -> bool:
|
||||
from ouroboros.platform_layer import file_lock_exclusive_nb
|
||||
|
||||
try:
|
||||
file_lock_exclusive_nb(fd)
|
||||
except OSError:
|
||||
return False
|
||||
return True
|
||||
|
||||
|
||||
class PresenceTurnLease:
|
||||
"""The gate resources one admitted turn holds; its owner releases them exactly once."""
|
||||
|
||||
def __init__(self, conversation_key: str, resources: ExitStack) -> None:
|
||||
self.conversation_key = conversation_key
|
||||
self._resources: ExitStack | None = resources
|
||||
self._lock = threading.Lock()
|
||||
|
||||
def release(self) -> None:
|
||||
with self._lock:
|
||||
resources, self._resources = self._resources, None
|
||||
if resources is not None:
|
||||
resources.close()
|
||||
|
||||
|
||||
class PresenceTurnGate:
|
||||
"""Cross-process cap plus one active turn for each conversation."""
|
||||
"""Cross-process cap plus one active turn for each conversation.
|
||||
|
||||
``run`` waits on a thread (the tool path); ``admit`` takes the same resources as a
|
||||
coroutine for the Host, whose queued turns must hold no thread while they wait.
|
||||
"""
|
||||
|
||||
def __init__(self, max_active: int = 2, *, state_root: Path | None = None) -> None:
|
||||
self._max_active = max(1, int(max_active))
|
||||
|
|
@ -165,11 +204,25 @@ class PresenceTurnGate:
|
|||
self._state_root = Path(state_root).resolve(strict=False) if state_root is not None else None
|
||||
self._claimed_slots: set[int] = set()
|
||||
|
||||
def _gate_file(self, name: str) -> Path:
|
||||
root = self._state_root / "presence_turn_gate"
|
||||
root.mkdir(parents=True, exist_ok=True)
|
||||
return root / name
|
||||
|
||||
def _conversation(self, conversation_key: str) -> tuple[str, threading.Lock]:
|
||||
key = str(conversation_key or "").strip()
|
||||
if not key:
|
||||
raise PresenceTurnError("presence_conversation_key_required", "conversation_key")
|
||||
with self._guard:
|
||||
return key, self._conversations.setdefault(key, threading.Lock())
|
||||
|
||||
def _conversation_file(self, key: str) -> Path:
|
||||
return self._gate_file(f"conversation-{hashlib.sha256(key.encode('utf-8')).hexdigest()}.lock")
|
||||
|
||||
@contextmanager
|
||||
def _file_lock(self, path: Path):
|
||||
from ouroboros.platform_layer import file_lock_exclusive, file_unlock
|
||||
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
fd = os.open(path, os.O_CREAT | os.O_RDWR, 0o600)
|
||||
try:
|
||||
file_lock_exclusive(fd)
|
||||
|
|
@ -180,58 +233,83 @@ class PresenceTurnGate:
|
|||
finally:
|
||||
os.close(fd)
|
||||
|
||||
@contextmanager
|
||||
def _file_slot(self):
|
||||
from ouroboros.platform_layer import file_lock_exclusive_nb, file_unlock
|
||||
|
||||
def _try_slot(self) -> Callable[[], None] | None:
|
||||
"""Claim one free active slot without waiting: its release, or None while all are taken."""
|
||||
if self._state_root is None:
|
||||
with self._slots:
|
||||
yield
|
||||
return
|
||||
slot_root = self._state_root / "presence_turn_gate"
|
||||
slot_root.mkdir(parents=True, exist_ok=True)
|
||||
while True:
|
||||
for index in range(self._max_active):
|
||||
return self._slots.release if self._slots.acquire(blocking=False) else None
|
||||
from ouroboros.platform_layer import file_unlock
|
||||
|
||||
for index in range(self._max_active):
|
||||
with self._guard:
|
||||
if index in self._claimed_slots:
|
||||
continue
|
||||
self._claimed_slots.add(index)
|
||||
try:
|
||||
fd = os.open(self._gate_file(f"slot-{index}.lock"), os.O_CREAT | os.O_RDWR, 0o600)
|
||||
except BaseException: # an unopenable slot is an error, never a slot left claimed
|
||||
with self._guard:
|
||||
if index in self._claimed_slots:
|
||||
continue
|
||||
self._claimed_slots.add(index)
|
||||
path = slot_root / f"slot-{index}.lock"
|
||||
fd = os.open(path, os.O_CREAT | os.O_RDWR, 0o600)
|
||||
self._claimed_slots.discard(index)
|
||||
raise
|
||||
if not _try_exclusive(fd): # another process holds it
|
||||
os.close(fd)
|
||||
with self._guard:
|
||||
self._claimed_slots.discard(index)
|
||||
continue
|
||||
|
||||
def release(index: int = index, fd: int = fd) -> None:
|
||||
try:
|
||||
file_lock_exclusive_nb(fd)
|
||||
except OSError:
|
||||
file_unlock(fd)
|
||||
finally:
|
||||
os.close(fd)
|
||||
with self._guard:
|
||||
self._claimed_slots.discard(index)
|
||||
continue
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
try:
|
||||
file_unlock(fd)
|
||||
finally:
|
||||
os.close(fd)
|
||||
with self._guard:
|
||||
self._claimed_slots.discard(index)
|
||||
return
|
||||
time.sleep(0.05)
|
||||
|
||||
return release
|
||||
return None
|
||||
|
||||
@contextmanager
|
||||
def _file_slot(self):
|
||||
while (release := self._try_slot()) is None:
|
||||
time.sleep(_GATE_POLL_SEC)
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
release()
|
||||
|
||||
def run(self, conversation_key: str, callback: Callable[[], PresenceTurnResult]) -> PresenceTurnResult:
|
||||
key = str(conversation_key or "").strip()
|
||||
if not key:
|
||||
raise PresenceTurnError("presence_conversation_key_required", "conversation_key")
|
||||
with self._guard:
|
||||
conversation_lock = self._conversations.setdefault(key, threading.Lock())
|
||||
key, conversation_lock = self._conversation(conversation_key)
|
||||
with conversation_lock:
|
||||
if self._state_root is None:
|
||||
with self._slots:
|
||||
return callback()
|
||||
with self._file_lock(self._conversation_file(key)):
|
||||
with self._file_slot():
|
||||
return callback()
|
||||
digest = hashlib.sha256(key.encode("utf-8")).hexdigest()
|
||||
conversation_path = self._state_root / "presence_turn_gate" / f"conversation-{digest}.lock"
|
||||
with self._file_lock(conversation_path):
|
||||
with self._file_slot():
|
||||
return callback()
|
||||
|
||||
async def admit(self, conversation_key: str) -> PresenceTurnLease:
|
||||
"""Take what ``run`` takes, in the same order, as a coroutine.
|
||||
|
||||
Every attempt is non-blocking and every wait is an asyncio sleep, so a turn queued
|
||||
behind its conversation or the active cap parks no thread. Cancellation releases
|
||||
whatever was already taken; success hands it all to the returned lease.
|
||||
"""
|
||||
from ouroboros.platform_layer import file_unlock
|
||||
|
||||
key, conversation_lock = self._conversation(conversation_key)
|
||||
with ExitStack() as held:
|
||||
while not conversation_lock.acquire(blocking=False):
|
||||
await asyncio.sleep(_GATE_POLL_SEC)
|
||||
held.callback(conversation_lock.release)
|
||||
if self._state_root is not None:
|
||||
fd = os.open(self._conversation_file(key), os.O_CREAT | os.O_RDWR, 0o600)
|
||||
held.callback(os.close, fd)
|
||||
while not _try_exclusive(fd):
|
||||
await asyncio.sleep(_GATE_POLL_SEC)
|
||||
held.callback(file_unlock, fd)
|
||||
while (release_slot := self._try_slot()) is None:
|
||||
await asyncio.sleep(_GATE_POLL_SEC)
|
||||
held.callback(release_slot)
|
||||
return PresenceTurnLease(key, held.pop_all())
|
||||
|
||||
|
||||
_GATES_LOCK = threading.Lock()
|
||||
|
|
@ -260,25 +338,90 @@ def _configured_gate(drive_root: Path | None = None) -> PresenceTurnGate:
|
|||
return _GATES.setdefault(key, PresenceTurnGate(limit, state_root=state_root))
|
||||
|
||||
|
||||
async def admit_configured_gate(drive_root: Path, conversation_key: str) -> PresenceTurnLease:
|
||||
"""Coroutine admission into the configured gate ``run_presence_turn`` would otherwise take."""
|
||||
return await _configured_gate(Path(drive_root)).admit(conversation_key)
|
||||
|
||||
|
||||
def _stable_numeric_id(prefix: str, value: str) -> int:
|
||||
digest = hashlib.sha256(f"{prefix}\0{value}".encode("utf-8")).digest()
|
||||
return (1 << 40) + (int.from_bytes(digest[:6], "big") & ((1 << 40) - 1))
|
||||
|
||||
|
||||
def _task_id(admission: PresenceAdmission, event: PresenceTurnEvent) -> str:
|
||||
digest = hashlib.sha256(
|
||||
f"{admission.binding_id}\0{event.source_event_id}".encode("utf-8")
|
||||
).hexdigest()
|
||||
def presence_turn_task_id(binding_id: str, source_event_id: str) -> str:
|
||||
"""The durable id of one turn; the Host joins a retry to its live execution by it."""
|
||||
digest = hashlib.sha256(f"{binding_id}\0{source_event_id}".encode("utf-8")).hexdigest()
|
||||
return f"presence-{digest[:24]}"
|
||||
|
||||
|
||||
def _task_id(admission: PresenceAdmission, event: PresenceTurnEvent) -> str:
|
||||
return presence_turn_task_id(admission.binding_id, event.source_event_id)
|
||||
|
||||
|
||||
def _stored_turn(drive_root: Path, task_id: str) -> dict[str, Any]:
|
||||
"""Read Presence authority without converting an unreadable row into a new event.
|
||||
|
||||
A generic fail-soft read can quarantine a malformed task result and return None.
|
||||
That is correct for a list, but a transport retry must not treat an already
|
||||
admitted event as never started. Quarantine keeps the original id occupied.
|
||||
"""
|
||||
from ouroboros.task_result_schema import TASK_RESULT_QUARANTINE_DIR
|
||||
|
||||
path = task_result_path(drive_root, task_id, create=False)
|
||||
quarantined = path.parent / TASK_RESULT_QUARANTINE_DIR
|
||||
try:
|
||||
if (quarantined / path.name).exists() or any(quarantined.glob(f"{task_id}.*.json")):
|
||||
raise PresenceTurnError("presence_result_unreadable", "source_event_id", turn_ref=task_id)
|
||||
return load_task_result(drive_root, task_id, strict=True) or {}
|
||||
except (OSError, ValueError) as exc:
|
||||
if isinstance(exc, PresenceTurnError):
|
||||
raise
|
||||
raise PresenceTurnError("presence_result_unreadable", "source_event_id", turn_ref=task_id) from exc
|
||||
|
||||
|
||||
def _cached_result(drive_root: Path, task_id: str) -> PresenceTurnResult | None:
|
||||
stored = load_task_result(drive_root, task_id) or {}
|
||||
stored = _stored_turn(drive_root, task_id)
|
||||
if str(stored.get("status") or "") not in {"completed", "failed"} or is_reconciled_presence_placeholder(stored):
|
||||
return None # a host-lost turn is not a result: the transport's retry runs it again
|
||||
return None # a host-lost turn is not a result; the later admission guard refuses regeneration
|
||||
return presence_result_from_stored(stored, task_id)
|
||||
|
||||
|
||||
def _notify_unresolved_turn(drive_root: Path, task_id: str) -> None:
|
||||
"""Tell the owner once without speaking to the correspondent or granting a retry.
|
||||
|
||||
The running result is the existing recovery record. A chat write precedes the
|
||||
best-effort notified stamp: a crash in between can repeat the question, never
|
||||
silently mark an undelivered question as delivered. This is not a new work queue.
|
||||
"""
|
||||
stored = load_task_result(drive_root, task_id) or {}
|
||||
if stored.get("presence_recovery_owner_notified"):
|
||||
return
|
||||
try:
|
||||
from ouroboros.contracts.chat_id_policy import WEB_UI_CHAT_ID
|
||||
from supervisor import message_bus
|
||||
|
||||
# Unit callers and unrelated installations must not write via a process-global
|
||||
# bridge bound to another data root. The real Host shares this canonical root.
|
||||
if message_bus.DATA_DIR is None or Path(message_bus.DATA_DIR).resolve() != drive_root.resolve():
|
||||
return
|
||||
if message_bus.try_get_bridge() is None:
|
||||
return
|
||||
message_bus.send_with_budget(
|
||||
WEB_UI_CHAT_ID,
|
||||
f"Presence turn {task_id} has an unconfirmed previous effect. Its original "
|
||||
"event remains with the transport; I will not start another model or tool "
|
||||
"attempt automatically. Please answer in Main after checking the prior "
|
||||
"operation and external delivery, so we can decide how to recover it.",
|
||||
role="system", system_type="presence_recovery_required", require_write=True,
|
||||
)
|
||||
from ouroboros.task_results import write_task_result
|
||||
|
||||
write_task_result(drive_root, task_id, str(stored["status"]), strict_existing_dict=True,
|
||||
presence_recovery_owner_notified=utc_now_iso())
|
||||
except Exception:
|
||||
log.warning("Presence recovery question could not be recorded for %s", task_id, exc_info=True)
|
||||
|
||||
|
||||
def _live_task_rows(drive_root: Path, task_id: str, conversation_key: str) -> list[dict[str, Any]]:
|
||||
"""This task's rows in the live chat generation, in order; an attempt starts with its inbound row.
|
||||
|
||||
|
|
@ -365,6 +508,12 @@ def _pointer_behind(drive_root: Path, conversation_key: str, task_id: str) -> bo
|
|||
pointer.get("task_id") != task_id and str(pointer.get("finished_at") or "") <= str(stored.get("ts") or ""))
|
||||
|
||||
|
||||
def presence_turn_replay(drive_root: Path, task_id: str, conversation_key: str) -> PresenceTurnResult | None:
|
||||
"""A settled turn's durable answer, returned without the gate; None when the turn must (re)run."""
|
||||
cached = _cached_result(Path(drive_root), task_id)
|
||||
return None if cached is None or _pointer_behind(Path(drive_root), conversation_key, task_id) else cached
|
||||
|
||||
|
||||
def _delivery_state(reporting_version: int, sends: Sequence[str] | None, text: str) -> str:
|
||||
"""One rule for the live write and the repair.
|
||||
|
||||
|
|
@ -571,12 +720,17 @@ def run_presence_turn(
|
|||
event_queue: Any = None,
|
||||
agent_factory: Callable[..., Any] | None = None,
|
||||
gate: PresenceTurnGate | None = None,
|
||||
admitted: PresenceTurnLease | None = None,
|
||||
) -> PresenceTurnResult:
|
||||
"""Run one bounded turn; adapters retain durable provider custody."""
|
||||
"""Run one bounded turn; adapters retain durable provider custody.
|
||||
|
||||
``admitted`` is a lease the caller already took for this conversation (``PresenceTurnGate.admit``):
|
||||
the turn runs under it instead of waiting on ``gate``, and the caller keeps releasing it.
|
||||
"""
|
||||
|
||||
task_id = _task_id(admission, event)
|
||||
cached = _cached_result(Path(drive_root), task_id)
|
||||
if cached is not None and not _pointer_behind(Path(drive_root), event.conversation_key, task_id):
|
||||
cached = presence_turn_replay(Path(drive_root), task_id, event.conversation_key)
|
||||
if cached is not None:
|
||||
return cached
|
||||
|
||||
def execute() -> PresenceTurnResult:
|
||||
|
|
@ -605,9 +759,22 @@ def run_presence_turn(
|
|||
def _execute_live() -> PresenceTurnResult:
|
||||
# Both locks are held: no other execution of this conversation runs, so a running or
|
||||
# interrupted row of this task (not yet reconciled) belongs to a lost attempt too.
|
||||
stored = load_task_result(Path(drive_root), task_id) or {}
|
||||
stored = _stored_turn(Path(drive_root), task_id)
|
||||
if str(stored.get("status") or "") == STATUS_CANCELLED:
|
||||
# An owner's Stop is not an absent turn. Never regenerate work after it,
|
||||
# nor acknowledge the original transport event as completed/silent.
|
||||
raise PresenceTurnError("presence_turn_cancelled", "source_event_id", turn_ref=task_id)
|
||||
lost_attempt = is_reconciled_presence_placeholder(stored) or str(stored.get("status") or "") in {
|
||||
STATUS_RUNNING, STATUS_INTERRUPTED}
|
||||
if lost_attempt:
|
||||
# A dead local stack is not evidence the provider or an external tool never acted.
|
||||
# In particular a dispatched operation administratively settled as "abandoned"
|
||||
# retains an unknown effect. Retain the original event with its transport, and
|
||||
# leave this row untouched: no new inbound row, agent, or paid generation. A
|
||||
# positively pre-thread failure wrote no running row and reaches the normal path;
|
||||
# a later canonical terminal result wins the cached-replay check above.
|
||||
_notify_unresolved_turn(Path(drive_root), task_id)
|
||||
raise PresenceTurnError("presence_attempt_outcome_unknown", "source_event_id", turn_ref=task_id)
|
||||
def chat_generation() -> tuple[int, int] | None: # rotation renames the live file: a new inode
|
||||
try:
|
||||
stat = os.stat(Path(drive_root) / "logs" / "chat.jsonl")
|
||||
|
|
@ -691,15 +858,161 @@ def run_presence_turn(
|
|||
delivery=_delivery_state(event.delivery_reporting_version, sends, result.text))
|
||||
return result
|
||||
|
||||
if admitted is not None:
|
||||
if admitted.conversation_key != str(event.conversation_key or "").strip():
|
||||
raise PresenceTurnError("presence_admission_conversation_mismatch", "conversation_key")
|
||||
return execute()
|
||||
return (gate or _configured_gate(Path(drive_root))).run(event.conversation_key, execute)
|
||||
|
||||
|
||||
class PresenceTurnNotStarted(RuntimeError):
|
||||
"""Nothing ran: admission was interrupted or no thread could start; retry the same event."""
|
||||
|
||||
|
||||
class PresenceTurnExecution:
|
||||
"""One live host turn: the result every waiter shares and the admission that starts it."""
|
||||
|
||||
__slots__ = ("turn_id", "result", "admission")
|
||||
|
||||
def __init__(self, turn_id: str) -> None:
|
||||
self.turn_id = turn_id
|
||||
self.result: concurrent.futures.Future = concurrent.futures.Future()
|
||||
# RUNNING from birth: a cancelled waiter's wrapper calls Future.cancel(), which a running
|
||||
# future refuses, so no HTTP waiter can cancel the result other waiters share.
|
||||
self.result.set_running_or_notify_cancel()
|
||||
self.admission: asyncio.Task | None = None
|
||||
|
||||
|
||||
class PresenceTurnExecutions:
|
||||
"""The Host's live presence turns by durable turn id: start one, or join the one already running.
|
||||
|
||||
A turn is work, not its request. It queues on the gate as a coroutine (no thread, shared or
|
||||
owned), then runs on its own daemon thread with the admitting request's context variables
|
||||
(settings, usage and wait scopes), as ``asyncio.to_thread`` carried them. ``reserve``d
|
||||
capacity returns only at true settlement: a cancelled or disconnected waiter leaves the turn
|
||||
and its capacity in place, and a retry of the same event joins it. A turn the process exits
|
||||
under stays host-lost for the transport's retry; the daemon thread never extends a drain.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._lock = threading.Lock()
|
||||
self._live: dict[str, PresenceTurnExecution] = {}
|
||||
|
||||
def live(self) -> list[str]:
|
||||
with self._lock:
|
||||
return list(self._live)
|
||||
|
||||
def start_or_join(
|
||||
self,
|
||||
turn_id: str,
|
||||
*,
|
||||
reserve: Callable[[], bool],
|
||||
release: Callable[[], None],
|
||||
admit: Callable[[], Any],
|
||||
run: Callable[[PresenceTurnLease], Any],
|
||||
) -> tuple[PresenceTurnExecution | None, bool]:
|
||||
"""``(execution, started)``; ``(None, False)`` when ``reserve`` refuses a new turn.
|
||||
|
||||
A live turn is joined before capacity is consulted: a retry never needs a new slot.
|
||||
Call from the event loop; ``admit`` is a coroutine function returning the lease.
|
||||
"""
|
||||
with self._lock:
|
||||
live = self._live.get(turn_id)
|
||||
if live is not None:
|
||||
return live, False
|
||||
if not reserve():
|
||||
return None, False
|
||||
execution = self._live[turn_id] = PresenceTurnExecution(turn_id)
|
||||
admission = self._admit_and_start(execution, admit, run, release)
|
||||
try:
|
||||
execution.admission = asyncio.get_running_loop().create_task(admission)
|
||||
except BaseException:
|
||||
admission.close()
|
||||
self._settle(execution, release, error=PresenceTurnNotStarted("presence turn admission did not start"))
|
||||
raise
|
||||
execution.admission.add_done_callback(lambda task: self._admission_ended(execution, release, task))
|
||||
return execution, True
|
||||
|
||||
def _admission_ended(self, execution: PresenceTurnExecution, release: Callable[[], None],
|
||||
task: asyncio.Task) -> None:
|
||||
# Cancellation lands only before the thread starts — in ``admit`` or before the first step,
|
||||
# where the coroutine's own handlers never run (e.g. a loop shutting down).
|
||||
if task.cancelled() and not execution.result.done():
|
||||
self._settle(execution, release, error=PresenceTurnNotStarted("presence turn admission was interrupted"))
|
||||
|
||||
async def _admit_and_start(self, execution: PresenceTurnExecution, admit: Callable[[], Any],
|
||||
run: Callable[[PresenceTurnLease], Any], release: Callable[[], None]) -> None:
|
||||
try:
|
||||
lease = await admit()
|
||||
except asyncio.CancelledError:
|
||||
raise # settled by _admission_ended
|
||||
except BaseException as exc:
|
||||
self._settle(execution, release, error=exc)
|
||||
if not isinstance(exc, Exception):
|
||||
raise
|
||||
return
|
||||
try:
|
||||
thread = threading.Thread(
|
||||
target=contextvars.copy_context().run, args=(self._execute, execution, run, lease, release),
|
||||
name=f"presence-turn-{execution.turn_id}", daemon=True,
|
||||
)
|
||||
thread.start()
|
||||
except BaseException as exc: # e.g. "can't start new thread": nothing ran, so nothing stays held
|
||||
self._release_lease(execution, lease)
|
||||
error = PresenceTurnNotStarted("presence turn thread could not start")
|
||||
error.__cause__ = exc
|
||||
self._settle(execution, release, error=error)
|
||||
if not isinstance(exc, Exception):
|
||||
raise
|
||||
|
||||
def _execute(self, execution: PresenceTurnExecution, run: Callable[[PresenceTurnLease], Any],
|
||||
lease: PresenceTurnLease, release: Callable[[], None]) -> None:
|
||||
outcome: Any = None
|
||||
error: BaseException | None = None
|
||||
try:
|
||||
outcome = run(lease)
|
||||
except BaseException as exc: # relayed unchanged to every waiter
|
||||
error = exc
|
||||
self._release_lease(execution, lease)
|
||||
self._settle(execution, release, result=outcome, error=error)
|
||||
|
||||
@staticmethod
|
||||
def _release_lease(execution: PresenceTurnExecution, lease: PresenceTurnLease) -> None:
|
||||
try:
|
||||
lease.release()
|
||||
except Exception: # every release callback still ran; settlement must not be lost with it
|
||||
log.warning("presence turn %s could not release its gate lease", execution.turn_id, exc_info=True)
|
||||
|
||||
def _settle(self, execution: PresenceTurnExecution, release: Callable[[], None], *,
|
||||
result: Any = None, error: BaseException | None = None) -> None:
|
||||
# Capacity first, then every waiter sees the outcome, then the id retires: a retry that
|
||||
# lands before retirement joins the settled result instead of paying for another turn.
|
||||
try:
|
||||
release()
|
||||
except Exception:
|
||||
log.warning("presence turn %s could not return its capacity", execution.turn_id, exc_info=True)
|
||||
if error is None:
|
||||
execution.result.set_result(result)
|
||||
else:
|
||||
execution.result.set_exception(error)
|
||||
with self._lock:
|
||||
if self._live.get(execution.turn_id) is execution:
|
||||
del self._live[execution.turn_id]
|
||||
|
||||
|
||||
__all__ = [
|
||||
"PresenceTurnError",
|
||||
"PresenceTurnEvent",
|
||||
"PresenceTurnExecution",
|
||||
"PresenceTurnExecutions",
|
||||
"PresenceTurnGate",
|
||||
"PresenceTurnLease",
|
||||
"PresenceTurnNotStarted",
|
||||
"PresenceTurnResult",
|
||||
"admit_configured_gate",
|
||||
"build_presence_result_event",
|
||||
"presence_turn_is_live",
|
||||
"presence_turn_replay",
|
||||
"presence_turn_task_id",
|
||||
"run_presence_turn",
|
||||
]
|
||||
|
|
|
|||
|
|
@ -1385,7 +1385,8 @@ def send_with_budget(chat_id: int, text: str, log_text: Optional[str] = None,
|
|||
progress_meta: Optional[Dict[str, Any]] = None,
|
||||
ts: Optional[str] = None,
|
||||
role: str = "", system_type: str = "",
|
||||
narration: Optional[bool] = None) -> None:
|
||||
narration: Optional[bool] = None,
|
||||
require_write: bool = False) -> None:
|
||||
"""Send one owner-visible message through the shared host seam.
|
||||
|
||||
``narration`` is the note's VOICE, the same typed fact the worker stamps on
|
||||
|
|
@ -1440,6 +1441,7 @@ def send_with_budget(chat_id: int, text: str, log_text: Optional[str] = None,
|
|||
task_id=task_id,
|
||||
record_type=system_type,
|
||||
message_meta=progress_meta,
|
||||
require_write=require_write,
|
||||
)
|
||||
|
||||
if _text.strip() in ("", "\u200b"):
|
||||
|
|
|
|||
|
|
@ -131,6 +131,10 @@ SCENARIOS = {
|
|||
"S30": ("exact budget pause -> owner Resume, MANAGED root: priced rounds hit the global ledger fence, the task parks nonterminal under its SAME id (durable paused row + PENDING _budget_pause carrier, no task_done, no paid wrap-up, /api/state budget_paused); Resume while exhausted is the typed 409; raising TOTAL_BUDGET alone wakes nothing (bounded watch); Resume after the increase mints ONE single-use grant, the loop consumes it, the task completes with cumulative rounds/spend and task_done exactly once; a repeated Resume is the typed 404", LANE_MOCK),
|
||||
"S31": ("exact budget pause -> owner Resume, DIRECT owner-chat turn: a WS chat turn hits its graceful in-task ceiling mid-turn, the live actor ends and its own record parks inline as the PENDING _budget_pause carrier with _is_direct_chat under the SAME id (no second actor, no task_done, no paid wrap-up; census direct_chat/budget_paused); a budget increase wakes nothing; Resume mints one grant, a pooled worker continues the checkpoint with the Q10 threshold refresh, and the turn concludes in the chat (durable chat row + keyed final frame over the same /ws); a repeated Resume is the typed 404", LANE_MOCK),
|
||||
"S32": ("exact budget pause survives the physical epoch: a paused MANAGED root rides a GRACEFUL server SIGTERM (the lifespan teardown's kill_workers ran: server_shutdown row) untouched — no task_done, no cancel, the same paused row and PENDING _budget_pause carrier in the final snapshot; the next boot parks the SAME id again unheld and un-dispatched (census budget_paused, public detail scheduled/budget_paused, Resume still the typed 409 while exhausted); after the increase one Resume grant completes the same id with cumulative rounds/spend", LANE_MOCK),
|
||||
# Presence resilience: Host responsiveness and turn custody under real waits.
|
||||
# FOUR tests: executor starvation alone, a held Host authentication alone,
|
||||
# both together, and the owner's Panic under the combined load.
|
||||
"S33": ("Presence waits keep owner controls answering: 12 events (2 held at the model, slot and same-conversation waits) on a 12-thread default executor and/or one held Host authentication; while held, /api/state, a v1 receipt and the owner's Stop (durable cancel intent, cancelled terminal) answer inside 10s/15s windows (health alone never passes); a disconnected turn and its retry are ONE model call and a later replay answers the same projection; Panic under the combined load ends the whole tree", LANE_MOCK),
|
||||
}
|
||||
|
||||
MOCK_SLUG = "openai-compatible::mock-model"
|
||||
|
|
|
|||
868
tests/system_e2e/test_presence_resilience.py
Normal file
868
tests/system_e2e/test_presence_resilience.py
Normal file
|
|
@ -0,0 +1,868 @@
|
|||
"""S33 — Presence under real waits keeps the owner's controls and receipts
|
||||
answering, and a retried or abandoned turn stays ONE physical turn, as a REAL
|
||||
consumer of the candidate: the working tree copied by ``candidate_checkout``,
|
||||
served by a keyless isolated server, driven over the SAME Host Service and
|
||||
main-app HTTP surfaces a transport and the SPA use, against the scripted
|
||||
loopback model.
|
||||
|
||||
THE TWO MECHANISMS (reproduced on this harness in the sprint's isolated repro):
|
||||
|
||||
1. EXECUTOR STARVATION. A Presence turn waiting on the model, on a Presence slot
|
||||
or on its conversation's turn kept an ``asyncio.to_thread`` worker, and the
|
||||
Host Service shares the main app's event loop and so its ONE default
|
||||
executor. Twelve waiting turns on a twelve-thread executor stalled
|
||||
``/api/state``, Host ``/presence/delivery`` receipts and the owner's Stop (no
|
||||
cancel intent was minted; the task later *completed*), while ``/api/health``
|
||||
— pure async — kept answering 200 in 2 ms.
|
||||
2. SYNCHRONOUS HOST AUTH ON THE LOOP. Token authentication, permission and
|
||||
admission ran synchronously ON the event loop, so one slow authentication
|
||||
froze every request of both apps, health included.
|
||||
|
||||
HOW THE WAITS ARE BUILT (event-gated, never a timed race):
|
||||
|
||||
* the server tree's default executor is pinned to 12 threads (the incident
|
||||
host's ``min(32, cpu + 4)``) by a child-only ``sitecustomize`` (``_HOOK_SOURCE``);
|
||||
* ``OUROBOROS_PRESENCE_MAX_ACTIVE=2``: e00/e01 are HELD at the loopback model by
|
||||
``ModelGate``s (the two active turns), e02-e05 open four more rooms (slot
|
||||
waits), e06-e11 are each room's second event (conversation waits);
|
||||
* e00's HTTP client DISCONNECTS while its turn is at the model and a
|
||||
transport-style retry of the SAME event follows (the 13th request);
|
||||
* the slow authentication is the same hook: it reports every
|
||||
``HostServiceContext.authenticate_token_payload`` to the test's loopback
|
||||
``HostAuthGate``, which HOLDS one skill's call until released (the repro needed
|
||||
4 000-file skill payloads for the same effect).
|
||||
|
||||
WHAT IS ASSERTED. ``/api/health`` answering is never a pass; it is only recorded
|
||||
as the diagnosis that tells a blocked loop from a starved executor.
|
||||
|
||||
* WHILE every wait is still held, inside the fixture's client windows (10 s
|
||||
reads, 15 s Stop): ``/api/state`` answers a ready snapshot and a fresh v1
|
||||
receipt is recorded — two rounds, so one lucky free thread cannot pass — and
|
||||
the owner's Stop of a held pooled task answers ``ok`` with its durable
|
||||
``requested`` cancel intent and a ``cancelled`` terminal;
|
||||
* with only the auth held: the same, plus another skill's complete turn;
|
||||
* after release: every event answers ``message`` with its OWN reply, the retry
|
||||
answers e00's reply, a later replay answers the identical projection, the model
|
||||
saw every event EXACTLY once, and every recorded receipt is a durable chat row;
|
||||
* the owner's Panic under the combined load answers inside the read window and
|
||||
the whole server tree is gone, with the panic exit code and flag.
|
||||
|
||||
KNOWN LIMITS. The knob resizes only DEFAULT-sized pools: a candidate that sizes
|
||||
the loop's default executor explicitly fails the premise by name instead of
|
||||
passing on a bigger pool. The hold models a slow AUTHENTICATION; permission and
|
||||
admission are exercised on the same request path but never held separately. The
|
||||
stubs prove Host wiring only, never a model's or a transport's behaviour.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import http.client
|
||||
import json
|
||||
import pathlib
|
||||
import re
|
||||
import socket
|
||||
import subprocess
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
import urllib.parse
|
||||
import uuid
|
||||
from dataclasses import dataclass
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
import pytest
|
||||
|
||||
from tests.candidate_checkout import CandidateError, candidate_checkout, require_candidate_interpreter
|
||||
from tests.system_e2e.harness import (
|
||||
LANE_MOCK,
|
||||
REPO_ROOT,
|
||||
ArtifactOracle,
|
||||
KeylessIsolatedServer,
|
||||
ModelGate,
|
||||
ScriptedStubModel,
|
||||
assert_settings_keyless,
|
||||
clone_repo,
|
||||
keyless_settings,
|
||||
message_text,
|
||||
require_lane,
|
||||
submit_running,
|
||||
wait_until,
|
||||
write_settings_file,
|
||||
)
|
||||
|
||||
S33_DEFAULT_EXECUTOR_WORKERS = 12 # the incident host; one held event per thread
|
||||
S33_MAX_ACTIVE = 2
|
||||
S33_READ_WINDOW_SEC = 10.0 # /api/state and a receipt must answer inside this
|
||||
S33_STOP_WINDOW_SEC = 15.0 # the owner's Stop must answer inside this
|
||||
S33_TURN_WINDOW_SEC = 60.0 # a complete unheld turn while only the auth is held
|
||||
S33_HOLD_SEC = 240.0 # every gate's bound: a broken run fails, never hangs
|
||||
S33_SETTLE_SEC = 180.0 # after release, every held request answers inside this
|
||||
S33_PANIC_EXIT_SEC = 60.0
|
||||
|
||||
PROVIDER = "s33chat"
|
||||
ACCOUNT = "s33-account"
|
||||
BEHAVIOR_SKILL = "s33-presence-profile"
|
||||
PROBE_SKILL = "s33-probe"
|
||||
SLOW_SKILL = "s33-slow"
|
||||
HELD_TRANSPORTS = tuple(f"s33-t{i}" for i in range(4)) # <= 4 in flight each (the Host caps 5)
|
||||
CONTROL_MARKER = "[S33-CONTROL-TASK]"
|
||||
# The product's default install ports: Panic sweeps its BOUND port, so the run
|
||||
# refuses to Panic unless the server provably bound somewhere else.
|
||||
LIVE_DEFAULT_PORTS = frozenset({8765, 8766, 8767})
|
||||
|
||||
# The host frames THIS turn's input with its exact source facts
|
||||
# (presence_context.frame_presence_user_content: "Source facts: {json}"); the
|
||||
# event id is ``<provider>:<run>:<label>``, so a model call names its event
|
||||
# exactly, and a reply or an earlier turn quoted into context never does.
|
||||
_SOURCE_EVENT_RE = re.compile(r'"source_event_id": "' + re.escape(PROVIDER) + r':([0-9a-f]{8}):([a-z0-9]+)"')
|
||||
|
||||
# Loaded ONLY into the isolated server tree (a ``sitecustomize`` on the child's
|
||||
# PYTHONPATH), never into this process. Each knob is off unless its env var is set:
|
||||
# S33_DEFAULT_EXECUTOR_WORKERS — every DEFAULT-sized ThreadPoolExecutor (the
|
||||
# loop's lazily created default executor: what ``asyncio.to_thread`` uses)
|
||||
# gets this many threads; explicitly sized pools are untouched; every
|
||||
# construction is logged to S33_EXECUTOR_LOG as proof of the applied size.
|
||||
# S33_AUTH_GATE_URL — ``HostServiceContext.authenticate_token_payload`` reports
|
||||
# each authenticated Host request (skill, and whether it ran ON the event-loop
|
||||
# thread) to the test's HostAuthGate, which may hold it: the event-gated
|
||||
# stand-in for a slow synchronous authentication.
|
||||
_HOOK_SOURCE = r'''"""S33 fixture hook: isolated server tree only (tests/system_e2e/test_presence_resilience.py)."""
|
||||
import os
|
||||
import sys
|
||||
|
||||
_WORKERS = os.environ.get("S33_DEFAULT_EXECUTOR_WORKERS", "").strip()
|
||||
if _WORKERS:
|
||||
import json
|
||||
import concurrent.futures.thread as _cft
|
||||
|
||||
_original_init = _cft.ThreadPoolExecutor.__init__
|
||||
|
||||
def _init(self, max_workers=None, *args, **kwargs):
|
||||
forced = max_workers is None
|
||||
_original_init(self, int(_WORKERS) if forced else max_workers, *args, **kwargs)
|
||||
log_path = os.environ.get("S33_EXECUTOR_LOG", "").strip()
|
||||
if log_path:
|
||||
try:
|
||||
with open(log_path, "a", encoding="utf-8") as fh:
|
||||
fh.write(json.dumps({"pid": os.getpid(), "forced_default": forced,
|
||||
"max_workers": self._max_workers}) + "\n")
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
_cft.ThreadPoolExecutor.__init__ = _init
|
||||
|
||||
_GATE = os.environ.get("S33_AUTH_GATE_URL", "").strip()
|
||||
if _GATE:
|
||||
import importlib.abc
|
||||
import importlib.machinery
|
||||
|
||||
def _report(skill):
|
||||
import asyncio
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
|
||||
try:
|
||||
asyncio.get_running_loop()
|
||||
on_loop = "1"
|
||||
except RuntimeError:
|
||||
on_loop = "0"
|
||||
query = urllib.parse.urlencode({"skill": skill, "on_loop": on_loop})
|
||||
opener = urllib.request.build_opener(urllib.request.ProxyHandler({}))
|
||||
try:
|
||||
opener.open(_GATE + "/auth?" + query, timeout=300).read()
|
||||
except Exception:
|
||||
pass # a vanished gate must never change what the product does
|
||||
|
||||
def _patch(module):
|
||||
context = getattr(module, "HostServiceContext", None)
|
||||
original = getattr(context, "authenticate_token_payload", None)
|
||||
if original is None:
|
||||
return
|
||||
|
||||
def authenticate_token_payload(self, raw_token):
|
||||
skill, payload = original(self, raw_token)
|
||||
_report(skill)
|
||||
return skill, payload
|
||||
|
||||
context.authenticate_token_payload = authenticate_token_payload
|
||||
|
||||
class _Finder(importlib.abc.MetaPathFinder):
|
||||
def find_spec(self, name, path=None, target=None):
|
||||
if name != "ouroboros.gateway.host_service":
|
||||
return None
|
||||
spec = importlib.machinery.PathFinder.find_spec(name, path)
|
||||
if spec is None or spec.loader is None:
|
||||
return spec
|
||||
exec_module = spec.loader.exec_module
|
||||
|
||||
def patched(module):
|
||||
exec_module(module)
|
||||
_patch(module)
|
||||
|
||||
spec.loader.exec_module = patched
|
||||
return spec
|
||||
|
||||
sys.meta_path.insert(0, _Finder())
|
||||
'''
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Loopback gates
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
class HostAuthGate:
|
||||
"""The test side of the auth hook: records every authenticated Host request
|
||||
and HOLDS the first one of ``hold_skill`` until ``release`` (bounded by
|
||||
``timeout``, like ``ModelGate``), so "authentication is slow right now" is a
|
||||
state the scenario controls instead of a latency it races."""
|
||||
|
||||
def __init__(self, *, timeout: float = S33_HOLD_SEC) -> None:
|
||||
self.timeout = float(timeout)
|
||||
self.hold_skill = ""
|
||||
self.arrived = threading.Event()
|
||||
self.release = threading.Event()
|
||||
self.timed_out = False
|
||||
self.calls: list = [] # (skill, ran_on_event_loop) in arrival order
|
||||
self._cv = threading.Condition()
|
||||
outer = self
|
||||
|
||||
class _Handler(BaseHTTPRequestHandler):
|
||||
def do_GET(self): # noqa: N802 - stdlib callback name
|
||||
query = urllib.parse.parse_qs(urllib.parse.urlsplit(self.path).query)
|
||||
skill = (query.get("skill") or [""])[0]
|
||||
hold = False
|
||||
with outer._cv:
|
||||
outer.calls.append((skill, (query.get("on_loop") or ["0"])[0] == "1"))
|
||||
if skill and skill == outer.hold_skill and not outer.arrived.is_set():
|
||||
hold = True
|
||||
outer.arrived.set()
|
||||
outer._cv.notify_all()
|
||||
if hold and not outer.release.wait(outer.timeout):
|
||||
outer.timed_out = True
|
||||
self.send_response(204)
|
||||
self.send_header("Content-Length", "0")
|
||||
self.end_headers()
|
||||
|
||||
def log_message(self, *_args):
|
||||
return
|
||||
|
||||
self._server = ThreadingHTTPServer(("127.0.0.1", 0), _Handler)
|
||||
self._thread = threading.Thread(target=self._server.serve_forever, daemon=True)
|
||||
|
||||
@property
|
||||
def url(self) -> str:
|
||||
return f"http://127.0.0.1:{self._server.server_address[1]}"
|
||||
|
||||
def count(self, skills) -> int:
|
||||
with self._cv:
|
||||
return sum(1 for skill, _ in self.calls if skill in skills)
|
||||
|
||||
def wait_count(self, skills, count: int, timeout: float) -> bool:
|
||||
with self._cv:
|
||||
return self._cv.wait_for(
|
||||
lambda: sum(1 for skill, _ in self.calls if skill in skills) >= count, timeout)
|
||||
|
||||
def on_loop(self) -> dict:
|
||||
with self._cv:
|
||||
return {skill: on for skill, on in self.calls}
|
||||
|
||||
def __enter__(self):
|
||||
self._thread.start()
|
||||
return self
|
||||
|
||||
def __exit__(self, *_exc) -> None:
|
||||
self.release.set()
|
||||
self._server.shutdown()
|
||||
self._server.server_close()
|
||||
self._thread.join(timeout=10)
|
||||
|
||||
|
||||
def _user_texts(body: dict) -> list:
|
||||
return [message_text(m) for m in (body.get("messages") or [])
|
||||
if isinstance(m, dict) and m.get("role") == "user"]
|
||||
|
||||
|
||||
def _event_labels(run: str, body: dict) -> list:
|
||||
"""The S33 events a model call serves, read from the host's framing of the
|
||||
turn's own input (never from the system prompt, which may list other work)."""
|
||||
return sorted({label for text in _user_texts(body)
|
||||
for rid, label in _SOURCE_EVENT_RE.findall(text) if rid == run})
|
||||
|
||||
|
||||
def _event_id(run: str, label: str) -> str:
|
||||
return f"{PROVIDER}:{run}:{label}"
|
||||
|
||||
|
||||
def _reply(run: str, label: str) -> str:
|
||||
return f"S33 reply to event {label} of run {run}."
|
||||
|
||||
|
||||
class _Arrivals:
|
||||
"""Every model call's S33 event labels, recorded at ARRIVAL — before any
|
||||
hold — so a held call is counted while it is still in flight."""
|
||||
|
||||
def __init__(self, run: str) -> None:
|
||||
self.run = run
|
||||
self.rows: list = []
|
||||
self._lock = threading.Lock()
|
||||
|
||||
def __call__(self, body: dict) -> None:
|
||||
with self._lock:
|
||||
self.rows.append(_event_labels(self.run, body))
|
||||
|
||||
def per_event(self) -> dict:
|
||||
counts: dict = {}
|
||||
with self._lock:
|
||||
for labels in self.rows:
|
||||
for label in labels:
|
||||
counts[label] = counts.get(label, 0) + 1
|
||||
return counts
|
||||
|
||||
|
||||
class _AllGates:
|
||||
def __init__(self, *gates) -> None:
|
||||
self.gates = gates
|
||||
|
||||
def __call__(self, body: dict) -> None:
|
||||
for gate in self.gates:
|
||||
gate(body)
|
||||
|
||||
|
||||
def _reply_step(run: str):
|
||||
def step(body: dict) -> dict:
|
||||
labels = _event_labels(run, body)
|
||||
return {"final": _reply(run, labels[0] if len(labels) == 1 else "unmarked")}
|
||||
return step
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Seeding: the owner-side facts a transport skill and a Presence profile need
|
||||
# (the same helpers tests/test_host_service_api.py and test_presence_admission.py use)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class _Transport:
|
||||
name: str
|
||||
access: str # the X-Skill-Token value: a fixture string, never a credential
|
||||
binding_id: str
|
||||
|
||||
|
||||
def _seed_behavior(data_root: pathlib.Path) -> None:
|
||||
from ouroboros.presence_capabilities import (
|
||||
PresenceSelection,
|
||||
PresenceState,
|
||||
PresenceToolTarget,
|
||||
presence_state_fingerprint,
|
||||
save_presence_state,
|
||||
)
|
||||
from ouroboros.presence_profile import parse_presence_profile, presence_request_fingerprint
|
||||
from ouroboros.skill_loader import SkillReviewState, load_skill, save_enabled, save_review_state
|
||||
|
||||
skill_dir = data_root / "skills" / "external" / BEHAVIOR_SKILL
|
||||
skill_dir.mkdir(parents=True)
|
||||
(skill_dir / "SKILL.md").write_text(
|
||||
"---\n"
|
||||
f"name: {BEHAVIOR_SKILL}\n"
|
||||
"description: Neutral S33 presence fixture.\n"
|
||||
"version: 0.1.0\n"
|
||||
"type: instruction\n"
|
||||
"presence:\n"
|
||||
" instructions: Participate helpfully in the selected room. Answer in one short sentence.\n"
|
||||
" context_topics: [community-guidelines]\n"
|
||||
" runtime_defaults:\n"
|
||||
" model_slot: main\n"
|
||||
" inline_max_rounds: 4\n"
|
||||
" capability_requests:\n"
|
||||
" - id: history\n"
|
||||
" kind: tool\n"
|
||||
" required: true\n"
|
||||
" purpose: Read relevant room history.\n"
|
||||
"---\n"
|
||||
"# S33 presence profile\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
loaded = load_skill(skill_dir, data_root)
|
||||
assert loaded is not None and not loaded.load_error, getattr(loaded, "load_error", "not loaded")
|
||||
save_enabled(data_root, loaded.name, True)
|
||||
save_review_state(data_root, loaded.name, SkillReviewState(status="pass", content_hash=loaded.content_hash))
|
||||
profile = parse_presence_profile(loaded.manifest, skill_dir)
|
||||
assert profile is not None
|
||||
save_presence_state(
|
||||
data_root, loaded.name,
|
||||
PresenceState((PresenceSelection(
|
||||
presence_request_fingerprint(profile.capability_requests[0]),
|
||||
PresenceToolTarget("builtin", "chat_history"),
|
||||
),)),
|
||||
expected_state_fingerprint=presence_state_fingerprint(PresenceState()),
|
||||
)
|
||||
|
||||
|
||||
def _seed_transport(data_root: pathlib.Path, name: str) -> _Transport:
|
||||
from ouroboros.gateway.host_service import AUTH_TOKEN_FILENAME
|
||||
from ouroboros.presence_bindings import (
|
||||
PresenceBinding,
|
||||
PresenceEndpoint,
|
||||
new_presence_binding_id,
|
||||
save_presence_binding,
|
||||
)
|
||||
from ouroboros.skill_loader import (
|
||||
SkillReviewState,
|
||||
compute_content_hash,
|
||||
save_enabled,
|
||||
save_review_state,
|
||||
save_skill_grants,
|
||||
)
|
||||
from ouroboros.utils import atomic_write_json
|
||||
|
||||
skill_dir = data_root / "skills" / "external" / name
|
||||
skill_dir.mkdir(parents=True)
|
||||
(skill_dir / "SKILL.md").write_text(
|
||||
f"---\nname: {name}\ndescription: S33 transport fixture\nversion: 0.1\n"
|
||||
"type: extension\nentry: plugin.py\npermissions: [presence]\nsubscribe_events: []\n---\n# S33 transport\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
(skill_dir / "plugin.py").write_text("def register(api):\n pass\n", encoding="utf-8")
|
||||
content_hash = compute_content_hash(skill_dir, manifest_entry="plugin.py")
|
||||
save_review_state(data_root, name, SkillReviewState(status="pass", content_hash=content_hash))
|
||||
save_enabled(data_root, name, True)
|
||||
save_skill_grants(data_root, name, [], content_hash=content_hash, requested_keys=[],
|
||||
granted_permissions=["presence"], requested_permissions=["presence"])
|
||||
access = f"s33-{name}-{uuid.uuid4().hex}"
|
||||
atomic_write_json(data_root / "state" / "skills" / name / AUTH_TOKEN_FILENAME,
|
||||
{"token": access, "content_hash": content_hash, "issued_at": "s33"})
|
||||
binding = save_presence_binding(data_root, PresenceBinding(
|
||||
new_presence_binding_id(), name, BEHAVIOR_SKILL,
|
||||
PresenceEndpoint(PROVIDER, ACCOUNT, "*", ""), # account-wide: any room
|
||||
PresenceEndpoint(PROVIDER, ACCOUNT, "s33-room", ""),
|
||||
))
|
||||
return _Transport(name, access, binding.binding_id)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Wire helpers (http.client: no proxy lookup, and an explicit disconnect)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _call(port: int, method: str, path: str, payload=None, *, headers=None, timeout: float) -> dict:
|
||||
started = time.monotonic()
|
||||
conn = http.client.HTTPConnection("127.0.0.1", int(port), timeout=timeout)
|
||||
try:
|
||||
data = None if payload is None else json.dumps(payload).encode("utf-8")
|
||||
conn.request(method, path, body=data,
|
||||
headers={**({"Content-Type": "application/json"} if data is not None else {}),
|
||||
**(headers or {})})
|
||||
response = conn.getresponse()
|
||||
raw = response.read()
|
||||
try:
|
||||
body = json.loads(raw.decode("utf-8")) if raw.strip() else {}
|
||||
except ValueError:
|
||||
body = {"raw": raw[:300].decode("utf-8", "replace")}
|
||||
return {"status": response.status, "body": body if isinstance(body, dict) else {"value": body},
|
||||
"error": "", "sec": round(time.monotonic() - started, 3)}
|
||||
except TimeoutError:
|
||||
return {"status": 0, "body": {}, "error": "timeout", "sec": round(time.monotonic() - started, 3)}
|
||||
except (OSError, http.client.HTTPException) as exc:
|
||||
return {"status": 0, "body": {}, "error": f"{type(exc).__name__}: {exc}",
|
||||
"sec": round(time.monotonic() - started, 3)}
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
|
||||
def _event_body(transport: _Transport, event_id: str, room: str, text: str) -> dict:
|
||||
return {
|
||||
"binding_id": transport.binding_id,
|
||||
"delivery_reporting_version": 0,
|
||||
"event": {
|
||||
"source_event_id": event_id, "provider": PROVIDER, "account_id": ACCOUNT,
|
||||
"conversation_id": room, "thread_id": "",
|
||||
"conversation_key": "caller-controlled-key-is-ignored",
|
||||
"actor": {"platform_actor_id": "user-7", "username": "alex", "display_name": "Alex"},
|
||||
"conversation": {"title": "S33 room"}, "message": {"message_id": event_id[-8:]},
|
||||
"text": text,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _receipt(delivery_id: str) -> dict:
|
||||
return {
|
||||
"schema_version": 1, "delivery_id": delivery_id, "part_id": "0", "state": "delivered",
|
||||
"provider": PROVIDER, "account_id": ACCOUNT, "conversation_id": "s33-probe-room", "thread_id": "",
|
||||
"text": "S33 probe receipt", "format": "markdown", "message": {"message_id": delivery_id[:8]},
|
||||
"origin": {"kind": "automatic", "task_id": "", "source_event_id": ""},
|
||||
}
|
||||
|
||||
|
||||
class _Turn(threading.Thread):
|
||||
"""One transport request for one event, answered whenever the Host answers it."""
|
||||
|
||||
def __init__(self, port: int, transport: _Transport, body: dict, label: str) -> None:
|
||||
super().__init__(name=f"s33-turn-{label}", daemon=True)
|
||||
self.port, self.transport, self.body, self.label = port, transport, body, label
|
||||
self.result: dict = {}
|
||||
|
||||
def run(self) -> None:
|
||||
self.result = _call(self.port, "POST", "/presence/turn", self.body,
|
||||
headers={"X-Skill-Token": self.transport.access},
|
||||
timeout=S33_HOLD_SEC + S33_SETTLE_SEC)
|
||||
|
||||
|
||||
class _HookedServer(KeylessIsolatedServer):
|
||||
"""``KeylessIsolatedServer`` whose child additionally carries the S33 hook env."""
|
||||
|
||||
def __init__(self, clone, data_root, settings_path, *, extra_env: dict) -> None:
|
||||
super().__init__(clone, data_root, settings_path)
|
||||
self.extra_env = dict(extra_env)
|
||||
|
||||
def _env(self) -> dict:
|
||||
env = super()._env()
|
||||
env.update(self.extra_env)
|
||||
return env
|
||||
|
||||
|
||||
def _jsonl_file(path: pathlib.Path) -> list:
|
||||
if not path.exists():
|
||||
return []
|
||||
rows = []
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
try:
|
||||
rows.append(json.loads(line))
|
||||
except ValueError:
|
||||
continue
|
||||
return [row for row in rows if isinstance(row, dict)]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# The scenario driver
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _run_s33(candidate, root: pathlib.Path, *, held_turns: bool, slow_auth: bool, panic: bool = False) -> None:
|
||||
run = uuid.uuid4().hex[:8]
|
||||
root = pathlib.Path(root)
|
||||
data_root = root / "data"
|
||||
data_root.mkdir(parents=True)
|
||||
(root / "home").mkdir()
|
||||
hook_dir = root / "hook"
|
||||
hook_dir.mkdir()
|
||||
(hook_dir / "sitecustomize.py").write_text(_HOOK_SOURCE, encoding="utf-8")
|
||||
executor_log = root / "executor_pools.jsonl"
|
||||
|
||||
arrivals = _Arrivals(run)
|
||||
# The control task's OWN input (other turns' system prompts may list it).
|
||||
control_gate = ModelGate(lambda body: bool(body.get("tools")) and CONTROL_MARKER in "".join(_user_texts(body)[:1]),
|
||||
timeout=S33_HOLD_SEC)
|
||||
held = {label: ModelGate(lambda body, label=label: _event_labels(run, body) == [label],
|
||||
timeout=S33_HOLD_SEC)
|
||||
for label in ("e00", "e01")}
|
||||
model_gates = (control_gate, *held.values())
|
||||
turns: dict = {}
|
||||
receipts: list = []
|
||||
probes: dict = {}
|
||||
server = None
|
||||
with HostAuthGate() as auth_gate, ScriptedStubModel(
|
||||
[_reply_step(run)] * 400, gate=_AllGates(arrivals, *model_gates)) as stub:
|
||||
try:
|
||||
settings = keyless_settings(stub, OUROBOROS_PRESENCE_MAX_ACTIVE=S33_MAX_ACTIVE,
|
||||
OUROBOROS_MAX_WORKERS=2)
|
||||
assert_settings_keyless(settings)
|
||||
write_settings_file(data_root / "settings.json", settings)
|
||||
_seed_behavior(data_root)
|
||||
transports = {name: _seed_transport(data_root, name)
|
||||
for name in (*HELD_TRANSPORTS, SLOW_SKILL, PROBE_SKILL)}
|
||||
server = _HookedServer(candidate, data_root, data_root / "settings.json", extra_env={
|
||||
"HOME": str(root / "home"),
|
||||
"PYTHONPATH": str(hook_dir),
|
||||
"S33_DEFAULT_EXECUTOR_WORKERS": str(S33_DEFAULT_EXECUTOR_WORKERS),
|
||||
"S33_EXECUTOR_LOG": str(executor_log),
|
||||
"S33_AUTH_GATE_URL": auth_gate.url,
|
||||
# Inherited by every descendant: the Panic survivor scan's membership.
|
||||
"S33_TREE": f"s33-tree-{run}",
|
||||
})
|
||||
server.start(ready_timeout=300)
|
||||
host_port = server.host_service_port
|
||||
oracle = ArtifactOracle(server.data_root)
|
||||
|
||||
# PREMISE: the server process's default executor really is 12 threads.
|
||||
pools = [row for row in _jsonl_file(executor_log) if row.get("pid") == server.proc.pid]
|
||||
assert any(row.get("forced_default") and row.get("max_workers") == S33_DEFAULT_EXECUTOR_WORKERS
|
||||
for row in pools), (
|
||||
"premise not established: the server never built a DEFAULT-sized executor at "
|
||||
f"{S33_DEFAULT_EXECUTOR_WORKERS} threads (an explicitly sized loop executor bypasses "
|
||||
f"the knob): {pools}")
|
||||
|
||||
def diagnosis() -> str:
|
||||
on_loop = auth_gate.on_loop()
|
||||
return json.dumps({
|
||||
"health": probes.get("health"),
|
||||
"auth_ran_on_event_loop": on_loop,
|
||||
"auth_calls": len(auth_gate.calls),
|
||||
"model_arrivals": arrivals.per_event(),
|
||||
"answered_early": {label: turn.result for label, turn in turns.items()
|
||||
if not turn.is_alive()},
|
||||
"executor_pools": pools,
|
||||
}, default=str)
|
||||
|
||||
# A REAL pooled task held mid model call: the target of the owner's Stop.
|
||||
control_id = submit_running(server, f"{CONTROL_MARKER} Say hello in one line and finish.")
|
||||
assert control_gate.arrived.wait(120), f"the control task never reached the model: {stub.kinds()}"
|
||||
|
||||
if held_turns:
|
||||
rooms = [f"s33-room-{i}" for i in range(6)]
|
||||
plan = [(f"e{i:02d}", rooms[i % 6], transports[HELD_TRANSPORTS[i % 4]])
|
||||
for i in range(S33_DEFAULT_EXECUTOR_WORKERS)]
|
||||
bodies = {label: _event_body(transport, _event_id(run, label), room,
|
||||
f"S33 event {label}: say one short sentence.")
|
||||
for label, room, transport in plan}
|
||||
by_label = {label: transport for label, _room, transport in plan}
|
||||
|
||||
# e00: sent on a connection this test will abandon mid-turn.
|
||||
abandoned = http.client.HTTPConnection("127.0.0.1", host_port, timeout=S33_HOLD_SEC)
|
||||
abandoned.request("POST", "/presence/turn", body=json.dumps(bodies["e00"]).encode("utf-8"),
|
||||
headers={"Content-Type": "application/json",
|
||||
"X-Skill-Token": by_label["e00"].access})
|
||||
assert held["e00"].arrived.wait(120), f"e00 never reached the model: {diagnosis()}"
|
||||
turns["e01"] = _Turn(host_port, by_label["e01"], bodies["e01"], "e01")
|
||||
turns["e01"].start()
|
||||
assert held["e01"].arrived.wait(120), f"e01 never reached the model: {diagnosis()}"
|
||||
# The transport gives up on e00 while its turn is at the model ...
|
||||
abandoned.sock.shutdown(socket.SHUT_RDWR)
|
||||
abandoned.close()
|
||||
# ... ten more events queue behind the two active turns, and e00 is retried.
|
||||
for label, _room, transport in plan[2:]:
|
||||
turns[label] = _Turn(host_port, transport, bodies[label], label)
|
||||
turns[label].start()
|
||||
turns["e00-retry"] = _Turn(host_port, by_label["e00"], bodies["e00"], "e00-retry")
|
||||
turns["e00-retry"].start()
|
||||
assert auth_gate.wait_count(HELD_TRANSPORTS, len(plan) + 1, 60), (
|
||||
f"the Host authenticated only {auth_gate.count(HELD_TRANSPORTS)} of {len(plan) + 1} "
|
||||
f"held turn requests (a blocked event loop, or the auth seam moved): {diagnosis()}")
|
||||
assert arrivals.per_event() == {"e00": 1, "e01": 1}, (
|
||||
f"exactly the two active turns may be at the model: {diagnosis()}")
|
||||
|
||||
if slow_auth:
|
||||
auth_gate.hold_skill = SLOW_SKILL
|
||||
turns["slow"] = _Turn(host_port, transports[SLOW_SKILL], _event_body(
|
||||
transports[SLOW_SKILL], _event_id(run, "slow"), "s33-slow-room",
|
||||
"S33 event slow: say one short sentence."), "slow")
|
||||
turns["slow"].start()
|
||||
assert auth_gate.arrived.wait(60), (
|
||||
"premise not established: the auth hook never saw the slow skill's Host "
|
||||
"authentication (HostServiceContext.authenticate_token_payload moved?): "
|
||||
f"{diagnosis()}")
|
||||
|
||||
if panic:
|
||||
_assert_panic_tears_down(server, oracle, f"S33_TREE=s33-tree-{run}", diagnosis)
|
||||
return
|
||||
|
||||
# Recorded, never asserted: a blocked loop times this out, a starved
|
||||
# executor answers it in milliseconds.
|
||||
probes["health"] = _call(server.port, "GET", "/api/health", timeout=S33_READ_WINDOW_SEC)
|
||||
for round_index in range(2):
|
||||
state = _call(server.port, "GET", "/api/state", timeout=S33_READ_WINDOW_SEC)
|
||||
probes[f"state-{round_index}"] = state
|
||||
assert state["status"] == 200 and state["body"].get("supervisor_ready") is True, (
|
||||
f"/api/state did not answer a ready snapshot inside {S33_READ_WINDOW_SEC:.0f}s "
|
||||
f"while Presence waited (round {round_index}): {state}; {diagnosis()}")
|
||||
delivery_id = f"s33-{run}-receipt-{round_index}"
|
||||
recorded = _call(host_port, "POST", "/presence/delivery", _receipt(delivery_id),
|
||||
headers={"X-Skill-Token": transports[PROBE_SKILL].access},
|
||||
timeout=S33_READ_WINDOW_SEC)
|
||||
probes[f"receipt-{round_index}"] = recorded
|
||||
assert 200 <= recorded["status"] < 300 and recorded["body"].get("ok") is True \
|
||||
and recorded["body"].get("duplicate") is not True, (
|
||||
f"a fresh receipt was not recorded inside {S33_READ_WINDOW_SEC:.0f}s while "
|
||||
f"Presence waited (round {round_index}): {recorded}; {diagnosis()}")
|
||||
receipts.append(delivery_id)
|
||||
|
||||
if slow_auth and not held_turns:
|
||||
# Another skill's whole turn completes while one authentication is held.
|
||||
fast = _call(host_port, "POST", "/presence/turn", _event_body(
|
||||
transports[HELD_TRANSPORTS[0]], _event_id(run, "fast"), "s33-fast-room",
|
||||
"S33 event fast: say one short sentence."),
|
||||
headers={"X-Skill-Token": transports[HELD_TRANSPORTS[0]].access},
|
||||
timeout=S33_TURN_WINDOW_SEC)
|
||||
assert fast["status"] == 200 and fast["body"].get("outcome") == "message" \
|
||||
and fast["body"].get("text") == _reply(run, "fast"), (
|
||||
f"another skill's turn did not complete while one Host authentication was "
|
||||
f"held: {fast}; {diagnosis()}")
|
||||
|
||||
stop = _call(server.port, "POST", f"/api/tasks/{control_id}/cancel", {},
|
||||
timeout=S33_STOP_WINDOW_SEC)
|
||||
# Accepted is enough here; the durable intent and the terminal below are the proof.
|
||||
assert 200 <= stop["status"] < 300 and stop["body"].get("ok") is True, (
|
||||
f"the owner's Stop did not answer inside {S33_STOP_WINDOW_SEC:.0f}s while Presence "
|
||||
f"waited: {stop}; {diagnosis()}")
|
||||
requested = wait_until(lambda: [
|
||||
row for row in oracle.supervisor_rows("cancel_intent")
|
||||
if row.get("task_id") == control_id and row.get("event") == "requested"], 10)
|
||||
assert requested, f"no durable cancel intent for {control_id}: {oracle.supervisor_rows('cancel_intent')}"
|
||||
cancelled = wait_until(
|
||||
lambda: oracle.task_result(control_id).get("status") == "cancelled", 30)
|
||||
assert cancelled, f"the stopped task did not settle cancelled: {oracle.task_result(control_id)}"
|
||||
|
||||
# Everything above happened INSIDE the hold: nothing had been released,
|
||||
# no held request has been answered, and no waiting event reached the model.
|
||||
assert not any(gate.release.is_set() for gate in model_gates) and not auth_gate.release.is_set()
|
||||
early = {label: turn.result for label, turn in turns.items() if not turn.is_alive()}
|
||||
assert not early, f"held Presence requests were answered before release: {early}"
|
||||
expected_arrivals = ({"e00": 1, "e01": 1} if held_turns else {}) | (
|
||||
{"fast": 1} if slow_auth and not held_turns else {})
|
||||
assert arrivals.per_event() == expected_arrivals, (
|
||||
f"a waiting event reached the model during the hold: {diagnosis()}")
|
||||
|
||||
auth_gate.release.set()
|
||||
for gate in model_gates:
|
||||
gate.release.set()
|
||||
deadline = time.monotonic() + S33_SETTLE_SEC
|
||||
for turn in turns.values():
|
||||
turn.join(timeout=max(0.0, deadline - time.monotonic()))
|
||||
unanswered = [label for label, turn in turns.items() if turn.is_alive()]
|
||||
assert not unanswered, f"held requests never answered after release: {unanswered}"
|
||||
assert not auth_gate.timed_out and not any(gate.timed_out for gate in model_gates)
|
||||
|
||||
refs: dict = {}
|
||||
for label, turn in turns.items():
|
||||
event = "e00" if label == "e00-retry" else label
|
||||
body = turn.result.get("body") or {}
|
||||
assert turn.result.get("status") == 200 and body.get("ok") is True \
|
||||
and body.get("outcome") == "message" and body.get("text") == _reply(run, event), (
|
||||
f"{label} did not answer its own reply after release: {turn.result}")
|
||||
refs[label] = str(body.get("turn_ref") or "")
|
||||
stored = oracle.task_result(refs[label])
|
||||
assert stored.get("status") == "completed", (label, stored)
|
||||
assert len(set(refs.values())) == len(refs), f"two events shared one turn: {refs}"
|
||||
|
||||
if held_turns:
|
||||
# The abandoned attempt and its retry were ONE turn; a later replay
|
||||
# answers the identical projection without another model call.
|
||||
replay = _call(host_port, "POST", "/presence/turn", bodies["e00"],
|
||||
headers={"X-Skill-Token": by_label["e00"].access}, timeout=60)
|
||||
assert replay["status"] == 200, replay
|
||||
assert {key: replay["body"].get(key) for key in ("outcome", "text", "turn_ref")} == {
|
||||
key: turns["e00-retry"].result["body"].get(key) for key in ("outcome", "text", "turn_ref")
|
||||
}, (replay, turns["e00-retry"].result)
|
||||
expected = {label: 1 for label in (
|
||||
*(bodies if held_turns else ()), *(("slow",) if slow_auth else ()),
|
||||
*(("fast",) if slow_auth and not held_turns else ()))}
|
||||
assert arrivals.per_event() == expected, (
|
||||
f"an event bought more (or fewer) than one model call: {arrivals.per_event()}")
|
||||
|
||||
recorded_ids = {
|
||||
str(((row.get("transport") or {}).get("delivery") or {}).get("delivery_id") or "")
|
||||
for row in (json.loads(line) for line in oracle.chat_bytes().decode("utf-8").splitlines()
|
||||
if '"presence_delivery"' in line)
|
||||
}
|
||||
assert set(receipts) <= recorded_ids, (receipts, recorded_ids)
|
||||
finally:
|
||||
auth_gate.release.set()
|
||||
for gate in model_gates:
|
||||
gate.release.set()
|
||||
for turn in turns.values():
|
||||
turn.join(timeout=S33_SETTLE_SEC)
|
||||
if server is not None:
|
||||
server.stop()
|
||||
for turn in turns.values():
|
||||
turn.join(timeout=10)
|
||||
|
||||
|
||||
def _assert_panic_tears_down(server, oracle, member: str, diagnosis) -> None:
|
||||
from ouroboros.config import PANIC_EXIT_CODE
|
||||
from ouroboros.process_containment import pids_with_env_marker
|
||||
|
||||
# Safety precondition, not the claim: Panic sweeps the port its server BOUND
|
||||
# and the Host port it reads from settings. Both must provably be this
|
||||
# server's own, or the sweep could reach the operator's live install.
|
||||
assert not {server.port, server.host_service_port} & LIVE_DEFAULT_PORTS
|
||||
assert oracle.server_port() == server.port, "the server's bound port is not provably its own"
|
||||
assert pids_with_env_marker(member), "the survivor scan cannot see the live server tree"
|
||||
answered = _call(server.port, "POST", "/api/command", {"cmd": "/panic"}, timeout=S33_READ_WINDOW_SEC)
|
||||
assert 200 <= answered["status"] < 300, (
|
||||
f"the owner's Panic was not accepted inside {S33_READ_WINDOW_SEC:.0f}s under held Presence "
|
||||
f"load: {answered}; {diagnosis()}")
|
||||
exited = wait_until(lambda: server.proc.poll() is not None, S33_PANIC_EXIT_SEC)
|
||||
assert exited, f"Panic did not end the server within {S33_PANIC_EXIT_SEC:.0f}s: {diagnosis()}"
|
||||
assert server.proc.returncode == PANIC_EXIT_CODE, server.proc.returncode
|
||||
assert (server.data_root / "state" / "panic_stop.flag").is_file()
|
||||
gone = wait_until(lambda: pids_with_env_marker(member) == [], 30)
|
||||
assert gone, f"Panic left members of the server tree alive: {pids_with_env_marker(member)}"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Mock lane
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def s33_candidate(tmp_path_factory):
|
||||
"""The tree every S33 server runs. Preferred: ONE byte-faithful
|
||||
``candidate_checkout`` of the working tree (uncommitted changes included),
|
||||
which needs a dependency-only interpreter. The nightly lane installs the
|
||||
checkout editable, which the candidate probe refuses; a CLEAN tree is then
|
||||
byte-identical to HEAD, so a HEAD clone serves the same bytes. A dirty tree
|
||||
under such an interpreter is refused, never silently tested as HEAD."""
|
||||
require_lane(LANE_MOCK)
|
||||
try:
|
||||
require_candidate_interpreter()
|
||||
except CandidateError as exc:
|
||||
dirty = subprocess.run(["git", "status", "--porcelain", "--untracked-files=all"], cwd=REPO_ROOT,
|
||||
check=True, capture_output=True, text=True).stdout
|
||||
if dirty.strip():
|
||||
pytest.fail(f"{exc} — and the working tree is dirty, so a HEAD clone would not be the candidate")
|
||||
yield clone_repo(tmp_path_factory.mktemp("s33"))
|
||||
return
|
||||
with candidate_checkout(REPO_ROOT, tmp_path_factory.mktemp("s33") / "candidate",
|
||||
origin_proof=True) as candidate:
|
||||
yield candidate
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.serial
|
||||
def test_s33_presence_waits_leave_state_receipts_and_stop_answering(s33_candidate, tmp_path):
|
||||
"""Mechanism 1 alone: 12 waiting events on a 12-thread default executor."""
|
||||
require_lane(LANE_MOCK)
|
||||
_run_s33(s33_candidate, tmp_path / "executor", held_turns=True, slow_auth=False)
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.serial
|
||||
def test_s33_held_host_authentication_leaves_the_event_loop_serving(s33_candidate, tmp_path):
|
||||
"""Mechanism 2 alone: one authentication held, nothing else waiting."""
|
||||
require_lane(LANE_MOCK)
|
||||
_run_s33(s33_candidate, tmp_path / "auth", held_turns=False, slow_auth=True)
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.serial
|
||||
def test_s33_presence_waits_and_held_authentication_together(s33_candidate, tmp_path):
|
||||
require_lane(LANE_MOCK)
|
||||
_run_s33(s33_candidate, tmp_path / "both", held_turns=True, slow_auth=True)
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.serial
|
||||
@pytest.mark.skipif(sys.platform == "win32", reason="the survivor scan is POSIX-only")
|
||||
def test_s33_owner_panic_tears_down_under_held_presence_load(s33_candidate, tmp_path):
|
||||
require_lane(LANE_MOCK)
|
||||
_run_s33(s33_candidate, tmp_path / "panic", held_turns=True, slow_auth=True, panic=True)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Default lane: the seams the fixture relies on
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def test_presence_resilience_fixture_seams_still_exist():
|
||||
"""The auth hook wraps a method by NAME and the scenario attributes model calls
|
||||
by marker; a drift in either must be a named failure here, not a silently
|
||||
vacuous hold in the mock lane."""
|
||||
from ouroboros.presence_context import frame_presence_user_content
|
||||
|
||||
compile(_HOOK_SOURCE, "sitecustomize.py", "exec")
|
||||
source = (REPO_ROOT / "ouroboros" / "gateway" / "host_service.py").read_text(encoding="utf-8")
|
||||
assert "class HostServiceContext" in source
|
||||
assert "def authenticate_token_payload(self, raw_token" in source
|
||||
run = "0123abcd"
|
||||
body = _event_body(_Transport("t", "a", "b"), _event_id(run, "e00"), "room", "hello")
|
||||
framed = frame_presence_user_content(
|
||||
{"_presence_turn": True, "metadata": {"presence": {"event": body["event"], "observed_text": "hello"}}},
|
||||
f"hello {_reply(run, 'e01')}")
|
||||
assert _event_labels(run, {"messages": [
|
||||
{"role": "system", "content": f'"source_event_id": "{_event_id(run, "e02")}"'},
|
||||
{"role": "user", "content": framed},
|
||||
]}) == ["e00"]
|
||||
|
|
@ -76,6 +76,10 @@ TERMINAL_WRITERS = {
|
|||
('ouroboros/mutation_attribution.py::capture_mutation_baseline', 'status'): 'dynamic',
|
||||
('ouroboros/mutation_attribution.py::record_terminal_mutation_candidates', 'status'): 'dynamic',
|
||||
('ouroboros/post_task_checkpoint.py::set_root_post_task_checkpoint', 'str(existing.get("status") or task.get("status") or STATUS_COMPLETED)'): 'terminal',
|
||||
# Presence recovery asks the owner only after a transport retry has found
|
||||
# an existing, unresolved turn. This field projection preserves the exact
|
||||
# stored status; it cannot terminalize or regenerate the lost attempt.
|
||||
('ouroboros/presence_runner.py::_notify_unresolved_turn', 'str(stored["status"])'): 'dynamic',
|
||||
# #1154: the compare-and-clear of a settled terminal-projection obligation.
|
||||
# It preserves the record's CURRENT status inside the projector and publishes
|
||||
# no lifecycle transition of its own; the status argument is only the
|
||||
|
|
|
|||
373
tests/test_host_service_responsiveness.py
Normal file
373
tests/test_host_service_responsiveness.py
Normal file
|
|
@ -0,0 +1,373 @@
|
|||
"""The Host Service keeps the owner's surfaces responsive while Presence turns wait (TZ3).
|
||||
|
||||
Before: token/grant/admission reads ran inline on the ASGI loop the owner's API shares, and
|
||||
each admitted turn — including its wait on the presence gate and its model/quota waits — ran
|
||||
on the shared default executor. These tests pin the replacement: blocking checks leave the
|
||||
loop, a queued turn holds no thread, an executing turn owns its thread, a cancelled HTTP wait
|
||||
neither stops the work nor frees its capacity early, and a retry of the same event joins it.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import contextvars
|
||||
import json
|
||||
import pathlib
|
||||
import threading
|
||||
import time
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
from ouroboros import presence_runner
|
||||
from ouroboros.gateway import host_service
|
||||
from ouroboros.gateway.host_service import create_host_service_app
|
||||
from ouroboros.presence_runner import PresenceTurnError, PresenceTurnGate
|
||||
from tests.test_host_service_api import _seed_presence_behavior, _seed_token
|
||||
|
||||
TOKEN = "presence-token"
|
||||
BUDGET = "telegram-bot:presence"
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _isolated_gates(monkeypatch):
|
||||
"""The configured-gate cache is process-global; each test gets its own, at two active turns."""
|
||||
monkeypatch.setattr(presence_runner, "_GATES", {})
|
||||
monkeypatch.setenv("OUROBOROS_PRESENCE_MAX_ACTIVE", "2")
|
||||
|
||||
|
||||
def _event(source_event_id: str, conversation_id: str = "room-1") -> dict:
|
||||
return {
|
||||
"source_event_id": source_event_id, "provider": "telegram", "account_id": "bot-1",
|
||||
"conversation_id": conversation_id, "thread_id": "topic-1", "conversation_key": "ignored",
|
||||
"actor": {"platform_actor_id": "user-7"}, "conversation": {}, "message": {}, "text": "Hello",
|
||||
}
|
||||
|
||||
|
||||
def _request(app, body=None):
|
||||
async def payload():
|
||||
return body
|
||||
|
||||
return SimpleNamespace(app=app, headers={"x-skill-token": TOKEN}, json=payload)
|
||||
|
||||
|
||||
def _turn(app, binding_id: str, source_event_id: str, conversation_id: str = "room-1"):
|
||||
return host_service._api_presence_turn(
|
||||
_request(app, {"binding_id": binding_id, "event": _event(source_event_id, conversation_id)}))
|
||||
|
||||
|
||||
def _presence_app(tmp_path: pathlib.Path, runner, *, account_wide: bool = False):
|
||||
_seed_token(tmp_path, skill="telegram-bot", token=TOKEN, permissions=["presence"],
|
||||
manifest_permissions=["presence"])
|
||||
binding_id = _seed_presence_behavior(tmp_path, account_wide=account_wide)
|
||||
app = create_host_service_app(tmp_path, presence_runner=runner)
|
||||
return app, binding_id, app.state.host_service_context
|
||||
|
||||
|
||||
def _answer(**kwargs):
|
||||
event_id = kwargs["event"].source_event_id
|
||||
return SimpleNamespace(outcome="message", text=f"answer {event_id}", task_id=event_id, work_ref="")
|
||||
|
||||
|
||||
async def _until(predicate, timeout: float = 5.0) -> None:
|
||||
deadline = time.monotonic() + timeout
|
||||
while not predicate():
|
||||
assert time.monotonic() < deadline, "condition not reached"
|
||||
await asyncio.sleep(0.01)
|
||||
|
||||
|
||||
def _turn_threads() -> list[threading.Thread]:
|
||||
return [thread for thread in threading.enumerate() if thread.name.startswith("presence-turn-")]
|
||||
|
||||
|
||||
def test_queued_and_executing_turns_leave_the_default_executor_free(tmp_path):
|
||||
"""Four live turns, two gate slots, ONE default-executor worker: the worker stays free.
|
||||
|
||||
Before, each turn waited inside ``asyncio.to_thread`` — the first parked turn was that
|
||||
worker, and the owner's reads (and every Host authentication) queued behind a model.
|
||||
"""
|
||||
entered, release = [], threading.Event()
|
||||
|
||||
def runner(**kwargs):
|
||||
entered.append(kwargs["event"].source_event_id)
|
||||
assert release.wait(10)
|
||||
return _answer(**kwargs)
|
||||
|
||||
app, binding_id, ctx = _presence_app(tmp_path, runner, account_wide=True)
|
||||
|
||||
async def scenario():
|
||||
shared = ThreadPoolExecutor(max_workers=1, thread_name_prefix="shared-default")
|
||||
asyncio.get_running_loop().set_default_executor(shared)
|
||||
turns = [asyncio.create_task(_turn(app, binding_id, f"e-{index}", f"room-{index}")) for index in range(4)]
|
||||
try:
|
||||
await _until(lambda: len(entered) == 2 and ctx._inflight[BUDGET] == 4)
|
||||
await asyncio.sleep(0.2)
|
||||
assert len(entered) == 2, "the gate admits two turns; two stay queued"
|
||||
assert len(ctx.presence_turns.live()) == 4
|
||||
assert len(_turn_threads()) == 2, "a queued turn holds no thread"
|
||||
assert await asyncio.wait_for(asyncio.to_thread(lambda: "free"), 2) == "free"
|
||||
identity = await asyncio.wait_for(host_service._api_identity(_request(app)), 5)
|
||||
assert identity.status_code == 200
|
||||
finally:
|
||||
release.set()
|
||||
responses = await asyncio.wait_for(asyncio.gather(*turns), 10)
|
||||
assert sorted(json.loads(response.body)["text"] for response in responses) == [
|
||||
f"answer e-{index}" for index in range(4)]
|
||||
shared.shutdown(wait=True)
|
||||
|
||||
asyncio.run(scenario())
|
||||
assert ctx.presence_turns.live() == [] and not any(ctx._inflight.values())
|
||||
|
||||
|
||||
def test_retries_join_turns_whose_waiters_were_cancelled_even_at_full_budget(tmp_path):
|
||||
"""Cancelled waiters (one executing turn, one queued) leave both turns running with their
|
||||
capacity; retries join them although the in-flight budget is full, and each event runs once."""
|
||||
calls, entered, release = [], threading.Event(), threading.Event()
|
||||
|
||||
def runner(**kwargs):
|
||||
calls.append(kwargs["event"].source_event_id)
|
||||
entered.set()
|
||||
assert release.wait(10)
|
||||
return _answer(**kwargs)
|
||||
|
||||
app, binding_id, ctx = _presence_app(tmp_path, runner)
|
||||
|
||||
async def scenario():
|
||||
# One conversation: e-0 executes, e-1..e-4 queue behind it; five turns fill the budget.
|
||||
waiters = {}
|
||||
for index in range(5):
|
||||
waiters[f"e-{index}"] = asyncio.create_task(_turn(app, binding_id, f"e-{index}"))
|
||||
await _until(lambda: ctx._inflight[BUDGET] == index + 1)
|
||||
if index == 0:
|
||||
await _until(entered.is_set)
|
||||
for event_id in ("e-0", "e-2"):
|
||||
waiters[event_id].cancel()
|
||||
with pytest.raises(asyncio.CancelledError):
|
||||
await waiters.pop(event_id)
|
||||
await asyncio.sleep(0.1)
|
||||
assert calls == ["e-0"]
|
||||
assert ctx._inflight[BUDGET] == 5, "a cancelled wait does not return its turn's capacity"
|
||||
assert len(ctx.presence_turns.live()) == 5
|
||||
|
||||
refused = await _turn(app, binding_id, "e-5")
|
||||
assert refused.status_code == 429, "a NEW event still meets the budget"
|
||||
retries = {event_id: asyncio.create_task(_turn(app, binding_id, event_id)) for event_id in ("e-0", "e-2")}
|
||||
await asyncio.sleep(0.2)
|
||||
assert not any(retry.done() for retry in retries.values()), "a retry joins; it is never refused"
|
||||
assert calls == ["e-0"] and ctx._inflight[BUDGET] == 5
|
||||
|
||||
release.set()
|
||||
for event_id, retry in retries.items():
|
||||
response = await asyncio.wait_for(retry, 10)
|
||||
assert response.status_code == 200
|
||||
assert json.loads(response.body)["text"] == f"answer {event_id}"
|
||||
for response in await asyncio.wait_for(asyncio.gather(*waiters.values()), 10):
|
||||
assert response.status_code == 200
|
||||
|
||||
asyncio.run(scenario())
|
||||
assert sorted(calls) == [f"e-{index}" for index in range(5)], "every event ran exactly once"
|
||||
assert ctx.presence_turns.live() == [] and not any(ctx._inflight.values())
|
||||
|
||||
|
||||
def test_turn_thread_carries_the_admitting_request_context(tmp_path):
|
||||
"""A raw thread starts with an empty context; the turn must see what to_thread carried."""
|
||||
from ouroboros.config import runtime_setting, task_settings_scope
|
||||
from ouroboros.settings_integrity import TaskSettingsSnapshot
|
||||
|
||||
probe = contextvars.ContextVar("tz3_probe", default="lost")
|
||||
seen = {}
|
||||
|
||||
def runner(**kwargs):
|
||||
seen.update(probe=probe.get(), setting=runtime_setting("OUROBOROS_TZ3_CONTEXT_PROBE"),
|
||||
thread=threading.current_thread().name)
|
||||
return _answer(**kwargs)
|
||||
|
||||
app, binding_id, _ctx = _presence_app(tmp_path, runner)
|
||||
|
||||
async def scenario():
|
||||
probe.set("carried")
|
||||
with task_settings_scope(TaskSettingsSnapshot(settings={}, environ={"OUROBOROS_TZ3_CONTEXT_PROBE": "task"})):
|
||||
return await _turn(app, binding_id, "e-ctx")
|
||||
|
||||
assert asyncio.run(scenario()).status_code == 200
|
||||
assert seen["thread"].startswith("presence-turn-")
|
||||
assert seen["probe"] == "carried" and seen["setting"] == "task"
|
||||
|
||||
|
||||
def test_a_thread_that_cannot_start_returns_capacity_gate_and_identity(tmp_path, monkeypatch):
|
||||
calls = []
|
||||
app, binding_id, ctx = _presence_app(tmp_path, lambda **kwargs: calls.append(1) or _answer(**kwargs))
|
||||
start = threading.Thread.start
|
||||
|
||||
def refuse_turn_threads(thread):
|
||||
if thread.name.startswith("presence-turn-"):
|
||||
raise RuntimeError("can't start new thread")
|
||||
return start(thread)
|
||||
|
||||
monkeypatch.setattr(threading.Thread, "start", refuse_turn_threads)
|
||||
refused = asyncio.run(_turn(app, binding_id, "e-1"))
|
||||
assert refused.status_code == 503
|
||||
assert json.loads(refused.body)["error"] == "presence turn thread could not start"
|
||||
assert calls == [] and ctx.presence_turns.live() == [] and not any(ctx._inflight.values())
|
||||
|
||||
monkeypatch.setattr(threading.Thread, "start", start)
|
||||
# The gate lease went back too: the same conversation admits at once.
|
||||
retried = asyncio.run(asyncio.wait_for(_turn(app, binding_id, "e-1"), 5))
|
||||
assert retried.status_code == 200 and calls == [1]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("failure_at", ["context", "constructor"])
|
||||
def test_thread_preparation_failure_releases_admitted_work(failure_at, monkeypatch):
|
||||
from contextlib import ExitStack
|
||||
from ouroboros.presence_runner import (
|
||||
PresenceTurnExecutions, PresenceTurnLease, PresenceTurnNotStarted,
|
||||
)
|
||||
|
||||
executions, capacity, gate = PresenceTurnExecutions(), [], []
|
||||
|
||||
async def admit():
|
||||
resources = ExitStack()
|
||||
gate.append("held")
|
||||
resources.callback(gate.remove, "held")
|
||||
return PresenceTurnLease("room", resources)
|
||||
|
||||
def fail(*args, **kwargs):
|
||||
raise RuntimeError("thread preparation refused")
|
||||
|
||||
async def scenario():
|
||||
# Install faults after asyncio's own Context preparation. Only the
|
||||
# admitted execution's thread preparation is under test.
|
||||
execution, _ = executions.start_or_join(
|
||||
"turn-preparation", reserve=lambda: capacity.append("held") or True,
|
||||
release=lambda: capacity.remove("held"), admit=admit,
|
||||
run=lambda lease: pytest.fail("no work may start"),
|
||||
)
|
||||
with monkeypatch.context() as patch:
|
||||
patch.setattr(contextvars if failure_at == "context" else threading,
|
||||
"copy_context" if failure_at == "context" else "Thread", fail)
|
||||
await execution.admission
|
||||
with pytest.raises(PresenceTurnNotStarted):
|
||||
execution.result.result()
|
||||
|
||||
asyncio.run(scenario())
|
||||
assert capacity == [] and gate == [] and executions.live() == []
|
||||
|
||||
|
||||
@pytest.mark.parametrize("admission_started", [False, True])
|
||||
def test_an_interrupted_admission_settles_waiters_and_capacity(admission_started):
|
||||
"""A loop that cancels a queued turn (shutdown) must not strand its waiters or its slot,
|
||||
including a cancel that lands before the admission coroutine ever ran."""
|
||||
from ouroboros.presence_runner import PresenceTurnExecutions, PresenceTurnNotStarted
|
||||
|
||||
executions, capacity = PresenceTurnExecutions(), []
|
||||
|
||||
async def never_admitted():
|
||||
await asyncio.Event().wait()
|
||||
|
||||
async def scenario():
|
||||
execution, started = executions.start_or_join(
|
||||
"turn-1", reserve=lambda: capacity.append("held") or True, release=lambda: capacity.remove("held"),
|
||||
admit=never_admitted, run=lambda _lease: pytest.fail("must not run"))
|
||||
assert started and capacity == ["held"]
|
||||
if admission_started:
|
||||
await asyncio.sleep(0.05)
|
||||
execution.admission.cancel()
|
||||
with pytest.raises(PresenceTurnNotStarted):
|
||||
await asyncio.wait_for(asyncio.wrap_future(execution.result), 2)
|
||||
|
||||
asyncio.run(scenario())
|
||||
assert capacity == [] and executions.live() == []
|
||||
|
||||
|
||||
def test_blocking_checks_run_off_the_event_loop(tmp_path, monkeypatch):
|
||||
"""Token discovery and presence admission read disk; the loop keeps ticking through both."""
|
||||
import ouroboros.presence_admission as presence_admission
|
||||
|
||||
app, binding_id, ctx = _presence_app(tmp_path, _answer)
|
||||
authenticate, admit = ctx.authenticate_token_payload, presence_admission.admit_presence_turn
|
||||
|
||||
def slow_authenticate(raw_token):
|
||||
time.sleep(0.3)
|
||||
return authenticate(raw_token)
|
||||
|
||||
def slow_admit(**kwargs):
|
||||
time.sleep(0.3)
|
||||
return admit(**kwargs)
|
||||
|
||||
monkeypatch.setattr(ctx, "authenticate_token_payload", slow_authenticate)
|
||||
monkeypatch.setattr(presence_admission, "admit_presence_turn", slow_admit)
|
||||
|
||||
async def ticks_during(awaitable):
|
||||
ticks, done = 0, asyncio.Event()
|
||||
|
||||
async def tick():
|
||||
nonlocal ticks
|
||||
while not done.is_set():
|
||||
ticks += 1
|
||||
await asyncio.sleep(0.01)
|
||||
|
||||
ticker = asyncio.create_task(tick())
|
||||
try:
|
||||
response = await awaitable
|
||||
finally:
|
||||
done.set()
|
||||
await ticker
|
||||
return response.status_code, ticks
|
||||
|
||||
async def scenario():
|
||||
return (await ticks_during(host_service._api_identity(_request(app))),
|
||||
await ticks_during(_turn(app, binding_id, "e-1")))
|
||||
|
||||
(identity_status, identity_ticks), (turn_status, turn_ticks) = asyncio.run(scenario())
|
||||
assert identity_status == 200 and turn_status == 200
|
||||
# Inline on the loop, each 0.3 s check would be one tick.
|
||||
assert identity_ticks > 10 and turn_ticks > 20
|
||||
|
||||
|
||||
def test_coroutine_admission_keeps_the_cross_process_gate_contract(tmp_path):
|
||||
"""``admit`` takes the same file locks ``run`` takes: another gate instance (another process
|
||||
in production) waits for the conversation and the slot cap, and cancellation leaks nothing."""
|
||||
host = PresenceTurnGate(1, state_root=tmp_path)
|
||||
other = PresenceTurnGate(1, state_root=tmp_path)
|
||||
entered = threading.Event()
|
||||
|
||||
async def scenario():
|
||||
lease = await host.admit("conversation-a")
|
||||
waiting = asyncio.create_task(host.admit("conversation-a"))
|
||||
await asyncio.sleep(0.1)
|
||||
waiting.cancel()
|
||||
with pytest.raises(asyncio.CancelledError):
|
||||
await waiting
|
||||
blocked = threading.Thread(target=other.run, args=("conversation-b", entered.set))
|
||||
blocked.start()
|
||||
try:
|
||||
await asyncio.sleep(0.3)
|
||||
assert not entered.is_set(), "the one cross-process slot is held by the lease"
|
||||
finally:
|
||||
lease.release()
|
||||
await asyncio.to_thread(blocked.join, 5)
|
||||
assert entered.is_set()
|
||||
# Nothing the cancelled admission took stayed held: the conversation admits at once.
|
||||
again = await asyncio.wait_for(host.admit("conversation-a"), 2)
|
||||
again.release()
|
||||
again.release() # idempotent
|
||||
|
||||
asyncio.run(scenario())
|
||||
|
||||
|
||||
def test_an_admitted_turn_runs_only_under_its_own_conversation(tmp_path):
|
||||
from ouroboros.presence_runner import PresenceTurnEvent, run_presence_turn
|
||||
|
||||
lease = asyncio.run(PresenceTurnGate(1).admit("telegram:bot-1:room-2:topic-1"))
|
||||
event = PresenceTurnEvent(
|
||||
source_event_id="e-1", provider="telegram", account_id="bot-1", conversation_id="room-1",
|
||||
thread_id="topic-1", conversation_key="telegram:bot-1:room-1:topic-1", actor={"id": "u"},
|
||||
conversation={}, message={}, text="Hello",
|
||||
)
|
||||
try:
|
||||
with pytest.raises(PresenceTurnError) as refused:
|
||||
run_presence_turn(admission=SimpleNamespace(binding_id="b"), event=event, repo_dir=tmp_path,
|
||||
drive_root=tmp_path, admitted=lease,
|
||||
agent_factory=lambda **_kwargs: pytest.fail("must not run"))
|
||||
assert refused.value.code == "presence_admission_conversation_mismatch"
|
||||
finally:
|
||||
lease.release()
|
||||
|
|
@ -434,13 +434,16 @@ def test_unknown_outcome_keeps_its_generation_limit_claim(setup):
|
|||
|
||||
def test_confirmed_provider_failure_settles_real_usage_before_raising(setup):
|
||||
root, gateway, client = setup
|
||||
# Auto asks the engine once more; it reselects the refused account, which ends rotation.
|
||||
gateway.results = [result(outcome="failed", cash=0.13, knowledge="exact", problem={
|
||||
"code": "subscription_window_exhausted", "message": "window exhausted", "retryable": True,
|
||||
"context": {"resetsAt": "2099-01-01T00:00:00Z", "httpStatus": 429},
|
||||
})]
|
||||
})] * 2
|
||||
gateway.dispatch = ["response_received"] * 2
|
||||
with pytest.raises(transport.ClaudexorModelError) as raised:
|
||||
client.chat([{"role": "user", "content": "hi"}], MODEL, model_role="vision")
|
||||
error = raised.value
|
||||
assert error.account_rotation["stop"] == "engine_reselected_refused_account"
|
||||
assert error.code == "subscription_window_exhausted" and error.model_role == "vision"
|
||||
assert error.reset_at == "2099-01-01T00:00:00Z"
|
||||
assert error.physical_attempt_capture.state == "settled"
|
||||
|
|
|
|||
|
|
@ -144,11 +144,11 @@ def _loop_tools(ctx, owner):
|
|||
return tools
|
||||
|
||||
|
||||
def test_main_wait_does_not_call_configured_api_fallback_before_owner_switch(main_call, monkeypatch):
|
||||
from ouroboros.gateways.claudexor import ClaudexorUnavailable
|
||||
def test_main_quota_calls_configured_api_fallback_before_any_owner_wait(main_call, monkeypatch):
|
||||
"""Owner order: Auto rotation, then the configured fallback, and only then the owner question."""
|
||||
from ouroboros.llm_attempt import _attempt_request, _candidate_before_dispatch
|
||||
|
||||
ctx, gateway, owner, events, decide, _observations = main_call
|
||||
ctx, gateway, owner, events, _decide, _observations = main_call
|
||||
tools = _loop_tools(ctx, owner)
|
||||
monkeypatch.setenv("OUROBOROS_MODEL_FALLBACKS", "openai::alternate")
|
||||
monkeypatch.setenv("OUROBOROS_TASK_REVIEW_MODE", "off")
|
||||
|
|
@ -170,13 +170,7 @@ def test_main_wait_does_not_call_configured_api_fallback_before_owner_switch(mai
|
|||
before_dispatch=_candidate_before_dispatch(request_body, request))
|
||||
|
||||
def catalog(*args, **kwargs):
|
||||
assert api_calls == [] and len(gateway.accepted_operations) == 1
|
||||
row = next(event for event in reversed(list(events.queue)) if event.get("type") == "task_model_wait")
|
||||
response = decide({"request_id": "switch-api", "decision_id": f"model_wait:task-one:{row['wait_id']}",
|
||||
"revision": row["revision"], "action": "switch", "model": "openai::alternate",
|
||||
"credential_profile_id": "", "use_local": False, "persist_role": False})
|
||||
assert response.status_code == 202
|
||||
raise ClaudexorUnavailable("subscription_window_exhausted", "still waiting")
|
||||
pytest.fail("the owner is asked only after every configured route of the round failed")
|
||||
|
||||
monkeypatch.setattr(ctx.llm, "_chat_remote", send)
|
||||
monkeypatch.setattr(ctx.llm, "claudexor_model_catalog", catalog)
|
||||
|
|
@ -184,6 +178,7 @@ def test_main_wait_does_not_call_configured_api_fallback_before_owner_switch(mai
|
|||
ctx.messages, tools, ctx.llm, ctx.drive_logs, lambda *_args, **_kwargs: None, queue.Queue(),
|
||||
task_id="task-one", drive_root=ctx.drive_root, event_queue=events)
|
||||
assert text == "Finished" and api_calls == ["openai"]
|
||||
assert not [event for event in list(events.queue) if event.get("type") == "task_model_wait"]
|
||||
assert usage["_model_route"] == {} and len(gateway.accepted_operations) == 1
|
||||
assert any(message.get("content") == "verified read A" for message in ctx.messages)
|
||||
assert any(message.get("content") == "completed review B" for message in ctx.messages)
|
||||
|
|
|
|||
|
|
@ -571,7 +571,10 @@ def scan_data_paths(root: pathlib.Path = REPO) -> frozenset[str]:
|
|||
# one rebuildable projection per conversation written by presence_runner at the end of an executed
|
||||
# turn; it has its own row in section 2.
|
||||
# 293 -> 295: the disposable test-environment caches (``cache/pip``, ``cache/uv``; test root only).
|
||||
EXPECTED_SCAN_PATHS = 295
|
||||
# 295 -> 296: Presence recovery inspects the retained quarantine members before
|
||||
# deciding whether an event ever started; a quarantined task id cannot become
|
||||
# a fresh model generation on a transport retry.
|
||||
EXPECTED_SCAN_PATHS = 296
|
||||
|
||||
# Scanned paths that must always be present — guards the scanner itself
|
||||
# against a silent regression that would shrink coverage while keeping counts
|
||||
|
|
|
|||
|
|
@ -4,6 +4,7 @@ from __future__ import annotations
|
|||
import asyncio
|
||||
import json
|
||||
import threading
|
||||
import time
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from contextlib import contextmanager
|
||||
from types import SimpleNamespace
|
||||
|
|
@ -418,7 +419,12 @@ def test_five_open_turns_leave_receipts_and_inject_admitted(tmp_path):
|
|||
with ThreadPoolExecutor(max_workers=5) as pool:
|
||||
turns = [pool.submit(_turn, client, binding, index) for index in range(5)]
|
||||
try:
|
||||
assert all(entered.acquire(timeout=10) for _ in range(5))
|
||||
# One conversation: the Host gate runs one turn; the other four queue, holding their budget.
|
||||
assert entered.acquire(timeout=10)
|
||||
deadline = time.monotonic() + 10
|
||||
while len(ctx.presence_turns.live()) < 5 and time.monotonic() < deadline:
|
||||
time.sleep(0.01)
|
||||
assert len(ctx.presence_turns.live()) == 5
|
||||
assert _turn(client, binding, 5).status_code == 429 # the turn budget itself still binds
|
||||
assert _post(client, _payload()).status_code == 200
|
||||
assert _inject(client).status_code == 202
|
||||
|
|
|
|||
|
|
@ -1,24 +1,42 @@
|
|||
"""An orphan-reconciled presence turn is neither spoken nor final: the retry answers it."""
|
||||
"""A lost presence turn is neither spoken nor final, and its retry never regenerates it.
|
||||
|
||||
Whether a lost attempt's model/tool work and sends took effect is unproven — absent receipts,
|
||||
logs or ledger rows are no proof of no effect — so a retry of an event whose persisted row is
|
||||
RUNNING, INTERRUPTED or the reconciler's placeholder fails closed before any agent exists
|
||||
(``presence_attempt_outcome_unknown``; the Host answers 409 ``retry``) until a canonical terminal
|
||||
replays. An event with no running row of its own (an attempt that died before its running write,
|
||||
a turn thread that never started, another event) still runs, and a real terminal still replays.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
import threading
|
||||
import time
|
||||
from dataclasses import replace
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
from starlette.testclient import TestClient
|
||||
|
||||
from ouroboros import agent_task_pipeline as pipeline
|
||||
from ouroboros.gateway import host_service
|
||||
from ouroboros.gateway.host_service import create_host_service_app
|
||||
from ouroboros.outcomes import infra_failed_axes
|
||||
from ouroboros.presence_context import build_presence_context_section
|
||||
from ouroboros.presence_runner import (
|
||||
PresenceTurnError,
|
||||
PresenceTurnGate,
|
||||
_cached_result,
|
||||
_live_task_rows,
|
||||
_previous_turn_path,
|
||||
_read_previous_turn,
|
||||
_stable_numeric_id,
|
||||
_task_id,
|
||||
_turn_sends,
|
||||
presence_result_from_stored,
|
||||
presence_turn_is_live,
|
||||
run_presence_turn,
|
||||
)
|
||||
from ouroboros.task_results import (
|
||||
|
|
@ -26,6 +44,7 @@ from ouroboros.task_results import (
|
|||
STATUS_FAILED,
|
||||
STATUS_INTERRUPTED,
|
||||
STATUS_RUNNING,
|
||||
is_reconciled_presence_placeholder,
|
||||
load_task_result,
|
||||
reopen_reconciled_presence_placeholder,
|
||||
task_result_path,
|
||||
|
|
@ -39,6 +58,7 @@ from tests.test_presence_runner import _admission, _event
|
|||
|
||||
_PRESENCE_METADATA = {"source": "presence", "presence": {"binding_id": "1" * 32, "delivery_reporting_version": 0}}
|
||||
_NOW = 1_800_000_000.0 # 2027-01-15T08:00:00Z, the fresh queue snapshot's own time
|
||||
_OUTCOME_UNKNOWN = "presence_attempt_outcome_unknown"
|
||||
|
||||
|
||||
def _sweep(tmp_path, monkeypatch, task_id, status=STATUS_RUNNING, *, seed=True, boot="2026-05-28T00:00:02+00:00",
|
||||
|
|
@ -132,6 +152,21 @@ def test_only_the_orphan_placeholder_of_a_presence_turn_reopens(tmp_path, monkey
|
|||
assert load_task_result(tmp_path, task_id)["status"] == STATUS_FAILED # sticky terminal unchanged
|
||||
|
||||
|
||||
def _complete(task, reply, drive_root):
|
||||
"""The real pipeline's terminal write for *task*: an accepted presence_finish, then emit_task_results."""
|
||||
task["_skip_post_task_synthesis"] = True
|
||||
ctx = SimpleNamespace(task_contract=task["task_contract"], task_metadata=task["metadata"])
|
||||
_finish_presence(ctx, "message", reply)
|
||||
ctx._presence_completion_accepted = True
|
||||
pending: list = []
|
||||
pipeline.emit_task_results(
|
||||
SimpleNamespace(drive_root=drive_root, repo_dir=drive_root), None, None, pending, task, reply,
|
||||
{"terminal_origin": "model_final"}, {"tool_calls": [], "reasoning_notes": []}, 0.0,
|
||||
drive_root / "logs", ctx=ctx,
|
||||
)
|
||||
return pending
|
||||
|
||||
|
||||
def _answering_agent(calls, reply, drive_root, *, during=None, lost=False):
|
||||
"""Real durable pipeline: the RUNNING start write, then the terminal write (or a lost worker)."""
|
||||
|
||||
|
|
@ -145,41 +180,73 @@ def _answering_agent(calls, reply, drive_root, *, during=None, lost=False):
|
|||
during(task)
|
||||
if lost:
|
||||
raise RuntimeError("worker lost")
|
||||
ctx = SimpleNamespace(task_contract=task["task_contract"], task_metadata=task["metadata"])
|
||||
_finish_presence(ctx, "message", reply)
|
||||
ctx._presence_completion_accepted = True
|
||||
pending: list = []
|
||||
pipeline.emit_task_results(
|
||||
SimpleNamespace(drive_root=drive_root, repo_dir=drive_root), None, None, pending, task, reply,
|
||||
{"terminal_origin": "model_final"}, {"tool_calls": [], "reasoning_notes": []}, 0.0,
|
||||
drive_root / "logs", ctx=ctx,
|
||||
)
|
||||
return pending
|
||||
return _complete(task, reply, drive_root)
|
||||
|
||||
return Agent()
|
||||
|
||||
|
||||
def test_reconciled_turn_is_not_cached_and_its_rerun_persists(tmp_path, monkeypatch):
|
||||
task_id = _task_id(_admission(), _event())
|
||||
_reconciled(tmp_path, monkeypatch, task_id)
|
||||
assert _cached_result(tmp_path, task_id) is None
|
||||
def _chat_rows(drive_root):
|
||||
chat = drive_root / "logs" / "chat.jsonl"
|
||||
return [json.loads(line) for line in chat.read_text(encoding="utf-8").splitlines()] if chat.exists() else []
|
||||
|
||||
|
||||
def _trace(drive_root, task_id, conversation_key=_event().conversation_key):
|
||||
"""What a refused retry must leave untouched: the turn's row, the conversation's pointer and dialogue."""
|
||||
row, pointer = task_result_path(drive_root, task_id), _previous_turn_path(drive_root, conversation_key)
|
||||
chat_id = _stable_numeric_id("presence-conversation", conversation_key)
|
||||
dialogue = [item for item in _chat_rows(drive_root)
|
||||
if item.get("task_id") == task_id or item.get("chat_id") == chat_id]
|
||||
return (row.read_bytes() if row.exists() else None, pointer.read_bytes() if pointer.exists() else None, dialogue)
|
||||
|
||||
|
||||
def _retry_kwargs(tmp_path, built, *, version=1, reply="Regenerated answer", **overrides):
|
||||
"""A retry whose agent factory records every build: a refused retry must fail before the first."""
|
||||
calls: list = []
|
||||
kwargs = dict(admission=_admission(), event=_event(), repo_dir=tmp_path, drive_root=tmp_path,
|
||||
agent_factory=lambda **_kw: _answering_agent(calls, "Real answer", tmp_path),
|
||||
gate=PresenceTurnGate(1))
|
||||
first = run_presence_turn(**kwargs)
|
||||
assert [task["id"] for task in calls] == [task_id] and first.outcome == "message" and first.text == "Real answer"
|
||||
|
||||
def factory(**kwargs):
|
||||
built.append(kwargs)
|
||||
return _answering_agent(calls, reply, tmp_path)
|
||||
|
||||
return dict(admission=_admission(), event=replace(_event(), delivery_reporting_version=version),
|
||||
repo_dir=tmp_path, drive_root=tmp_path, agent_factory=factory, gate=PresenceTurnGate(1), **overrides)
|
||||
|
||||
|
||||
def _refused(kwargs):
|
||||
"""The typed refusal: the event stays with its transport, named by the turn it could not settle."""
|
||||
with pytest.raises(PresenceTurnError) as raised:
|
||||
run_presence_turn(**kwargs)
|
||||
assert (raised.value.code, raised.value.field, getattr(raised.value, "turn_ref", None)) == (
|
||||
_OUTCOME_UNKNOWN, "source_event_id", _task_id(kwargs["admission"], kwargs["event"]))
|
||||
|
||||
|
||||
def _assert_fails_closed(tmp_path, task_id, kwargs, built):
|
||||
"""Two retries, each refused before any agent exists, and neither leaves a trace to inherit."""
|
||||
before = _trace(tmp_path, task_id, kwargs["event"].conversation_key)
|
||||
for _second in (False, True): # a second retry is refused exactly like the first, never promoted
|
||||
_refused(kwargs)
|
||||
assert built == [] and not presence_turn_is_live(task_id)
|
||||
assert _trace(tmp_path, task_id, kwargs["event"].conversation_key) == before
|
||||
assert _cached_result(tmp_path, task_id) is None # still no result to replay
|
||||
|
||||
|
||||
def _prior_sends(tmp_path, task_id):
|
||||
"""The receipt facts a lost attempt left in the live generation, settled by the runner's own rule."""
|
||||
return _turn_sends(_live_task_rows(tmp_path, task_id, _event().conversation_key))
|
||||
|
||||
|
||||
@pytest.mark.parametrize("version", [0, 1])
|
||||
@pytest.mark.parametrize("status", [STATUS_RUNNING, STATUS_INTERRUPTED])
|
||||
def test_reconciled_turn_fails_closed_and_is_never_regenerated(tmp_path, monkeypatch, status, version):
|
||||
"""Formerly the placeholder re-ran with 'what the lost attempt sent' as model context. That context
|
||||
cannot prove the lost model/tool work had no effect, so the retry is refused and the host mark stays."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
row = _reconciled(tmp_path, monkeypatch, task_id, status)
|
||||
assert _cached_result(tmp_path, task_id) is None # the placeholder is still not a result
|
||||
built: list = []
|
||||
_assert_fails_closed(tmp_path, task_id, _retry_kwargs(tmp_path, built, version=version), built)
|
||||
stored = load_task_result(tmp_path, task_id)
|
||||
assert stored["status"] == STATUS_COMPLETED and stored["terminal_origin"] == "model_final"
|
||||
assert "status_reconciled_from" not in stored and "TASK_ORPHAN_RECONCILED" not in stored["result"]
|
||||
assert stored["superseded_placeholder"]["reason_code"] == "orphaned_running_after_worker_restart"
|
||||
# A v0 transport reports no receipts: what the lost attempt sent is unknown, and the model is told so.
|
||||
attempt = calls[0]["metadata"]["presence"]["previous_attempt"]
|
||||
assert attempt == {"delivered_count": None, "delivered": None, "uncertain_count": 0}
|
||||
section = build_presence_context_section(tmp_path, calls[0]["metadata"]["presence"])
|
||||
assert "whether it already sent anything is unknown" in section
|
||||
# The persisted answer is now the cached result: a further retry does not re-run.
|
||||
assert run_presence_turn(**kwargs) == first and len(calls) == 1
|
||||
assert is_reconciled_presence_placeholder(stored) and stored == row # never reopened, never moved aside
|
||||
assert "superseded_placeholder" not in stored
|
||||
|
||||
|
||||
def test_reconciler_repersist_keeps_the_placeholder(tmp_path, monkeypatch):
|
||||
|
|
@ -225,65 +292,205 @@ def test_reconciled_non_presence_task_keeps_the_sticky_terminal(tmp_path, monkey
|
|||
assert row["status"] == STATUS_FAILED and row["status_reconciled_from"] == STATUS_RUNNING
|
||||
|
||||
|
||||
def test_host_retry_reruns_a_lost_turn_once_and_then_replays(tmp_path, monkeypatch):
|
||||
"""Through POST /presence/turn: a lost v1 attempt, the reconciler, the re-run, a replay."""
|
||||
_HOST_EVENT_ID = "telegram:bot-1:42"
|
||||
|
||||
|
||||
def _host(tmp_path, agents):
|
||||
"""The real Host over the real runner; each turn builds the next queued agent (an exception raises)."""
|
||||
_seed_token(tmp_path, skill="telegram-bot", token="presence-token",
|
||||
permissions=["presence"], manifest_permissions=["presence"])
|
||||
binding_id = _seed_presence_behavior(tmp_path)
|
||||
task_id = "presence-" + hashlib.sha256(f"{binding_id}\0telegram:bot-1:42".encode("utf-8")).hexdigest()[:24]
|
||||
agents: list = []
|
||||
|
||||
def factory(**_kw):
|
||||
agent = agents.pop(0)
|
||||
if isinstance(agent, BaseException):
|
||||
raise agent
|
||||
return agent
|
||||
|
||||
def run_real_presence(**kwargs):
|
||||
return run_presence_turn(repo_dir=tmp_path, drive_root=tmp_path, agent_factory=lambda **_kw: agents.pop(0),
|
||||
return run_presence_turn(repo_dir=tmp_path, drive_root=tmp_path, agent_factory=factory,
|
||||
gate=PresenceTurnGate(1), **kwargs)
|
||||
|
||||
client = TestClient(create_host_service_app(tmp_path, presence_runner=run_real_presence))
|
||||
recorder = client.app.state.host_service_context.presence_deliveries
|
||||
app = create_host_service_app(tmp_path, presence_runner=run_real_presence)
|
||||
return app, binding_id
|
||||
|
||||
def early_send(task): # the lost attempt had already delivered one transport message
|
||||
recorder.record("telegram-bot", {
|
||||
"schema_version": 1, "delivery_id": "send:early", "part_id": "0", "state": "delivered",
|
||||
"provider": "telegram", "account_id": "bot-1", "conversation_id": "room-1", "thread_id": "topic-1",
|
||||
"text": "Early part", "format": "markdown", "message": {"provider_message_id": "501"},
|
||||
"origin": {"kind": "tool", "task_id": task["id"], "source_event_id": "telegram:bot-1:42"},
|
||||
})
|
||||
|
||||
def post():
|
||||
return client.post("/presence/turn", headers={"X-Skill-Token": "presence-token"}, json={
|
||||
"binding_id": binding_id, "delivery_reporting_version": 1, "event": {
|
||||
"source_event_id": "telegram:bot-1:42", "provider": "telegram", "account_id": "bot-1",
|
||||
"conversation_id": "room-1", "thread_id": "topic-1", "conversation_key": "ignored",
|
||||
"actor": {"platform_actor_id": "user-7"}, "conversation": {}, "message": {"message_id": "42"},
|
||||
"text": "Hello",
|
||||
}})
|
||||
def _host_body(binding_id, source_event_id=_HOST_EVENT_ID):
|
||||
return {"binding_id": binding_id, "delivery_reporting_version": 1, "event": {
|
||||
"source_event_id": source_event_id, "provider": "telegram", "account_id": "bot-1",
|
||||
"conversation_id": "room-1", "thread_id": "topic-1", "conversation_key": "ignored",
|
||||
"actor": {"platform_actor_id": "user-7"}, "conversation": {}, "message": {"message_id": "42"},
|
||||
"text": "Hello",
|
||||
}}
|
||||
|
||||
calls: list = []
|
||||
agents.append(_answering_agent(calls, "", tmp_path, during=early_send, lost=True))
|
||||
assert post().status_code == 500 and load_task_result(tmp_path, task_id)["status"] == STATUS_RUNNING
|
||||
_reconciled(tmp_path, monkeypatch, task_id, seed=False)
|
||||
|
||||
def _host_post(tmp_path, agents):
|
||||
app, binding_id = _host(tmp_path, agents)
|
||||
client = TestClient(app)
|
||||
|
||||
def post(source_event_id=_HOST_EVENT_ID):
|
||||
return client.post("/presence/turn", headers={"X-Skill-Token": "presence-token"},
|
||||
json=_host_body(binding_id, source_event_id))
|
||||
|
||||
task_id = "presence-" + hashlib.sha256(f"{binding_id}\0{_HOST_EVENT_ID}".encode("utf-8")).hexdigest()[:24]
|
||||
return app.state.host_service_context, post, task_id
|
||||
|
||||
|
||||
def _host_refused(response, task_id):
|
||||
body = response.json()
|
||||
assert response.status_code == 409 and body["ok"] is False and not body.get("text"), body
|
||||
assert (body["code"], body["disposition"], body["field"], body["turn_ref"]) == (
|
||||
_OUTCOME_UNKNOWN, "retry", "source_event_id", task_id)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("prior_send", [None, "delivered", "uncertain"])
|
||||
def test_host_retry_of_a_lost_turn_is_a_409_retry_never_a_rerun(tmp_path, monkeypatch, prior_send):
|
||||
"""Through POST /presence/turn: a lost v1 attempt, retries before and after the reconciler, a second
|
||||
retry — all refused, whatever the lost attempt's receipts said. Formerly the placeholder re-ran with
|
||||
the prior sends as context; a confirmed send, an unconfirmed one or none proves nothing about the rest."""
|
||||
agents: list = []
|
||||
ctx, post, task_id = _host_post(tmp_path, agents)
|
||||
recorder = ctx.presence_deliveries
|
||||
sweeps = []
|
||||
|
||||
def sweep_during_rerun(_task): # stale by clock and by a later boot, but executing in this process
|
||||
sweeps.append(_sweep(tmp_path, monkeypatch, task_id, seed=False, boot="2027-01-15T07:59:00+00:00")[0])
|
||||
assert load_task_result(tmp_path, task_id)["status"] == STATUS_RUNNING
|
||||
def lost_attempt_effects(task):
|
||||
if prior_send: # the lost attempt had already reported one transport send
|
||||
recorder.record("telegram-bot", {
|
||||
"schema_version": 1, "delivery_id": "send:early", "part_id": "0", "state": prior_send,
|
||||
"provider": "telegram", "account_id": "bot-1", "conversation_id": "room-1", "thread_id": "topic-1",
|
||||
"text": "Early part", "format": "markdown", "message": {"provider_message_id": "501"},
|
||||
"origin": {"kind": "tool", "task_id": task["id"], "source_event_id": _HOST_EVENT_ID},
|
||||
})
|
||||
# stale by clock and by a later boot, but executing in this process: the live set protects it
|
||||
sweeps.append(_sweep(tmp_path, monkeypatch, task["id"], seed=False, boot="2027-01-15T07:59:00+00:00")[0])
|
||||
assert load_task_result(tmp_path, task["id"])["status"] == STATUS_RUNNING
|
||||
|
||||
agents.append(_answering_agent(calls, "Real answer", tmp_path, during=sweep_during_rerun))
|
||||
calls: list = []
|
||||
agents.append(_answering_agent(calls, "", tmp_path, during=lost_attempt_effects, lost=True))
|
||||
assert post().status_code == 500 and load_task_result(tmp_path, task_id)["status"] == STATUS_RUNNING
|
||||
assert sweeps == [0] and len(calls) == 1
|
||||
must_not_run = _answering_agent(calls, "Regenerated answer", tmp_path)
|
||||
agents.append(must_not_run)
|
||||
before = _trace(tmp_path, task_id)
|
||||
_host_refused(post(), task_id) # the unreconciled RUNNING row: nothing of this turn is live here any more
|
||||
assert _trace(tmp_path, task_id) == before
|
||||
_reconciled(tmp_path, monkeypatch, task_id, seed=False)
|
||||
before = _trace(tmp_path, task_id)
|
||||
for _second in (False, True):
|
||||
_host_refused(post(), task_id)
|
||||
assert agents == [must_not_run] and len(calls) == 1 and _trace(tmp_path, task_id) == before
|
||||
assert ctx.presence_turns.live() == [] and not any(ctx._inflight.values())
|
||||
assert is_reconciled_presence_placeholder(load_task_result(tmp_path, task_id))
|
||||
assert _prior_sends(tmp_path, task_id) == ((["Early part"] if prior_send == "delivered" else []),
|
||||
int(prior_send == "uncertain"))
|
||||
inbound = [row for row in _chat_rows(tmp_path) if row.get("direction") == "in" and row.get("task_id") == task_id]
|
||||
assert len(inbound) == 1 # the refusals never log the correspondent's message again
|
||||
# Fail-closed is per turn, and the refusals kept no capacity: another event of the room still runs.
|
||||
agents[:] = [_answering_agent(calls, "Fresh answer", tmp_path)]
|
||||
fresh = post("telegram:bot-1:43")
|
||||
assert fresh.status_code == 200 and fresh.json()["text"] == "Fresh answer" and agents == []
|
||||
assert "previous_attempt" not in calls[-1]["metadata"]["presence"]
|
||||
_host_refused(post(), task_id)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("status", [STATUS_RUNNING, STATUS_INTERRUPTED])
|
||||
def test_host_replays_the_late_canonical_terminal_of_a_refused_turn(tmp_path, status):
|
||||
"""An older worker still finishing the turn writes its terminal after a refusal: the retry answers
|
||||
from that durable row, once, and never builds an agent."""
|
||||
agents: list = []
|
||||
_ctx, post, task_id = _host_post(tmp_path, agents)
|
||||
calls: list = []
|
||||
agents.append(_answering_agent(calls, "", tmp_path, lost=True))
|
||||
assert post().status_code == 500
|
||||
if status == STATUS_INTERRUPTED:
|
||||
write_task_result(tmp_path, task_id, STATUS_INTERRUPTED, result="Task interrupted.")
|
||||
must_not_run = _answering_agent(calls, "Regenerated answer", tmp_path)
|
||||
agents.append(must_not_run)
|
||||
_host_refused(post(), task_id)
|
||||
_complete(calls[0], "Late answer", tmp_path) # the lost attempt's own terminal lands after all
|
||||
first = post()
|
||||
assert first.status_code == 200 and first.json()["text"] == "Real answer" and first.json()["turn_ref"] == task_id
|
||||
assert sweeps == [0] and len(calls) == 2
|
||||
assert calls[1]["metadata"]["presence"]["previous_attempt"] == {
|
||||
"delivered_count": 1, "delivered": ["Early part"], "uncertain_count": 0}
|
||||
section = build_presence_context_section(tmp_path, calls[1]["metadata"]["presence"])
|
||||
assert 'already delivered 1 message(s): "Early part"' in section
|
||||
stored = load_task_result(tmp_path, task_id)
|
||||
assert stored["status"] == STATUS_COMPLETED and "status_reconciled_from" not in stored
|
||||
assert stored["superseded_placeholder"]["status"] == STATUS_FAILED
|
||||
rows = [json.loads(line) for line in (tmp_path / "logs" / "chat.jsonl").read_text(encoding="utf-8").splitlines()]
|
||||
inbound = [row for row in rows if row.get("direction") == "in" and row.get("client_message_id") == "telegram:bot-1:42"]
|
||||
assert len(inbound) == 1 # the re-run does not log the correspondent's message twice
|
||||
assert first.status_code == 200 and first.json()["text"] == "Late answer" and first.json()["turn_ref"] == task_id
|
||||
assert agents == [must_not_run] and len(calls) == 1
|
||||
assert _read_previous_turn(tmp_path, _event().conversation_key)["task_id"] == task_id # repaired on replay
|
||||
replay = post()
|
||||
assert replay.status_code == 200 and replay.json() == first.json() and len(calls) == 2
|
||||
assert replay.status_code == 200 and replay.json() == first.json() and agents == [must_not_run]
|
||||
stored = load_task_result(tmp_path, task_id)
|
||||
assert stored["status"] == STATUS_COMPLETED and "superseded_placeholder" not in stored
|
||||
|
||||
|
||||
def test_host_retry_of_a_live_turn_joins_it_instead_of_refusing(tmp_path):
|
||||
"""A RUNNING row whose execution is live in this Host is not a lost attempt: the retry joins it."""
|
||||
agents: list = []
|
||||
app, binding_id = _host(tmp_path, agents)
|
||||
ctx = app.state.host_service_context
|
||||
started, release, calls, joins = threading.Event(), threading.Event(), [], []
|
||||
|
||||
def hold(_task):
|
||||
started.set()
|
||||
assert release.wait(5)
|
||||
|
||||
agents.append(_answering_agent(calls, "Real answer", tmp_path, during=hold))
|
||||
real_start_or_join = ctx.presence_turns.start_or_join
|
||||
|
||||
async def scenario():
|
||||
second_joined = asyncio.Event()
|
||||
|
||||
def recording(turn_id, **kwargs):
|
||||
outcome = real_start_or_join(turn_id, **kwargs)
|
||||
joins.append(outcome[1])
|
||||
if len(joins) == 2:
|
||||
second_joined.set()
|
||||
return outcome
|
||||
|
||||
ctx.presence_turns.start_or_join = recording
|
||||
|
||||
async def json_body():
|
||||
return _host_body(binding_id)
|
||||
|
||||
def request():
|
||||
return SimpleNamespace(app=app, headers={"x-skill-token": "presence-token"}, json=json_body)
|
||||
|
||||
first = asyncio.ensure_future(host_service._api_presence_turn(request()))
|
||||
assert await asyncio.to_thread(started.wait, 5)
|
||||
assert load_task_result(tmp_path, calls[0]["id"])["status"] == STATUS_RUNNING
|
||||
second = asyncio.ensure_future(host_service._api_presence_turn(request()))
|
||||
await asyncio.wait_for(second_joined.wait(), 5)
|
||||
release.set()
|
||||
return await asyncio.wait_for(asyncio.gather(first, second), 5)
|
||||
|
||||
first, second = asyncio.run(scenario())
|
||||
assert joins == [True, False] and len(calls) == 1
|
||||
assert first.status_code == second.status_code == 200 and first.body == second.body
|
||||
assert json.loads(first.body)["text"] == "Real answer"
|
||||
|
||||
|
||||
def test_host_turn_that_never_started_runs_on_its_retry(tmp_path, monkeypatch):
|
||||
"""No thread, or no agent before the running write: nothing of the turn ran, so the retry answers it."""
|
||||
agents: list = []
|
||||
ctx, post, task_id = _host_post(tmp_path, agents)
|
||||
start = threading.Thread.start
|
||||
|
||||
def refuse_turn_threads(thread):
|
||||
if thread.name.startswith("presence-turn-"):
|
||||
raise RuntimeError("can't start new thread")
|
||||
return start(thread)
|
||||
|
||||
monkeypatch.setattr(threading.Thread, "start", refuse_turn_threads)
|
||||
refused = post()
|
||||
assert refused.status_code == 503 and (refused.json()["code"], refused.json()["disposition"]) == (
|
||||
"presence_turn_not_started", "retry")
|
||||
assert load_task_result(tmp_path, task_id) is None and _chat_rows(tmp_path) == []
|
||||
assert ctx.presence_turns.live() == [] and not any(ctx._inflight.values())
|
||||
monkeypatch.setattr(threading.Thread, "start", start)
|
||||
calls: list = []
|
||||
agents[:] = [RuntimeError("worker died before the running write"), _answering_agent(calls, "Real answer", tmp_path)]
|
||||
assert post().status_code == 500 and load_task_result(tmp_path, task_id) is None
|
||||
first = post()
|
||||
assert first.status_code == 200 and first.json()["text"] == "Real answer" and len(calls) == 1
|
||||
assert "previous_attempt" not in calls[0]["metadata"]["presence"]
|
||||
assert [row["direction"] for row in _chat_rows(tmp_path) if row.get("task_id") == task_id].count("in") == 1
|
||||
replay = post()
|
||||
assert replay.status_code == 200 and replay.json() == first.json() and len(calls) == 1
|
||||
|
||||
|
||||
def _lost_v1_attempt(tmp_path, task_id, *, chat_id):
|
||||
|
|
@ -300,122 +507,176 @@ def _lost_v1_attempt(tmp_path, task_id, *, chat_id):
|
|||
|
||||
|
||||
def _v1_kwargs(tmp_path, calls, reply="Real answer", **overrides):
|
||||
from dataclasses import replace
|
||||
|
||||
return dict(admission=_admission(), event=replace(_event(), delivery_reporting_version=1), repo_dir=tmp_path,
|
||||
drive_root=tmp_path, agent_factory=lambda **_kw: _answering_agent(calls, reply, tmp_path),
|
||||
gate=PresenceTurnGate(1), **overrides)
|
||||
|
||||
|
||||
def test_stale_running_row_is_a_lost_attempt_before_the_reconciler_runs(tmp_path):
|
||||
"""An adapter retry inside the reconciler's grace window still learns what the dead attempt sent."""
|
||||
from dataclasses import replace
|
||||
def _lose_attempt(tmp_path, monkeypatch, kind):
|
||||
"""A real v1 attempt: inbound row, running write, one confirmed tool send, then its worker is lost.
|
||||
|
||||
task_id = _task_id(_admission(), _event())
|
||||
chat_id = 0
|
||||
``kind`` is how the lost row persists: still RUNNING, marked INTERRUPTED, or ``reopened`` — a
|
||||
placeholder an earlier release had already moved back to RUNNING for a re-run that was lost too.
|
||||
"""
|
||||
calls: list = []
|
||||
kwargs = _v1_kwargs(tmp_path, calls)
|
||||
fresh = run_presence_turn(**{**kwargs, "event": replace(kwargs["event"], source_event_id="telegram:bot-1:1")})
|
||||
assert fresh.text == "Real answer" and "previous_attempt" not in calls[0]["metadata"]["presence"]
|
||||
chat_id = calls[0]["chat_id"]
|
||||
_lost_v1_attempt(tmp_path, task_id, chat_id=chat_id)
|
||||
assert _cached_result(tmp_path, task_id) is None
|
||||
|
||||
def confirmed_send(task):
|
||||
append_jsonl(tmp_path / "logs" / "chat.jsonl", {
|
||||
"type": "presence_delivery", "direction": "out", "chat_id": task["chat_id"], "text": "Early part",
|
||||
"task_id": task["id"], "transport": {"conversation_key": _event().conversation_key, "delivery": {
|
||||
"state": "delivered", "delivery_id": "send:early", "part_id": "0"}}})
|
||||
|
||||
with pytest.raises(RuntimeError, match="worker lost"):
|
||||
run_presence_turn(**{**_v1_kwargs(tmp_path, calls), "agent_factory": lambda **_kw: _answering_agent(
|
||||
calls, "", tmp_path, during=confirmed_send, lost=True)})
|
||||
task_id = calls[0]["id"]
|
||||
if kind == STATUS_INTERRUPTED:
|
||||
write_task_result(tmp_path, task_id, STATUS_INTERRUPTED, result="Task interrupted.")
|
||||
elif kind == "reopened":
|
||||
_reconciled(tmp_path, monkeypatch, task_id, seed=False)
|
||||
assert reopen_reconciled_presence_placeholder(tmp_path, task_id) is True
|
||||
assert load_task_result(tmp_path, task_id)["status"] == (STATUS_INTERRUPTED if kind == STATUS_INTERRUPTED
|
||||
else STATUS_RUNNING)
|
||||
return calls[0]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("kind", [STATUS_RUNNING, STATUS_INTERRUPTED, "reopened"])
|
||||
def test_stale_running_row_fails_closed_before_the_reconciler_runs(tmp_path, monkeypatch, kind):
|
||||
"""An adapter retry inside the reconciler's grace window finds the dead attempt's row: formerly it re-ran
|
||||
with the confirmed send as context, now it is refused until the attempt's own terminal replays."""
|
||||
lost = _lose_attempt(tmp_path, monkeypatch, kind)
|
||||
task_id = lost["id"]
|
||||
assert _cached_result(tmp_path, task_id) is None and _prior_sends(tmp_path, task_id) == (["Early part"], 0)
|
||||
built: list = []
|
||||
kwargs = _retry_kwargs(tmp_path, built)
|
||||
_assert_fails_closed(tmp_path, task_id, kwargs, built)
|
||||
_complete(lost, "Late answer", tmp_path) # a later canonical terminal
|
||||
first = run_presence_turn(**kwargs)
|
||||
assert first.text == "Real answer" and [task["id"] for task in calls] == [calls[0]["id"], task_id]
|
||||
assert calls[1]["metadata"]["presence"]["previous_attempt"] == {
|
||||
"delivered_count": 1, "delivered": ["Early part"], "uncertain_count": 0}
|
||||
assert (first.outcome, first.text, first.task_id) == ("message", "Late answer", task_id) and built == []
|
||||
assert run_presence_turn(**kwargs) == first and built == []
|
||||
stored = load_task_result(tmp_path, task_id)
|
||||
assert stored["status"] == STATUS_COMPLETED and "superseded_placeholder" not in stored # nothing to reopen
|
||||
rows = [json.loads(line) for line in (tmp_path / "logs" / "chat.jsonl").read_text(encoding="utf-8").splitlines()]
|
||||
assert sum(1 for row in rows if row.get("direction") == "in" and row.get("task_id") == task_id) == 1
|
||||
assert run_presence_turn(**kwargs) == first and len(calls) == 2
|
||||
assert stored["status"] == STATUS_COMPLETED and stored["terminal_origin"] == "model_final"
|
||||
assert [row["direction"] for row in _chat_rows(tmp_path) if row.get("task_id") == task_id].count("in") == 1
|
||||
|
||||
|
||||
def test_rotation_between_the_lost_attempt_and_its_retry_makes_prior_sends_unknown(tmp_path):
|
||||
"""Receipts in a rotated archive are not counted as zero: the model is told the count is unknown."""
|
||||
def test_a_lost_attempt_blocks_only_its_own_event(tmp_path, monkeypatch):
|
||||
"""Another event of the same conversation has no running row of its own: it runs as a fresh turn,
|
||||
and neither touches the lost row nor unblocks it."""
|
||||
lost = _lose_attempt(tmp_path, monkeypatch, STATUS_RUNNING)
|
||||
lost_row = task_result_path(tmp_path, lost["id"]).read_bytes()
|
||||
built: list = []
|
||||
fresh_kwargs = _retry_kwargs(tmp_path, built, reply="Fresh answer")
|
||||
fresh_kwargs["event"] = replace(fresh_kwargs["event"], source_event_id="telegram:bot-1:43")
|
||||
fresh = run_presence_turn(**fresh_kwargs)
|
||||
assert fresh.text == "Fresh answer" and fresh.task_id != lost["id"] and len(built) == 1
|
||||
assert task_result_path(tmp_path, lost["id"]).read_bytes() == lost_row
|
||||
built.clear()
|
||||
_assert_fails_closed(tmp_path, lost["id"], _retry_kwargs(tmp_path, built), built)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("reconciled", [False, True])
|
||||
def test_a_refused_retry_asks_the_owner_once_and_never_the_correspondent(tmp_path, monkeypatch, reconciled):
|
||||
"""With the installation's bridge bound, the refusal raises one owner-side question: a second retry
|
||||
does not repeat it, the correspondent's conversation never hears it, its once-stamp is the only row
|
||||
change, and an event with no running row of its own asks nothing."""
|
||||
from ouroboros.contracts.chat_id_policy import WEB_UI_CHAT_ID
|
||||
from supervisor import message_bus
|
||||
|
||||
lost = _lose_attempt(tmp_path, monkeypatch, STATUS_RUNNING)
|
||||
if reconciled:
|
||||
_reconciled(tmp_path, monkeypatch, lost["id"], seed=False)
|
||||
before, dialogue = load_task_result(tmp_path, lost["id"]), _trace(tmp_path, lost["id"])[2]
|
||||
notices: list = []
|
||||
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
|
||||
monkeypatch.setattr(message_bus, "try_get_bridge", lambda: object())
|
||||
monkeypatch.setattr(message_bus, "send_with_budget", lambda *args, **kwargs: notices.append((args, kwargs)))
|
||||
built: list = []
|
||||
kwargs = _retry_kwargs(tmp_path, built)
|
||||
for _second in (False, True):
|
||||
_refused(kwargs)
|
||||
assert built == [] and len(notices) == 1 and _trace(tmp_path, lost["id"])[2] == dialogue
|
||||
args, meta = notices[0]
|
||||
assert args[0] == WEB_UI_CHAT_ID != lost["chat_id"] and lost["id"] in args[1]
|
||||
assert meta["system_type"] == "presence_recovery_required"
|
||||
after, stamp = load_task_result(tmp_path, lost["id"]), {"presence_recovery_owner_notified", "updated_at"}
|
||||
assert after["presence_recovery_owner_notified"] and is_reconciled_presence_placeholder(after) is reconciled
|
||||
assert {k: v for k, v in after.items() if k not in stamp} == {k: v for k, v in before.items() if k not in stamp}
|
||||
|
||||
class Quiet:
|
||||
def handle_task(self, _task):
|
||||
return [{"type": "presence_result", "outcome": "silent", "text": "", "work_ref": ""}]
|
||||
|
||||
fresh = run_presence_turn(**{**kwargs, "event": replace(kwargs["event"], source_event_id="telegram:bot-1:43"),
|
||||
"agent_factory": lambda **_kw: Quiet()})
|
||||
assert fresh.outcome == "silent" and len(notices) == 1
|
||||
|
||||
|
||||
def test_rotation_between_the_lost_attempt_and_its_retry_keeps_prior_sends_unknown(tmp_path):
|
||||
"""Receipts in a rotated archive are not counted as zero, and a refused retry never re-logs the
|
||||
message into the live generation (formerly the re-run was told the count was unknown)."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
chat = _lost_v1_attempt(tmp_path, task_id, chat_id=7)
|
||||
(tmp_path / "archive").mkdir()
|
||||
chat.rename(tmp_path / "archive" / "chat_20260528T000100.jsonl") # the live generation rotated
|
||||
calls: list = []
|
||||
first = run_presence_turn(**_v1_kwargs(tmp_path, calls))
|
||||
assert first.text == "Real answer" and len(calls) == 1
|
||||
assert calls[0]["metadata"]["presence"]["previous_attempt"] == {
|
||||
"delivered_count": None, "delivered": None, "uncertain_count": 0}
|
||||
section = build_presence_context_section(tmp_path, calls[0]["metadata"]["presence"])
|
||||
assert "whether it already sent anything is unknown" in section and "delivered 0 message" not in section
|
||||
live = [json.loads(line) for line in chat.read_text(encoding="utf-8").splitlines()] if chat.exists() else []
|
||||
assert not [row for row in live if row.get("task_id") == task_id and row.get("direction") == "in"] # never re-logged
|
||||
built: list = []
|
||||
_assert_fails_closed(tmp_path, task_id, _retry_kwargs(tmp_path, built), built)
|
||||
assert not chat.exists() and _prior_sends(tmp_path, task_id) == (None, 0)
|
||||
|
||||
|
||||
def test_a_second_death_after_the_rotation_keeps_the_count_unknown(tmp_path):
|
||||
"""The retry that follows a rotation must not leave a fresh inbound row that a later retry mistakes
|
||||
for complete receipt coverage: the archived send stays unknown, never zero."""
|
||||
def test_a_second_refusal_after_the_rotation_keeps_the_count_unknown(tmp_path):
|
||||
"""No retry may leave a fresh inbound row a later reader mistakes for complete receipt coverage:
|
||||
formerly a second death proved it for the re-run; now neither refusal writes into the live file."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
chat = _lost_v1_attempt(tmp_path, task_id, chat_id=7)
|
||||
(tmp_path / "archive").mkdir()
|
||||
chat.rename(tmp_path / "archive" / "chat_20260528T000100.jsonl")
|
||||
calls: list = []
|
||||
with pytest.raises(RuntimeError): # retry A dies after its running write
|
||||
run_presence_turn(**{**_v1_kwargs(tmp_path, calls),
|
||||
"agent_factory": lambda **_kw: _answering_agent(calls, "", tmp_path, lost=True)})
|
||||
second = run_presence_turn(**_v1_kwargs(tmp_path, calls)) # retry B
|
||||
assert second.text == "Real answer" and len(calls) == 2
|
||||
assert [task["metadata"]["presence"]["previous_attempt"] for task in calls] == [
|
||||
{"delivered_count": None, "delivered": None, "uncertain_count": 0}] * 2
|
||||
chat.touch() # the rotator leaves a fresh live generation behind, as in production
|
||||
built: list = []
|
||||
kwargs = _retry_kwargs(tmp_path, built)
|
||||
for _attempt in ("A", "B"):
|
||||
_refused(kwargs)
|
||||
assert built == [] and chat.read_bytes() == b"" and _prior_sends(tmp_path, task_id) == (None, 0)
|
||||
|
||||
|
||||
def test_rejected_build_leaves_the_placeholder_for_the_next_retry(tmp_path, monkeypatch):
|
||||
"""The host mark moves aside only when the turn actually runs; a rejected build keeps it."""
|
||||
from ouroboros.presence_runner import PresenceTurnError
|
||||
from ouroboros.task_results import is_reconciled_presence_placeholder
|
||||
|
||||
def test_rejected_build_leaves_the_placeholder_and_the_next_retry_still_fails_closed(tmp_path, monkeypatch):
|
||||
"""The host mark never moves aside. Formerly a rejected build kept it for the re-run that followed; now
|
||||
the lost attempt is refused before any build or staging, whatever files the retry carries."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
_reconciled(tmp_path, monkeypatch, task_id)
|
||||
calls: list = []
|
||||
kwargs = dict(admission=_admission(), event=_event(), repo_dir=tmp_path, drive_root=tmp_path,
|
||||
agent_factory=lambda **_kw: _answering_agent(calls, "Real answer", tmp_path), gate=PresenceTurnGate(1))
|
||||
with pytest.raises(PresenceTurnError):
|
||||
run_presence_turn(**kwargs, staged_files=[tmp_path / "missing.png"])
|
||||
stored = load_task_result(tmp_path, task_id)
|
||||
assert is_reconciled_presence_placeholder(stored) and calls == [] # still the placeholder, not a phantom running row
|
||||
first = run_presence_turn(**kwargs)
|
||||
assert first.text == "Real answer" and "previous_attempt" in calls[0]["metadata"]["presence"]
|
||||
stored = load_task_result(tmp_path, task_id)
|
||||
assert stored["status"] == STATUS_COMPLETED and stored["superseded_placeholder"]["status"] == STATUS_FAILED
|
||||
row = _reconciled(tmp_path, monkeypatch, task_id)
|
||||
built: list = []
|
||||
kwargs = _retry_kwargs(tmp_path, built, version=0)
|
||||
_refused({**kwargs, "staged_files": [tmp_path / "missing.png"]})
|
||||
assert load_task_result(tmp_path, task_id) == row and built == [] # still the placeholder
|
||||
_assert_fails_closed(tmp_path, task_id, kwargs, built)
|
||||
assert load_task_result(tmp_path, task_id) == row
|
||||
|
||||
|
||||
def test_uncertain_receipts_make_the_prior_count_a_floor(tmp_path):
|
||||
"""A timed-out send the provider never confirmed is neither counted nor forgotten."""
|
||||
def test_uncertain_receipts_fail_closed_and_stay_a_floor(tmp_path):
|
||||
"""A timed-out send the provider never confirmed is neither counted nor forgotten — and it is exactly
|
||||
the effect a regenerated turn could duplicate, so the retry is refused (formerly it re-ran)."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
chat = _lost_v1_attempt(tmp_path, task_id, chat_id=7)
|
||||
append_jsonl(chat, {"type": "presence_delivery", "direction": "system", "chat_id": 7, "text": "Maybe part",
|
||||
"task_id": task_id, "transport": {"conversation_key": _event().conversation_key, "delivery": {
|
||||
"state": "uncertain", "delivery_id": "send:late", "part_id": "0"}}})
|
||||
calls: list = []
|
||||
run_presence_turn(**_v1_kwargs(tmp_path, calls))
|
||||
attempt = calls[0]["metadata"]["presence"]["previous_attempt"]
|
||||
assert attempt == {"delivered_count": 1, "delivered": ["Early part"], "uncertain_count": 1}
|
||||
section = build_presence_context_section(tmp_path, calls[0]["metadata"]["presence"])
|
||||
assert 'delivered at least 1 message(s): "Early part"; 1 more part(s) may have landed' in section
|
||||
built: list = []
|
||||
_assert_fails_closed(tmp_path, task_id, _retry_kwargs(tmp_path, built), built)
|
||||
assert _prior_sends(tmp_path, task_id) == (["Early part"], 1)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("states, uncertain", [(("uncertain", "failed"), 0), (("failed", "uncertain"), 1)])
|
||||
def test_a_part_settles_by_its_latest_receipt(tmp_path, states, uncertain):
|
||||
"""A timed-out part the provider later refused is neither delivered nor possibly landed; a refused
|
||||
part whose retry timed out may have landed after all."""
|
||||
part whose retry timed out may have landed after all. Either way the lost turn is not regenerated."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
chat = _lost_v1_attempt(tmp_path, task_id, chat_id=7)
|
||||
for state in states:
|
||||
append_jsonl(chat, {"type": "presence_delivery", "direction": "system", "chat_id": 7, "text": "Maybe part",
|
||||
"task_id": task_id, "transport": {"conversation_key": _event().conversation_key, "delivery": {
|
||||
"state": state, "delivery_id": "send:late", "part_id": "0"}}})
|
||||
calls: list = []
|
||||
run_presence_turn(**_v1_kwargs(tmp_path, calls))
|
||||
attempt = calls[0]["metadata"]["presence"]["previous_attempt"]
|
||||
assert attempt == {"delivered_count": 1, "delivered": ["Early part"], "uncertain_count": uncertain}
|
||||
section = build_presence_context_section(tmp_path, calls[0]["metadata"]["presence"])
|
||||
assert ("may have landed" in section) is bool(uncertain)
|
||||
built: list = []
|
||||
_assert_fails_closed(tmp_path, task_id, _retry_kwargs(tmp_path, built), built)
|
||||
assert _prior_sends(tmp_path, task_id) == (["Early part"], uncertain)
|
||||
|
||||
|
||||
def test_an_attempt_that_died_before_its_running_write_still_logs_the_message_once(tmp_path):
|
||||
|
|
@ -439,16 +700,16 @@ def test_an_attempt_that_died_before_its_running_write_still_logs_the_message_on
|
|||
|
||||
|
||||
def test_a_confirmed_part_later_refused_is_not_delivered(tmp_path):
|
||||
"""One latest-state rule for confirmed and uncertain parts alike."""
|
||||
"""One latest-state rule for confirmed and uncertain parts alike; zero confirmed sends is still no proof
|
||||
that the lost model/tool work had no effect, so the retry is refused (formerly it re-ran)."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
chat = _lost_v1_attempt(tmp_path, task_id, chat_id=7)
|
||||
append_jsonl(chat, {"type": "presence_delivery", "direction": "system", "chat_id": 7, "text": "Early part",
|
||||
"task_id": task_id, "transport": {"conversation_key": _event().conversation_key, "delivery": {
|
||||
"state": "failed", "delivery_id": "send:early", "part_id": "0"}}})
|
||||
calls: list = []
|
||||
run_presence_turn(**_v1_kwargs(tmp_path, calls))
|
||||
assert calls[0]["metadata"]["presence"]["previous_attempt"] == {
|
||||
"delivered_count": 0, "delivered": [], "uncertain_count": 0}
|
||||
built: list = []
|
||||
_assert_fails_closed(tmp_path, task_id, _retry_kwargs(tmp_path, built), built)
|
||||
assert _prior_sends(tmp_path, task_id) == ([], 0)
|
||||
|
||||
|
||||
def test_a_receipt_addressed_to_another_conversation_does_not_count(tmp_path):
|
||||
|
|
@ -458,10 +719,9 @@ def test_a_receipt_addressed_to_another_conversation_does_not_count(tmp_path):
|
|||
append_jsonl(chat, {"type": "presence_delivery", "direction": "out", "chat_id": 8, "text": "Early part",
|
||||
"task_id": task_id, "transport": {"conversation_key": "telegram:bot-1:other-room:0", "delivery": {
|
||||
"state": "delivered", "delivery_id": "send:elsewhere", "part_id": "0"}}})
|
||||
calls: list = []
|
||||
run_presence_turn(**_v1_kwargs(tmp_path, calls))
|
||||
assert calls[0]["metadata"]["presence"]["previous_attempt"] == {
|
||||
"delivered_count": 1, "delivered": ["Early part"], "uncertain_count": 0}
|
||||
built: list = []
|
||||
_assert_fails_closed(tmp_path, task_id, _retry_kwargs(tmp_path, built), built)
|
||||
assert _prior_sends(tmp_path, task_id) == (["Early part"], 0)
|
||||
|
||||
|
||||
def test_reconciler_skips_a_row_whose_retry_went_live_after_the_decision(tmp_path, monkeypatch):
|
||||
|
|
@ -517,7 +777,6 @@ def test_reconciler_settles_nothing_when_the_row_was_requeued_after_the_decision
|
|||
def test_an_unwritable_inbound_row_fails_the_turn_before_the_model_runs(tmp_path, monkeypatch):
|
||||
"""The no-re-log rule assumes the inbound row landed; a failed append is a failed turn, then a retry logs it."""
|
||||
from ouroboros import presence_runner
|
||||
from ouroboros.presence_runner import PresenceTurnError
|
||||
|
||||
calls: list = []
|
||||
kwargs = _v1_kwargs(tmp_path, calls)
|
||||
|
|
@ -546,27 +805,22 @@ def _sending_agent(calls, reply, drive_root, *, part, rotate_first=False):
|
|||
return _answering_agent(calls, reply, drive_root, during=send)
|
||||
|
||||
|
||||
def test_a_retry_after_a_rotation_still_knows_its_own_sends(tmp_path):
|
||||
"""The lost attempt's archived sends stay unknown to the retry, but the retry's own receipts, which
|
||||
all landed in the generation live when it started, are its pointer's confirmed sends."""
|
||||
from ouroboros.presence_runner import _read_previous_turn
|
||||
|
||||
def test_a_refused_retry_after_a_rotation_claims_no_sends_of_its_own(tmp_path):
|
||||
"""Formerly the re-run's own receipts in the fresh live generation became its pointer's confirmed sends.
|
||||
A refused retry sends nothing and names no turn: the conversation's pointer is never written."""
|
||||
task_id = _task_id(_admission(), _event())
|
||||
chat = _lost_v1_attempt(tmp_path, task_id, chat_id=7)
|
||||
(tmp_path / "archive").mkdir()
|
||||
chat.rename(tmp_path / "archive" / "chat_20260528T000100.jsonl")
|
||||
chat.touch() # the rotator leaves a fresh live generation behind, as in production
|
||||
calls: list = []
|
||||
kwargs = _v1_kwargs(tmp_path, calls)
|
||||
run_presence_turn(**{**kwargs, "agent_factory": lambda **_kw: _sending_agent(calls, "Real answer", tmp_path, part="Retry part")})
|
||||
assert calls[0]["metadata"]["presence"]["previous_attempt"]["delivered_count"] is None # the archived send
|
||||
pointer = _read_previous_turn(tmp_path, kwargs["event"].conversation_key)
|
||||
assert (pointer["transport_sends"], pointer["delivery"]) == (["Retry part"], "partly confirmed")
|
||||
kwargs = {**_v1_kwargs(tmp_path, calls), "agent_factory": lambda **_kw: _sending_agent(
|
||||
calls, "Real answer", tmp_path, part="Retry part")}
|
||||
_refused(kwargs)
|
||||
assert calls == [] and chat.read_bytes() == b"" and _read_previous_turn(tmp_path, _event().conversation_key) is None
|
||||
|
||||
|
||||
def test_a_rotation_during_the_turn_leaves_its_sends_unknown(tmp_path):
|
||||
from ouroboros.presence_runner import _read_previous_turn
|
||||
|
||||
calls: list = []
|
||||
kwargs = _v1_kwargs(tmp_path, calls)
|
||||
run_presence_turn(**{**kwargs, "agent_factory": lambda **_kw: _sending_agent(
|
||||
|
|
@ -575,43 +829,39 @@ def test_a_rotation_during_the_turn_leaves_its_sends_unknown(tmp_path):
|
|||
assert (pointer["transport_sends"], pointer["delivery"]) == ([], "unknown")
|
||||
|
||||
|
||||
def test_a_presence_placeholder_owes_no_terminal_projection_but_its_rerun_does(tmp_path, monkeypatch):
|
||||
"""The target's terminal-projection sweep must not post the placeholder's failure into the room; the
|
||||
re-run's own completion originates the room's terminal row, even when a stale marker was inherited."""
|
||||
from ouroboros.terminal_projection import reconcile_terminal_projections
|
||||
|
||||
from ouroboros.terminal_projection import SETTLEMENT_NONE, settle_terminal_projection
|
||||
def test_a_presence_placeholder_owes_no_terminal_projection_even_when_its_retry_is_refused(tmp_path, monkeypatch):
|
||||
"""The target's terminal-projection sweep must not post the placeholder's failure into the room, and a
|
||||
refused retry originates no terminal of its own (formerly the re-run's completion did), even when a
|
||||
stale marker was inherited."""
|
||||
from ouroboros.terminal_projection import (
|
||||
SETTLEMENT_NONE,
|
||||
reconcile_terminal_projections,
|
||||
settle_terminal_projection,
|
||||
)
|
||||
|
||||
task_id = _task_id(_admission(), _event())
|
||||
_reconciled(tmp_path, monkeypatch, task_id)
|
||||
assert reconcile_terminal_projections(tmp_path) == 0 # nothing owed for a placeholder
|
||||
chat = tmp_path / "logs" / "chat.jsonl"
|
||||
rows = [json.loads(line) for line in chat.read_text(encoding="utf-8").splitlines()] if chat.exists() else []
|
||||
assert not [row for row in rows if row.get("type") == "task_summary"]
|
||||
assert not [row for row in _chat_rows(tmp_path) if row.get("type") == "task_summary"]
|
||||
# An install on the previous release already recorded readiness for the placeholder (its chat append
|
||||
# failed): the direct settlement path, as startup recovery calls it, must not publish it either.
|
||||
write_task_result(tmp_path, task_id, STATUS_FAILED, canonical_terminal_projection_ready={
|
||||
"summary_id": f"task-terminal:{task_id}", "token": "stale", "attempt": {}, "task_done_ts": "2026-05-28T00:00:05+00:00",
|
||||
"chat_id": 7})
|
||||
assert settle_terminal_projection(tmp_path, task_id) == SETTLEMENT_NONE
|
||||
rows = [json.loads(line) for line in chat.read_text(encoding="utf-8").splitlines()] if chat.exists() else []
|
||||
assert not [row for row in rows if row.get("type") == "task_summary"]
|
||||
assert not [row for row in _chat_rows(tmp_path) if row.get("type") == "task_summary"]
|
||||
# An install that ran the sweep before this rule left a failed marker on the placeholder.
|
||||
write_task_result(tmp_path, task_id, STATUS_FAILED, canonical_terminal_projection={
|
||||
"summary_id": f"task-terminal:{task_id}", "summary_kind": "terminal_root_projection", "attempt": {}})
|
||||
calls: list = []
|
||||
first = run_presence_turn(admission=_admission(), event=_event(), repo_dir=tmp_path, drive_root=tmp_path,
|
||||
agent_factory=lambda **_kw: _answering_agent(calls, "Real answer", tmp_path),
|
||||
gate=PresenceTurnGate(1))
|
||||
built: list = []
|
||||
_assert_fails_closed(tmp_path, task_id, _retry_kwargs(tmp_path, built, version=0), built)
|
||||
stored = load_task_result(tmp_path, task_id)
|
||||
assert first.text == "Real answer" and stored["status"] == STATUS_COMPLETED
|
||||
assert {"canonical_terminal_projection", "canonical_terminal_projection_ready"} <= set(stored["superseded_placeholder"])
|
||||
assert reconcile_terminal_projections(tmp_path) == 1 # the re-run's completion owes and gets its row
|
||||
rows = [json.loads(line) for line in chat.read_text(encoding="utf-8").splitlines()]
|
||||
summaries = [row for row in rows if row.get("type") == "task_summary" and row.get("task_id") == task_id]
|
||||
assert [row["status"] for row in summaries] == ["completed"] and summaries[0]["presence_provenance"]["provider"] == "telegram"
|
||||
assert is_reconciled_presence_placeholder(stored) and "superseded_placeholder" not in stored
|
||||
assert {"canonical_terminal_projection", "canonical_terminal_projection_ready"} <= set(stored) # not moved aside
|
||||
assert reconcile_terminal_projections(tmp_path) == 0
|
||||
assert not [row for row in _chat_rows(tmp_path) if row.get("type") == "task_summary"]
|
||||
# An ordinary orphaned root (no presence identity) still gets its failed terminal row.
|
||||
_sweep(tmp_path, monkeypatch, "plain-orphan", metadata={"source": "chat"})
|
||||
assert reconcile_terminal_projections(tmp_path) == 1
|
||||
rows = [json.loads(line) for line in chat.read_text(encoding="utf-8").splitlines()]
|
||||
assert [row["status"] for row in rows if row.get("type") == "task_summary" and row.get("task_id") == "plain-orphan"] == ["failed"]
|
||||
assert [row["status"] for row in _chat_rows(tmp_path)
|
||||
if row.get("type") == "task_summary" and row.get("task_id") == "plain-orphan"] == ["failed"]
|
||||
|
|
|
|||
192
tests/test_presence_refusals.py
Normal file
192
tests/test_presence_refusals.py
Normal file
|
|
@ -0,0 +1,192 @@
|
|||
"""Recovery disposition is a producer fact, not a guess from an HTTP status."""
|
||||
import asyncio
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from ouroboros.gateway import host_service
|
||||
from ouroboros.presence_admission import PresenceAdmissionError
|
||||
from ouroboros.presence_runner import PresenceTurnError
|
||||
from ouroboros.presence_runner import PresenceTurnGate, presence_turn_task_id, run_presence_turn
|
||||
from ouroboros.presence_runner import _notify_unresolved_turn
|
||||
from ouroboros.task_results import load_task_result, task_result_path, write_task_result
|
||||
from tests.test_host_service_responsiveness import _answer, _event, _presence_app, _request, _turn
|
||||
from tests.test_presence_delivery import _payload
|
||||
|
||||
|
||||
def test_auth_block_and_origin_rejection_share_status_not_disposition(tmp_path):
|
||||
app, binding, _ctx = _presence_app(tmp_path, _answer)
|
||||
request = _request(app, {"binding_id": binding, "event": _event("event")})
|
||||
request.headers = {"x-skill-token": "invalid"}
|
||||
auth = asyncio.run(host_service._api_presence_turn(request))
|
||||
event = _event("event")
|
||||
event["provider"] = "wrong-provider"
|
||||
origin = asyncio.run(host_service._api_presence_turn(
|
||||
_request(app, {"binding_id": binding, "event": event})))
|
||||
assert auth.status_code == origin.status_code == 403
|
||||
assert json.loads(auth.body)["disposition"] == "blocked"
|
||||
assert json.loads(auth.body)["code"] == "presence_auth_blocked"
|
||||
assert json.loads(origin.body)["disposition"] == "rejected"
|
||||
assert json.loads(origin.body)["code"] == "presence_origin_mismatch"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("code,disposition", [
|
||||
("chat_log_unwritable", "retry"),
|
||||
("presence_result_missing", "retry"),
|
||||
("presence_attachment_admission_rejected", "rejected"),
|
||||
("presence_admission_conversation_mismatch", "rejected"),
|
||||
])
|
||||
def test_turn_refusals_keep_legacy_facts_and_semantic_disposition(tmp_path, code, disposition):
|
||||
manifest = [{"status": "rejected", "label": "input"}]
|
||||
|
||||
def refuse(**kwargs):
|
||||
raise PresenceTurnError(code, "staged_files", attachment_manifest=manifest)
|
||||
|
||||
app, binding, _ctx = _presence_app(tmp_path, refuse)
|
||||
response = asyncio.run(_turn(app, binding, "event"))
|
||||
assert response.status_code == 409
|
||||
body = json.loads(response.body)
|
||||
assert body == {"ok": False, "error": f"{code}: staged_files", "code": code,
|
||||
"field": "staged_files", "attachment_manifest": manifest,
|
||||
"disposition": disposition}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("code,disposition", [
|
||||
("presence_behavior_skill_disabled", "blocked"),
|
||||
("presence_required_capability_unavailable", "blocked"),
|
||||
("presence_bindings_unreadable", "retry"),
|
||||
("presence_binding_wrong_transport", "rejected"),
|
||||
])
|
||||
def test_admission_dispositions(tmp_path, monkeypatch, code, disposition):
|
||||
app, binding, _ctx = _presence_app(tmp_path, _answer)
|
||||
|
||||
def refuse(*args):
|
||||
raise PresenceAdmissionError(code, "binding_id")
|
||||
|
||||
monkeypatch.setattr(host_service, "_admit_presence", refuse)
|
||||
response = asyncio.run(_turn(app, binding, "event"))
|
||||
body = json.loads(response.body)
|
||||
assert response.status_code == 409
|
||||
assert body["code"] == code and body["field"] == "binding_id"
|
||||
assert body["disposition"] == disposition
|
||||
|
||||
|
||||
def test_receipt_write_retry_and_conflict_do_not_share_disposition(tmp_path, monkeypatch):
|
||||
app, _binding, ctx = _presence_app(tmp_path, _answer)
|
||||
payload = _payload()
|
||||
first = asyncio.run(host_service._api_presence_delivery(_request(app, payload)))
|
||||
assert first.status_code == 200
|
||||
changed = dict(payload, text="different facts")
|
||||
conflict = asyncio.run(host_service._api_presence_delivery(_request(app, changed)))
|
||||
assert conflict.status_code == 409
|
||||
assert json.loads(conflict.body)["disposition"] == "rejected"
|
||||
|
||||
def unwritable(*args):
|
||||
raise OSError("unreadable archive")
|
||||
|
||||
monkeypatch.setattr(ctx.presence_deliveries, "record", unwritable)
|
||||
unavailable = asyncio.run(host_service._api_presence_delivery(_request(app, payload)))
|
||||
assert unavailable.status_code == 503
|
||||
assert json.loads(unavailable.body)["disposition"] == "retry"
|
||||
assert json.loads(unavailable.body)["code"] == "presence_receipt_unwritable"
|
||||
|
||||
|
||||
def test_missing_work_is_rejected_without_losing_original_error(tmp_path):
|
||||
app, binding, _ctx = _presence_app(tmp_path, _answer)
|
||||
request = _request(app)
|
||||
request.path_params = {"work_ref": "missing-work"}
|
||||
request.query_params = {"binding_id": binding}
|
||||
response = asyncio.run(host_service._api_presence_work(request))
|
||||
assert response.status_code == 404
|
||||
assert json.loads(response.body) == {
|
||||
"ok": False, "error": "presence work reference not found",
|
||||
"code": "presence_work_not_found", "disposition": "rejected",
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("status,code,disposition", [
|
||||
("running", "presence_attempt_outcome_unknown", "retry"),
|
||||
("interrupted", "presence_attempt_outcome_unknown", "retry"),
|
||||
("cancelled", "presence_turn_cancelled", "blocked"),
|
||||
])
|
||||
def test_lost_or_cancelled_turn_never_regenerates_and_releases_host_capacity(tmp_path, status, code, disposition):
|
||||
attempted = []
|
||||
|
||||
def runner(**kwargs):
|
||||
return run_presence_turn(
|
||||
repo_dir=tmp_path, drive_root=tmp_path, gate=PresenceTurnGate(1),
|
||||
agent_factory=lambda **_kw: attempted.append("agent") or None, **kwargs,
|
||||
)
|
||||
|
||||
app, binding, ctx = _presence_app(tmp_path, runner)
|
||||
turn_id = presence_turn_task_id(binding, "event")
|
||||
write_task_result(tmp_path, turn_id, status, result="prior attempt unconfirmed",
|
||||
metadata={"source": "presence", "presence": {"binding_id": binding}})
|
||||
original = load_task_result(tmp_path, turn_id)
|
||||
for _ in range(2):
|
||||
response = asyncio.run(_turn(app, binding, "event"))
|
||||
body = json.loads(response.body)
|
||||
assert response.status_code == 409
|
||||
assert body["ok"] is False and body["code"] == code
|
||||
assert body["disposition"] == disposition and body["turn_ref"] == turn_id
|
||||
assert not body.get("text") and attempted == []
|
||||
assert ctx.presence_turns.live() == [] and not any(ctx._inflight.values())
|
||||
assert load_task_result(tmp_path, turn_id) == original
|
||||
|
||||
|
||||
def test_unresolved_turn_asks_owner_once_and_never_addresses_correspondent(tmp_path, monkeypatch):
|
||||
from supervisor import message_bus
|
||||
|
||||
task_id = "presence-unresolved"
|
||||
write_task_result(tmp_path, task_id, "running", metadata={"source": "presence"})
|
||||
deliveries = []
|
||||
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
|
||||
monkeypatch.setattr(message_bus, "try_get_bridge", lambda: object())
|
||||
monkeypatch.setattr(message_bus, "send_with_budget", lambda *args, **kwargs: deliveries.append((args, kwargs)))
|
||||
_notify_unresolved_turn(tmp_path, task_id)
|
||||
_notify_unresolved_turn(tmp_path, task_id)
|
||||
assert len(deliveries) == 1
|
||||
args, kwargs = deliveries[0]
|
||||
assert args[0] == 1 and task_id in args[1]
|
||||
assert kwargs == {"role": "system", "system_type": "presence_recovery_required", "require_write": True}
|
||||
assert load_task_result(tmp_path, task_id)["presence_recovery_owner_notified"]
|
||||
|
||||
|
||||
def test_unrecorded_owner_question_never_gets_a_notified_stamp(tmp_path, monkeypatch):
|
||||
from supervisor import message_bus
|
||||
|
||||
task_id = "presence-unrecorded"
|
||||
write_task_result(tmp_path, task_id, "running", metadata={"source": "presence"})
|
||||
monkeypatch.setattr(message_bus, "DATA_DIR", tmp_path)
|
||||
monkeypatch.setattr(message_bus, "try_get_bridge", lambda: object())
|
||||
monkeypatch.setattr(message_bus, "get_bridge", lambda: pytest.fail("unrecorded message was sent"))
|
||||
monkeypatch.setattr(message_bus, "log_chat", lambda *args, **kwargs: (_ for _ in ()).throw(OSError("disk")))
|
||||
_notify_unresolved_turn(tmp_path, task_id)
|
||||
assert not load_task_result(tmp_path, task_id).get("presence_recovery_owner_notified")
|
||||
assert not (tmp_path / "logs" / "chat.jsonl").exists()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("quarantined", [False, True])
|
||||
def test_unreadable_or_quarantined_result_never_becomes_a_new_turn(tmp_path, quarantined):
|
||||
invoked = []
|
||||
|
||||
def runner(**kwargs):
|
||||
return run_presence_turn(
|
||||
repo_dir=tmp_path, drive_root=tmp_path, gate=PresenceTurnGate(1),
|
||||
agent_factory=lambda **_kw: invoked.append("agent") or None, **kwargs,
|
||||
)
|
||||
|
||||
app, binding, ctx = _presence_app(tmp_path, runner)
|
||||
task_id = presence_turn_task_id(binding, "event")
|
||||
result_path = task_result_path(tmp_path, task_id)
|
||||
if quarantined:
|
||||
result_path = result_path.parent / "quarantine" / result_path.name
|
||||
result_path.parent.mkdir(parents=True)
|
||||
result_path.write_text("{broken", encoding="utf-8")
|
||||
for _ in range(2):
|
||||
response = asyncio.run(_turn(app, binding, "event"))
|
||||
body = json.loads(response.body)
|
||||
assert response.status_code == 409 and body["code"] == "presence_result_unreadable"
|
||||
assert body["disposition"] == "retry" and body["turn_ref"] == task_id
|
||||
assert invoked == [] and ctx.presence_turns.live() == [] and not any(ctx._inflight.values())
|
||||
assert result_path.read_text(encoding="utf-8") == "{broken"
|
||||
424
tests/test_quota_rotation_fallback.py
Normal file
424
tests/test_quota_rotation_fallback.py
Normal file
|
|
@ -0,0 +1,424 @@
|
|||
"""Quota recovery order: Auto account rotation, then the configured fallback, then the owner.
|
||||
|
||||
Engine evidence, serving bundle ``state/cx/3.14.0-8849e922607b/claudexord.bundle.cjs``:
|
||||
``resolve38`` resolves ONE account per model operation before dispatch and
|
||||
``ModelOperations.execute`` invokes it once ("A model operation may send inference
|
||||
only once"). A vendor HTTP quota refusal is a settled, failed result naming that
|
||||
account; ``observe`` ingests its ``resetsAt``/``retryAfterMs`` as the account's
|
||||
cooldown. A refusal before dispatch is the engine's own verdict, and only its
|
||||
``poolCause=quota`` says every compatible account is quota-blocked.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import queue
|
||||
import time
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
from ouroboros import fallback_cooldown, loop, loop_llm_call, model_wait
|
||||
from ouroboros import llm_claudexor as transport
|
||||
from ouroboros import usage_accounting as ua
|
||||
from ouroboros.loop_llm_call import RETRY_WALL_EXHAUSTED_KEY, call_llm_with_retry, provider_no_call_source
|
||||
from ouroboros.loop_model_call import RESOURCE_REFUSAL_KEY
|
||||
from ouroboros.loop_transport import provider_recovery_hint
|
||||
from ouroboros.model_slots import MODEL_ACCOUNTS_KEY
|
||||
from ouroboros.presence_runner import _presence_delivery
|
||||
from tests.test_llm_claudexor import MODEL, ROUTE, ledger, result
|
||||
from tests.test_llm_claudexor import setup as gateway_fixture
|
||||
from tests.test_model_wait import live_wait as wait_fixture
|
||||
from tests.test_model_wait_controls import _loop_tools
|
||||
from tests.test_subscription_main_wait import main_call as main_call_fixture
|
||||
|
||||
setup = gateway_fixture
|
||||
live_wait = wait_fixture
|
||||
main_call = main_call_fixture
|
||||
ROUTE_B = {**ROUTE, "credentialProfileId": "account-b", "accountFingerprint": "fingerprint-b"}
|
||||
RESET = "2099-01-01T00:00:00Z"
|
||||
FALLBACK = "openai::alternate"
|
||||
|
||||
|
||||
def _vendor_quota(route=ROUTE, *, reset=True, http=True):
|
||||
"""The vendor refused this account over HTTP: dispatched, settled, nothing generated."""
|
||||
context = {"vendorCode": "usage_limit_reached", **({"httpStatus": 429} if http else {}),
|
||||
**({"resetsAt": RESET} if reset else {})}
|
||||
return result(outcome="failed", route=route, problem={
|
||||
"code": "subscription_window_exhausted", "message": "Codex model request was refused (HTTP 429).",
|
||||
"retryable": True, "context": context})
|
||||
|
||||
|
||||
def _pool_quota():
|
||||
"""The engine's own verdict before dispatch: every compatible account is quota-blocked."""
|
||||
return result(outcome="failed", route={}, problem={
|
||||
"code": "subscription_window_exhausted", "retryable": False,
|
||||
"message": "Every available model account is blocked by subscription quota",
|
||||
"context": {"source": "codex", "poolCause": "quota", "resetsAt": RESET}})
|
||||
|
||||
|
||||
def _final(text):
|
||||
row = result()
|
||||
row["message"] = {"role": "assistant", "content": text}
|
||||
return row
|
||||
|
||||
|
||||
def _history():
|
||||
"""A conversation whose last same-route answer came from account-a."""
|
||||
return [{"role": "user", "content": "go"}, result()["message"],
|
||||
{"role": "tool", "tool_call_id": "a", "content": "A"},
|
||||
{"role": "tool", "tool_call_id": "b", "content": "B"}]
|
||||
|
||||
|
||||
def _rows(root, kind):
|
||||
path = root / "logs" / "events.jsonl"
|
||||
rows = [json.loads(line) for line in path.read_text().splitlines()] if path.exists() else []
|
||||
return [row for row in rows if row.get("type") == kind]
|
||||
|
||||
|
||||
def _waits(events):
|
||||
return [event for event in list(events.queue) if event.get("type") == "task_model_wait"]
|
||||
|
||||
|
||||
# -- account rotation (transport) ---------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.parametrize("asynchronous", [False, True])
|
||||
def test_unpinned_vendor_quota_reasks_the_same_payload_and_the_engine_picks_the_next_account(setup, asynchronous):
|
||||
root, gateway, client = setup
|
||||
gateway.results, gateway.dispatch = [_vendor_quota(), result(route=ROUTE_B)], ["response_received"] * 2
|
||||
call = client.chat_async if asynchronous else client.chat
|
||||
answer = call(_history(), MODEL, None, "high", model_role="main", cache_affinity="affinity")
|
||||
message, usage = asyncio.run(answer) if asynchronous else answer
|
||||
assert message["content"] == result()["message"]["content"]
|
||||
assert usage["claudexor"]["route"]["credentialProfileId"] == "account-b"
|
||||
assert usage["claudexor"]["account_rotation"] == {"refused_accounts": ["account-a"]}
|
||||
first, second = (payload for payload, _key in gateway.uploads)
|
||||
assert first["account"] == {"mode": "auto", "preferredProfileId": "account-a"}
|
||||
# Same model, effort, messages and tools: only the refused account's preference is gone.
|
||||
assert second == {**first, "account": {"mode": "auto"}}
|
||||
assert len(gateway.accepted_operations) == 2 and [row["state"] for row in ledger(root)].count("settled") == 2
|
||||
assert [(row["account"], row["disposition"]) for row in _rows(root, "model_account_rotation")] == [
|
||||
("account-a", "reask")]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("dispatch", ["response_received", "not_started"])
|
||||
def test_a_pinned_route_is_never_rotated(setup, monkeypatch, dispatch):
|
||||
_root, gateway, client = setup
|
||||
monkeypatch.setenv(MODEL_ACCOUNTS_KEY, json.dumps({"main": "account-a"}))
|
||||
gateway.results, gateway.dispatch = [_vendor_quota()], [dispatch]
|
||||
with pytest.raises(transport.ClaudexorModelError) as refused:
|
||||
client.chat(_history(), MODEL, model_role="main", cache_affinity="affinity")
|
||||
assert [payload["account"] for payload, _key in gateway.uploads] == [{"mode": "pin", "profileId": "account-a"}]
|
||||
assert refused.value.account_rotation == {
|
||||
"refused_accounts": ["account-a"], "stop": "pinned_account", "pool_exhausted": False}
|
||||
assert len(gateway.accepted_operations) == 1
|
||||
|
||||
|
||||
@pytest.mark.parametrize("second,stop,pool_exhausted", [
|
||||
(_vendor_quota(), "engine_reselected_refused_account", False),
|
||||
(_pool_quota(), "pool_exhausted", True),
|
||||
])
|
||||
def test_rotation_ends_on_engine_evidence_and_claims_exhaustion_only_from_the_pool_verdict(
|
||||
setup, second, stop, pool_exhausted):
|
||||
"""A re-selected refused account is not asked a third time, and proves nothing about the pool."""
|
||||
root, gateway, client = setup
|
||||
gateway.results = [_vendor_quota(), second]
|
||||
gateway.dispatch = ["response_received", "response_received" if second["route"] else "not_started"]
|
||||
with pytest.raises(transport.ClaudexorModelError) as refused:
|
||||
client.chat(_history(), MODEL, model_role="main", cache_affinity="affinity")
|
||||
assert len(gateway.accepted_operations) == 2
|
||||
assert refused.value.account_rotation == {
|
||||
"refused_accounts": ["account-a"], "stop": stop, "pool_exhausted": pool_exhausted}
|
||||
assert [row["disposition"] for row in _rows(root, "model_account_rotation")] == ["reask", stop]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("refusal,dispatch,stop", [
|
||||
(_pool_quota(), "not_started", "pool_exhausted"),
|
||||
(_vendor_quota(http=False), "response_received", "generation_not_excluded"),
|
||||
])
|
||||
def test_an_engine_verdict_or_a_stream_level_refusal_is_never_asked_again(setup, refusal, dispatch, stop):
|
||||
root, gateway, client = setup
|
||||
gateway.results, gateway.dispatch = [refusal], [dispatch]
|
||||
with pytest.raises(transport.ClaudexorModelError) as refused:
|
||||
client.chat(_history(), MODEL, model_role="main")
|
||||
assert refused.value.account_rotation["stop"] == stop and len(gateway.accepted_operations) == 1
|
||||
|
||||
|
||||
def test_an_unknown_outcome_is_never_rotated_or_resent(setup):
|
||||
root, gateway, client = setup
|
||||
gateway.results, gateway.dispatch = [result(outcome="unknown")], ["unknown"]
|
||||
with pytest.raises(transport.ClaudexorModelError) as unknown:
|
||||
client.chat(_history(), MODEL, model_role="main")
|
||||
assert unknown.value.code == "model_outcome_unknown" and len(gateway.accepted_operations) == 1
|
||||
assert not _rows(root, "model_account_rotation")
|
||||
|
||||
|
||||
def test_a_spent_send_budget_ends_rotation(setup):
|
||||
_root, gateway, client = setup
|
||||
gateway.results, gateway.dispatch = [_vendor_quota(), result(route=ROUTE_B)], ["response_received"] * 2
|
||||
with ua.physical_attempt_limit(1), pytest.raises(transport.ClaudexorModelError) as refused:
|
||||
client.chat(_history(), MODEL, model_role="main")
|
||||
assert refused.value.account_rotation["stop"] == "send_budget_spent"
|
||||
assert len(gateway.accepted_operations) == 1
|
||||
|
||||
|
||||
# -- the round: configured fallback before the owner, and the owner from the retained refusal ------
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def one_fallback(monkeypatch):
|
||||
monkeypatch.setenv("OUROBOROS_MODEL_FALLBACKS", FALLBACK)
|
||||
monkeypatch.setenv("OUROBOROS_TASK_REVIEW_MODE", "off")
|
||||
monkeypatch.setattr(fallback_cooldown, "is_cooling_down", lambda *_args: False)
|
||||
|
||||
|
||||
def _api_fallback(ctx, monkeypatch, *, answer):
|
||||
"""The configured API fallback answers or fails; the subscription primary stays real."""
|
||||
calls, remote = [], ctx.llm._chat_remote
|
||||
|
||||
def send(target, messages, schemas, *args, **kwargs):
|
||||
if target["provider"] == "claudexor":
|
||||
return remote(target, messages, schemas, *args, **kwargs)
|
||||
calls.append(target["provider"])
|
||||
if not answer:
|
||||
raise RuntimeError("fixture API fallback failed")
|
||||
return {"role": "assistant", "content": "Finished by the fallback"}, {
|
||||
"prompt_tokens": 1, "completion_tokens": 1, "cost": 0.0, "provider": "openai"}
|
||||
|
||||
monkeypatch.setattr(ctx.llm, "_chat_remote", send)
|
||||
return calls
|
||||
|
||||
|
||||
def _run(ctx, tools, events):
|
||||
return loop.run_llm_loop(ctx.messages, tools, ctx.llm, ctx.drive_logs, lambda *_args, **_kwargs: None,
|
||||
queue.Queue(), task_id="task-one", drive_root=ctx.drive_root, event_queue=events)
|
||||
|
||||
|
||||
def test_after_the_fallback_fails_the_owner_wait_opens_from_the_retained_refusal(main_call, one_fallback, monkeypatch):
|
||||
"""No primary generation opens the owner question; the primary sends again only after it resolves."""
|
||||
ctx, gateway, owner, events, _decide, _observations = main_call
|
||||
tools = _loop_tools(ctx, owner)
|
||||
gateway.results = [_pool_quota(), _final("Finished after the owner wait")]
|
||||
gateway.dispatch = ["not_started", "response_received"]
|
||||
api_calls = _api_fallback(ctx, monkeypatch, answer=False)
|
||||
at_owner_question = []
|
||||
catalog = ctx.llm.claudexor_model_catalog
|
||||
|
||||
def observed_catalog(*args, **kwargs):
|
||||
at_owner_question.append((len(gateway.accepted_operations), list(api_calls),
|
||||
len({event["wait_id"] for event in _waits(events)})))
|
||||
return catalog(*args, **kwargs)
|
||||
|
||||
monkeypatch.setattr(ctx.llm, "claudexor_model_catalog", observed_catalog)
|
||||
text, _usage, _trace = _run(ctx, tools, events)
|
||||
assert text == "Finished after the owner wait"
|
||||
# The wait was already open, after the fallback, with only the ORIGINAL refusal sent.
|
||||
assert at_owner_question[0] == (1, ["openai"], 1)
|
||||
waits = _waits(events)
|
||||
assert len({event["wait_id"] for event in waits}) == 1 and waits[-1]["resolution"] == "resource_available"
|
||||
assert waits[0]["reason"] == "quota" and waits[0]["reset_at"] == RESET
|
||||
assert waits[0]["account_rotation"] == {"refused_accounts": [], "stop": "pool_exhausted", "pool_exhausted": True}
|
||||
assert len(gateway.accepted_operations) == 2
|
||||
|
||||
|
||||
def test_a_fallback_answer_leaves_no_owner_question_and_no_refusal(main_call, one_fallback, monkeypatch):
|
||||
ctx, gateway, owner, events, _decide, _observations = main_call
|
||||
tools = _loop_tools(ctx, owner)
|
||||
gateway.results, gateway.dispatch = [_pool_quota()], ["not_started"]
|
||||
api_calls = _api_fallback(ctx, monkeypatch, answer=True)
|
||||
monkeypatch.setattr(ctx.llm, "claudexor_model_catalog", lambda *_a, **_kw: pytest.fail("no owner question"))
|
||||
text, usage, _trace = _run(ctx, tools, events)
|
||||
assert text == "Finished by the fallback" and api_calls == ["openai"]
|
||||
assert not _waits(events) and RESOURCE_REFUSAL_KEY not in usage and len(gateway.accepted_operations) == 1
|
||||
|
||||
|
||||
def test_an_unknown_fallback_outcome_asks_no_owner_and_sends_nothing_more(main_call, one_fallback, monkeypatch):
|
||||
ctx, gateway, owner, _events, _decide, _observations = main_call
|
||||
gateway.results, gateway.dispatch = [_pool_quota()], ["not_started"]
|
||||
assert loop._call_round_model(ctx)[0] is None and ctx.tools._ctx._deferred_resource_refusal is not None
|
||||
original = loop._call_round_model
|
||||
|
||||
def candidate(round_call):
|
||||
assert round_call.defer_resource_wait is True # the primary's owner question would follow
|
||||
round_call.accumulated_usage["_last_llm_error_kind"] = "provider_outcome_unknown"
|
||||
return None, 0.0, "max"
|
||||
|
||||
monkeypatch.setattr(loop, "_call_round_model", candidate)
|
||||
message, *_rest = loop._run_cross_model_fallback_chain(
|
||||
llm=ctx.llm, ctx=ctx.tools._ctx, tools=ctx.tools, messages=ctx.messages, active_model=ctx.active_model,
|
||||
active_use_local=False, tool_schemas=[], active_effort="medium", max_retries=3, drive_logs=ctx.drive_logs,
|
||||
task_id=ctx.task_id, round_idx=1, event_queue=_events, accumulated_usage=ctx.accumulated_usage,
|
||||
task_type="task", emit_progress=lambda *_a, **_kw: None, context_fit_plan=ctx.context_fit_plan,
|
||||
active_context_mode="max")
|
||||
monkeypatch.setattr(loop, "_call_round_model", original)
|
||||
assert message is None and not _waits(_events) and len(gateway.accepted_operations) == 1
|
||||
refusal = ctx.accumulated_usage[RESOURCE_REFUSAL_KEY]
|
||||
assert (refusal["owner_wait"], refusal["fallbacks_tried"]) == ("not_asked", [FALLBACK])
|
||||
# The unknown-outcome fence still outranks: nothing is resent, not even a forced final.
|
||||
assert provider_no_call_source(ctx.accumulated_usage, False)[0] == "provider_outcome_unknown_no_resend"
|
||||
assert provider_no_call_source({RESOURCE_REFUSAL_KEY: refusal}, True) == ("resource_refusal_no_resend", True)
|
||||
|
||||
|
||||
def test_a_candidate_defers_only_while_a_later_route_or_the_owner_question_follows(main_call, monkeypatch):
|
||||
ctx, _gateway, _owner, _events, _decide, _observations = main_call
|
||||
monkeypatch.setenv("OUROBOROS_MODEL_FALLBACKS", "openai::one,openai::two")
|
||||
monkeypatch.setattr(fallback_cooldown, "is_cooling_down", lambda *_args: False)
|
||||
monkeypatch.setattr(loop, "_rebind_context_fit_plan", lambda plan, *_a, **_kw: (plan, "max"))
|
||||
seen = []
|
||||
monkeypatch.setattr(loop, "_call_round_model",
|
||||
lambda call: seen.append((call.active_model, call.defer_resource_wait)) or (None, 0.0, "max"))
|
||||
loop._run_cross_model_fallback_chain(
|
||||
llm=ctx.llm, ctx=ctx.tools._ctx, tools=ctx.tools, messages=ctx.messages, active_model=ctx.active_model,
|
||||
active_use_local=False, tool_schemas=[], active_effort="medium", max_retries=1, drive_logs=ctx.drive_logs,
|
||||
task_id=ctx.task_id, round_idx=1, event_queue=None, accumulated_usage={}, task_type="task",
|
||||
emit_progress=lambda *_a, **_kw: None, context_fit_plan=ctx.context_fit_plan, active_context_mode="max")
|
||||
assert seen == [("openai::one", True), ("openai::two", False)]
|
||||
|
||||
|
||||
# -- Presence: never waits, typed temporary refusal, no speech ------------------------------------
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def presence(main_call, monkeypatch):
|
||||
ctx, gateway, owner, events, _decide, _observations = main_call
|
||||
owner.task["_presence_turn"] = True
|
||||
|
||||
def never(*_args, **_kwargs):
|
||||
pytest.fail("a Presence turn must not wait for quota or the owner")
|
||||
|
||||
monkeypatch.setattr(model_wait, "time", SimpleNamespace(monotonic=time.monotonic, time=time.time, sleep=never))
|
||||
monkeypatch.setattr(loop_llm_call, "_sleep_within_deadline", never)
|
||||
monkeypatch.setattr(ctx.llm, "claudexor_model_catalog", never)
|
||||
return main_call
|
||||
|
||||
|
||||
@pytest.mark.parametrize("configured_fallback", [True, False])
|
||||
def test_presence_quota_turn_never_waits_and_ends_in_a_typed_temporary_refusal_without_speech(
|
||||
presence, monkeypatch, configured_fallback):
|
||||
ctx, gateway, owner, events, _decide, _observations = presence
|
||||
if configured_fallback:
|
||||
monkeypatch.setenv("OUROBOROS_MODEL_FALLBACKS", FALLBACK)
|
||||
monkeypatch.setattr(fallback_cooldown, "is_cooling_down", lambda *_args: False)
|
||||
monkeypatch.setenv("OUROBOROS_TASK_REVIEW_MODE", "off")
|
||||
api_calls = _api_fallback(ctx, monkeypatch, answer=False)
|
||||
gateway.results, gateway.dispatch = [_pool_quota()], ["not_started"]
|
||||
text, usage, _trace = _run(ctx, _loop_tools(ctx, owner), events)
|
||||
assert not _waits(events) and len(gateway.accepted_operations) == 1
|
||||
assert api_calls == (["openai"] if configured_fallback else [])
|
||||
refusal = usage[RESOURCE_REFUSAL_KEY]
|
||||
assert {key: refusal[key] for key in ("reason", "role", "model", "fallbacks_tried", "owner_wait", "temporary")} == {
|
||||
"reason": "quota", "role": "main", "model": MODEL, "owner_wait": "not_allowed", "temporary": True,
|
||||
"fallbacks_tried": [FALLBACK] if configured_fallback else []}
|
||||
assert refusal["account_rotation"]["pool_exhausted"] is True and refusal["reset_at"] == RESET
|
||||
assert provider_no_call_source(usage, False) == ("resource_refusal_no_resend", True)
|
||||
assert usage["execution_status"] == "infra_failed"
|
||||
# Host-authored terminal: the correspondent hears nothing.
|
||||
assert _presence_delivery("message", text, str(usage.get("terminal_origin") or "")) == ("silent", "")
|
||||
|
||||
|
||||
def test_presence_sign_in_refusal_is_typed_without_an_owner_wait(presence, monkeypatch):
|
||||
ctx, gateway, _owner, events, _decide, _observations = presence
|
||||
refusal = _pool_quota()
|
||||
refusal["problem"] = {**refusal["problem"], "code": "auth_required", "context": {"poolCause": "auth"}}
|
||||
gateway.results, gateway.dispatch = [refusal], ["not_started"]
|
||||
assert loop._call_round_model(ctx)[0] is None
|
||||
assert not _waits(events) and ctx.accumulated_usage[RESOURCE_REFUSAL_KEY]["reason"] == "auth"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("role", ["light", "vision"])
|
||||
def test_presence_helper_roles_such_as_safety_refuse_at_once_instead_of_waiting(presence, role):
|
||||
"""Safety keeps its own fail-closed refusal; it only stops waiting inside a Presence turn."""
|
||||
ctx, gateway, _owner, events, _decide, _observations = presence
|
||||
gateway.results, gateway.dispatch = [_pool_quota()], ["not_started"]
|
||||
with pytest.raises(transport.ClaudexorModelNotDispatched):
|
||||
ctx.llm.chat([{"role": "user", "content": "check"}], MODEL, model_role=role)
|
||||
assert not _waits(events) and len(gateway.accepted_operations) == 1
|
||||
|
||||
|
||||
# -- no timer, no same-request resend; the owner wait needs evidence, not the same refused account --
|
||||
|
||||
|
||||
def test_main_retry_never_sleeps_to_a_reset_nor_resends_a_spent_window(tmp_path, monkeypatch):
|
||||
monkeypatch.setattr(loop_llm_call, "_sleep_within_deadline", lambda *_a, **_kw: pytest.fail("slept to a reset"))
|
||||
refusal = transport.ClaudexorModelError(_vendor_quota()["problem"], route=ROUTE)
|
||||
refusal.physical_attempt_capture = SimpleNamespace(state="settled", attempt_id="a-1", provider="claudexor",
|
||||
route_is_loopback=False)
|
||||
calls = []
|
||||
|
||||
class Client:
|
||||
def chat(self, **kwargs):
|
||||
calls.append(kwargs)
|
||||
raise refusal
|
||||
|
||||
usage = {}
|
||||
message, _cost = call_llm_with_retry(Client(), [{"role": "user", "content": "go"}], MODEL, None, "medium", 3,
|
||||
tmp_path / "logs", "task-one", 1, None, usage, deadline_ts=None)
|
||||
assert message is None and len(calls) == 1 and usage[RETRY_WALL_EXHAUSTED_KEY] is True
|
||||
|
||||
|
||||
@pytest.mark.parametrize("presence", [False, True])
|
||||
def test_an_unproven_dated_pool_keeps_its_reset_timer_only_where_waiting_is_allowed(tmp_path, monkeypatch, presence):
|
||||
slept = []
|
||||
monkeypatch.setattr(loop_llm_call, "_sleep_within_deadline", lambda seconds, *_a, **_kw: slept.append(seconds) and False)
|
||||
dated = transport.ClaudexorModelNotDispatched({"code": "credential_pool_exhausted", "message": "no account",
|
||||
"context": {"poolCause": "unavailable", "resetsAt": RESET}})
|
||||
dated.physical_attempt_capture = SimpleNamespace(state="released", attempt_id="a-1", provider="claudexor",
|
||||
route_is_loopback=False)
|
||||
|
||||
class Client:
|
||||
def chat(self, **_kwargs):
|
||||
raise dated
|
||||
|
||||
with model_wait.task_model_wait_scope(task={"id": "task-one", "_presence_turn": presence}, drive_root=tmp_path,
|
||||
event_queue=None, worker_slot_held=False):
|
||||
call_llm_with_retry(Client(), [{"role": "user", "content": "go"}], MODEL, None, "medium", 3,
|
||||
tmp_path / "logs", "task-one", 1, None, {}, deadline_ts=None)
|
||||
assert (slept == []) is presence and all(seconds > 3600 for seconds in slept)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("code,context,outage", [
|
||||
("credential_pool_exhausted", {"poolCause": "mixed"}, False),
|
||||
("model_account_unavailable", {}, True),
|
||||
])
|
||||
def test_an_engine_resource_verdict_is_not_a_network_outage(code, context, outage):
|
||||
refusal = transport.ClaudexorModelNotDispatched({"code": code, "message": code, "context": context})
|
||||
refusal.physical_attempt_capture = SimpleNamespace(state="released", provider="claudexor", route_is_loopback=False)
|
||||
assert (loop_llm_call.classify_llm_exception(refusal).kind == "transport_unavailable") is outage
|
||||
|
||||
|
||||
def test_an_owner_wait_without_reset_evidence_does_not_resend_to_the_same_refused_account(live_wait, monkeypatch):
|
||||
root, gateway, client, _owner, events, _decide = live_wait
|
||||
gateway.results = [_vendor_quota(reset=False), _vendor_quota(reset=False), result(route=ROUTE_B)]
|
||||
gateway.dispatch = ["response_received"] * 3
|
||||
clock = SimpleNamespace(now=time.monotonic())
|
||||
monkeypatch.setattr(model_wait, "time", SimpleNamespace(
|
||||
monotonic=lambda: clock.now, time=time.time, sleep=lambda seconds: setattr(clock, "now", clock.now + seconds)))
|
||||
chosen = iter(["account-a", "account-b"])
|
||||
seen = []
|
||||
|
||||
def catalog(*_args, **_kwargs):
|
||||
seen.append((next(chosen), len(gateway.accepted_operations)))
|
||||
return {"source": "codex", "credentialProfileId": seen[-1][0], "models": [{"id": "exact-model"}]}
|
||||
|
||||
monkeypatch.setattr(client, "claudexor_model_catalog", catalog)
|
||||
answer, usage = client.chat(_history(), MODEL, model_role="main", cache_affinity="affinity")
|
||||
assert usage["claudexor"]["route"]["credentialProfileId"] == "account-b"
|
||||
# The engine's choice of the refused account, with no cooldown behind it, resent nothing.
|
||||
assert seen == [("account-a", 2), ("account-b", 2)] and len(gateway.accepted_operations) == 3
|
||||
waits = _waits(events)
|
||||
assert waits[0]["account_rotation"]["stop"] == "engine_reselected_refused_account"
|
||||
assert waits[0]["account_rotation"]["pool_exhausted"] is False
|
||||
assert waits[-1]["resolution"] == "resource_available"
|
||||
|
||||
|
||||
def test_the_terminal_hint_claims_every_account_only_from_the_pool_verdict():
|
||||
base = {"_last_llm_error_kind": "subscription_window_exhausted", "_last_llm_reset_at": RESET}
|
||||
proven = provider_recovery_hint({**base, RESOURCE_REFUSAL_KEY: {
|
||||
"account_rotation": {"stop": "pool_exhausted", "pool_exhausted": True}, "fallbacks_tried": [FALLBACK]}})
|
||||
unproven = provider_recovery_hint({**base, RESOURCE_REFUSAL_KEY: {
|
||||
"account_rotation": {"stop": "engine_reselected_refused_account", "pool_exhausted": False}}})
|
||||
assert "every compatible account blocked" in proven and FALLBACK in proven
|
||||
assert "every compatible account" not in unproven and "unproven" in unproven
|
||||
assert "scheduled" not in proven + unproven and "nothing sleeps" in unproven
|
||||
Loading…
Add table
Add a link
Reference in a new issue