Merge pull request #1264 from ouroboros-agent/claude/claudexor-3.14.0-pin-20260924

Pin managed Claudexor to 3.14.0, undated pool refusals, alias labels in choosers
This commit is contained in:
Anton Razzhigaev 2026-09-24 22:48:55 +03:00 • committed by GitHub
commit 06ebf8a377
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
10 changed files with 163 additions and 55 deletions

View file

@ -311,7 +311,7 @@ A definite refusal needs a typed cause, no custody handle and either producer `d
**Configured dispatch is exact** (`subagent_runtime.resolve_configured_actor_dispatch`, resolved at dispatch, the last moment route availability is current): a bad choice returns a typed reason, current alternatives and any reset time with `host_fallback: false`; the host never ranks alternatives, converts session work to API work or picks the first healthy row. The WHY is recorded so no redesign undoes it: selecting an `agent_session` row IS the parent's typed LLM decision that this work executes on the harness, the FLOOR the host hardcodes (truth, money, authorship), while topology, decomposition and supervision judgment remain the model's CEILING (BIBLE P13/P5). At completion `actual_substrate` derives only from custody rows; a typed startup refusal never authorizes work on a different substrate: another session route or API fallback is an explicit LLM-selected start.
**Route health.** `subagents.route_health` is the one health reader for dispatcher, nanny and review slots, and it does not guess admission: the doctor `status` describes only the DEFAULT credential store while real accounts live in the engine's credential-profile pool, so admission is the engine's, whose start POST answers an empty pool with its own typed refusal (`credential_pool_exhausted` + earliest reset) at $0. `enabled` is honored as `route_disabled` (the owner's switch, not an observation); the other structural refusals are `route_not_in_capability_catalog`, an access-profile mismatch, `engine_rejects_delegated_marker` and positively proven quota exhaustion. A fully-used ratio needs a valid future reset before it can refuse a route; an incomplete reading delegates admission to the engine and never certifies available quota; explicit active cooldowns bind. Review slots inherit "the engine decides", never a silent fallback onto metered API spend: on the `auto` lane a dead daemon keeps the native fallback with its visible marker, while "daemon alive, pool empty" is discovered at the engine and disclosed. The model sees a semi-stable facts-only catalog (`model_visible_subagent_catalog`); invalid, list-disabled or all-rows-off configuration projects nothing; dispatch is the live check.
**Route health.** `subagents.route_health` is the one health reader for dispatcher, nanny and review slots, and it does not guess admission: the doctor `status` describes only the DEFAULT credential store while real accounts live in the engine's credential-profile pool, so admission is the engine's, whose start POST answers an empty pool with its own typed refusal (`credential_pool_exhausted` + earliest reset) at $0; the gateway types it as a timer-healing window, scheduled on its reset like a spent subscription window, only when `resetsAt` is dated; an undated (structural) pool stays a plain `ClaudexorUnavailable`. `enabled` is honored as `route_disabled` (the owner's switch, not an observation); the other structural refusals are `route_not_in_capability_catalog`, an access-profile mismatch, `engine_rejects_delegated_marker` and positively proven quota exhaustion. A fully-used ratio needs a valid future reset before it can refuse a route; an incomplete reading delegates admission to the engine and never certifies available quota; explicit active cooldowns bind. Review slots inherit "the engine decides", never a silent fallback onto metered API spend: on the `auto` lane a dead daemon keeps the native fallback with its visible marker, while "daemon alive, pool empty" is discovered at the engine and disclosed. The model sees a semi-stable facts-only catalog (`model_visible_subagent_catalog`); invalid, list-disabled or all-rows-off configuration projects nothing; dispatch is the live check.
Historical helper observations use `state/subagent_last_delegation.json`, owned by `subagent_history.py`. Task finalization joins configured API actors to their own attempts, preserving requested pins and observed routes/accounts; fallback success never certifies the failed route, nor does code/test failure condemn a provider. Sparse pre-response failures stay in usage/history/raw events; served-call trace references retain received response-backed calls, including incomplete responses. Foreground/recovered session terminals and typed start failures carry original effort/processing through `RunCustody`/`STARTED`/`START_FAILED` facts, without history-only archive reads; missing options stay unknown, distinct from captured defaults. Pre-invocation bootstrap refusals keep their source; reviewers keep separate receipts. Known occurrence and first observation differ: replay neither refreshes age nor replaces newer dated evidence; definite evidence may resolve an unknown start. `MAX_CONFIGURED_SUBAGENTS` bounds latest actor rows, with old single receipts readable. Dynamic Runtime context and owner UI consume history without changing catalog, admission, quota, dispatch or ranking.

View file

@ -1,12 +1,12 @@
{
"schema_version": 1,
"release": {
"version": "3.13.0",
"build_sha": "cc34a19ee1b0b41bc4fb3627a73ca40e1bf72e7a",
"version": "3.14.0",
"build_sha": "8849e922607b5cd741f4996264b026ed778181e5",
"protocol_major": 3,
"archive_url": "https://github.com/razzant/claudexor/releases/download/v3.13.0/claudexor-runtime-3.13.0.tar.gz",
"sha256": "91248839a11bf2e342605bc5a1cd217f75ce2201007299820483c9b5203a6ce2",
"size_bytes": 7851937,
"archive_url": "https://github.com/razzant/claudexor/releases/download/v3.14.0/claudexor-runtime-3.14.0.tar.gz",
"sha256": "b1e915e1bd87398259fa8a76e240284fbf0a44ff8e0fc26c610dcfd41662cdd1",
"size_bytes": 7869179,
"node_version": "24.16.0",
"node_artifacts": {
"darwin-arm64": {

View file

@ -54,22 +54,6 @@ async def run_sync_to_completion(function, /, *args, **kwargs):
raise
def _read_jsonl_segment_with_gaps(
path: pathlib.Path,
*,
tail_bytes: int | None = None,
) -> tuple[list, set[str]]:
"""Gateway wrapper over ``jsonl_tail.read_jsonl_segment_with_gaps``.
The parser is THIS module's ``iter_jsonl_objects`` name, resolved at call
time, so the gateway tests that monkeypatch it keep governing every gateway
read (``tests/test_gateway_history.py``).
"""
from ouroboros.jsonl_tail import read_jsonl_segment_with_gaps
return read_jsonl_segment_with_gaps(path, tail_bytes=tail_bytes, iter_objects=iter_jsonl_objects)
def read_rotated_jsonl_entries(
live: pathlib.Path,
archive_dir: pathlib.Path,

View file

@ -123,8 +123,11 @@ class ClaudexorUnavailable(RuntimeError):
# Cross-repo contract (B1): the engine's window-exhausted RunFailure codes. A
# newer engine (rotation PR-A) reports a spent credential POOL under its own
# code; both heal on a timer, so both map onto the exhausted class below — with
# the ORIGINAL code preserved. Any other code stays a generic
# code. A spent subscription window always heals on a timer (the engine always
# dates it); a pool heals on a timer only when the engine DATED it (`resetsAt`),
# an undated pool (every enabled account refused the model, every account
# disabled) is structural. `_window_exhausted_refusal` is the one reader; the
# ORIGINAL code is preserved either way. Any other code stays a generic
# ClaudexorUnavailable: fail-open, old engines emitting code:null included.
WINDOW_EXHAUSTED_CODES = ("subscription_window_exhausted", "credential_pool_exhausted")
@ -138,6 +141,20 @@ class ClaudexorSubscriptionWindowExhausted(ClaudexorUnavailable):
self.reset_at = str(reset_at or "")
def _window_exhausted_refusal(code: str, message: str, resets_at: Any, *,
status_code: int = 0) -> Optional[ClaudexorSubscriptionWindowExhausted]:
"""The timer-healing class for a window-exhausted code, or None for the caller's
plain refusal. Keyed on the structured ``resetsAt`` field only, never on prose: a
pool maps only when the engine dated it (a non-empty string), so an undated pool
never masquerades as a quota timer."""
dated = isinstance(resets_at, str) and bool(resets_at.strip())
if code not in WINDOW_EXHAUSTED_CODES or (
code != "subscription_window_exhausted" and not dated):
return None
return ClaudexorSubscriptionWindowExhausted(
message, reset_at=str(resets_at or ""), status_code=status_code, code=code)
REPORTED_CAUSE_CHARS = 512 # strict bound of the record field, omission marker included
@ -161,10 +178,8 @@ def run_failure_error(run_id: str, run_state: str, failure: Any) -> ClaudexorUna
message = (f"delegated review session {run_id} ended {run_state or 'unknown'}"
+ (f": {json.dumps(failure, ensure_ascii=False)}" if failure else ""))
code = str(failure.get("code") or "")
exc = (ClaudexorSubscriptionWindowExhausted(
message, reset_at=str(failure.get("resetsAt") or ""), code=code)
if code in WINDOW_EXHAUSTED_CODES
else ClaudexorUnavailable(code or f"run_{run_state or 'unknown'}", message))
exc = (_window_exhausted_refusal(code, message, failure.get("resetsAt"))
or ClaudexorUnavailable(code or f"run_{run_state or 'unknown'}", message))
exc.reported_cause = run_failure_cause(failure)
return exc
@ -471,19 +486,17 @@ class ClaudexorGateway:
if isinstance(item, str) and item
)
# The CODE decides, exactly as every other classification on this seam does.
# Sniffing `context` for a reset key instead had it both ways: no producer puts
# `resetsAt`/`resets_at`/`cooldownUntil` in a ControlProblem context (a spent
# window is reported as a run-detail RunFailure, and `cooldown_until` lives in a
# quota snapshot), so the transient class was unreachable — while any unrelated
# refusal that happened to carry one, an `idempotency_conflict` say, would have
# been announced as a spent subscription window and retried on a timer.
if code in WINDOW_EXHAUSTED_CODES:
return ClaudexorSubscriptionWindowExhausted(
message, reset_at=str(context.get("resetsAt") or ""),
status_code=response.status_code, code=code,
)
return ClaudexorUnavailable(code, message, status_code=response.status_code,
required_actions=required_actions)
# Sniffing `context` for a reset key on ANY code had it both ways: an unrelated
# refusal that happened to carry one, an `idempotency_conflict` say, was announced
# as a spent subscription window and retried on a timer. So a reset key is read
# only inside a window code, and there only the pool's own `resetsAt` says whether
# it is a timer at all. At engine 3.14.0 the daemon serializes no `resetsAt` into a
# pool ControlProblem context (the dated producer is the run-detail RunFailure, and
# `cooldown_until` lives in a quota snapshot), so this seam yields the plain class.
return (_window_exhausted_refusal(code, message, context.get("resetsAt"),
status_code=response.status_code)
or ClaudexorUnavailable(code, message, status_code=response.status_code,
required_actions=required_actions))
# -- operations ------------------------------------------------------------

View file

@ -722,13 +722,12 @@ def classify_llm_exception(exc: Exception, safe_error: str = "") -> LlmErrorClas
if isinstance(exc, LocalContextTooLargeError):
return LlmErrorClassification("context_overflow", False)
# Structured fact, not a keyword scan (Bible P5): a transport that KNOWS its
# window is spent carries the typed code plus the reset instant.
if str(getattr(exc, "code", "") or "") == SUBSCRIPTION_WINDOW_EXHAUSTED:
reset_at = str(getattr(exc, "reset_at", "") or "")
return LlmErrorClassification(
SUBSCRIPTION_WINDOW_EXHAUSTED, True, _exception_status_code(exc),
"", seconds_until(reset_at), reset_at,
)
# window is spent carries the typed code plus the reset instant; a DATED credential
# pool heals on the same timer (its code stays as evidence), an undated one never gets here.
wcode, reset_at = str(getattr(exc, "code", "") or ""), str(getattr(exc, "reset_at", "") or "")
if wcode == SUBSCRIPTION_WINDOW_EXHAUSTED or (wcode == "credential_pool_exhausted" and reset_at):
return LlmErrorClassification(SUBSCRIPTION_WINDOW_EXHAUSTED, True, _exception_status_code(exc),
"" if wcode == SUBSCRIPTION_WINDOW_EXHAUSTED else wcode, seconds_until(reset_at), reset_at)
status_code = _exception_status_code(exc)
provider_code = _exception_provider_code(exc, safe)
# Typed refusal (llm_attempt.ProviderPolicyRefusal): nothing upstream answered,

View file

@ -154,7 +154,7 @@ def test_tracked_pin_names_a_cli_capable_release():
of this proposal converged on while the pin was still pre-CLI 3.6.0)."""
pin = runtime.load_runtime_pin()
assert pin is not None
assert pin.version == "3.13.0"
assert pin.version == "3.14.0"
assert pin.cli_entrypoint == "claudexor.bundle.cjs"

View file

@ -307,6 +307,82 @@ def test_the_window_class_is_transient_and_scheduled_by_its_reset():
assert classification.reset_at == "2030-01-01T00:00:00Z"
def _control_problem(code: str, context: dict, status: int = 409) -> cx.ClaudexorUnavailable:
def handler(_request: httpx.Request) -> httpx.Response:
return httpx.Response(status, json={
"code": code, "message": f"{code} refusal", "retryable": False, "context": context})
with _gateway(handler) as gateway:
with pytest.raises(cx.ClaudexorUnavailable) as excinfo:
gateway.get_run("run-1")
return excinfo.value
def test_a_pool_heals_on_a_timer_only_when_the_engine_dated_it():
"""Owner decision: a credential pool exhausted for a STRUCTURAL reason (every enabled
account refused the model, every account disabled: the engine names no reset) must not
masquerade as a quota timer. Both producer seams (run failure, ControlProblem) key on
the structured `resetsAt` through one helper; nothing reads the prose."""
reset = "2030-01-01T00:00:00Z"
dated_failure = {"code": "credential_pool_exhausted", "resetsAt": reset,
"safeMessage": "every account is cooling down"}
for dated in (_control_problem("credential_pool_exhausted", {"resetsAt": reset}),
cx.run_failure_error("run-1", "failed", dated_failure)):
assert isinstance(dated, cx.ClaudexorSubscriptionWindowExhausted)
assert (dated.code, dated.reset_at) == ("credential_pool_exhausted", reset)
dated_run = cx.run_failure_error("run-1", "failed", dated_failure)
assert dated_run.reported_cause == "every account is cooling down"
# A dated pool heals on the same timer as a spent window, so it is SCHEDULED
# against its reset (not the short backoff) and keeps its own code as evidence.
dated_class = classify_llm_exception(dated_run)
assert (dated_class.kind, dated_class.retry_same_request) == (SUBSCRIPTION_WINDOW_EXHAUSTED, True)
assert dated_class.provider_code == "credential_pool_exhausted"
assert dated_class.reset_at == reset and dated_class.retry_after_sec is not None
assert dated_class.retry_after_sec > 60.0
# Other direction: an absent, empty or null reset is structural: a plain refusal
# under the SAME code, the engine's words still carried beside it.
undated_failure = {"code": "credential_pool_exhausted",
"safeMessage": "every enabled account refused the model"}
undated = [
_control_problem("credential_pool_exhausted", {}),
_control_problem("credential_pool_exhausted", {"resetsAt": ""}),
_control_problem("credential_pool_exhausted", {"resetsAt": None}),
cx.run_failure_error("run-1", "failed", undated_failure),
cx.run_failure_error("run-1", "failed", {**undated_failure, "resetsAt": ""}),
]
for exc in undated:
assert type(exc) is cx.ClaudexorUnavailable
assert exc.code == "credential_pool_exhausted"
run_undated = undated[3]
assert run_undated.reported_cause == "every enabled account refused the model"
# It classifies exactly as a plain refusal of the same code: no timer, no instant.
classification = classify_llm_exception(run_undated)
assert classification == classify_llm_exception(
cx.ClaudexorUnavailable("credential_pool_exhausted", str(run_undated)))
assert classification.kind == "provider_error"
assert classification.kind != SUBSCRIPTION_WINDOW_EXHAUSTED
assert (classification.retry_after_sec, classification.reset_at) == (None, "")
# A spent subscription window keeps its unconditional mapping (the engine always
# dates it): even an undated one stays the window class.
for context in ({"resetsAt": reset}, {}):
window = _control_problem("subscription_window_exhausted", context, status=429)
assert isinstance(window, cx.ClaudexorSubscriptionWindowExhausted)
assert classify_llm_exception(window).kind == SUBSCRIPTION_WINDOW_EXHAUSTED
def test_an_invalid_request_stays_a_permanent_plain_refusal():
"""The code decides: a stray reset instant never turns a request refusal into a
timer, and the engine's 400 stays non-retryable."""
exc = _control_problem("invalid_request", {"resetsAt": "2030-01-01T00:00:00Z"}, status=400)
assert type(exc) is cx.ClaudexorUnavailable and exc.code == "invalid_request"
classification = classify_llm_exception(exc)
assert classification.kind == "bad_request"
assert classification.retry_same_request is False
assert (classification.retry_after_sec, classification.reset_at) == (None, "")
def test_a_billing_refusal_stays_permanently_classified():
classification = classify_llm_exception(RuntimeError("402 payment required"))
assert classification.kind == "quota_exhausted"

View file

@ -666,8 +666,8 @@ def test_session_is_never_restarted_for_format_repair(tmp_path, fake_route, monk
def test_a_pool_exhausted_terminal_is_typed_like_a_spent_window(tmp_path, fake_route):
"""Cross-repo forward-compat (B1): a newer engine reports a spent credential POOL
with its own RunFailureCode. Same timer-healing semantics, same exception class —
with the ORIGINAL code preserved, never relabelled. An unknown code stays the
with its own RunFailureCode. A DATED pool gets the same timer-healing semantics and
exception class — with the ORIGINAL code preserved, never relabelled. An unknown code stays the
generic typed refusal (fail-open: old engines emit code:null and behave as today)."""
from ouroboros.gateways.claudexor import (
ClaudexorSubscriptionWindowExhausted, ClaudexorUnavailable)

View file

@ -201,18 +201,27 @@ function routeCatalogItems(route, items = []) {
/**
* One suggestion per model: the label names the model and makes no account claim
* (DESIGN.md §7). Availability, the reading account and its observation time are
* account facts, so they never travel on a model option.
* account facts, so they never travel on a model option. Two row facts do: what an
* alias resolves to (`resolved_model`, shown only while every supplying row agrees)
* and a row known only from the engine's frozen list (`origin: "hint"` on every
* supplying row; absent origin is live). The value stays the row id either way.
*/
export function catalogModelOptions(items = []) {
const values = new Map();
for (const item of items) {
const value = String(item?.value || item?.id || item);
const name = String(item?.name || item?.label || '');
if (!values.has(value)) values.set(value, { value, label: value, named: false, live: false, resolved: new Set() });
const current = values.get(value);
if (!current) values.set(value, { value, label: name || value, named: Boolean(name) });
else if (name && !current.named) Object.assign(current, { label: name, named: true });
if (name && !current.named) Object.assign(current, { label: name, named: true });
if (item?.origin !== 'hint') current.live = true;
const resolved = item?.resolved_model;
if (typeof resolved === 'string' && resolved && resolved !== value) current.resolved.add(resolved);
}
return [...values.values()].map(({ value, label }) => ({ value, label }));
return [...values.values()].map(({ value, label, live, resolved }) => {
const named = resolved.size === 1 ? `${value} → ${[...resolved][0]}` : label;
return { value, label: live ? named : `${named} (shipped list)` };
});
}
/** Suggestions carry the model alone; the source select already names the provider. */

View file

@ -53,6 +53,33 @@ test('an option label is byte-identical with one account and with eighteen', ()
assert.equal(many[0].label.length, 'Same model'.length);
});
test('an alias names its resolution and a frozen-list row says so; the value stays the row id', () => {
const alias = { id: 'default', label: 'Default (recommended)', origin: 'live', resolved_model: 'claude-opus-5-5[1m]' };
const hint = { id: 'claude-sonnet-5', label: null, origin: 'hint', resolved_model: null };
const [aliasOption, hintOption] = catalogModelOptions([alias, hint]);
assert.deepEqual(aliasOption, { value: 'default', label: 'default → claude-opus-5-5[1m]' });
assert.deepEqual(hintOption, { value: 'claude-sonnet-5', label: 'claude-sonnet-5 (shipped list)' });
// The one projection feeds the session chooser too: the value is what travels.
assert.deepEqual(sessionModelOptions({ models: [alias, hint] }, 'default').slice(1), [aliasOption, hintOption]);
// Other direction: a row without the fields (a 3.13 engine), a live row, a resolution
// equal to the id and an empty one all label exactly as before.
for (const row of [{ id: 'gpt-5.6-sol', name: 'GPT-5.6 Sol' }, { id: 'gpt-5.6-sol', name: 'GPT-5.6 Sol', origin: 'live' },
{ id: 'gpt-5.6-sol', name: 'GPT-5.6 Sol', resolved_model: 'gpt-5.6-sol' }, { id: 'gpt-5.6-sol', name: 'GPT-5.6 Sol', resolved_model: '' }]) {
assert.deepEqual(catalogModelOptions([row]), [{ value: 'gpt-5.6-sol', label: 'GPT-5.6 Sol' }]);
}
assert.deepEqual(catalogModelOptions(['plain-id']), [{ value: 'plain-id', label: 'plain-id' }]);
// Account-free: eighteen supplying accounts label byte-identically to one; the suffix
// needs every supplier on the frozen list, and a disputed resolution is not shown.
for (const row of [alias, hint]) {
const eighteen = Array.from({ length: 18 }, (_, index) => ({ ...row, credential_profile_id: `acct-${index + 1}`,
availability: index % 2 ? 'unavailable' : 'available', observed_at: '2026-09-24T00:00:00Z' }));
assert.equal(catalogModelOptions(eighteen)[0].label, catalogModelOptions([eighteen[0]])[0].label);
assert.doesNotMatch(catalogModelOptions(eighteen)[0].label, /acct|18|available|2026-09-24/);
}
assert.equal(catalogModelOptions([hint, { ...hint, origin: 'live' }])[0].label, 'claude-sonnet-5');
assert.equal(catalogModelOptions([alias, { ...alias, resolved_model: 'claude-sonnet-5' }])[0].label, 'Default (recommended)');
});
test('a nameless first duplicate yields to a later name, and otherwise the value is the label', () => {
const nameless = { value: 'same', id: 'same', name: undefined, label: undefined, credential_profile_id: 'personal' };
assert.equal(catalogModelOptions([nameless, { ...nameless, name: 'Named' }])[0].label, 'Named');