Reviewer P2: a lost response proves no Host-executed tool or correspondent action from the failed attempt, not that nothing ran anywhere — priced read-only provider-owned retrieval (server-side web search) may rerun on the repeat. The 409 and owner notice apply when the round ends unresolved without a further permitted repeat (deadline or a finalize control can refuse the first repeat). Test comment on the owner-notice writer corrected; chapter budgets untouched.
Regression drives the real round dispatcher with a ledger-backed fake provider: one dispatched ReadError, one repeat, two attempts, no tool run or send between them, no forced final, no presence_unknown_outcome. test_provider_terminal_notice converted from the pre-TZ3 200-deferred expectation to the 409 presence_attempt_outcome_unknown contract with admitted child custody preserved. Ch.06/ch.12 state the contract.
presence_runner._terminal_refusal keys on a missing canonical terminal or the new durable presence_unknown_outcome marker (stamped by agent_task_pipeline from the loop no-resend predicate); confirmed outages and overflows keep the TZ2 deferred projection with admitted child custody. Regressions drive the real emit_task_results pipeline over host-fallback, salvaged-draft and death-record shapes.
Independent exact-range review (Astra, 55ad78213..15dbc7533) found two P1 paths
where the Host acknowledged a Presence event the model never answered:
- a quota-refused primary whose fallback died with an unknown outcome is worded
provider_unavailable by the forced rail (the unknown fence outranks the
refusal source), so a guard keyed on resource_refusal_no_resend alone let the
host-authored terminal project to completed/silent (HTTP 200) live and on
replay, and the adapter could drop the event;
- a terminal write that fails after the durable start barrier is swallowed by
the pipeline, and the in-memory presence_result envelope was returned as the
answer while the durable row still said RUNNING.
One helper, _terminal_refusal, now owns the verdict the durable row demands
before acknowledgement: a confirmed resource refusal keeps the event
(presence_resources_unavailable); no canonical terminal (RUNNING/INTERRUPTED,
reconciled placeholder, lost terminal write) or any infra_failed execution
axis is presence_attempt_outcome_unknown — 409, empty external speech, the
admitted work_ref preserved (durable metadata, else the execution's own
handoff fact), one serialized owner-only notice, never a retry certificate.
Both guards (live and cached replay) consume it. Two consumer regressions
through the real Host app cover both paths; test fakes that returned an
envelope without writing a terminal now write the durable row the real
pipeline writes. Chapter 12 records the rule.
Four independent reviewers (Fable 5.1, Opus 5.5, Codex gpt-6-sol, Grok 4.7)
ran the Ouroboros scope-review brief on the reworked candidate; all four
returned "merge after small fixes" and no round-1 blocker survived. The
accepted findings, each mirroring the MiniMax region pattern:
- The one-click Colab collector never asked for ZAI_API_KEY, so a Z.ai-only
notebook was prompted for OpenRouter (colab_bootstrap.provider_keys).
- ZAI_PLAN is a transport choice, not a credential: it no longer satisfies the
onboarding "has a provider" checks (settings_setup_contract, the wizard's
two lists), and an unknown plan value is refused at settings save and by the
provider Test instead of silently selecting pay-as-you-go.
- The Capability Evidence route readers (gateway/settings._active_main_route,
reviewer_window.reviewer_route) resolve the plan's endpoint, so the two
Z.ai plans no longer share one route fingerprint.
- The task loop classifies Z.ai's HTTP 429 code 1113 as quota_exhausted
(typed code, exact match) instead of retrying it as a transient rate limit.
- OpenRouter's GLM namespace is z-ai/, so the catalog label key follows it.
- Docs: Z.ai joins the exclusive-direct-provider list in 02-startup-onboarding
and the forbidden skill settings list in CREATING_SKILLS; the 02-naming
sentence names the two provider-specific projections instead of implying
direct OpenAI carries no effort; glm ids join the onboarding suggestions.
Declined as disproportionate or out of scope, with the reason recorded in
the review ledger: the CHECKLISTS.md prose parenthetical (protected file;
the runtime deny list is authoritative), the GAIA/Terminal-Bench launcher
key lists (benchmark-only), a recorded live wire fixture (no key), and
GLM-5.2's skip-thinking on none/minimal (owner-accepted, disclosed in the
external-fact inventory).
zai:: joins the direct providers exactly the way deepseek:: did: prefix and
credential registry, ZAI_API_KEY plus a ZAI_PLAN endpoint selector
(provider_models.resolve_zai_base_url: empty/payg = api.z.ai/api/paas/v4,
coding = the Coding Plan endpoint), the routing target, live catalog fetch,
provider Test, settings card, onboarding contract, review-fallback roles,
single-provider startup and review detection, secret masking, benchmark
env hygiene, and docs.
Reasoning effort now reaches Z.ai. The provider serves an ABSENT
reasoning_effort at its maximum tier, so every call on the old generic
compatible route was billed at max regardless of the configured effort.
The canonical scale is projected onto Z.ai's own low/high/max enum
(ZAI_REASONING_EFFORT_ALIASES: none/minimal -> low, medium -> high,
xhigh/ultra -> max), disclosed as reasoning_effort_clamped when the tier
changes; GLM-5.3 rejects every other value and cannot disable thinking
(HTTP 400 code 1210), and forced tool_choice works with thinking on, so
there is no DeepSeek-style suppression arm. The projection is keyed on the
provider id the owner configured, never on a model name: a GLM served
from an owner's own OpenAI-compatible endpoint keeps today's behavior.
The provider port is the contributor's own work from the closed PR #1194,
narrowed to Z.ai (the DashScope and Moonshot lanes were not measured and
stay out). 07-configuration gains two settings rows and one route
paragraph (budget 37300 -> 38400), 02-naming records the dated Z.ai
probe in the external-fact inventory, and the onboarding bootstrap
fixture and data-layout inventory are regenerated.
Co-authored-by: josephsteuerjr <josephsteuerjr@gmail.com>
Revert the route-descriptor registry, the descriptor-driven effort pickers,
the Anthropic builder indirection and the effort_not_carried plumbing back
to the target's state, and restore DeepSeek's documented projection
(xhigh -> high per api-docs.deepseek.com Thinking Mode).
What the maintainers found on the combined tree: nine target-green tests
went red (five llm_golden compatible/cloudru/minimax cases, three
host-message-metadata cases, the provider_models/config leaf guard), the
07-configuration chapter left its byte budget, the picker hid saved and
working tiers (xhigh/max/none) and matched the catalog by bare model id,
the reviewer-slot wiring passed a fourth argument that the local wrapper
dropped, and the new usage fact never reached events.jsonl. The z.ai GLM
effort carriage itself is kept and lands as a first-class provider in the
next commit.
Kept from this change: Z.ai's HTTP 429 code 1113 ("Insufficient balance")
is a billing fact, so the provider Test maps it to "No credits" instead of
"Rate limited" (typed code only; the status copy stays provider-neutral).
Expose a body-only edit mode with one exact anchor and revision CAS. Preserve frontmatter, capture the complete old/new source, and report an honest host-derived delta without selecting the semantic writer. Add consumer and refusal regressions; keep release carriers unchanged for contributor integration.
Keep the host-derived own-binding fence across scheduling and supervisor steering without turning children into Presence speakers. Add real scheduling and fake-model consumer regressions.
Keep Host auth work from saturating control threads, refuse uncertain JSONL boundaries, and retain quota-exhausted events without speaking draft replies. This is an intermediate contributor commit pending the remaining no-effect retry decision and exact-diff review.
Recheck saved authority before Safety only for exact hits, keep misses pure, and confine catalog omissions to callable names. Bound ambiguous identity guidance and cover resource-blocked builtin discovery.
The moved target (829164b90, via the headless lineage-read PR) edited the
data-layout tree in docs/architecture/01-high-level-architecture.md without
regenerating docs/inventories/DATA_LAYOUT_INVENTORY.md, so
tests/test_generated_inventories.py::test_data_layout_inventory_is_byte_identical
is red on the target itself. Only the source SHA line changes; the entries are
identical. Regenerated with scripts/regenerate_inventories.py.
Findings from the simulated triad (gpt-6-astra, grok-4.7, fable-5.1), one batch:
- The budget line's unbounded-budget branch still rendered the full
usage_breakdown; it now reads usage_writer_snapshot like /api/state, so the
chapter 05 sentence that both read the slim shapes is true on every install.
- A restart that lands right after a drain now flushes that turn's dirty budget
projection before leaving the loop, so the last drained llm_usage events
still reach state.json.
- The events batch time budget is checked between handlers; the docstring,
the constant comment and chapter 05 now say a running handler can overrun it
instead of promising a hard elapsed-time bound.
- The socket-open project refresh forces its read: the in-flight read predates
the socket and was just invalidated, so joining it would bring nothing fresh.
- The update_budget_from_usage docstring distinguishes the loop's once-per-turn
llm_usage path from direct callers; plan labels removed from test docstrings.
Provide scoped MCP raw-to-wire discovery, pure pre-safety name resolution, and a shared policy-filtered refusal path for tool namespaces. Preserve exact dispatch and extension adoption; cover real registry consumers and classification.
The chat header's 3 s refresh and the projects-nav refresh both read
/api/state through the shared snapshot sequencer, which ordered responses
but never limited how many reads were in flight. With a slow server answer
the polls piled up (5 concurrent from the loop, 6 with the socket's own SHA
read) and the project panel's history read queued behind them for minutes
(issue #1102 / #1195 class: header polls stacked behind slow /api/state).
createStateSnapshotSequencer gains gate(force): one read in flight page-wide;
a periodic tick during a read starts nothing and resolves once that read
settles; forced refreshes that land mid-flight coalesce into exactly one
follow-up read that starts after the in-flight one settles (success, HTTP
error, thrown fetch), and every forced caller resolves after it. begin()
and the generation ordering are unchanged; a gated request settles through
apply/fail of the same request object. chat.js and app.js read through the
gate; the boot prefetch, socket-open nav refresh and 20 s nav poll join an
in-flight read instead of forcing a second one. Test doubles learn gate().
Every llm_usage worker event made the supervisor loop render the FULL usage
breakdown (five grouped axes plus the per-root projection over the whole
ledger) to persist a handful of totals, take STATE_LOCK and rewrite
state.json, and the loop drained the event queue to empty before it read the
owner's bridge; at one settled attempt per second the events phase never
returned to intake and the chat went silent while the header said Online.
The compatibility writer now reads a slim render-cached snapshot
(usage_writer_snapshot: totals, the ordering marker, the OpenRouter bucket its
drift check compares, the totals-only projection) computed from the same
validated rows as usage_breakdown, so the marker and the money still come from
one read; /api/state and the budget line read the same slim shapes.
_projection_from_final moves verbatim into the _usage_rows leaf and is
re-exported. An llm_usage event only marks the loop context dirty (its row and
live frame say "deferred"); the loop flushes one write per turn AFTER bridge
intake, keeps the flag on a refused or failed write and retries no more often
than BUDGET_PROJECTION_RETRY_SEC, and the OpenRouter ground-truth check fires on
crossing each multiple of 50 physical calls (a coalesced write may jump 49 to
51). One events pass drains at most SUPERVISOR_EVENT_BATCH_MAX_EVENTS or
SUPERVISOR_EVENT_BATCH_MAX_SEC, FIFO, restart requests and event-lag
observation preserved, the remainder next turn; a turn that hit its bound skips
the idle sleep. The drain and the flush live in server_liveness beside the
other loop-phase helpers, so server.py shrinks.
Tests pin the snapshot against the full breakdown on every key the writer reads
with the full render spied out, byte-identical state.json from either render,
the 49 to 51 crossing, N events to one write after intake with "deferred" rows,
the dirty-flag retry contract including a corrupt ledger, the count and time
bounds with intake reached while events remain queued, and /api/state serving
the same values for both budget branches without the full breakdown.