Centralize post-admission drive settlement, preserve captured identities and complete input closures, make metadata reads pure, serve confined nested files and directory archives, and keep maintenance off the supervisor loop. Preserve generation fences at actual mutation boundaries and truthful queued forwarding receipts.
Claudexor platform gate (API keys — subscription auth NOT covered) / toolchain · ubuntu-latest · real managed Node/npm, no ambient Node, no harness install (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / toolchain · macos-latest · real managed Node/npm, no ambient Node, no harness install (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / toolchain · windows-latest · real managed Node/npm, no ambient Node, no harness install (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / consumer · windows-latest · real pinned Codex local install + doctor, no login, no model task (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · windows-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · ubuntu-latest · claude · API key only, subscription NOT covered (push) Waiting to run
Claudexor platform gate (API keys — subscription auth NOT covered) / live · macos-latest · codex · API key only, subscription NOT covered (push) Waiting to run
Adapt the structural intent of PR #1291 (aafad5c5a6) onto #1295/#1299/#1311. Share exact body-span compilation; allow an explicit summary update in the same source-locked transaction while preserving other YAML fields. Keep one outcome per nomination, host provenance, manual edits and supported deletion/reorganization.
Version-neutral contributor candidate. Independent exact-range review and publication are still pending; Stage0 writer-choice experiment remains separate.
Await durable web acceptance in the existing off-loop completion seam; test responsiveness with a held ingress lock. Diagnose long edit needles honestly, and update old access-matrix assertions while preserving child secret/write refusals. Version-neutral contributor fixup.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Upstream advanced 86d4e6298 -> 4f8dce1f62
(v7.5.0: PR #1300 single Dockerfile + extra CA bundle, #1207 Z.ai provider,
#1309 OpenAI-family cache layout, TZ2 Presence work), which left PR #1150
CONFLICTING on GitHub.
One textual conflict, tests/test_reference_book_budgets.py: upstream's
chapter-08 budget (22000, measured 21921 on its own tree) does not cover the
merged chapter, measured 22454 = upstream's Docker subsection + this PR's
533-byte platform-gate sentence. Upstream's rationale block is kept and the
budget is raised minimally to 22500 with the PR's rationale re-stated on top.
docs/architecture/08-git-branching-ci-and-build.md itself auto-merged
(different paragraphs, no text touched).
No upstream change since 86d4e6298 touched claudexor-platform-gate.yml,
scripts/claudexor_toolchain_smoke.py, ouroboros/claudexor_runtime.py, the
runtime pin, size_ratchet_manifest.py or the claudexor tests; those files are
byte-identical to a5fa2c998 here. DOMAIN_MAP.md and the generated inventories
already match the merged tree, so nothing was regenerated.
Timer follow-ups carry the same objective_author stamp a promote carries (B3). The host_task_facts row and the cancel receipt state files_rescued as a stat-only positive/zero/unknown count with hash_computed:false, walking the canonical and child-drive artifact stores (C2 fallback per owner correction; B5 receipt stays NOT_DONE pending the TZ-1 landed API). New real-server Chromium test for a zero-option question card (B1). Chapter budgets bumped with rationale.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Close tool-rights affordances and bounded edit diagnostics, retain per-message batch transport with live-process tail handback, preserve accepted web input evidence through quiz responses, and show Starting/Input saved without implying a running supervisor. This version-neutral contributor slice deliberately defers artifact custody, directory ZIP, and V10 to a separate structural PR.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
tests/test_persistence_inventory.py: EXPECTED_SCAN_PATHS measured at 297 on the merged tree (base 294; upstream +2 for state/extra-ca-bundle and its *.pem glob; ours +1 for task_results/quarantine/*), both comment histories kept.
ouroboros/llm_claudexor.py: upstream's declared-prefix projection and wire_layout stamp (#1309) and our quota re-ask / all-operations-not-started proof are disjoint hunks; the re-ask reuses the projected payload and changes only its account field, so mutable context stays out of the cache unit.
supervisor/message_bus.py: quiz-card host_facts (#1302) and our fail-closed ensure_record_boundary plumbing are separate functions.
docs/architecture/06-agent-core.md: 314156 -> 313972 bytes under the unchanged 314000 budget by tightening the quota-wait paragraph only; no fact removed.
docs/PERSISTENCE.md, tests/test_reference_book_budgets.py: auto-merge kept, rows and budget entries disjoint.
ouroboros/size_ratchet_manifest.py: band rationale for llm_claudexor.py, which enters the 1001-line band only on the merged tree.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Round-3 repairs of the exact-range review of 193a948e:
- era_retry dispatch key is read AFTER the paid era call, so an owner switch
inside the wait rebinding the Light role keys the record to the binding that
answered (regression switches the override inside the fake LLM call).
- reflection skip events carry input_ref (the retained task-input prompt /
bound entry copy), not a misnamed reflection_ref; docs say what is retained.
- _compact_chronicle docstring: a meta-less pass reads/writes no durable retry
metadata; every run is still paid for.
Repairs authored by a Claude Fable-5.1 subscription session (run-f7ef06af47ab).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Freshness checks bind a finish only; a stop is recorded as author_stop without buying a panel. The task row reason slot carries the author rationale (project_dialogue + log_events twin + parity fixture). Authored by a delegated Claude Fable run.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Reviewer P2: a lost response proves no Host-executed tool or correspondent action from the failed attempt, not that nothing ran anywhere — priced read-only provider-owned retrieval (server-side web search) may rerun on the repeat. The 409 and owner notice apply when the round ends unresolved without a further permitted repeat (deadline or a finalize control can refuse the first repeat). Test comment on the owner-notice writer corrected; chapter budgets untouched.
Closes the six findings of the exact-diff review of a1f651878..844107a0: model-supplied _nomination_route (and every _-prefixed host key) is stripped in bind_entries and stamped only at the host nomination seam; era_retry is keyed on the effective Light dispatch binding (lane, account pin, model-wait override) shared with dispatch; validator rejections (unsupported_type/empty_content/missing_topic) emit reflection_memory_action_skipped on the real generate_reflection path and cannot abort the batch; observed_route_stamp reads physical facts only (partial -> per-field unknown) and a usage merge forwards only the last call stamp; chapter 06 and PERSISTENCE distinguish ordinary vs pressure era pass and mark calendar recompression deferred; the per-nomination route test covers three rooms with an unknown correction route outranking the block stamp.
Repairs authored by Claude Code (claude-fable-5-1) under Ouroboros supervision.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Regression drives the real round dispatcher with a ledger-backed fake provider: one dispatched ReadError, one repeat, two attempts, no tool run or send between them, no forced final, no presence_unknown_outcome. test_provider_terminal_notice converted from the pre-TZ3 200-deferred expectation to the 409 presence_attempt_outcome_unknown contract with admitted child custody preserved. Ch.06/ch.12 state the contract.
supervisor.check_openrouter_ground_truth is a provider call made through
urllib and still verified against the default store (review round 2,
Codex): hand it the same trust context. The merged file is written through
utils.write_bytes_atomic instead of a private tmp/replace, and a superseded
bundle is pruned only after a day, so a task that still holds the earlier
setting keeps its file instead of two processes re-materializing each
other's bundle.
The merged bundle lived at one stable path, and the SSL context and the
provider clients were keyed on that path: after the owner corrected the
PEM the file on disk changed but every cached context and client kept
trusting the first one until a restart (review round 2, Grok, reproduced
against a loopback CA). Name the merged file by its digest under
state/extra-ca-bundle/, remove stale siblings, and let the path itself be
the cache identity. Also: the Settings copy names the machine that runs
Ouroboros rather than the device showing the page, DEPLOYMENT lists the
web-search scraper among the tools with their own trust store, and the
persistence inventory tracks the new directory.
Repair batch authored by a delegated Claude Fable run; reproduced each finding via a failing test through the real seam.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
OpenAI's public API (direct and via OpenRouter) and the Codex backend reuse a
prompt cache for a NEW conversation only up to the end of the leading system
section / input item, and only under one routing key (measured 2026-09-25/26).
Main's single system message carried governance + memory + dynamic context, so
every new task, child, wake, direct turn and presence event paid the whole
prompt cold (~$1.96 per event on a ~393k-token prompt; 7.7% first-round cache
on a Codex install).
- context_fit.ContextFitProjection.system_message declares the stable prefix
(_stable_prefix_blocks: 1, host-only, popped from every send copy).
- llm_messages.split_leading_system_prefix projects a declared leading system
message into [system: block 0] + one [SYSTEM NOTICE] user message (byte-stable
provenance header + memory + dynamic context) before the task; pure function
of the canonical messages, so the prospective wrap-up candidate and the send
agree and round N+1 extends round N. project_declared_system_prefix stamps the
per-call target with wire_layout, copied onto usage by the response normalizer.
- Applied inside llm_openai_compatible._build_remote_kwargs for OpenAI-family
routes (llm_attempt.openai_family_route) and inside llm_claudexor._request for
every Claudexor model source. Undeclared systems (reviews, safety, light
calls) and every other family send byte-identical wire.
- llm_routing._openrouter_session_identity: the OpenAI family shares one sticky
session per model and governance prefix; other families keep the
conversation-stable derivation; explicit affinity and reroute rotation win.
Measured: the next conversation's first round read 198,797 of 393,676 tokens
from cache ($1.05 instead of $1.96); Codex shares 213,888 tokens instead of
33,024. Replay of 8 real events x 3 layouts x 2 samples: 12/16 first actions
matched production with this layout, 9/16 with today's, 8/16 with a
developer-after-task variant.
Docs: ARCHITECTURE §6 prompt-caching paragraph, DEVELOPMENT §6 cache-friendliness
invariant and notice rule, DEVELOPMENT §2 inventory row; chapter budgets raised
with reasons; domain manifest regenerated (drift predates this change).
Tests: tests/test_openai_system_prefix_split.py (new), test_prompt_cache_v664,
test_wrapup_real_send_parity, test_handover_native_reset, test_cache_optimization,
golden fixtures (two new cases, one deliberate re-record).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Add `OUROBOROS_EXTRA_CA_BUNDLE`: a PEM file whose CA certificates are merged
over certifi once (`state/extra-ca-bundle.pem`, rewritten only on content
change) and handed to every first-party HTTP client through
`net_transport` — the httpx transports and the proxy-routed Default client,
the Anthropic requests lane and its SDK web-search client, the GigaChat
SDK, the catalog, onboarding, Provider Test, pricing and capability
probes — so an endpoint behind a national or corporate CA (GigaChat's, a
TLS-inspecting proxy) works without editing a system store the no-proxy
lanes never consult. Unset keeps every client byte-identical; an
unreadable or malformed file is a typed `ExtraCaBundleError`, never a
silent fall-back. Settings -> Advanced gains the field, and DESIGN records
the rule that controls a typical owner never needs live under Advanced.
Docs: ARCHITECTURE §1, §6, §7, DEPLOYMENT.
Fold the base/app split back into a single Dockerfile. The base tag was
never published, so a plain `docker build -t ouroboros-web .` failed on
every fresh machine, and a runtime `UV_PROJECT_ENVIRONMENT` let an agent's
`uv sync` in a task workspace rewrite the runtime venv. What the split was
for stays: the Playwright browsers pinned to the locked version are
installed from an ephemeral uvx tool environment into `/ms-playwright`
above `COPY pyproject.toml uv.lock` (every release rewrites the lock), and
the BuildKit cache mounts keep the uv and apt caches out of the layers
(4.82 GB -> 4.12 GB). The tag-only docker lanes stop forcing
`PLAYWRIGHT_BROWSERS_PATH=0`, which pointed at an empty package tree and
made them download browsers mid-test; the portable path-length test reads
the shipped browser path. The Russian Trusted CA files leave the image:
the next commit gives deployments an owner-side setting instead.
Docs: ARCHITECTURE §8 Docker, README.
presence_runner._terminal_refusal keys on a missing canonical terminal or the new durable presence_unknown_outcome marker (stamped by agent_task_pipeline from the loop no-resend predicate); confirmed outages and overflows keep the TZ2 deferred projection with admitted child custody. Regressions drive the real emit_task_results pipeline over host-fallback, salvaged-draft and death-record shapes.
Semantic merge of origin/ouroboros a1f651878 (PR #1302 and earlier) into the
writer-neutral TZ-3 PR-1 candidate. Textually clean; both edits to
ouroboros/loop_llm_call.py kept (the `_observed_route` stamp from this branch,
the Z.ai "1113" billing marker from upstream). Generated inventories were
regenerated with their scripts, not hand-merged: docs/DOMAIN_MAP.md and
ouroboros/domains.toml gain the lazy-only pair D05->D08, which upstream
a1f651878 already drifts by on its own tree (ouroboros/tools/core_artifacts.py
lazy imports); size-ratchet manifest and docs/inventories were already in sync.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Writer-neutral memory step of TZ-3 PR-1 on top of PR #1299. It does not choose
the semantic memory writer (owner quiz 1C stays open).
Review findings closed (Astra review of a1a2a505, child 2cf6834a):
- F1 later-era compression: the era run is the OLDEST run of up to
ERA_COMPRESS_COUNT summary blocks anywhere before the newest block, bounded
by gaps and earlier eras (never an era of an era); the official tree scanned
only all_blocks[:ERA_COMPRESS_COUNT], so once those were eras no later
summary was ever compressed again (latent bug in origin/ouroboros).
- F2 era_retry persisted per run across pressure passes: _compact_chronicle
reads and records the SAME era_retry in dialogue_meta.json as the ordinary
run, so a not-shorter era on an unchanged run and route is not paid for
again by the next pressure pass.
- F3 route stamp = what ANSWERED: knowledge.observed_route_stamp builds the
history `route` from the returned usage (provider, resolved model,
Claudexor serving credentialProfileId + accountFingerprint), never from the
configured route a model-wait override rebinds; the loop stamps
_observed_route on accumulated usage for the turn writer.
- F4 project reflection locator readable through read_file: apply_memory_actions
names an exact task-source copy of the reflection (retained through
retain_memory_source) on each action; the canonical Project row is only a
pointer, and an unavailable source is an explicit `source_unavailable`.
- F5 skip-event logging outside the per-action handler: a failed
reflection_memory_action_skipped audit write is logged and cannot abort the
remaining independent actions of the batch.
- N1 per-run era_retry: dict keyed by source_sha256 (dispatch route + observed
route), newest ERA_RETRY_MAX_RUNS=16 kept, legacy single record still reads
(_era_retry_runs); Health shows up to three withheld runs.
- N2 per-nomination route: room_consolidation stamps _nomination_route on each
bound entry from the correction call that produced it; _write_knowledge_entries
honours it over the block-level stamp.
- N3 size ratchet: loop_llm_call.py back under the giant band (1599 lines);
manifest regenerated with the reflection.py band rationale.
- N4 chapter 06 budget: the three paragraphs this diff touches were compressed
in place (310774 bytes under the 310800 budget already carried by this diff).
Typed events: consolidation_skipped_locked, era_not_shorter (attempted or
not), scratchpad_consolidation (every exit names its outcome),
reflection_memory_action_skipped. knowledge_history.jsonl source_capture rows
carry writer/route/writer_input_ref and old_chars/new_chars; the reader-less
knowledge_journal.jsonl size telemetry is no longer written.
Known local red (pre-existing on clean fa741a7f, unrelated to this diff):
tests/test_post_task_model_wait.py controlled-worker tests fail on this Mac
because the monkeypatched llm_claudexor.model_sources lambda lacks the
processing_view keyword.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Independent exact-range review (Astra, 55ad78213..15dbc7533) found two P1 paths
where the Host acknowledged a Presence event the model never answered:
- a quota-refused primary whose fallback died with an unknown outcome is worded
provider_unavailable by the forced rail (the unknown fence outranks the
refusal source), so a guard keyed on resource_refusal_no_resend alone let the
host-authored terminal project to completed/silent (HTTP 200) live and on
replay, and the adapter could drop the event;
- a terminal write that fails after the durable start barrier is swallowed by
the pipeline, and the in-memory presence_result envelope was returned as the
answer while the durable row still said RUNNING.
One helper, _terminal_refusal, now owns the verdict the durable row demands
before acknowledgement: a confirmed resource refusal keeps the event
(presence_resources_unavailable); no canonical terminal (RUNNING/INTERRUPTED,
reconciled placeholder, lost terminal write) or any infra_failed execution
axis is presence_attempt_outcome_unknown — 409, empty external speech, the
admitted work_ref preserved (durable metadata, else the execution's own
handoff fact), one serialized owner-only notice, never a retry certificate.
Both guards (live and cached replay) consume it. Two consumer regressions
through the real Host app cover both paths; test fakes that returned an
envelope without writing a terminal now write the durable row the real
pipeline writes. Chapter 12 records the rule.
F1 era window scans the oldest summary run anywhere before the newest block (eras ahead no longer hide later summaries); F2 pressure chronicle consults/records the same era_retry and tolerates a recorded refusal; F3 history route stamp is the route that ANSWERED (knowledge.observed_route_stamp over returned usage, forwarded by usage merges, loop records _observed_route) never the configured Light route; F4 reflection actions bind an exact retained actor-readable source; F5 skip-event logging cannot abort later actions.
A root's owner card now carries one host-written sentence under the
question: the asking task, how its run started (read from the typed
run_origin provenance of the task record: the owner's message and its
time, a scheduled follow-up of task X, background consciousness, a
promotion, a schedule, or origin unknown with the recorded marker), and
when the owner last wrote in the card's chat, read from the chat log
tail. It is computed from host records only, never from the question
text, and an unrecorded fact is written as unknown.
The sentence is stored in the owner_quiz block (record_asked host_facts),
forwarded by the send_quiz event handler, and carried by the live quiz
frame, the chat.quiz host event, the chat row, history replay (with the
block as the source for rows logged without it), the Main question
pointer, the activity census and the browser mirror. The web card renders
it as a muted plain-text line (.chat-quiz-host-facts), absent when empty.
The chat.quiz event also carries the Project name for a Project card;
the browser wire does not.
QuizOutbound and ChatOutbound gain the optional host_facts field in both
language mirrors; the frozen-contract row and DESIGN.md describe it. No
gateway version change.
Interpret existing typed memory-operation errors at the post-task stage boundary while preserving ordinary consolidation partial-chunk semantics. Stop later paid stages after budget refusal or unresolved provider attempt, retaining free task facts and a degraded checkpoint. RED/GREEN regression covers actual consolidation callback and later stages.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Four independent reviewers (Fable 5.1, Opus 5.5, Codex gpt-6-sol, Grok 4.7)
ran the Ouroboros scope-review brief on the reworked candidate; all four
returned "merge after small fixes" and no round-1 blocker survived. The
accepted findings, each mirroring the MiniMax region pattern:
- The one-click Colab collector never asked for ZAI_API_KEY, so a Z.ai-only
notebook was prompted for OpenRouter (colab_bootstrap.provider_keys).
- ZAI_PLAN is a transport choice, not a credential: it no longer satisfies the
onboarding "has a provider" checks (settings_setup_contract, the wizard's
two lists), and an unknown plan value is refused at settings save and by the
provider Test instead of silently selecting pay-as-you-go.
- The Capability Evidence route readers (gateway/settings._active_main_route,
reviewer_window.reviewer_route) resolve the plan's endpoint, so the two
Z.ai plans no longer share one route fingerprint.
- The task loop classifies Z.ai's HTTP 429 code 1113 as quota_exhausted
(typed code, exact match) instead of retrying it as a transient rate limit.
- OpenRouter's GLM namespace is z-ai/, so the catalog label key follows it.
- Docs: Z.ai joins the exclusive-direct-provider list in 02-startup-onboarding
and the forbidden skill settings list in CREATING_SKILLS; the 02-naming
sentence names the two provider-specific projections instead of implying
direct OpenAI carries no effort; glm ids join the onboarding suggestions.
Declined as disproportionate or out of scope, with the reason recorded in
the review ledger: the CHECKLISTS.md prose parenthetical (protected file;
the runtime deny list is authoritative), the GAIA/Terminal-Bench launcher
key lists (benchmark-only), a recorded live wire fixture (no key), and
GLM-5.2's skip-thinking on none/minimal (owner-accepted, disclosed in the
external-fact inventory).
zai:: joins the direct providers exactly the way deepseek:: did: prefix and
credential registry, ZAI_API_KEY plus a ZAI_PLAN endpoint selector
(provider_models.resolve_zai_base_url: empty/payg = api.z.ai/api/paas/v4,
coding = the Coding Plan endpoint), the routing target, live catalog fetch,
provider Test, settings card, onboarding contract, review-fallback roles,
single-provider startup and review detection, secret masking, benchmark
env hygiene, and docs.
Reasoning effort now reaches Z.ai. The provider serves an ABSENT
reasoning_effort at its maximum tier, so every call on the old generic
compatible route was billed at max regardless of the configured effort.
The canonical scale is projected onto Z.ai's own low/high/max enum
(ZAI_REASONING_EFFORT_ALIASES: none/minimal -> low, medium -> high,
xhigh/ultra -> max), disclosed as reasoning_effort_clamped when the tier
changes; GLM-5.3 rejects every other value and cannot disable thinking
(HTTP 400 code 1210), and forced tool_choice works with thinking on, so
there is no DeepSeek-style suppression arm. The projection is keyed on the
provider id the owner configured, never on a model name: a GLM served
from an owner's own OpenAI-compatible endpoint keeps today's behavior.
The provider port is the contributor's own work from the closed PR #1194,
narrowed to Z.ai (the DashScope and Moonshot lanes were not measured and
stay out). 07-configuration gains two settings rows and one route
paragraph (budget 37300 -> 38400), 02-naming records the dated Z.ai
probe in the external-fact inventory, and the onboarding bootstrap
fixture and data-layout inventory are regenerated.
Co-authored-by: josephsteuerjr <josephsteuerjr@gmail.com>