Phase B of the poltergeist delegation sprint, squash-landed onto the v6.98.0
cancellation release. The nanny contract rides host-authored run instructions
(strict truncate_within_limit budgets), a proportional dual-axis reminder
replaces the permanent post-success silence (wait rounds never reset the cost
axis), the executor is resolved before the model lane with auto->light on
harness dispatch and recorded lane provenance, delegate_wait holds a typed
idle-rail-only external-wait lease (window 1800 < tool-kill 2100 < lease 2400),
and the new delegate_answer verb lets a run's own nanny answer its pending
interactive questions through the engine's interaction API with typed
outcomes — registered across the workspace surface and both child profiles
wherever the other three verbs live, with the subset invariant re-pinned
against the post-workspace-authority tool surface. Merge resolution grafts
phase A's containment/evidence logic (A3 two-fact breach rule, nested-home
disclosure, confinement telemetry) into B's restructured delegate module.
The formal six-lane exact-SHA gate plus two verified fix rounds (BR1: honest
cancel outcomes with completion-wins in the hosted-review poller, an import-
cycle break through the delegate_shared leaf, the conditional $0 reminder,
R2-vocabulary prose sync; BR2: a discovered success is never lost to a
transport blip — cancel_and_verify carries its terminal detail additively,
with a typed succeeded-but-unretrievable refusal — and natural terminals are
attributed to the run itself, never to the host's cancel) are squashed into
this landing.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Triad gate #3 (same panel, restaged diff incl. F/S fix rounds); fable-5 0
critical, gemini clean, sol 6 FAIL dispositioned:
- G3-4 (real, W1 core): reserve_attempt/_transition now re-stash the root
accounting AFTER appending the row, so the loop's 120s-trusted snapshot
includes the current in-flight hold; regression proves the reservation
participates in the graceful stop before the hard fence.
- G3-5 (real): reservation pricing carries the finalizer's applied prompt-cache
TTL (AttemptRequest.prompt_cache_ttl; payload-free sites fall back to
config.resolve_prompt_cache_ttl) instead of hardcoding 1h — the owner's
cheaper tier can no longer flip an affordable call to BudgetExceeded; TTL
consumer pin + docstrings updated (three named consumers).
- G3-6 (real, both bind points): plan-claim whitespace is preserved
byte-for-byte through wave freeze (_bounded_wave_acceptance_claims) AND
read-time bind (_bounded_claim_text) — only the disclosed
truncate_review_artifact bound applies; regression with quoted whitespace.
- G3-3 (split): sites delegate_custody/review_evidence_refs/task_tree_ledger
REJECTED with byte-level base precedents (full text adjacent in the same
row/dict — house idiom, not P1 cognitive artifacts); build_review_context
retired[:5] ellipsis was a real deviation -> exact (+N more) count.
- G3-2: estimate_cost_optional cache/knob params now keyword-only — a stale
positional caller raises instead of silently dropping cache accounting.
- G3-7: DEVELOPMENT.md context-mode section rewritten to the implemented
D-ARCH/D-DEV contract (ARCH always full in max for every task class; DEV
keyed on the active-repo binding).
- G3-8: ARCHITECTURE §11.1 frozen gateway inventory row for TaskCostBreakdown
+ TaskDetailResponse with parity anchors.
Module gate held by extracting the pure row-math leaf _usage_rows.py
(usage_accounting 1508; re-exports keep import sites unchanged; ARCH row
added). G3-1 (vision-inspected Settings render) deferred to the final live
E2E pass by design.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The submarine shape still ran the summarizer + full-transcript rewrite in 34 of
35 rounds (measured with the REAL compactor on the wave3 numbers: low mode, an
irreducible ~630K-char frame, an agent that keeps calling tools). The rearm
criterion was the COMPACTABLE region — which is exactly what a pass collapses,
so the pass's own output cleared the 1.2x growth bar on the very next tool
round. The shipped pin missed it because it stubbed the compactor to return the
transcript UNCHANGED, the one shape where the bar is hard to cross.
Judge utility on the FLOOR a pass could reach (frozen frame + the spans
_emergency_keep_recent must keep, summaries only add to it). While the floor is
over the trigger no pass can help, so region growth cannot rearm and only the
N-round window refires — which still bounds the transcript. A futile pass now
arms on the region it ACTUALLY HANDLED, not on the remainder it just collapsed.
Measured after: 3/35 rounds, 4 disclosures (was 34/35, 35 disclosures).
Also in this area, from the same adversarial wave:
- the emergency trigger's real magnitude is now stated in context_budget.py and
pinned per density/tool-envelope (max, density 1.7, 37K schema tokens fires at
~559K chars — 2.1x earlier than the 1.2M constant reads); the eagerness itself
is the owner's "necessity = total calibrated pressure" decision and is left
alone, but it can no longer move silently.
- the finalizer no longer claims the global prompt-cache TTL is "consumed HERE
and only here" while review_helpers.cached_prompt_blocks reads it too; a pin
derives the readers from the call sites and requires both docstrings to agree.
- cache_horizon_note documents its honest reachability (at the shipped 1h only
wait_tasks can emit it; wait_task sits exactly on 3600s, delegate_wait's 2100s
ceiling cannot reach it; at 5m all three do), pinned against the real clamps.
- a golden covers the per-tier cache-write split on the OpenRouter lane: the
harvest is passthrough-dependent, not structurally dead.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
`_CACHE_TTL_SECONDS` mapped the reported "default" to 300s, and the wait tools
showed that number to the agent as a fact. "default" records ONE thing: a BARE
ephemeral marker went out. It is not a proven provider TTL, and the report is
route-blind — the finalizer returns "default" for ANY marker-carrying payload,
including routes it never normalizes. Verified live: a Gemini OpenRouter target
(`_route_normalizes_cache_breakpoints` False, bare markers kept because Gemini's
explicit cache documents no ttl field) reports "default" and therefore claimed a
300-second horizon that Gemini never established. That is exactly the "second TTL
truth" the design forbids: readers consume the RECORDED applied fact and never
re-derive a prediction from a route the fact does not carry.
- cache_ttl_seconds now answers ONLY for the explicitly stamped 5m/1h tiers;
"default" joins empty/unknown as None.
- The wait tools (wait_task/wait_tasks/delegate_wait) therefore emit NO horizon
line on a bare marker. Silence beats a fabricated number (BIBLE P1) — a wrong
horizon is worse than none, because the agent re-plans its wait windows on it.
- Under the owner's global default ('1h') the disclosure is unaffected: every
Anthropic-family send stamps a named tier and still reports a true horizon.
Docs: ARCHITECTURE now states the wait-tool disclosure contract and why a bare
marker is silent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
W3 wait-TTL disclosure (plan v2: minimal factual line, NO token-count
predictions — codex disposition; compatible with owner 7=NO which was about a
nanny model LANE, not factual disclosures). wait_task appends the line to its
result, wait_tasks carries it as cache_horizon_note in the JSON envelope, and
delegate_wait attaches it on both its terminal payload and its expired-window
render: 'configured prompt-cache horizon (<ttl>, <N>s) elapsed during this
wait (<M>s); the next model send may be cold.'
The threshold is cache_ttl_seconds() over the RECORDED applied TTL of the
task's latest send (accumulated_usage._last_prompt_cache_ttl, the finalizer's
reported fact) — the loop publishes the accumulated-usage reference on the
tool ctx (the established dynamic-ctx pattern: _pending_compaction,
_swarm_handoff_attempt), so no route-level second predictor exists. No
recorded fact (no cached send yet, marker-less route) -> no line, never an
invented horizon.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Facts of the round only (no dollar accumulator — that number is
counterfactual: the next send may reroute, compact, or hit a live cache):
- llm_round events gain cache_cold_restart (round>1 AND cached < 0.2×prompt
AND cache_write > 0.5×prompt — a prompt that was almost entirely re-written)
and gap_since_prev_round_sec (monotonic gap since this task's previous
successful round, the datum that separates a TTL expiry from a prefix
rewrite). Live evidence: the submarine waves burned ~$21 of ~$94 on 8 such
restarts after wait_tasks blocks, invisible without joining timestamps by
hand.
- accumulated_usage records _last_prompt_cache_ttl — the APPLIED TTL of the
latest send (the finalizer's reported fact), which the wait-tool
cache-horizon disclosure reads instead of re-deriving a route-level
prediction.
Owner decision (2026-08-08, batch #2 Q2=A): a single global setting
OUROBOROS_PROMPT_CACHE_TTL ('default' | '5m' | '1h', default '1h') replaces the
scattered TTL policy. Consumed at exactly ONE point — the send-time cache
finalizer llm._normalize_payload_cache_ttl — which stamps the value onto every
EXISTING cache_control breakpoint of the Anthropic-normalizing family before
the promotion rule runs, so TTL ordering stays legal by construction (the
176567b every-call-400 class). It never creates markers (the d32f703d
empty-block 400 class) and never touches non-Anthropic wire formats (the
v5.30.0 Gemini ttl-field class); OpenAI/Grok/DeepSeek stay implicit and no
prompt_cache_retention is sent.
- config.py: SETTINGS_DEFAULTS entry + PROMPT_CACHE_TTL_SCALE +
resolve_prompt_cache_ttl() (resolve_effort-style validation; unknown values
fall back to the shipped default). apply_settings_to_env picks the key up
automatically via settings_env_keys().
- review_helpers.py: REVIEW_CACHE_TTL collapsed into the global getter —
cached_prompt_blocks(ttl=None) projects the setting, so an owner-selected
'5m' honestly lowers review lanes too; 'default' emits the bare marker.
- llm.py: finalizer reports the strongest APPLIED TTL ('1h' > '5m' >
'default') so usage metadata carries the fact readers consume — no second
route-level predictor exists. Anthropic's per-tier write split
(usage.cache_creation ephemeral_5m/1h counters) is harvested into
usage.cache_write_tokens_by_ttl.
- pricing.py: estimate_cost_optional bills only the genuine 1h write share at
the documented 2x/1.25 tier ratio when the split is reported; absent the
split, the pre-split behavior stands (never a loosened ratio).
- settings UI: one segmented Prompt Cache TTL control (default/5m/1h).
- docs: ARCHITECTURE settings-table row; DEVELOPMENT cache-friendliness
invariant extended (builders keep declaring bare markers; the finalizer owns
the value).
- tests: golden finalizer pins extended, not weakened — legacy pins now run
under an explicit 'default' global; new goldens cover default-1h stamping,
5m override of a caller-declared 1h, marker-creation and cross-route no-ops,
and per-tier write pricing.
Six reviewed phases land as one release.
Admission is the outer boundary: every migrated launcher records a manifest before
it can touch the filesystem, and finalizes a typed outcome on every path — success,
refusal, crash, and the real exit status. A structural audit enforces that boundary
across all fourteen launchers, together with confinement computed from the active
checkout and a single manifest publisher, judging by effect rather than by callee
name and failing closed on any write form it cannot resolve.
Harness exit codes are no longer trusted as run status: inspect returns zero for an
eval that raised and harbor returns zero for a job whose trials all errored, so the
launchers now read the harness's own artefact and keep "the harness failed", "it
scored nothing" and "it scored honest zeros" distinguishable.
The acceptance dialogue reconciles receipts through one typed identity that is an
equivalence by construction, so a passing check can no longer clear a red it never
addressed. Prompt caching is normalized at every send site and cached calls stop
under-reporting their input. The owner's context mode becomes explicit and
fail-closed, with one enforcement point for every writer of a disk-authored setting.
Deliberate limits are disclosed in each bench's METHODOLOGY.md rather than implied
by silence. Isolated benchmark egress and the multi-lane script generator are
deferred to a later release with restoration patches and carry-forward notes.