Findings from the simulated triad (gpt-6-astra, grok-4.7, fable-5.1), one batch:
- The budget line's unbounded-budget branch still rendered the full
usage_breakdown; it now reads usage_writer_snapshot like /api/state, so the
chapter 05 sentence that both read the slim shapes is true on every install.
- A restart that lands right after a drain now flushes that turn's dirty budget
projection before leaving the loop, so the last drained llm_usage events
still reach state.json.
- The events batch time budget is checked between handlers; the docstring,
the constant comment and chapter 05 now say a running handler can overrun it
instead of promising a hard elapsed-time bound.
- The socket-open project refresh forces its read: the in-flight read predates
the socket and was just invalidated, so joining it would bring nothing fresh.
- The update_budget_from_usage docstring distinguishes the loop's once-per-turn
llm_usage path from direct callers; plan labels removed from test docstrings.
Every llm_usage worker event made the supervisor loop render the FULL usage
breakdown (five grouped axes plus the per-root projection over the whole
ledger) to persist a handful of totals, take STATE_LOCK and rewrite
state.json, and the loop drained the event queue to empty before it read the
owner's bridge; at one settled attempt per second the events phase never
returned to intake and the chat went silent while the header said Online.
The compatibility writer now reads a slim render-cached snapshot
(usage_writer_snapshot: totals, the ordering marker, the OpenRouter bucket its
drift check compares, the totals-only projection) computed from the same
validated rows as usage_breakdown, so the marker and the money still come from
one read; /api/state and the budget line read the same slim shapes.
_projection_from_final moves verbatim into the _usage_rows leaf and is
re-exported. An llm_usage event only marks the loop context dirty (its row and
live frame say "deferred"); the loop flushes one write per turn AFTER bridge
intake, keeps the flag on a refused or failed write and retries no more often
than BUDGET_PROJECTION_RETRY_SEC, and the OpenRouter ground-truth check fires on
crossing each multiple of 50 physical calls (a coalesced write may jump 49 to
51). One events pass drains at most SUPERVISOR_EVENT_BATCH_MAX_EVENTS or
SUPERVISOR_EVENT_BATCH_MAX_SEC, FIFO, restart requests and event-lag
observation preserved, the remainder next turn; a turn that hit its bound skips
the idle sleep. The drain and the flush live in server_liveness beside the
other loop-phase helpers, so server.py shrinks.
Tests pin the snapshot against the full breakdown on every key the writer reads
with the full render spied out, byte-identical state.json from either render,
the 49 to 51 crossing, N events to one write after intake with "deferred" rows,
the dirty-flag retry contract including a corrupt ledger, the count and time
bounds with intake reached while events remain queued, and /api/state serving
the same values for both budget branches without the full breakdown.