Keep the same task and live browser while required owner input lends active worker capacity. Restore acknowledged planned continuations without replaying effects or changing the original budget threshold. Align visible Publish admission, root questions, ongoing replies, display fallback names and definite publication failures with existing owners.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Keep independent native Main and Project turns addressable through post-task
work and model-access waits, with Project operations bound to their folder.
Separate web-tool restrictions from network access and preserve owner controls.
Deliver complete chosen delegation inputs and keep full operative plan sources
outside the bounded review-state index, retaining paid and legacy custody.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
A planned restart (self-restart after self-modification, managed update
or its rollback) ends the owned Claudexor daemon when the landed checkout
pins another engine version or build, so the next generation starts on
the pinned engine instead of attaching to the old one. The check runs in
the server lifespan teardown of a requested restart, the one seam both
planned restarts share: load_runtime_pin from the landed checkout against
the serving engine's handshake over the attach-only read_owned_gateway.
An unchanged pin, an unpublished or unreadable pin and an unreachable,
foreign or already stopped daemon leave the existing handoff untouched.
The stop is the manual Restart's attested stop, now shared as
_stop_owned_daemon: an unconfirmed or raising stop is the same critical
diagnostic and process_stop_unconfirmed row, custody retained, and the
restart proceeds. Delegated runs in flight end with the old daemon and
the next generation closes them as absent.
The owner's Restart (/restart, the chat Restart button) keeps its
checkout-first order: _safe_restart_serialized (update lock, strict
managed-update transaction gate, then checkout, dependency sync, import
test and stable fallback) still runs before anything is stopped, so every
refusal there - the assisted-update "restart was deferred" included -
leaves the server intact. Past the gate and the durable no-resume flags the
restart always follows: server_restart._stop_owned_work records one cancel
intent per RUNNING task, live direct/ephemeral activity and in-flight
post-task synthesis (cancel_intents.request_cancel), calls
kill_workers(force=True, terminal_status="cancelled") with Panic's
reconcile_delegate_custody=False plus the managed-update preservation
kwargs, cancels delegated runs through the public owner-gone seam
reconcile_orphaned_runs(running_task_ids=set()) over read_owned_gateway
(nothing between the cancel intents and the daemon stop may
ensure_owned_gateway - it would start a dead daemon), and makes the
attested owned-daemon stop exactly as Panic does; then the owner's stop
notice and the re-exec.
Nothing after the gate is a veto or a deferral. An unconfirmed or raising
worker shutdown and an unconfirmed daemon stop each leave a critical
diagnostic (the daemon's own process_stop_unconfirmed supervisor row, or the
same row when the stop raises), custody stays retained and the restart
proceeds; a stop with nothing to stop is quiet. OwnedClaudexorDaemon gains
the typed stop_outcome() ("stopped" / "nothing_to_stop" / "unconfirmed") so
no caller reads private manager state; stop() keeps its boolean contract.
The next generation does the rest with existing machinery: owner restart is
a no-resume cause, the startup custody sweep reconciles open delegated runs
against the daemon it reaches (absent -> close_absent_run, no invented
spend), and after an unconfirmed stop it attaches to the still-live daemon
instead of spawning a second one (pinned).
Resource change: the server lifespan makes one background
ensure_owned_gateway() (warm_owned_daemon) for an already provisioned owned
home, so the first delegation after a Restart finds the daemon serving. No
new invalidation: a Stop retires it through the existing start generation.
Not added: drain, execution leases, producer manifests, model-operation
census, new custody event kinds, early daemon start in review execution, a
"deferred" restart state. Planned self-restart, managed-update handoff and
Panic are unchanged. Version carriers untouched.
Residuals: on Windows the ledger identity is unmeasurable, so the stop ends
only the self-started child there (selective preservation of an attached
daemon is not implemented); the stop is installation-wide, so a Restart
ends every run the owned daemon serves, including runs another process
started against this installation's home; owner restart is no-resume, so
nothing interrupted is adopted by the next generation.
Closes#604.
Preserve the shared decision ingress, received HTTP status with byte payloads, both mailbox custody holds, visible ephemeral work and numeric owners. The dependency release and final public-byte pin remain pending.
Add caller-owned Claudexor model calls, unified subscription/API setup, explicit role accounts and live quota/auth continuation. Preserve prompt/tool ownership, physical-attempt custody and manual assignments. Keep the existing task process through waits and post-work.
The reviewed source checkpoint retains the current dependency pin. Published signed Claudexor bytes, cross-platform CI and the agreed merge ordering remain delivery prerequisites. No Ouroboros version, tag or release is created.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
An install upgraded with the retired reviewer comma-lists still in
settings.json (OUROBOROS_REVIEW_MODELS and friends) silently runs the
shipped default reviewer panel: config.normalize_settings_raw drops the
keys and reports the loss only on the module logger, which an owner who
never opens the Logs panel does not see (docs/archive/v7next/
LEDGER_CORRECTIONS.md "the notice is a LOG line, not a typed event").
The smallest honest mechanism on existing seams:
- settings_defaults.retired_setting_keys_notice is the ONE sentence about
a dropped retired-key set, moved out of config.normalize_settings_raw
next to the two retirement tables it reads. It now names the panel that
is ACTIVE (the authored OUROBOROS_REVIEWER_SLOTS, or the shipped default
until one is authored) instead of always claiming the shipped default.
config.retired_key_sets_seen exposes the sets the read seam already
records, so the boot reads no second copy of the document and keeps no
second table.
- server_maintenance._startup_retired_settings_notice posts one system
row (system_type retired_settings_notice) to the bound owner chat through
the existing send_with_budget path, from server._run_supervisor right
after the queue-restore notice. Dedupe is durable per retired-key set in
state.json:retired_settings_notified via update_state, so a restart or a
supervisor revival never repeats it; with no owner chat bound nothing is
sent and nothing marked, so the first boot that can deliver it does.
Pins: tests/test_retired_settings_chat_notice.py (emitted once for a
document with a retired key, not without, not repeated in a fresh process,
withheld and unmarked before an owner chat is bound, an authored slots
setting is named as the active panel in chat and log, and the boot
wiring). tests/test_config_extraction.py maps the moved sentence to its
leaf; the facade inventory is regenerated for the two new re-export
bindings; ARCHITECTURE 11.4 states the notice.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit c19759c0216a038efbc90cc219ae9085373334b7)
Second absorption of the frozen upstream line (23ab428f..db6d7cf8: 89
commits, 47 files) on top of the rc.10 hotfix tip, by the F2 rules (S1
upstream body in the owning leaf, S2 hand-merge, S3 only with proof;
retired 7.0 surfaces never return):
- PR #609 net-resilience: interactive transport-wait episodes bounded by
the task idle timeout with the typed task_incident/toast_once pair, a
bounded paid repeat after a typed post-dispatch transport death with a
round-keyed record that fences every other send, the shutdown-aware
supervisor crash counter and bounded lifespan join, Darwin keepalive
tuning. loop_transport/transport_custody/net_transport/loop_llm_call
land verbatim (same shapes on both sides); the loop.py deltas are
relocated into the v7 leaves (loop_round_limits, loop_model_call,
loop_delivery, loop_forced_finalization, loop_nudges, loop_messages,
loop_budget) with bodies AST-equal to upstream modulo the call-time
handles; _emit_overflow_retry_skipped stays a public helper (the v7
facade contract) and upstream's nested _skipped delegates to it.
- PR #614 update letter: ouroboros/update_letter.py and its web module
land verbatim; the new OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC key and
get_update_letter_timeout_sec live in their v7 owners
(settings_defaults.py, runtime_limits.py, re-exported by config.py);
_supervisor_stop lives in server_process.py beside the restart events;
docs/PERSISTENCE.md gains the state/update_letter.json row and the
inventory pin moves to 286.
- Tests: the relocated run_llm_loop tests take the emit_progress
incident keyword (every one-argument progress fake in tests/ swept, a
gap upstream itself left in test_tree_cost_ceiling); the official-update
runtime-section test lands in tests/test_context.py; _MOVED_OWNERS
registers the relocated getter.
- Docs: ARCHITECTURE/DEVELOPMENT hunks land on the upstream text; the two
legacy timeout rows upstream's context still carries stay retired (7.0).
- Size ratchet regenerated; the band rationale for tests/test_update_letter.py
is carried verbatim from upstream.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Final whole-diff review by gpt-5.6-sol (FIX FIRST) and gate round on db809e69:
- "Applied" had two proofs — git ancestry on the panel, a zero-behind heuristic
in the Runtime fact — and they disagreed in reachable cases (a kept older
letter under a merge that also consumed a newer target). The CHECK now
records one fact where git is, `target_in_head`, for the letter that will
be shown; both readers relate the letter through that fact and the head the
check described. No git runs on either reading path.
- The letter preemptively swapped Max for Low when the light route's known
window looked too small. DEVELOPMENT "Context mode": predicted pressure
never swaps in Low — only an actual provider overflow may use the task-local
Low projection, once, on the same route. The owner's mode is sent; one Low
retry follows a typed overflow, inside the same ceiling, both attempts
accounted. (Refines the owner's decision D-CODE-6 to the mechanism the
repository's own rule allows; the outcome he asked for — a letter on a small
light window — is kept.)
- A reply stopped by the output budget (finish_reason length / max_tokens) was
stored as ready: a partial cognitive artifact presented as the letter. It is
a typed `output_truncated` failure now, with the previous good letter kept.
- After the owner switches update channel, the cached check describes the
other channel; the Runtime fact read it as current. It is `unchecked` for
the active channel now, the letter kept as history — the same answer the
panel already gave.
- Commit bodies were fetched for the whole range and then dropped; only the
newest 200 are read now, subjects still for every commit.
- Boot starts the local model server before the check that may need it.
- The panel notes an applied letter written about an older version than the
one running.
Declined, with reasons in the sprint plan: streaming git output (no
reachable memory case at the sizes involved); a readiness wait on the local
model (a typed failure the next check repairs, disclosed instead).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
The PR target moved from 85c1e386 to f437408b (PRs #565, #590: the task
acceptance surface moved into acceptance_dialogue.py, deep self-review runs
on the configured reviewer row, the finalization nudges fold into one
_inject helper, and the project last-task-result reader grew in server.py).
One file conflicted: ouroboros/size_ratchet_manifest.py. Upstream shrank
ouroboros/loop.py to 272,905 bytes while this branch held 284,335, so the
BYTE_DEBT row was the only unmerged hunk. Resolved by regenerating the
manifest on the merged tree: BYTE_DEBT["ouroboros/loop.py"] = 272,805, which
is the merged file's exact size (upstream's 272,905 minus this branch's own
100-byte shrink) and below both parents' debts. Upstream's six new 1001-1500
band rationales and its shrunken tests/test_devtools_benchmarks.py debt ride
along unchanged.
Every other shared file auto-merged on disjoint regions: the transport-death
and wait-episode seams in loop.py, the incident= keyword on
OuroborosAgent._emit_progress, the supervisor stop event and lifespan
teardown in server.py, and the transport paragraphs in ARCHITECTURE and
DEVELOPMENT are all preserved beside the upstream hunks.
The bounded join now runs right after the shutdown notice, before workers,
bridge and event bus are torn down, so a tick still in flight cannot
respawn a worker the teardown just killed; its bound drops to 2 s, well
inside the launcher-managed 5 s force-exit budget that a 5 s join could
have consumed. A fresh lifespan clears the stop flag before it starts a
generation, symmetric with the teardown that sets it. Tests patch the
module's bound names rather than process-wide time/threading and restore
the supervisor error they touch.
A graceful shutdown (window close, SIGTERM) tore the event bus/Manager down
under the running supervisor tick, so the loop met BrokenPipe/EOF three
times in a row, declared "Supervisor loop died after 3 consecutive
crashes" and posted an owner alarm while the process was exiting on
purpose. The lifespan teardown now sets a process-local stop event as its
very first statement and joins the loop (bounded, 5 s) before the bridge
and event bus shut down; the loop checks the event in its while condition
and in its crash handler (an exception raised while a stop or restart is
in progress exits at info level without counting, erroring or alarming),
waits its crash backoff on that event instead of time.sleep so a shutdown
is prompt, and reaches the one shared exit (watchdog stop, thread cleared)
on every path. Three genuine consecutive crashes keep their visible death.
A fetching managed-update check (boot, and the Updates panel's Check for updates) now
also writes an update letter: one accounted LIGHT-slot call with the ordinary task
context (build_llm_messages routed at the light model), fed the first-parent commits
between the running base and the official target plus the README Version History rows
those commits added, recovered from the commit diffs because the table is capped and
untagged releases have no tag to look up.
The letter lives in state/update_letter.json under its exact (base, target, channel,
ref) key and is never deleted after the update lands; one projection relates it to the
live HEAD (pending / applied / superseded / other) for both the Updates payload
(`letter`) and the Runtime context, whose new `official_update` fact tells every task
and every consciousness cycle whether the body is current with its official source as
of the last check. No wake is forced, no new endpoint or WebSocket type, no model
substitution when the light slot has no credentials (typed failure instead), the last
good letter survives a failed rewrite, and cost stays in the ledger by attempt id.
CONSCIOUSNESS.md gains one guideline (6682 -> 6927 bytes): mentioning an update is the
mind's judgment, not a duty.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Owner batch #6 item 3=A ordered two SSOT fixes; only the first (the flat
timeout pair, D04) landed. supervisor/workers.py declared its own
TOTAL_BUDGET_LIMIT and workers.init accepted a total_budget_limit to write
it, and nothing read the result — not the module, not a caller, not a test.
An accepted argument that is written to a global nobody reads is worse than
a dead constant: it reads as a live tunable to anyone grepping for the
budget, and server.py fed it the real TOTAL_BUDGET on every boot.
The two copies that ARE read stay untouched: supervisor.state (the authority
budget_remaining and every dispatch gate read, hot-reloaded through
set_budget_limit) and supervisor.message_bus (the reporting plane).
Pin: tests/test_settings_budget_hotreload.py — the module carries no such
attribute, init no longer asks for the value, the source no longer names it,
and the one authority still feeds budget_remaining. Red before this commit
on the attribute assertion.
Focused review of 4cf13aa0 (scope critical, sol major): a task-result file
whose stat raised OSError was ordered last and excluded from every tie group,
so when it was the project's NEWER result the fallback returned the older row
AND persisted it as the durable pointer — hiding the newer file even after
the transient error cleared. The added regressions were also neutralized by
the production name tie-break (their listing orders sorted back into the
favourable order), and the docstring omitted the `updated_at` fallback.
Any stat or JSON read uncertainty now answers best-effort for THIS call only
and blocks the pointer write-back, so a later lookup sees the file once it
is readable again. The regressions choose file NAMES whose sorted order is
adverse (foreign a00.., then y_old, z_new at the window boundary and in the
self-heal tail; a key-less row sorting first), assert the persisted pointer
after each order, and pin the uncertainty rule with a readable file whose
stat fails. The docstring names the key exactly: `ts`, then `updated_at`,
then the task id. Owner decision R33.
Focused review of 5624c77f (sol major, scope critical): the durable-ts
tie-break was truncated at the 64-entry search window (a tie group crossing
entry 64 returned the older row), absent from the self-heal tail, listing-
dependent for equal or absent `ts`, abandoned on a stat OSError, and a
sort-time OSError fell back to name order where equal mtimes are no longer
contiguous — and the wrong choice could be persisted as the durable pointer.
The fallback order is now TOTAL and deterministic: every file's mtime is
stamped once (an unreadable stat orders last and joins no tie group); the
list is sorted newest first with the file name as the stable tie-break, so
equal mtimes are contiguous by construction; the first match's whole
equal-mtime group is read to its end — across the 64-entry window, which
bounds the SEARCH and never cuts a group whose order the mtime cannot settle
(disclosed in the docstring) — and inside the group the durable `ts`
(`updated_at` as its fallback) decides, the task id last. Regressions cover a
group crossing the window, a group met in the self-heal tail, equal `ts` and
absent `ts`, each in both listing orders. Owner decision R32.
Task results finalized within one clock tick share a file mtime (1 ms at
HZ=1000; 32 of 40 warm triple-writes tied in a probe), so the pointer-absent
fallback scan — newest-first by mtime, first project match wins — returned
whichever tied file the directory listed first. The routing test
`test_project_room_decision_turn_carries_last_task_result_ground_truth` was
red in every full parallel run of the phase tip while green in isolation and
on the base; the code path itself is untouched by the phase.
Inside an exact-mtime tie group the durable `ts` now decides; only the tied
files are opened, so the bounded-read contract
(`test_project_last_task_result_lookup_is_bounded`) holds and results with
distinct mtimes behave exactly as before. The regression forces the adverse
listing order through the directory listing (a real filesystem's order is
not controllable) with mtimes forced equal, and fails without the fix. The
routing suite enters the 1001–1500 band with its rationale (the
last-task-result contracts stay with their sibling routing tests). Owner
decision R29.
The read seam carries the vocabulary normalization only; the provider
normalization (apply_runtime_provider_defaults) is a separate, never-persisted
derivation, and _owner_read_settings_raw does not make it. The context-fit route
resolver and its failed-route fallback read the owner-raw document, so on a
direct-provider install with no explicit model they probed the OpenRouter-form
shipped default while the loop ran the anthropic::/openai:: route: known_window
and fits were computed against a route the loop never used. The base tree only
agreed because the retired boot write had persisted the normalization. Both
resolvers now read apply_runtime_provider_defaults(load_settings()) — the same
derivation the task-start projection and the settings GET already make. The
boot comment in server.py and ARCHITECTURE now say which normalization the read
seam carries and that the route consumers re-derive the other.
Pin: tests/test_settings_read_seam.py::test_the_context_fit_route_is_the_provider_normalized_effective_route
(red on the pre-fix tree: the owner-raw openrouter route).
Spec 4.3.5, MIGRATION rows 913-917 and 1080-1081 (owner batch 11 2=A), re-derived
on tip bytes rather than replayed from the frozen reference.
The read. `load_settings` ran the raw-stage migrations inline before merging the
shipped defaults; `_owner_read_settings_raw` - the reader behind the four
single-decision owner endpoints, the generic save and the context-fit route
resolver - merged defaults over the RAW document and got none of them. One
auto-grant POST therefore destroyed every owner customization written under a
renamed key and re-persisted retired ghosts, permanently. `config.normalize_settings_raw`
is now that step, pure and idempotent, applied by both readers before defaults;
the loader keeps the settings_integrity verified read path under it. The
load-bearing order on this tree is pass-count-before-purge and purge-before-rename
(the reference's singular scope-pin clause is dead here: both spellings are retired).
The write. `_owner_update_settings(transform, expected_digest)` reads, changes and
persists one document inside one settings lock, with the persistence prologue
under that lock; a transform returning None writes nothing; a digest mismatch
refuses before the transform runs. `_owner_write_settings` keeps its name and
contract as the whole-document caller. `settings_document_digest` moves from
onboarding beside the primitive (byte-identical); onboarding calls it directly.
The serialization. `serialize_settings` is the one on-disk spelling for the config
saver, the owner writer's atomic helper and the packaged bootstrap saver, which
also routes through the persistence prologue and is now visible to its tripwire.
Boot. The server lifespan applies the provider normalization in-process and
persists nothing, like the launcher half that landed at a4481521.
Red-first: 11 of the 18 tests in tests/test_settings_read_seam.py observed
failing on 1072a317 before the fix; the guard-string pins in
test_onboarding_host / test_server_runtime flip to the no-boot-write form, the
prologue tripwire names the primitive and the packaged saver, and the moved
unreadable-sentinel pin follows the digest.
OUROBOROS_SOFT_TIMEOUT_SEC and OUROBOROS_HARD_TIMEOUT_SEC stopped terminating
anything when the activity model (idle window + subtree liveness + absolute
ceiling) replaced them. What survived was five surfaces discussing a value none
of them obeyed: SETTINGS_DEFAULTS offered it, the Settings UI accepted a number,
the save response apologised for it, queue.init compared the caller's value
against the constant it then wrote anyway and logged a deprecation row, and
/status printed "legacy_timeouts_ignored: soft=600s, hard=1800s" on every
request. A knob discussed everywhere and obeyed nowhere reads as a live tunable.
Retired through the existing idiom - RETIRED_SETTING_KEYS, stripped on load. No
successor knob (the activity model already governs), so nothing to seed. Gone
with them: both globals and init parameters in queue/workers, the
_emit_timeout_deprecation_once emitter and its latch, the gateway's
_RETIRED_NO_EFFECT_KEYS bucket (a retired key cannot reach an effect bucket at
all, so _effect_buckets no longer needs the warnings parameter), the status_text
parameters and legacy line, the server reads and ctx fields, the bench settings
carriers and the TB forwarded-env allowlist, and the two ARCHITECTURE rows.
rc_audit's `since` stopped being a one-key special case: RETIRED_IN_THIS_ABI
names the distinction, so an upgrading install still learns the difference
between "stopped working in THIS upgrade" and "was already inert".
Pins: tests/test_legacy_timeout_retirement.py (10 cases, incl. a grep-class
sweep and the auditor's since/behavior). The N-1 fixture carries the pair at its
DEFAULT values - a default-valued ghost is the one nobody looks for - so the
rc_audit fixture suite now pins that both produce a retired-setting finding.
Two tests that asserted the old no-op semantics are reshaped, not deleted.
Disclosed: saving the key through POST /api/settings no longer returns an
explicit "Retired setting(s) saved" warning; it is merged away silently like
every other retired key. Restoring it would mean reading the raw body for keys
the merge deliberately never looks at.
events/tools/supervisor/task_reflections now rotate on the supervisor tick
with the existing rotator; agent_stdout.log is size-capped in the launcher
copy thread (2 MB x 4, mirroring server.log). events.jsonl rotation lands
BEHIND its custody readers: delegate_custody replay/fault-scan/timing/
invocation records, complete_custody_rows, the settled-terminal cursor
(now a monotonic chain offset), the legacy usage import snapshot, the
swarm-fanout rollup and the worker boot verify all read the rotated
archive chain (utils.jsonl_chain_handles: open-live-first + inode dedup,
rotation-race-safe). memory.read_jsonl_tail backfills bounded tails from
the newest archive segments. Hot-store tripwires: events/tools thresholds
become 8 MB rotation-regression tripwires, supervisor/task_reflections
gain rows, and a 100 MB events-chain watch inherits the replay-degradation
signal. skill_review_history readers are byte-bounded (CPL4-C12,
find_history_job_bounded idiom); lifecycle terminal rows already persist
their ordinals, so bounded counters stay exact inside the window.
Integrates the moved target without rewrite. Two parallel sprints fixed the
same two defect classes independently; each conflict is resolved onto ONE
coherent mechanism, deferring to landed owner decisions and keeping this
branch's reviewed engine internals:
- Delivery-control (O4 x custody-absorption Q4=A): the target's structure
and semantics stand — delivery_protocol.py leaf module, protocol-intent
CONTAINMENT (an embedded trailing object is an attempt, repaired or
degraded-preserved, never honored and never published raw), the
latch-off passthrough pin, and the child-absorption hold. Inside that
contract, parse_delivery_control_body's per-brace raw_decode walk
(O(n*braces), ~10s on a code-heavy forced answer) is replaced with this
branch's O(n) string-aware extractor (fences peeled, duplicate protocol
keys flagged, RecursionError degraded, bounded line-anchor retries after
unbalanced prose); the protocol-key judgment stays in the body parser,
so an ordinary trailing JSON object and a nested protocol object remain
prose. This branch's honor-the-trailing-directive resolver semantics and
its tests are dropped as superseded by Q4=A; two attempted 'improvements'
to the landed residuals were reverted on discovering their deliberate
pins.
- Terminal-custody refresh (O-B/R11 x custody-absorption Q2=B/D1, owner
option A 2026-08-31): both machineries stay, split by surface. The
target's trigger/snapshot/no-churn/verified-write refresh, boot backfill
(reverse join over stored unreconciled disclosures) and kill-path clears
carry the disclosure class; this branch's evidence heal carries the
current-truth class — actual_substrate and the subagent_envelope
evidence mirror (where readers take subscription_cost_usd) are rewritten
from live custody, while the top-level delegated_runs_* counters stay
the frozen historical snapshot (filtered out of the mirror write). The
settled-rows cursor pass finds what neither sweep nor backfill can see
(a stale-counter row with an empty unreconciled list marks itself
nowhere); at boot it runs AFTER the backfill so same-generation heals
keep their pinned boot_backfill attribution.
- supervisor/events.py: both handler tables merge (**_CDE chat delivery +
**_TELEMETRY passthrough); the registry AST scan learns
chat_delivery_events.py and confirms the four #407 media types are
registered.
- loop.py stays under its shrink-only byte debt by moving the pure
extract_plain_text_from_content into delivery_protocol (historical name
re-exported); review_evidence.py returns under its 1600-line cap.
- ARCHITECTURE.md: merged prose for events/delegate_terminal/delivery
reflects exactly the mechanisms above.
Combined-tree checks: 12038 parallel + 629 serial + 723 web green; ruff -F
clean; size-ratchet pairwise vs the target base green (chat.js/manifest
values settle with this commit — the transition validator reads the
committed manifest).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Module side. 34 D11 owners classified: 13 byte-identical across
tip/merge-base/reference (client_surface, gateway __init__/files/logs/mcp/
onboarding_host/schedules/task_events/task_hurry/ui_preferences, server_auth,
server_entrypoint, server_web); 17 pure upstream drift - tip bytes stand
(gateway _helpers/claudexor_accounts/contracts/control/extensions/history/
host_service/marketplace/models/presence_settings/projects/router/
skill_publish/state/tasks/ws, server_runtime); 3 carry only the D03
settings-seam / D04 retired-knob reference deltas - HOT-DEFERRED with that
seam (gateway owner_settings/onboarding/settings, D12/D17 precedent);
server.py split.
The split: 6 leaves, 43 moved spans (rows 1034-1078 + 3948-3949), every span
transplant-tool proof-green against git show HEAD:server.py (ast=tokens=
byte-roundtrip on every symbol, leaf_invariants=[], zero declared names - the
reference design homes shared rebindable state in server_process, so all six
are projection-only leaves, no handles). server.py 3191->1640: server_process
(drive root, logger, restart signals), server_routing_context (13 owner-turn
projections), server_owner_routing (attachment staging, mailbox delivery,
routing receipt, owner-message dispatch, /evolve off), server_liveness (WS3
wedge watchdog), server_maintenance (startup sweeps + periodic cadences),
server_restart (live census, teardown args, update guards, bus shutdown).
Facade = tip parent - moved spans + reference-style top import block (module-
level PORT_FILE/logging reads force top imports), facade audit green: every
kept span byte-identical to tip, every moved name re-exported by identity.
Drift-probe first per leaf: reference leaves byte-true except 10 spans
falsified by upstream drift (attachment-report train, OB-03 monotonic clock,
child-ref promotion, planned-handoff train) - re-emitted from tip bytes, no
oracle semantics replayed over drift.
HOT-DEFERRED with evidence: rows 1070/1072/1073/1074 (_pending_restart,
_handle_restart_in_supervisor, _check_pending_restart_drain,
_perform_supervisor_restart) - the upstream delegation train coupled the
restart performer to main() through the written module global
_planned_delegate_restart_transaction_id; byte-preserving relocation would
fork that state (D09-class second answer about ownership); inventory pinned
in test_server_extraction._SERVER_OWNED. Rows 1080-1081 (lifespan, D03
settings-seam server half) - reader halves 913-917 are hot-deferred by
D12/D17 (settings_integrity rewrite); landing the boot half alone would leave
normalization neither persisted nor re-derived. Gateway ABI/alias rows = F3;
web/ untouched.
LIVE delta landed: the same-qualname FUNCTION_DEBT relocation rule (row 1033,
delta id D11) replayed byte-identical from the reference into the tip-shaped
validate_manifest_transition (ouroboros/review.py is NOT a protected file;
the reference-only MODULE_DEBT_1500 layer NOT replayed - Q11=B). Pin renamed
per the row with reference bytes. This unblocks the D08 lane's row 2016
deferral (FUNCTION_DEBT relocation of _handle_schedule_task).
Test side: pin suite test_server_extraction (6, reference-adapted: deferred
restart rows moved to the _SERVER_OWNED work order, prune-sweep rows 3948-3949
added to _MOVED_OWNERS, facade bound 1700 while the restart organ is
deferred); rows 1192/1259 landed as the D11 slice of the reconciliation theme
split (the two server_maintenance-bound tests re-homed - the byte-debt
ratchet refused the +40-byte in-place retarget, giant 320340->318310); owner
retargets mirrored path-keyed (run_isolation DATA_DIR, phase3c maintenance
owner, project_routing_v664 routing_context, client_surface joined-text pin,
ws3 fake clock -> server_liveness, panic-sweep leaf floor 5->11); patches of
facade-resident readers deliberately NOT retargeted (they ride the deferral).
All touched test files lossless (the one rename is ledger row 1033); no new
ast-identical dup bodies.
size-ratchet manifest regenerated with the official tool (one byte-debt
shrink); ratchet lane 4 passed + 1 pre-existing base red reproduced bit-for-
bit on pristine a56bb76a (parent-pair manifest/tree mismatch at 7d2dca49,
documented in LEDGER_CORRECTIONS 9) - the (a56bb76a -> this commit) pair is
consistent. ruff check . --select F clean. CI-shape battery on this tree:
parallel 11983 passed rc=0, serial 609 passed rc=0. Import smoke
green (identity re-exports, shared Events single home).
docs/v7next/LEDGER_CORRECTIONS.md: D11 lane section (10 entries).
(cherry picked from commit 849f90be0c3b554753b2c580d14ac70450b2c58c)
Phase O-B of the harness-health sprint:
- task_status.find_child_tasks: filter raw disk rows on lineage BEFORE the
materializing projection — with the default it was not a read (artifact
copy2 + manifest rewrites + double hashing of every unrelated task with a
child drive, 4-5x per finalization). Rows carrying a retry pointer still
materialize (the retry chain projects the retry's lineage); lineage-less
live rows keep arriving through the queue-snapshot overlay. One shared
_EventsTailIndex per call replaces N tail parses.
- task_status.wait_for_effective_tasks: poll ticks read status/cancel_state
projections only (a 600s wait over 5 children was ~1500 spurious
materializations); the single materializing read runs after the loop on
every exit path so wait consumers keep their place in the child-result
sha economy.
- delegate_terminal.refresh_terminal_reconciliation: heals the second stale
class PR #402 missed — stored substrate counters, actual_substrate and
subscription_cost_usd that disagree with live custody survive a
successful refresh (empirically reproduced: succeeded=1 in custody,
0/'harness_attempted'/None stored forever). The refresh now rewrites the
envelope evidence and the top-level mirror through the same producers
(actual_substrate/substrate_result_fields); unreadable custody never
rewrites, undelegated tasks never gain a fabricated substrate block.
- delegate_terminal.refresh_recently_settled_terminals + sweep wiring: a
durable byte-offset cursor over the append-only custody log finds tasks
whose runs settled OUTSIDE the sweep's reconcile outcomes
(terminal-boundary settlements, prior generations) — bounded per tick,
never a full replay.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
A task's terminal result freezes delegated_runs_unreconciled at its terminal
write, and the sweep-side refresh (nanny-leaf S1) reaches only task ids named
in the CURRENT reconcile pass's outcomes — a settlement from a previous server
generation left the stored projection lying forever, and no kill path ever
cleared a stale non-empty list.
D1a — boot backfill: once per generation, directly after the startup orphan
reconcile, backfill_terminal_reconciliations scans stored TERMINAL results
still carrying a non-empty disclosure (reverse join over the self-clearing
set; never a replay-driven scan) and re-audits each through the existing
refresh seam with trigger=boot_backfill. One custody_audit_snapshot — one
replay() pass plus one pending-invocations pass — serves every audit in the
batch; open_runs/undisposed_patches/owned_project_registrations gained an
optional pre-replayed state parameter (same seam as replay(rows=...)), passed
only when a snapshot is really shared so existing call shapes stay identical.
Each row is fail-soft, mirroring the sweep's own per-task guard.
Refresh/recorder are now change-gated: an audit that matches the stored
disclosure (unreconciled list + deferred retirements; never the envelope's
trigger/outcomes provenance) performs NO write and NO custody emit, so a
permanently-unreconcilable row (an undisposed patch) does not rewrite the
result or grow events.jsonl on every boot — the second boot is a byte-level
no-op, pinned by test. The recorder gate subsumes the old empty-over-empty
guard and additionally stops identical-content rewrites on every kill/boot.
refresh_terminal_reconciliation gained trigger= (default sweep_refresh for
the existing server caller) so the envelope distinguishes
sweep_refresh/boot_backfill/kill_path_clear.
D1b — kill-path clear, fable-#6 shape (chosen by the mandatory pin 'killing a
fresh task must not mint a row via the recorder's STATUS_RUNNING fallback and
must not add a second write per kill'): the four conditional-kwargs terminal
writes (task_lifecycle running-kill cancel write; task_reaper cancel-suppressed
retry, interrupted-retry and no-retry failed writes) now carry the audit
UNCONDITIONALLY, so a clean audit writes [] and clears a stale list inside the
caller's own single terminal write — no recorder call, no extra write, no mint.
The fast already-settled cancel lane, which performs no terminal write of its
own, runs the guarded change-gated refresh (trigger=kill_path_clear): it
touches only an existing row with a non-empty stored list, so an ordinary
fast-lane kill pays zero task-result writes (pinned) and an absent row can
never be minted (pinned). The grok-#7 shape (one recorder call inside
_reconcile_delegated_runs_on_kill) was rejected against the pin: it would add
a second write per disclosing kill on every lane that already writes its own
terminal result, and reaches the recorder's STATUS_RUNNING fallback on lanes
that run before any durable row exists.
Owner decisions honored (not re-opened): Q2=B — the delegated_runs_* counters
stay a historical snapshot at the original terminal write, never recomputed
(a healed row may read 'unreconciled: []' beside 'settled: 0'; the
delegate_terminal_reconciliation envelope with trigger + open_run_ids is the
current-liveness surface); Q5=A — reason_code is never rewritten. Patch debt
survives every refresh as patch:<run_id>, never a blind clear (bug-report
boundary, pinned). Known disclosed residual (pre-existing, unchanged): the
finalize-on-miss cancel write threads the audit into delivery only, never onto
the row; a stale settled row raced into that lane heals on the next boot's
backfill.
Docs: ARCHITECTURE §5 startup-reconciliation paragraph and the
delegated-subagents custody section document the generation-crossing refresh,
the frozen-counters owner decision and the envelope as the liveness surface;
the delegate_terminal.py module-map entry follows.
Tests: tests/test_delegate_sweep_refresh.py — settled-before-boot backfill,
undisposed-patch debt preserved, byte-level second-boot no-op, single shared
replay snapshot, lock-timeout fail-soft + next-boot heal, same-generation
sweep-settlement ordering through the real _startup_custody_sweep, absent-row
mint pin, fast-lane stale clear (trigger + frozen counters) and fast-lane
zero-write pin; tests/test_cancel_intents_phase_a.py — running-kill clean
audit clears a stale list on the cancelled result.
Cancellation/custody sits on the F2 border and the drift probe drew the
line sharply. Quiet side, landed: the D07 Emergency Stop delta on
ouroboros/server_control.py - execute_panic_stop takes bound_port as a
keyword instead of lazy-importing the composition root, server.py hands
down _actual_bound_port() - applied in place onto tip bytes with an exact
two-way byte proof (vs tip: only the D07 hunks; vs the reference leaf:
only upstream's dc4c0204 reconcile_delegate_custody hunk). Four reference
pins ride along: test_panic_stop_port_sweep.py and
test_owner_stop_fences_s6.py re-pinned to tip bytes where upstream
genuinely moved (kill_workers kwargs, 5 host leaves not 11, the bounded
re-sweep's two token-less calls - each noted in-line),
test_cascade_chatless_residual_s6.py and
test_preflight_process_containment.py verbatim-green.
Hot side, HOT-DEFERRED with evidence in docs/v7next/LEDGER_CORRECTIONS.md
(entries 7-12): the 16-symbol cancel_custody extraction (upstream 65b5d19f
re-decomposed the same ownership differently), the cancel_intents D08
corruption rule (upstream absorbed it for request_cancel/claim_intent,
four mutators still fail open), the S7b test split bound to the deferred
extraction, and two D09-listed pins that belong to D07's module delta and
the delegation organ (delegate_start now refuses selectorless starts).
The E-suite stays with F2/F4 per the lane charter; cross-domain test
adaptations (events_task_done, loop_delivery, shell_process/git_ops_reset)
are left for their owning lanes. scripts/v7next_transplant.py +
tests/test_v7next_transplant.py byte-synced from the hardened def681bd (no
tool-emitted leaves in this lane; the in-place proof above is the byte
gate). 204 tests green in isolation at the lane base; size-ratchet
manifest regeneration is a no-op and lane 5 passed; ruff F clean.
(cherry picked from commit f1a37d3b413b66a9167fcf7e1a6fee5e229e60be)
Incident class: a configured-session nanny's metered round dies
provider_outcome_unknown (dispatched request, no terminal provider fact —
never resent, per custody doctrine) while its one physical delegated leaf
is alive; terminalizing the nanny let the cause-blind terminal cleanup
cancel the healthy leaf.
Main fix (D1-min): the round gate latches a durable hold and the next
round top parks the task in the same $0 supervised_wait the nanny would
have chosen. Eligibility is narrow and fail-closed: exact-route configured
sessions, exactly one open run, no pending invocations, no open
containment fault, and a READ-ONLY engine poll proving a live
non-terminal state. A meaningful leaf wake resumes with a NEW round whose
transcript carries the wake receipt (bound to the unknown attempt id);
owner dialogue drained at the round top resumes the same way. One wake =
one dispatch: an unacknowledgeable wake fails closed to the no-resend
terminal with its receipt removed. Control wakes (Stop, deadline,
finalize_now — re-checked at the source), daemon refusals, round-limit
boundaries, and budget exits all close the hold into no-call terminals,
never a paid dial; in-process loop exits clear the latch (a worker crash
preserves it for recovery). Repeated unknown cycles re-latch behind a
bounded backoff floor kept under the idle-rail minimum.
Custody companions: every provider-death arm of the rail stamps
terminal_origin=host_salvage through one wrapper (deadline grace finals
and scheduled swarm handoffs keep their legacy shape; budget/round-limit
rails stay untouched); durable llm_api_error events bind the physical
attempt (capture resolved through the explicit cause chain) and the
bounded transport cause type, which also leads the unresolved-attempt
reason; and the periodic sweep, after settling runs, re-runs the
read-only terminal custody audit so a stale delegated_runs_unreconciled
disclosure heals instead of lying forever (retry-lineage projections read
the original row live). Doctrine docs updated to state the resend
boundary precisely.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Three class-level completions of the parent commits, plus the reflow
cleanup:
- Downgrade-safe fingerprint pair: the unversioned `start_time` field
keeps the legacy `ps` spelling an N-1 reader compares correctly (a
rollback that read a boot-qualified token there mismatched every row
and pruned it WITHOUT killing - a permanently orphaned process on a
supported managed-update flow), while the current reader prefers the
new `start_time_boot` sibling, omitted when it would duplicate the
legacy spelling (macOS/BSD). One `ps` per SPAWN (the write path always
paid it before /proc-first); the sweep's hot path stays
subprocess-free, and a SIBLING-FREE mismatched row falls back to one
`ps` comparison of the legacy field, so a mid-generation boot-id
readability change never prunes a live owned row. A row that DOES
carry a boot sibling is judged by boot evidence alone: a mismatched
sibling is positive proof of another boot and prunes without the
legacy fallback, whose ps-less degradation to bare ticks would
otherwise re-qualify a cross-boot recycled pid (review finding,
pinned by test).
- A bare tick never authorizes a kill: the tick-half shortcut accepted a
recorded boot-id-less tick against the tick half of a live
boot-qualified token - the exact cross-boot collision (ticks recur,
pids recycle, command hashes repeat) the boot id exists to refuse, and
it sat on the KILL path where the pre-change reader only pruned. The
compatibility helper is now a strict legacy-spelling comparison; the
one host class that ever mints bare ticks (no usable `ps`) still
resolves PRE-UPGRADE sibling-free rows through the legacy helper's
own tick fallback - the disclosed one-migration-window residual,
alongside the "<ticks>." cross-boot twin. The boot id is
stored in full (the 8-hex truncation bought nothing but 32-bit
collision odds).
- One clock for the whole watchdog: the direct-chat heartbeat
(`agent._last_activity_ts`) joins the supervisor tick on
`time.monotonic()`, the wedge half compares on the same `now`, and the
two-clock split (plus its `now_wall` plumbing) is deleted - a forward
wall step past the 90s deadline recommended a needless /restart for a
healthy long turn, and a backward step masked a real wedge. The test
that pinned the wall-clock implementation is replaced by a
jump-resilience test symmetric to the stall one, plus a source-level
pin of all four production stamp writers. The stall toast key gains
the server pid, since a monotonic uptime stamp alone can repeat across
generations against the browser's restart-surviving dedupe set.
- The unrelated docstring re-wraps in platform_layer.py are dropped:
they bought band headroom that no longer exists on the target (the
file left the 1001-1500 band there), and they buried the real change.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Symptom: an ordinary wall-clock step — an NTP correction, a DST or timezone
change, a manual set, a resumed VM — made the supervisor liveness watchdog
report a phantom "Supervisor loop STALLED ~Ns", log a durable
supervisor_loop_stall event and send the owner an alert about a loop that had
never missed a tick. A backward step did the reverse and could hide a real
stall for as long as the skew lasted.
Root cause: the liveness tick (server.py `_loop_liveness`) and the watchdog's
comparison both read `time.time()`. That value is only ever consumed as an
elapsed gap, so it never needed to be a wall stamp — and a wall stamp is
exactly what a clock jump invalidates.
Fix shape: the tick and the stall half of the watchdog move to
`time.monotonic()`. The watchdog's two halves do NOT share one `now`: the
chat-turn wedge half compares against `agent._last_activity_ts`, which is a
wall stamp written by agent.py, so a single monotonic `now` would make
`now - turn_ts` hugely negative and silently disable wedge detection forever.
`now` (monotonic) therefore feeds the stall half and a new `now_wall` feeds the
wedge half, with a comment stating why they must not be collapsed back. The
owner stall alert path is untouched; its `toast_once` key becomes a monotonic
number, still unique per stall episode within a generation.
Tests: tests/test_ws3_wedge_resilience.py seeds the liveness list on the same
clock the tick now uses. Two tests cover the genuinely new failure modes — a
wall-clock jump neither fabricates nor masks a stall, and the wedge half still
fires on a wall-stale turn under a healthy monotonic tick. Both drive the
watchdog through a controllable fake `server.time` while the harness keeps real
clocks for its own timeouts, and both fail against the pre-fix shared-clock
code. A shared `_stop_watchdog` helper joins the thread before monkeypatch
teardown so an in-flight iteration cannot run against a restored clock: the
watchdog re-reads its stop token only at the TOP of the loop, so an iteration
already in flight still completes its clock reads and checks. All four tests
that start a watchdog route their teardown through it, including the two
pre-existing ones — they run immediately before the fake-clock tests, and a
leaked iteration there would read the real `time` module against a fake
liveness stamp and fire a phantom alert into whatever the next test patched.
Compatibility: no persisted value carries the liveness tick, so nothing on disk
or on the wire changes; the tick is process-local and re-seeded per supervisor
generation.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Project task trees leaked into Main Chat as empty Working... cards (issue
#296 residual): a Project child's diagnostics carry only their own task_id,
so the frame went out unaddressed (chat 0) and Main minted an unresolvable
card. The server now addresses every task-scoped live event at one seam
(supervisor/log_addressing.py): lineage from the RUNNING row, Project
binding wins, an explicit chat_id is preserved (0 is the real Skill Review
session; None is absence), then the task row's chat, then the
DirectActivityRegistry entry; direct/ephemeral turns additionally carry
their chat BY VALUE via the turn-scoped TurnEventQueue proxy, and A2A
frames are suppressed at the push_log choke while durable rows keep the
honest audience. The addressing runs at supervisor ingress, in the
server-process append sink (make_server_log_sink replaces the raw
bridge.push_log), and at every handler that owns a suppressed type's
explicit push.
Exactly-once: append_jsonl streams only logs/*.jsonl (never chat.jsonl or
state/memory/receipt stores); each process suppresses the types whose live
delivery has a dedicated owner (WORKER_/SERVER_LOG_SINK_SUPPRESSED_TYPES);
llm_usage gains its explicit addressed push; emit_review_cycles_exhausted
guarantees a live sibling even for a queue-less server-process caller.
project_chat_for_task_tree resolves all three lineage probes from one
mtime/size-cached bindings lens.
tests/test_log_forwarding.py is rewritten production-shaped (the real sink
installed — the old suite stubbed it out and asserted an exactly-once the
production wiring violated) and pins the leak repro, the addressing
precedence, the A2A end-to-end drop, the by-value turn proxy, and the
exactly-once contract. ARCHITECTURE.md documents the new module and both
contracts in the same change; the size-ratchet manifest carries the
message_bus band entry. Version carriers untouched (version-neutral
contributor-style PR).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Bind applied settings to the isolated child, keep the snapshot immutable, and make archive publication and cleanup descriptor-safe across replacement races. Add focused regression coverage and document the three-task smoke contract.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Keep owner text attempt-aware until durable settlement, preserve attachment provenance through task admission, steering, retry, and child inheritance, and make initial task admission atomic while keeping running-task steering LLM-first.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Carry exact task contracts, owner context, predecessor authority, and plan state through canonical roots, forks, subagents, and delegated work orders. Make archive-aware recovery truthful and fail before model work when a named authority source cannot be resolved.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Conflict resolution decisions:
- The sprint's spawn-time rotation deferral (_admit_spawned/run_deferred_rotation)
is SUPERSEDED by the mainline's reconcile_rotation, which rides every
ensure_owned_gateway (spawn and attach), is idempotent/conditional, and treats
the recovery-window 503 as retry-next-ensure — the same incident class closed
without spawn-time state. The bounded admission wait stays; reconcile runs
after admission, and an expired admission skips it until the next ensure.
- ReviewRouteUnavailable raises keep the sprint's typed .code; the mainline's
ClaudexorSubscriptionWindowExhausted branch stays ahead of them.
- Admission tests re-pin the rotation TIMING against reconcile_rotation
(fires only after admission, attach included); mainline gateway/route_health
stubs accept the sprint's additive timeout_sec/pinned_profile kwargs.
- size ratchet manifest regenerated over the merged tree.
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
Ouroboros advised a desktop-app owner to reload 'the browser tab' because the
runtime had no fact about the client UI: the launcher never exported the
retired OUROBOROS_DESKTOP_MODE flag, and messages carried no sending-surface
provenance at all.
Two additive facts, no device taxonomy (the model classifies raw observables):
- launcher.py exports OUROBOROS_PRESENTATION (desktop_window|browser_fallback;
absent=web) — the process posture, rendered as runtime_env.presentation.
- The SPA measures raw observables AT SEND TIME (pywebview bridge, ua,
viewport, matchMedia booleans, captured_at) and attaches client_surface to
each chat frame; the gateway normalizes it through the closed-key bounded
ouroboros/client_surface.py SSOT, stamps received_at, carries it in
task_metadata, persists it as an optional chat.jsonl column, and renders it
as the owner_client context fact. Non-web ingress gets a host-stamped
{channel: <source>} fallback; promotion/steering/mailbox carry it, and the
loop notes a mid-task surface change only when the surface IDENTITY differs
(viewport resize is not a device change), with a neutral note for the first
observed fact. SYSTEM.md documents the semantics and the pywebview product
facts (no Cmd+R, SHA auto-reload).
Adversarial waves 1-2 are folded in: Infinity viewport crash at ws ingress,
strict booleans, no-identity facts never mint change notes, provenance-honest
prompt wording, behavioral producer tests, send-site and received_at pins.
client_surface helpers live in their own module (message_bus/loop/chat.js stay
inside their ratchet sizes).