Commit graph

149 commits

Author SHA1 Message Date
Anton Razzhigaev
3640d668bc Merge PR #750 into the v7 integration candidate
Preserve latest owner-wait budget and Telegram delivery fixes. Local ancestry checkpoint; combined runtime repairs and exact-candidate verification remain pending.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-08 22:18:03 +03:00
Ouroboros
f3ed795ea1 Preserve owner dialogue and required waiting
Keep the same task and live browser while required owner input lends active worker capacity. Restore acknowledged planned continuations without replaying effects or changing the original budget threshold. Align visible Publish admission, root questions, ongoing replies, display fallback names and definite publication failures with existing owners.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-08 20:24:08 +03:00
Ouroboros
b9bb1fe311 Restore ordinary conversation capabilities and complete chosen inputs
Keep independent native Main and Project turns addressable through post-task
work and model-access waits, with Project operations bound to their folder.
Separate web-tool restrictions from network access and preserve owner controls.

Deliver complete chosen delegation inputs and keep full operative plan sources
outside the bounded review-state index, retaining paid and legacy custody.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-08 19:26:53 +03:00
Ouroboros
08797bba11 fix(restart): suppress post-stop custody reconciliation 2026-09-08 14:05:37 +00:00
Ouroboros
25a7bd7636 Planned restart applies a changed Claudexor engine pin
A planned restart (self-restart after self-modification, managed update
or its rollback) ends the owned Claudexor daemon when the landed checkout
pins another engine version or build, so the next generation starts on
the pinned engine instead of attaching to the old one. The check runs in
the server lifespan teardown of a requested restart, the one seam both
planned restarts share: load_runtime_pin from the landed checkout against
the serving engine's handshake over the attach-only read_owned_gateway.
An unchanged pin, an unpublished or unreadable pin and an unreachable,
foreign or already stopped daemon leave the existing handoff untouched.
The stop is the manual Restart's attested stop, now shared as
_stop_owned_daemon: an unconfirmed or raising stop is the same critical
diagnostic and process_stop_unconfirmed row, custody retained, and the
restart proceeds. Delegated runs in flight end with the old daemon and
the next generation closes them as absent.
2026-09-08 11:29:05 +00:00
Ouroboros
4a75c311fb Manual Restart: stop owned work and the owned daemon before re-exec
The owner's Restart (/restart, the chat Restart button) keeps its
checkout-first order: _safe_restart_serialized (update lock, strict
managed-update transaction gate, then checkout, dependency sync, import
test and stable fallback) still runs before anything is stopped, so every
refusal there - the assisted-update "restart was deferred" included -
leaves the server intact. Past the gate and the durable no-resume flags the
restart always follows: server_restart._stop_owned_work records one cancel
intent per RUNNING task, live direct/ephemeral activity and in-flight
post-task synthesis (cancel_intents.request_cancel), calls
kill_workers(force=True, terminal_status="cancelled") with Panic's
reconcile_delegate_custody=False plus the managed-update preservation
kwargs, cancels delegated runs through the public owner-gone seam
reconcile_orphaned_runs(running_task_ids=set()) over read_owned_gateway
(nothing between the cancel intents and the daemon stop may
ensure_owned_gateway - it would start a dead daemon), and makes the
attested owned-daemon stop exactly as Panic does; then the owner's stop
notice and the re-exec.

Nothing after the gate is a veto or a deferral. An unconfirmed or raising
worker shutdown and an unconfirmed daemon stop each leave a critical
diagnostic (the daemon's own process_stop_unconfirmed supervisor row, or the
same row when the stop raises), custody stays retained and the restart
proceeds; a stop with nothing to stop is quiet. OwnedClaudexorDaemon gains
the typed stop_outcome() ("stopped" / "nothing_to_stop" / "unconfirmed") so
no caller reads private manager state; stop() keeps its boolean contract.
The next generation does the rest with existing machinery: owner restart is
a no-resume cause, the startup custody sweep reconciles open delegated runs
against the daemon it reaches (absent -> close_absent_run, no invented
spend), and after an unconfirmed stop it attaches to the still-live daemon
instead of spawning a second one (pinned).

Resource change: the server lifespan makes one background
ensure_owned_gateway() (warm_owned_daemon) for an already provisioned owned
home, so the first delegation after a Restart finds the daemon serving. No
new invalidation: a Stop retires it through the existing start generation.

Not added: drain, execution leases, producer manifests, model-operation
census, new custody event kinds, early daemon start in review execution, a
"deferred" restart state. Planned self-restart, managed-update handoff and
Panic are unchanged. Version carriers untouched.

Residuals: on Windows the ledger identity is unmeasurable, so the stop ends
only the self-started child there (selective preservation of an attached
daemon is not implemented); the stop is installation-wide, so a Restart
ends every run the owned daemon serves, including runs another process
started against this installation's home; owner restart is no-resume, so
nothing interrupted is adopted by the next generation.

Closes #604.
2026-09-08 11:29:05 +00:00
Ouroboros
4ec7840952 Integrate subscriptions with current task and service ownership
Preserve the shared decision ingress, received HTTP status with byte payloads, both mailbox custody holds, visible ephemeral work and numeric owners. The dependency release and final public-byte pin remain pending.
2026-09-07 16:26:28 +00:00
Anton Razzhigaev
f6711d3369 feat: integrate subscription accounts into model setup and execution
Add caller-owned Claudexor model calls, unified subscription/API setup, explicit role accounts and live quota/auth continuation. Preserve prompt/tool ownership, physical-attempt custody and manual assignments. Keep the existing task process through waits and post-work.

The reviewed source checkpoint retains the current dependency pin. Published signed Claudexor bytes, cross-platform CI and the agreed merge ordering remain delivery prerequisites. No Ouroboros version, tag or release is created.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-07 10:54:51 +00:00
Ouroboros
aed9809b49 Tell the owner in chat when retired settings keys are not honored (D-07)
An install upgraded with the retired reviewer comma-lists still in
settings.json (OUROBOROS_REVIEW_MODELS and friends) silently runs the
shipped default reviewer panel: config.normalize_settings_raw drops the
keys and reports the loss only on the module logger, which an owner who
never opens the Logs panel does not see (docs/archive/v7next/
LEDGER_CORRECTIONS.md "the notice is a LOG line, not a typed event").

The smallest honest mechanism on existing seams:

- settings_defaults.retired_setting_keys_notice is the ONE sentence about
  a dropped retired-key set, moved out of config.normalize_settings_raw
  next to the two retirement tables it reads. It now names the panel that
  is ACTIVE (the authored OUROBOROS_REVIEWER_SLOTS, or the shipped default
  until one is authored) instead of always claiming the shipped default.
  config.retired_key_sets_seen exposes the sets the read seam already
  records, so the boot reads no second copy of the document and keeps no
  second table.
- server_maintenance._startup_retired_settings_notice posts one system
  row (system_type retired_settings_notice) to the bound owner chat through
  the existing send_with_budget path, from server._run_supervisor right
  after the queue-restore notice. Dedupe is durable per retired-key set in
  state.json:retired_settings_notified via update_state, so a restart or a
  supervisor revival never repeats it; with no owner chat bound nothing is
  sent and nothing marked, so the first boot that can deliver it does.

Pins: tests/test_retired_settings_chat_notice.py (emitted once for a
document with a retired key, not without, not repeated in a fresh process,
withheld and unmarked before an owner chat is bound, an authored slots
setting is named as the active panel in chat and log, and the boot
wiring). tests/test_config_extraction.py maps the moved sentence to its
leaf; the facade inventory is regenerated for the two new re-export
bindings; ARCHITECTURE 11.4 states the notice.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
(cherry picked from commit c19759c0216a038efbc90cc219ae9085373334b7)
2026-09-05 09:43:36 +00:00
Ouroboros
62f87cc94c Merge upstream ouroboros db6d7cf8 into the v7 line: absorb PR #609 net-resilience and PR #614 update letter
Second absorption of the frozen upstream line (23ab428f..db6d7cf8: 89
commits, 47 files) on top of the rc.10 hotfix tip, by the F2 rules (S1
upstream body in the owning leaf, S2 hand-merge, S3 only with proof;
retired 7.0 surfaces never return):

- PR #609 net-resilience: interactive transport-wait episodes bounded by
  the task idle timeout with the typed task_incident/toast_once pair, a
  bounded paid repeat after a typed post-dispatch transport death with a
  round-keyed record that fences every other send, the shutdown-aware
  supervisor crash counter and bounded lifespan join, Darwin keepalive
  tuning. loop_transport/transport_custody/net_transport/loop_llm_call
  land verbatim (same shapes on both sides); the loop.py deltas are
  relocated into the v7 leaves (loop_round_limits, loop_model_call,
  loop_delivery, loop_forced_finalization, loop_nudges, loop_messages,
  loop_budget) with bodies AST-equal to upstream modulo the call-time
  handles; _emit_overflow_retry_skipped stays a public helper (the v7
  facade contract) and upstream's nested _skipped delegates to it.
- PR #614 update letter: ouroboros/update_letter.py and its web module
  land verbatim; the new OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC key and
  get_update_letter_timeout_sec live in their v7 owners
  (settings_defaults.py, runtime_limits.py, re-exported by config.py);
  _supervisor_stop lives in server_process.py beside the restart events;
  docs/PERSISTENCE.md gains the state/update_letter.json row and the
  inventory pin moves to 286.
- Tests: the relocated run_llm_loop tests take the emit_progress
  incident keyword (every one-argument progress fake in tests/ swept, a
  gap upstream itself left in test_tree_cost_ceiling); the official-update
  runtime-section test lands in tests/test_context.py; _MOVED_OWNERS
  registers the relocated getter.
- Docs: ARCHITECTURE/DEVELOPMENT hunks land on the upstream text; the two
  legacy timeout rows upstream's context still carries stay retired (7.0).
- Size ratchet regenerated; the band rationale for tests/test_update_letter.py
  is carried verbatim from upstream.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 22:53:15 +00:00
Ouroboros
a97793adcb Merge managed/ouroboros (3881b899) into the network-resilience branch 2026-09-04 15:57:47 +00:00
Anton Razzhigaev
3496f3a39f Update letter: the check records the ancestry fact once; no predicted Low; a cut reply is typed
Final whole-diff review by gpt-5.6-sol (FIX FIRST) and gate round on db809e69:

- "Applied" had two proofs — git ancestry on the panel, a zero-behind heuristic
  in the Runtime fact — and they disagreed in reachable cases (a kept older
  letter under a merge that also consumed a newer target). The CHECK now
  records one fact where git is, `target_in_head`, for the letter that will
  be shown; both readers relate the letter through that fact and the head the
  check described. No git runs on either reading path.
- The letter preemptively swapped Max for Low when the light route's known
  window looked too small. DEVELOPMENT "Context mode": predicted pressure
  never swaps in Low — only an actual provider overflow may use the task-local
  Low projection, once, on the same route. The owner's mode is sent; one Low
  retry follows a typed overflow, inside the same ceiling, both attempts
  accounted. (Refines the owner's decision D-CODE-6 to the mechanism the
  repository's own rule allows; the outcome he asked for — a letter on a small
  light window — is kept.)
- A reply stopped by the output budget (finish_reason length / max_tokens) was
  stored as ready: a partial cognitive artifact presented as the letter. It is
  a typed `output_truncated` failure now, with the previous good letter kept.
- After the owner switches update channel, the cached check describes the
  other channel; the Runtime fact read it as current. It is `unchecked` for
  the active channel now, the letter kept as history — the same answer the
  panel already gave.
- Commit bodies were fetched for the whole range and then dropped; only the
  newest 200 are read now, subjects still for every commit.
- Boot starts the local model server before the check that may need it.
- The panel notes an applied letter written about an older version than the
  one running.

Declined, with reasons in the sprint plan: streaming git output (no
reachable memory case at the sizes involved); a readiness wait on the local
model (a typed failure the next check repairs, disclosed instead).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 11:25:59 +00:00
Ouroboros
5da2a81acd Merge managed/ouroboros (f437408b) into the network-resilience branch
The PR target moved from 85c1e386 to f437408b (PRs #565, #590: the task
acceptance surface moved into acceptance_dialogue.py, deep self-review runs
on the configured reviewer row, the finalization nudges fold into one
_inject helper, and the project last-task-result reader grew in server.py).

One file conflicted: ouroboros/size_ratchet_manifest.py. Upstream shrank
ouroboros/loop.py to 272,905 bytes while this branch held 284,335, so the
BYTE_DEBT row was the only unmerged hunk. Resolved by regenerating the
manifest on the merged tree: BYTE_DEBT["ouroboros/loop.py"] = 272,805, which
is the merged file's exact size (upstream's 272,905 minus this branch's own
100-byte shrink) and below both parents' debts. Upstream's six new 1001-1500
band rationales and its shrunken tests/test_devtools_benchmarks.py debt ride
along unchanged.

Every other shared file auto-merged on disjoint regions: the transport-death
and wait-episode seams in loop.py, the incident= keyword on
OuroborosAgent._emit_progress, the supervisor stop event and lifespan
teardown in server.py, and the transport paragraphs in ARCHITECTURE and
DEVELOPMENT are all preserved beside the upstream hunks.
2026-09-04 04:22:14 +00:00
Ouroboros
195a275162 server: join the supervisor loop before killing workers, inside the force-exit budget
The bounded join now runs right after the shutdown notice, before workers,
bridge and event bus are torn down, so a tick still in flight cannot
respawn a worker the teardown just killed; its bound drops to 2 s, well
inside the launcher-managed 5 s force-exit budget that a 5 s join could
have consumed. A fresh lifespan clears the stop flag before it starts a
generation, symmetric with the teardown that sets it. Tests patch the
module's bound names rather than process-wide time/threading and restore
the supervisor error they touch.
2026-09-04 02:00:41 +00:00
Ouroboros
f059e25fb6 server: make the supervisor crash counter shutdown-aware
A graceful shutdown (window close, SIGTERM) tore the event bus/Manager down
under the running supervisor tick, so the loop met BrokenPipe/EOF three
times in a row, declared "Supervisor loop died after 3 consecutive
crashes" and posted an owner alarm while the process was exiting on
purpose. The lifespan teardown now sets a process-local stop event as its
very first statement and joins the loop (bounded, 5 s) before the bridge
and event bus shut down; the loop checks the event in its while condition
and in its crash handler (an exception raised while a stop or restart is
in progress exits at info level without counting, erroring or alarming),
waits its crash backoff on that event instead of time.sleep so a shutdown
is prompt, and reaches the one shared exit (watchdog stop, thread cleared)
on every path. Three genuine consecutive crashes keep their visible death.
2026-09-04 02:00:40 +00:00
Anton Razzhigaev
a93e05cba2 Merge remote-tracking branch 'managed/ouroboros' into claude/update-letter-20260903
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 01:36:18 +00:00
Anton Razzhigaev
d89bba3d4e Update letter: Ouroboros writes one paragraph about a pending official update
A fetching managed-update check (boot, and the Updates panel's Check for updates) now
also writes an update letter: one accounted LIGHT-slot call with the ordinary task
context (build_llm_messages routed at the light model), fed the first-parent commits
between the running base and the official target plus the README Version History rows
those commits added, recovered from the commit diffs because the table is capped and
untagged releases have no tag to look up.

The letter lives in state/update_letter.json under its exact (base, target, channel,
ref) key and is never deleted after the update lands; one projection relates it to the
live HEAD (pending / applied / superseded / other) for both the Updates payload
(`letter`) and the Runtime context, whose new `official_update` fact tells every task
and every consciousness cycle whether the body is current with its official source as
of the last check. No wake is forced, no new endpoint or WebSocket type, no model
substitution when the light slot has no credentials (typed failure instead), the last
good letter survives a failed rewrite, and cost stays in the ledger by attempt id.

CONSCIOUSNESS.md gains one guideline (6682 -> 6927 bytes): mentioning an update is the
mind's judgment, not a duty.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-03 19:20:55 +00:00
Ouroboros
7ae0dbe3fd budget: drop the third TOTAL_BUDGET_LIMIT copy from the worker pool (batch #6 item 3=A)
Owner batch #6 item 3=A ordered two SSOT fixes; only the first (the flat
timeout pair, D04) landed. supervisor/workers.py declared its own
TOTAL_BUDGET_LIMIT and workers.init accepted a total_budget_limit to write
it, and nothing read the result — not the module, not a caller, not a test.
An accepted argument that is written to a global nobody reads is worse than
a dead constant: it reads as a live tunable to anyone grepping for the
budget, and server.py fed it the real TOTAL_BUDGET on every boot.

The two copies that ARE read stay untouched: supervisor.state (the authority
budget_remaining and every dispatch gate read, hot-reloaded through
set_budget_limit) and supervisor.message_bus (the reporting plane).

Pin: tests/test_settings_budget_hotreload.py — the module carries no such
attribute, init no longer asks for the value, the source no longer names it,
and the one authority still feeds budget_remaining. Red before this commit
on the attribute assertion.
2026-09-02 15:15:53 +00:00
Ouroboros
3b90af8d60 server: read uncertainty never persists the last-task-result pointer; adverse tie-order regressions
Focused review of 4cf13aa0 (scope critical, sol major): a task-result file
whose stat raised OSError was ordered last and excluded from every tie group,
so when it was the project's NEWER result the fallback returned the older row
AND persisted it as the durable pointer — hiding the newer file even after
the transient error cleared. The added regressions were also neutralized by
the production name tie-break (their listing orders sorted back into the
favourable order), and the docstring omitted the `updated_at` fallback.

Any stat or JSON read uncertainty now answers best-effort for THIS call only
and blocks the pointer write-back, so a later lookup sees the file once it
is readable again. The regressions choose file NAMES whose sorted order is
adverse (foreign a00.., then y_old, z_new at the window boundary and in the
self-heal tail; a key-less row sorting first), assert the persisted pointer
after each order, and pin the uncertainty rule with a readable file whose
stat fails. The docstring names the key exactly: `ts`, then `updated_at`,
then the task id. Owner decision R33.
2026-09-02 13:24:58 +00:00
Ouroboros
4cf13aa085 server: total, deterministic order for the project last-task-result fallback
Focused review of 5624c77f (sol major, scope critical): the durable-ts
tie-break was truncated at the 64-entry search window (a tie group crossing
entry 64 returned the older row), absent from the self-heal tail, listing-
dependent for equal or absent `ts`, abandoned on a stat OSError, and a
sort-time OSError fell back to name order where equal mtimes are no longer
contiguous — and the wrong choice could be persisted as the durable pointer.

The fallback order is now TOTAL and deterministic: every file's mtime is
stamped once (an unreadable stat orders last and joins no tie group); the
list is sorted newest first with the file name as the stable tie-break, so
equal mtimes are contiguous by construction; the first match's whole
equal-mtime group is read to its end — across the 64-entry window, which
bounds the SEARCH and never cuts a group whose order the mtime cannot settle
(disclosed in the docstring) — and inside the group the durable `ts`
(`updated_at` as its fallback) decides, the task id last. Regressions cover a
group crossing the window, a group met in the self-heal tail, equal `ts` and
absent `ts`, each in both listing orders. Owner decision R32.
2026-09-02 13:00:51 +00:00
Ouroboros
5624c77f10 server: project last-task-result fallback breaks an mtime tie by the durable ts
Task results finalized within one clock tick share a file mtime (1 ms at
HZ=1000; 32 of 40 warm triple-writes tied in a probe), so the pointer-absent
fallback scan — newest-first by mtime, first project match wins — returned
whichever tied file the directory listed first. The routing test
`test_project_room_decision_turn_carries_last_task_result_ground_truth` was
red in every full parallel run of the phase tip while green in isolation and
on the base; the code path itself is untouched by the phase.

Inside an exact-mtime tie group the durable `ts` now decides; only the tied
files are opened, so the bounded-read contract
(`test_project_last_task_result_lookup_is_bounded`) holds and results with
distinct mtimes behave exactly as before. The regression forces the adverse
listing order through the directory listing (a real filesystem's order is
not controllable) with mtimes forced equal, and fails without the fix. The
routing suite enters the 1001–1500 band with its rationale (the
last-task-result contracts stay with their sibling routing tests). Owner
decision R29.
2026-09-02 12:41:03 +00:00
Ouroboros
f30219bc27 context_fit: resolve the route from the provider-normalized effective settings
The read seam carries the vocabulary normalization only; the provider
normalization (apply_runtime_provider_defaults) is a separate, never-persisted
derivation, and _owner_read_settings_raw does not make it. The context-fit route
resolver and its failed-route fallback read the owner-raw document, so on a
direct-provider install with no explicit model they probed the OpenRouter-form
shipped default while the loop ran the anthropic::/openai:: route: known_window
and fits were computed against a route the loop never used. The base tree only
agreed because the retired boot write had persisted the normalization. Both
resolvers now read apply_runtime_provider_defaults(load_settings()) — the same
derivation the task-start projection and the settings GET already make. The
boot comment in server.py and ARCHITECTURE now say which normalization the read
seam carries and that the route consumers re-derive the other.

Pin: tests/test_settings_read_seam.py::test_the_context_fit_route_is_the_provider_normalized_effective_route
(red on the pre-fix tree: the owner-raw openrouter route).
2026-09-02 11:18:19 +00:00
Ouroboros
d8e76b6e30 v7next D03: one settings normalization, one locked update, one digest, one serializer
Spec 4.3.5, MIGRATION rows 913-917 and 1080-1081 (owner batch 11 2=A), re-derived
on tip bytes rather than replayed from the frozen reference.

The read. `load_settings` ran the raw-stage migrations inline before merging the
shipped defaults; `_owner_read_settings_raw` - the reader behind the four
single-decision owner endpoints, the generic save and the context-fit route
resolver - merged defaults over the RAW document and got none of them. One
auto-grant POST therefore destroyed every owner customization written under a
renamed key and re-persisted retired ghosts, permanently. `config.normalize_settings_raw`
is now that step, pure and idempotent, applied by both readers before defaults;
the loader keeps the settings_integrity verified read path under it. The
load-bearing order on this tree is pass-count-before-purge and purge-before-rename
(the reference's singular scope-pin clause is dead here: both spellings are retired).

The write. `_owner_update_settings(transform, expected_digest)` reads, changes and
persists one document inside one settings lock, with the persistence prologue
under that lock; a transform returning None writes nothing; a digest mismatch
refuses before the transform runs. `_owner_write_settings` keeps its name and
contract as the whole-document caller. `settings_document_digest` moves from
onboarding beside the primitive (byte-identical); onboarding calls it directly.

The serialization. `serialize_settings` is the one on-disk spelling for the config
saver, the owner writer's atomic helper and the packaged bootstrap saver, which
also routes through the persistence prologue and is now visible to its tripwire.

Boot. The server lifespan applies the provider normalization in-process and
persists nothing, like the launcher half that landed at a4481521.

Red-first: 11 of the 18 tests in tests/test_settings_read_seam.py observed
failing on 1072a317 before the fix; the guard-string pins in
test_onboarding_host / test_server_runtime flip to the no-boot-write form, the
prologue tripwire names the primitive and the packaged saver, and the moved
unreadable-sentinel pin follows the digest.
2026-09-02 11:18:18 +00:00
Ouroboros
5b1767fadd D04: retire the flat wall-clock timeout pair (owner 1B)
OUROBOROS_SOFT_TIMEOUT_SEC and OUROBOROS_HARD_TIMEOUT_SEC stopped terminating
anything when the activity model (idle window + subtree liveness + absolute
ceiling) replaced them. What survived was five surfaces discussing a value none
of them obeyed: SETTINGS_DEFAULTS offered it, the Settings UI accepted a number,
the save response apologised for it, queue.init compared the caller's value
against the constant it then wrote anyway and logged a deprecation row, and
/status printed "legacy_timeouts_ignored: soft=600s, hard=1800s" on every
request. A knob discussed everywhere and obeyed nowhere reads as a live tunable.

Retired through the existing idiom - RETIRED_SETTING_KEYS, stripped on load. No
successor knob (the activity model already governs), so nothing to seed. Gone
with them: both globals and init parameters in queue/workers, the
_emit_timeout_deprecation_once emitter and its latch, the gateway's
_RETIRED_NO_EFFECT_KEYS bucket (a retired key cannot reach an effect bucket at
all, so _effect_buckets no longer needs the warnings parameter), the status_text
parameters and legacy line, the server reads and ctx fields, the bench settings
carriers and the TB forwarded-env allowlist, and the two ARCHITECTURE rows.

rc_audit's `since` stopped being a one-key special case: RETIRED_IN_THIS_ABI
names the distinction, so an upgrading install still learns the difference
between "stopped working in THIS upgrade" and "was already inert".

Pins: tests/test_legacy_timeout_retirement.py (10 cases, incl. a grep-class
sweep and the auditor's since/behavior). The N-1 fixture carries the pair at its
DEFAULT values - a default-valued ghost is the one nobody looks for - so the
rc_audit fixture suite now pins that both produce a retired-setting finding.
Two tests that asserted the old no-op semantics are reshaped, not deleted.

Disclosed: saving the key through POST /api/settings no longer returns an
explicit "Retired setting(s) saved" warning; it is merged away silently like
every other retired key. Restoring it would mean reading the raw body for keys
the merge deliberately never looks at.
2026-09-01 23:01:30 +00:00
Ouroboros
9723251f98 persistence: rotate the remaining hot logs behind chain-aware readers (CPL4-C1..C5, C12)
events/tools/supervisor/task_reflections now rotate on the supervisor tick
with the existing rotator; agent_stdout.log is size-capped in the launcher
copy thread (2 MB x 4, mirroring server.log). events.jsonl rotation lands
BEHIND its custody readers: delegate_custody replay/fault-scan/timing/
invocation records, complete_custody_rows, the settled-terminal cursor
(now a monotonic chain offset), the legacy usage import snapshot, the
swarm-fanout rollup and the worker boot verify all read the rotated
archive chain (utils.jsonl_chain_handles: open-live-first + inode dedup,
rotation-race-safe). memory.read_jsonl_tail backfills bounded tails from
the newest archive segments. Hot-store tripwires: events/tools thresholds
become 8 MB rotation-regression tripwires, supervisor/task_reflections
gain rows, and a 100 MB events-chain watch inherits the replay-degradation
signal. skill_review_history readers are byte-bounded (CPL4-C12,
find_history_job_bounded idiom); lifecycle terminal rows already persist
their ordinals, so bounded counters stay exact inside the window.
2026-09-01 17:35:37 +00:00
Anton Razzhigaev
d98942a13b merge: ouroboros target (PRs #367 #398 #404 #407 #408) into harness-health
Integrates the moved target without rewrite. Two parallel sprints fixed the
same two defect classes independently; each conflict is resolved onto ONE
coherent mechanism, deferring to landed owner decisions and keeping this
branch's reviewed engine internals:

- Delivery-control (O4 x custody-absorption Q4=A): the target's structure
  and semantics stand — delivery_protocol.py leaf module, protocol-intent
  CONTAINMENT (an embedded trailing object is an attempt, repaired or
  degraded-preserved, never honored and never published raw), the
  latch-off passthrough pin, and the child-absorption hold. Inside that
  contract, parse_delivery_control_body's per-brace raw_decode walk
  (O(n*braces), ~10s on a code-heavy forced answer) is replaced with this
  branch's O(n) string-aware extractor (fences peeled, duplicate protocol
  keys flagged, RecursionError degraded, bounded line-anchor retries after
  unbalanced prose); the protocol-key judgment stays in the body parser,
  so an ordinary trailing JSON object and a nested protocol object remain
  prose. This branch's honor-the-trailing-directive resolver semantics and
  its tests are dropped as superseded by Q4=A; two attempted 'improvements'
  to the landed residuals were reverted on discovering their deliberate
  pins.
- Terminal-custody refresh (O-B/R11 x custody-absorption Q2=B/D1, owner
  option A 2026-08-31): both machineries stay, split by surface. The
  target's trigger/snapshot/no-churn/verified-write refresh, boot backfill
  (reverse join over stored unreconciled disclosures) and kill-path clears
  carry the disclosure class; this branch's evidence heal carries the
  current-truth class — actual_substrate and the subagent_envelope
  evidence mirror (where readers take subscription_cost_usd) are rewritten
  from live custody, while the top-level delegated_runs_* counters stay
  the frozen historical snapshot (filtered out of the mirror write). The
  settled-rows cursor pass finds what neither sweep nor backfill can see
  (a stale-counter row with an empty unreconciled list marks itself
  nowhere); at boot it runs AFTER the backfill so same-generation heals
  keep their pinned boot_backfill attribution.
- supervisor/events.py: both handler tables merge (**_CDE chat delivery +
  **_TELEMETRY passthrough); the registry AST scan learns
  chat_delivery_events.py and confirms the four #407 media types are
  registered.
- loop.py stays under its shrink-only byte debt by moving the pure
  extract_plain_text_from_content into delivery_protocol (historical name
  re-exported); review_evidence.py returns under its 1600-line cap.
- ARCHITECTURE.md: merged prose for events/delegate_terminal/delivery
  reflects exactly the mechanisms above.

Combined-tree checks: 12038 parallel + 629 serial + 723 web green; ruff -F
clean; size-ratchet pairwise vs the target base green (chat.js/manifest
values settle with this commit — the transition validator reads the
committed manifest).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-31 10:01:52 +00:00
Ouroboros
d1c8fca451 v7next F1: domain D11 quiet part - server.py six-leaf split from tip bytes; restart transaction hot-deferred; same-qualname ratchet delta landed
Module side. 34 D11 owners classified: 13 byte-identical across
tip/merge-base/reference (client_surface, gateway __init__/files/logs/mcp/
onboarding_host/schedules/task_events/task_hurry/ui_preferences, server_auth,
server_entrypoint, server_web); 17 pure upstream drift - tip bytes stand
(gateway _helpers/claudexor_accounts/contracts/control/extensions/history/
host_service/marketplace/models/presence_settings/projects/router/
skill_publish/state/tasks/ws, server_runtime); 3 carry only the D03
settings-seam / D04 retired-knob reference deltas - HOT-DEFERRED with that
seam (gateway owner_settings/onboarding/settings, D12/D17 precedent);
server.py split.

The split: 6 leaves, 43 moved spans (rows 1034-1078 + 3948-3949), every span
transplant-tool proof-green against git show HEAD:server.py (ast=tokens=
byte-roundtrip on every symbol, leaf_invariants=[], zero declared names - the
reference design homes shared rebindable state in server_process, so all six
are projection-only leaves, no handles). server.py 3191->1640: server_process
(drive root, logger, restart signals), server_routing_context (13 owner-turn
projections), server_owner_routing (attachment staging, mailbox delivery,
routing receipt, owner-message dispatch, /evolve off), server_liveness (WS3
wedge watchdog), server_maintenance (startup sweeps + periodic cadences),
server_restart (live census, teardown args, update guards, bus shutdown).
Facade = tip parent - moved spans + reference-style top import block (module-
level PORT_FILE/logging reads force top imports), facade audit green: every
kept span byte-identical to tip, every moved name re-exported by identity.
Drift-probe first per leaf: reference leaves byte-true except 10 spans
falsified by upstream drift (attachment-report train, OB-03 monotonic clock,
child-ref promotion, planned-handoff train) - re-emitted from tip bytes, no
oracle semantics replayed over drift.

HOT-DEFERRED with evidence: rows 1070/1072/1073/1074 (_pending_restart,
_handle_restart_in_supervisor, _check_pending_restart_drain,
_perform_supervisor_restart) - the upstream delegation train coupled the
restart performer to main() through the written module global
_planned_delegate_restart_transaction_id; byte-preserving relocation would
fork that state (D09-class second answer about ownership); inventory pinned
in test_server_extraction._SERVER_OWNED. Rows 1080-1081 (lifespan, D03
settings-seam server half) - reader halves 913-917 are hot-deferred by
D12/D17 (settings_integrity rewrite); landing the boot half alone would leave
normalization neither persisted nor re-derived. Gateway ABI/alias rows = F3;
web/ untouched.

LIVE delta landed: the same-qualname FUNCTION_DEBT relocation rule (row 1033,
delta id D11) replayed byte-identical from the reference into the tip-shaped
validate_manifest_transition (ouroboros/review.py is NOT a protected file;
the reference-only MODULE_DEBT_1500 layer NOT replayed - Q11=B). Pin renamed
per the row with reference bytes. This unblocks the D08 lane's row 2016
deferral (FUNCTION_DEBT relocation of _handle_schedule_task).

Test side: pin suite test_server_extraction (6, reference-adapted: deferred
restart rows moved to the _SERVER_OWNED work order, prune-sweep rows 3948-3949
added to _MOVED_OWNERS, facade bound 1700 while the restart organ is
deferred); rows 1192/1259 landed as the D11 slice of the reconciliation theme
split (the two server_maintenance-bound tests re-homed - the byte-debt
ratchet refused the +40-byte in-place retarget, giant 320340->318310); owner
retargets mirrored path-keyed (run_isolation DATA_DIR, phase3c maintenance
owner, project_routing_v664 routing_context, client_surface joined-text pin,
ws3 fake clock -> server_liveness, panic-sweep leaf floor 5->11); patches of
facade-resident readers deliberately NOT retargeted (they ride the deferral).
All touched test files lossless (the one rename is ledger row 1033); no new
ast-identical dup bodies.

size-ratchet manifest regenerated with the official tool (one byte-debt
shrink); ratchet lane 4 passed + 1 pre-existing base red reproduced bit-for-
bit on pristine a56bb76a (parent-pair manifest/tree mismatch at 7d2dca49,
documented in LEDGER_CORRECTIONS 9) - the (a56bb76a -> this commit) pair is
consistent. ruff check . --select F clean. CI-shape battery on this tree:
parallel 11983 passed rc=0, serial 609 passed rc=0. Import smoke
green (identity re-exports, shared Events single home).
docs/v7next/LEDGER_CORRECTIONS.md: D11 lane section (10 entries).

(cherry picked from commit 849f90be0c3b554753b2c580d14ac70450b2c58c)
2026-08-31 00:17:05 +00:00
Anton Razzhigaev
601dcc8ba1 fix(delegate,status): heal stale terminal evidence; stop mutating unrelated trees on child reads
Phase O-B of the harness-health sprint:

- task_status.find_child_tasks: filter raw disk rows on lineage BEFORE the
  materializing projection — with the default it was not a read (artifact
  copy2 + manifest rewrites + double hashing of every unrelated task with a
  child drive, 4-5x per finalization). Rows carrying a retry pointer still
  materialize (the retry chain projects the retry's lineage); lineage-less
  live rows keep arriving through the queue-snapshot overlay. One shared
  _EventsTailIndex per call replaces N tail parses.
- task_status.wait_for_effective_tasks: poll ticks read status/cancel_state
  projections only (a 600s wait over 5 children was ~1500 spurious
  materializations); the single materializing read runs after the loop on
  every exit path so wait consumers keep their place in the child-result
  sha economy.
- delegate_terminal.refresh_terminal_reconciliation: heals the second stale
  class PR #402 missed — stored substrate counters, actual_substrate and
  subscription_cost_usd that disagree with live custody survive a
  successful refresh (empirically reproduced: succeeded=1 in custody,
  0/'harness_attempted'/None stored forever). The refresh now rewrites the
  envelope evidence and the top-level mirror through the same producers
  (actual_substrate/substrate_result_fields); unreadable custody never
  rewrites, undelegated tasks never gain a fabricated substrate block.
- delegate_terminal.refresh_recently_settled_terminals + sweep wiring: a
  durable byte-offset cursor over the append-only custody log finds tasks
  whose runs settled OUTSIDE the sweep's reconcile outcomes
  (terminal-boundary settlements, prior generations) — bounded per tick,
  never a full replay.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-30 20:23:59 +00:00
Anton Razzhigaev
c754310c0b Refresh stale stored delegated-custody disclosures across generations and kill paths (D1)
A task's terminal result freezes delegated_runs_unreconciled at its terminal
write, and the sweep-side refresh (nanny-leaf S1) reaches only task ids named
in the CURRENT reconcile pass's outcomes — a settlement from a previous server
generation left the stored projection lying forever, and no kill path ever
cleared a stale non-empty list.

D1a — boot backfill: once per generation, directly after the startup orphan
reconcile, backfill_terminal_reconciliations scans stored TERMINAL results
still carrying a non-empty disclosure (reverse join over the self-clearing
set; never a replay-driven scan) and re-audits each through the existing
refresh seam with trigger=boot_backfill. One custody_audit_snapshot — one
replay() pass plus one pending-invocations pass — serves every audit in the
batch; open_runs/undisposed_patches/owned_project_registrations gained an
optional pre-replayed state parameter (same seam as replay(rows=...)), passed
only when a snapshot is really shared so existing call shapes stay identical.
Each row is fail-soft, mirroring the sweep's own per-task guard.

Refresh/recorder are now change-gated: an audit that matches the stored
disclosure (unreconciled list + deferred retirements; never the envelope's
trigger/outcomes provenance) performs NO write and NO custody emit, so a
permanently-unreconcilable row (an undisposed patch) does not rewrite the
result or grow events.jsonl on every boot — the second boot is a byte-level
no-op, pinned by test. The recorder gate subsumes the old empty-over-empty
guard and additionally stops identical-content rewrites on every kill/boot.
refresh_terminal_reconciliation gained trigger= (default sweep_refresh for
the existing server caller) so the envelope distinguishes
sweep_refresh/boot_backfill/kill_path_clear.

D1b — kill-path clear, fable-#6 shape (chosen by the mandatory pin 'killing a
fresh task must not mint a row via the recorder's STATUS_RUNNING fallback and
must not add a second write per kill'): the four conditional-kwargs terminal
writes (task_lifecycle running-kill cancel write; task_reaper cancel-suppressed
retry, interrupted-retry and no-retry failed writes) now carry the audit
UNCONDITIONALLY, so a clean audit writes [] and clears a stale list inside the
caller's own single terminal write — no recorder call, no extra write, no mint.
The fast already-settled cancel lane, which performs no terminal write of its
own, runs the guarded change-gated refresh (trigger=kill_path_clear): it
touches only an existing row with a non-empty stored list, so an ordinary
fast-lane kill pays zero task-result writes (pinned) and an absent row can
never be minted (pinned). The grok-#7 shape (one recorder call inside
_reconcile_delegated_runs_on_kill) was rejected against the pin: it would add
a second write per disclosing kill on every lane that already writes its own
terminal result, and reaches the recorder's STATUS_RUNNING fallback on lanes
that run before any durable row exists.

Owner decisions honored (not re-opened): Q2=B — the delegated_runs_* counters
stay a historical snapshot at the original terminal write, never recomputed
(a healed row may read 'unreconciled: []' beside 'settled: 0'; the
delegate_terminal_reconciliation envelope with trigger + open_run_ids is the
current-liveness surface); Q5=A — reason_code is never rewritten. Patch debt
survives every refresh as patch:<run_id>, never a blind clear (bug-report
boundary, pinned). Known disclosed residual (pre-existing, unchanged): the
finalize-on-miss cancel write threads the audit into delivery only, never onto
the row; a stale settled row raced into that lane heals on the next boot's
backfill.

Docs: ARCHITECTURE §5 startup-reconciliation paragraph and the
delegated-subagents custody section document the generation-crossing refresh,
the frozen-counters owner decision and the envelope as the liveness surface;
the delegate_terminal.py module-map entry follows.

Tests: tests/test_delegate_sweep_refresh.py — settled-before-boot backfill,
undisposed-patch debt preserved, byte-level second-boot no-op, single shared
replay snapshot, lock-timeout fail-soft + next-boot heal, same-generation
sweep-settlement ordering through the real _startup_custody_sweep, absent-row
mint pin, fast-lane stale clear (trigger + frozen counters) and fast-lane
zero-write pin; tests/test_cancel_intents_phase_a.py — running-kill clean
audit clears a stale list on the cancelled result.
2026-08-30 18:15:57 +00:00
Ouroboros
88479fa756 v7next F1: domain D09 - the quiet edge transplanted, the hot core deferred with proof
Cancellation/custody sits on the F2 border and the drift probe drew the
line sharply. Quiet side, landed: the D07 Emergency Stop delta on
ouroboros/server_control.py - execute_panic_stop takes bound_port as a
keyword instead of lazy-importing the composition root, server.py hands
down _actual_bound_port() - applied in place onto tip bytes with an exact
two-way byte proof (vs tip: only the D07 hunks; vs the reference leaf:
only upstream's dc4c0204 reconcile_delegate_custody hunk). Four reference
pins ride along: test_panic_stop_port_sweep.py and
test_owner_stop_fences_s6.py re-pinned to tip bytes where upstream
genuinely moved (kill_workers kwargs, 5 host leaves not 11, the bounded
re-sweep's two token-less calls - each noted in-line),
test_cascade_chatless_residual_s6.py and
test_preflight_process_containment.py verbatim-green.
Hot side, HOT-DEFERRED with evidence in docs/v7next/LEDGER_CORRECTIONS.md
(entries 7-12): the 16-symbol cancel_custody extraction (upstream 65b5d19f
re-decomposed the same ownership differently), the cancel_intents D08
corruption rule (upstream absorbed it for request_cancel/claim_intent,
four mutators still fail open), the S7b test split bound to the deferred
extraction, and two D09-listed pins that belong to D07's module delta and
the delegation organ (delegate_start now refuses selectorless starts).
The E-suite stays with F2/F4 per the lane charter; cross-domain test
adaptations (events_task_done, loop_delivery, shell_process/git_ops_reset)
are left for their owning lanes. scripts/v7next_transplant.py +
tests/test_v7next_transplant.py byte-synced from the hardened def681bd (no
tool-emitted leaves in this lane; the in-place proof above is the byte
gate). 204 tests green in isolation at the lane base; size-ratchet
manifest regeneration is a no-op and lane 5 passed; ruff F clean.

(cherry picked from commit f1a37d3b413b66a9167fcf7e1a6fee5e229e60be)
2026-08-30 17:27:39 +00:00
Anton Razzhigaev
508c27ceea Hold on the live delegated leaf across unknown provider outcomes; carry custody provenance
Incident class: a configured-session nanny's metered round dies
provider_outcome_unknown (dispatched request, no terminal provider fact —
never resent, per custody doctrine) while its one physical delegated leaf
is alive; terminalizing the nanny let the cause-blind terminal cleanup
cancel the healthy leaf.

Main fix (D1-min): the round gate latches a durable hold and the next
round top parks the task in the same $0 supervised_wait the nanny would
have chosen. Eligibility is narrow and fail-closed: exact-route configured
sessions, exactly one open run, no pending invocations, no open
containment fault, and a READ-ONLY engine poll proving a live
non-terminal state. A meaningful leaf wake resumes with a NEW round whose
transcript carries the wake receipt (bound to the unknown attempt id);
owner dialogue drained at the round top resumes the same way. One wake =
one dispatch: an unacknowledgeable wake fails closed to the no-resend
terminal with its receipt removed. Control wakes (Stop, deadline,
finalize_now — re-checked at the source), daemon refusals, round-limit
boundaries, and budget exits all close the hold into no-call terminals,
never a paid dial; in-process loop exits clear the latch (a worker crash
preserves it for recovery). Repeated unknown cycles re-latch behind a
bounded backoff floor kept under the idle-rail minimum.

Custody companions: every provider-death arm of the rail stamps
terminal_origin=host_salvage through one wrapper (deadline grace finals
and scheduled swarm handoffs keep their legacy shape; budget/round-limit
rails stay untouched); durable llm_api_error events bind the physical
attempt (capture resolved through the explicit cause chain) and the
bounded transport cause type, which also leads the unresolved-attempt
reason; and the periodic sweep, after settling runs, re-runs the
read-only terminal custody audit so a stale delegated_runs_unreconciled
disclosure heals instead of lying forever (retry-lineage projections read
the original row live). Doctrine docs updated to state the resend
boundary precisely.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-30 12:03:03 +00:00
Anton Razzhigaev
4a295255af
Merge pull request #385 from razzant/rework/pr348-supervisor-custody
Supervisor process custody: boot-qualified fingerprints, downgrade-safe token pair, one monotonic clock
2026-08-30 10:40:49 +03:00
Anton Razzhigaev
1eb7e2a0d6 Close the custody-identity classes: downgrade-safe token pair, no bare-tick kills, one monotonic clock
Three class-level completions of the parent commits, plus the reflow
cleanup:

- Downgrade-safe fingerprint pair: the unversioned `start_time` field
  keeps the legacy `ps` spelling an N-1 reader compares correctly (a
  rollback that read a boot-qualified token there mismatched every row
  and pruned it WITHOUT killing - a permanently orphaned process on a
  supported managed-update flow), while the current reader prefers the
  new `start_time_boot` sibling, omitted when it would duplicate the
  legacy spelling (macOS/BSD). One `ps` per SPAWN (the write path always
  paid it before /proc-first); the sweep's hot path stays
  subprocess-free, and a SIBLING-FREE mismatched row falls back to one
  `ps` comparison of the legacy field, so a mid-generation boot-id
  readability change never prunes a live owned row. A row that DOES
  carry a boot sibling is judged by boot evidence alone: a mismatched
  sibling is positive proof of another boot and prunes without the
  legacy fallback, whose ps-less degradation to bare ticks would
  otherwise re-qualify a cross-boot recycled pid (review finding,
  pinned by test).

- A bare tick never authorizes a kill: the tick-half shortcut accepted a
  recorded boot-id-less tick against the tick half of a live
  boot-qualified token - the exact cross-boot collision (ticks recur,
  pids recycle, command hashes repeat) the boot id exists to refuse, and
  it sat on the KILL path where the pre-change reader only pruned. The
  compatibility helper is now a strict legacy-spelling comparison; the
  one host class that ever mints bare ticks (no usable `ps`) still
  resolves PRE-UPGRADE sibling-free rows through the legacy helper's
  own tick fallback - the disclosed one-migration-window residual,
  alongside the "<ticks>." cross-boot twin. The boot id is
  stored in full (the 8-hex truncation bought nothing but 32-bit
  collision odds).

- One clock for the whole watchdog: the direct-chat heartbeat
  (`agent._last_activity_ts`) joins the supervisor tick on
  `time.monotonic()`, the wedge half compares on the same `now`, and the
  two-clock split (plus its `now_wall` plumbing) is deleted - a forward
  wall step past the 90s deadline recommended a needless /restart for a
  healthy long turn, and a backward step masked a real wedge. The test
  that pinned the wall-clock implementation is replaced by a
  jump-resilience test symmetric to the stall one, plus a source-level
  pin of all four production stamp writers. The stall toast key gains
  the server pid, since a monotonic uptime stamp alone can repeat across
  generations against the browser's restart-surviving dedupe set.

- The unrelated docstring re-wraps in platform_layer.py are dropped:
  they bought band headroom that no longer exists on the target (the
  file left the 1001-1500 band there), and they buried the real change.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-30 05:49:20 +00:00
Andrei Kaznacheev
e856a3010e Measure supervisor liveness on the monotonic clock
Symptom: an ordinary wall-clock step — an NTP correction, a DST or timezone
change, a manual set, a resumed VM — made the supervisor liveness watchdog
report a phantom "Supervisor loop STALLED ~Ns", log a durable
supervisor_loop_stall event and send the owner an alert about a loop that had
never missed a tick. A backward step did the reverse and could hide a real
stall for as long as the skew lasted.

Root cause: the liveness tick (server.py `_loop_liveness`) and the watchdog's
comparison both read `time.time()`. That value is only ever consumed as an
elapsed gap, so it never needed to be a wall stamp — and a wall stamp is
exactly what a clock jump invalidates.

Fix shape: the tick and the stall half of the watchdog move to
`time.monotonic()`. The watchdog's two halves do NOT share one `now`: the
chat-turn wedge half compares against `agent._last_activity_ts`, which is a
wall stamp written by agent.py, so a single monotonic `now` would make
`now - turn_ts` hugely negative and silently disable wedge detection forever.
`now` (monotonic) therefore feeds the stall half and a new `now_wall` feeds the
wedge half, with a comment stating why they must not be collapsed back. The
owner stall alert path is untouched; its `toast_once` key becomes a monotonic
number, still unique per stall episode within a generation.

Tests: tests/test_ws3_wedge_resilience.py seeds the liveness list on the same
clock the tick now uses. Two tests cover the genuinely new failure modes — a
wall-clock jump neither fabricates nor masks a stall, and the wedge half still
fires on a wall-stale turn under a healthy monotonic tick. Both drive the
watchdog through a controllable fake `server.time` while the harness keeps real
clocks for its own timeouts, and both fail against the pre-fix shared-clock
code. A shared `_stop_watchdog` helper joins the thread before monkeypatch
teardown so an in-flight iteration cannot run against a restored clock: the
watchdog re-reads its stop token only at the TOP of the loop, so an iteration
already in flight still completes its clock reads and checks. All four tests
that start a watchdog route their teardown through it, including the two
pre-existing ones — they run immediately before the fake-clock tests, and a
leaked iteration there would read the real `time` module against a fake
liveness stamp and fire a phantom alert into whatever the next test patched.

Compatibility: no persisted value carries the liveness tick, so nothing on disk
or on the wire changes; the tick is process-local and re-seeded per supervisor
generation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-30 05:49:20 +00:00
Anton Razzhigaev
c510de72d1 WIP: Claude-runtime retirement (Q4) + shared reviewer-row UI vocabulary (5A)
Engine: delete gateways/claude_code.py and the /api/claude-code/* endpoints,
launcher verify block, platform-layer SDK/CLI probes, CLAUDE_CODE_* dispatch
branches, CLAUDE_CODE_MODEL setting, claude-agent-sdk dependency; advisory
module loses the dead SDK budget/paid-suspect blocks (event renamed
advisory_suspect_result), error formatter drops retired probe fields.

Web: Claude-runtime status/repair panel and onboarding card removed; reviewer
rows get the Direct model / Configured subagent source picker with read-only
derived disclosure; advisory editor on the shared api_chat vocabulary (legacy
kind 'api' is parse-only); GET /api/reviewer-slots now round-trips
subagent_id references (resolved route as disclosure only).

Bench devtools: CLAUDE_CODE_MODEL dropped from forward lists, templates,
manifest active projection (read vocabulary keeps it), cybergym snapshot,
TB metadata roles; advisory bench rationale is comparability, not
impossibility.

Tests: suites repinned to the successor contracts (native episode caps,
in-process usage scope, credentials-based availability, endpoint round-trip);
size-ratchet manifest regenerated with band rationales.
2026-08-29 22:29:46 +00:00
Anton Razzhigaev
f67ed07164
Merge pull request #353 from razzant/fix/main-chat-leak-20260829
fix: task-scoped live log events carry their explicit audience; exactly-once live delivery
2026-08-29 12:01:39 +03:00
Anton Razzhigaev
ce862fc7ac fix: task-scoped live log events carry their explicit audience; live delivery of persisted rows is exactly-once
Project task trees leaked into Main Chat as empty Working... cards (issue
#296 residual): a Project child's diagnostics carry only their own task_id,
so the frame went out unaddressed (chat 0) and Main minted an unresolvable
card. The server now addresses every task-scoped live event at one seam
(supervisor/log_addressing.py): lineage from the RUNNING row, Project
binding wins, an explicit chat_id is preserved (0 is the real Skill Review
session; None is absence), then the task row's chat, then the
DirectActivityRegistry entry; direct/ephemeral turns additionally carry
their chat BY VALUE via the turn-scoped TurnEventQueue proxy, and A2A
frames are suppressed at the push_log choke while durable rows keep the
honest audience. The addressing runs at supervisor ingress, in the
server-process append sink (make_server_log_sink replaces the raw
bridge.push_log), and at every handler that owns a suppressed type's
explicit push.

Exactly-once: append_jsonl streams only logs/*.jsonl (never chat.jsonl or
state/memory/receipt stores); each process suppresses the types whose live
delivery has a dedicated owner (WORKER_/SERVER_LOG_SINK_SUPPRESSED_TYPES);
llm_usage gains its explicit addressed push; emit_review_cycles_exhausted
guarantees a live sibling even for a queue-less server-process caller.
project_chat_for_task_tree resolves all three lineage probes from one
mtime/size-cached bindings lens.

tests/test_log_forwarding.py is rewritten production-shaped (the real sink
installed — the old suite stubbed it out and asserted an exactly-once the
production wiring violated) and pins the leak repro, the addressing
precedence, the A2A end-to-end drop, the by-value turn proxy, and the
exactly-once contract. ARCHITECTURE.md documents the new module and both
contracts in the same change; the size-ratchet manifest carries the
message_bus band entry. Version carriers untouched (version-neutral
contributor-style PR).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-29 03:45:43 +00:00
Anton Razzhigaev
db910f5fc3 Harden CyberGym isolated settings and archive publication
Bind applied settings to the isolated child, keep the snapshot immutable, and make archive publication and cleanup descriptor-safe across replacement races. Add focused regression coverage and document the three-task smoke contract.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-28 04:49:41 +00:00
Ouroboros
68bd9761f7 fix: make presence admission and routing receipts honest
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-21 23:35:27 +03:00
Ouroboros
72b502228f fix: converge child reference promotion recovery
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-21 19:54:54 +03:00
Codex
37e22deb7e feat: preserve project display identity in main UI
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-21 17:17:20 +03:00
Codex
23aaa3accd fix: preserve complete authority through routed follow-ups
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-21 16:00:24 +03:00
Codex
ac2d82adb4 fix: preserve steering and attachment authority across retries
Keep owner text attempt-aware until durable settlement, preserve attachment provenance through task admission, steering, retry, and child inheritance, and make initial task admission atomic while keeping running-task steering LLM-first.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-21 15:29:33 +03:00
Codex
3459dd12e4 fix: preserve canonical authority across task boundaries
Carry exact task contracts, owner context, predecessor authority, and plan state through canonical roots, forks, subagents, and delegated work orders. Make archive-aware recovery truthful and fail before model work when a named authority source cannot be resolved.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-21 15:28:26 +03:00
Ouroboros
dc4c02047c fix: harden delegated nanny recovery and custody
Require explicit configured-session starts, persist event-only wakes until transcript delivery, bind planned restart adoption to an exact normal-exit transaction, and reconcile every non-panic terminal custody obligation.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-20 02:44:17 +03:00
Ouroboros
dca6ff9e1c Implement configured subagent runtime lifecycle
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-20 02:44:17 +03:00
Ouroboros
71190ca558 Merge managed/ouroboros (rotation-visibility mainline) into the delegation sprint
Conflict resolution decisions:
- The sprint's spawn-time rotation deferral (_admit_spawned/run_deferred_rotation)
  is SUPERSEDED by the mainline's reconcile_rotation, which rides every
  ensure_owned_gateway (spawn and attach), is idempotent/conditional, and treats
  the recovery-window 503 as retry-next-ensure — the same incident class closed
  without spawn-time state. The bounded admission wait stays; reconcile runs
  after admission, and an expired admission skips it until the next ensure.
- ReviewRouteUnavailable raises keep the sprint's typed .code; the mainline's
  ClaudexorSubscriptionWindowExhausted branch stays ahead of them.
- Admission tests re-pin the rotation TIMING against reconcile_rotation
  (fires only after admission, attach included); mainline gateway/route_health
  stubs accept the sprint's additive timeout_sec/pinned_profile kwargs.
- size ratchet manifest regenerated over the merged tree.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-18 06:46:08 +03:00
Ouroboros
98ee316858 Integrate owner-surface-fact over the repaired mainline (size-ratchet audit stays green)
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-18 06:16:57 +03:00
Ouroboros
cbd88e40ee Repair size-ratchet fallout from the #243 merge combination
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-18 06:03:34 +03:00
Ouroboros
d6bfb1f598 feat(chat): per-message owner surface fact and process presentation posture
Ouroboros advised a desktop-app owner to reload 'the browser tab' because the
runtime had no fact about the client UI: the launcher never exported the
retired OUROBOROS_DESKTOP_MODE flag, and messages carried no sending-surface
provenance at all.

Two additive facts, no device taxonomy (the model classifies raw observables):
- launcher.py exports OUROBOROS_PRESENTATION (desktop_window|browser_fallback;
  absent=web) — the process posture, rendered as runtime_env.presentation.
- The SPA measures raw observables AT SEND TIME (pywebview bridge, ua,
  viewport, matchMedia booleans, captured_at) and attaches client_surface to
  each chat frame; the gateway normalizes it through the closed-key bounded
  ouroboros/client_surface.py SSOT, stamps received_at, carries it in
  task_metadata, persists it as an optional chat.jsonl column, and renders it
  as the owner_client context fact. Non-web ingress gets a host-stamped
  {channel: <source>} fallback; promotion/steering/mailbox carry it, and the
  loop notes a mid-task surface change only when the surface IDENTITY differs
  (viewport resize is not a device change), with a neutral note for the first
  observed fact. SYSTEM.md documents the semantics and the pywebview product
  facts (no Cmd+R, SHA auto-reload).

Adversarial waves 1-2 are folded in: Infinity viewport crash at ws ingress,
strict booleans, no-identity facts never mint change notes, provenance-honest
prompt wording, behavioral producer tests, send-site and received_at pins.
client_surface helpers live in their own module (message_bus/loop/chat.js stay
inside their ratchet sizes).
2026-08-18 05:54:03 +03:00