Commit graph

45 commits

Author SHA1 Message Date
Ouroboros
cbf33022bc fix: settle child file and mailbox custody before cleanup
Centralize post-admission drive settlement, preserve captured identities and complete input closures, make metadata reads pure, serve confined nested files and directory archives, and keep maintenance off the supervisor loop. Preserve generation fences at actual mutation boundaries and truthful queued forwarding receipts.
2026-09-26 15:04:49 +03:00
Ouroboros
d49273bee4 merge: preserve TZ2 dialogue semantics over TZ1 ingress 2026-09-26 07:12:29 +03:00
Ouroboros
80fdcee156 fix: receive emergency controls while ordered chat acceptance settles 2026-09-26 06:12:30 +03:00
Ouroboros
192a5780a1 fix: settle late quiz forwarding off the ASGI loop 2026-09-26 05:42:05 +03:00
Ouroboros
9878b66ec9 Merge current ouroboros into focused TZ1 ingress and access candidate 2026-09-26 05:14:48 +03:00
Ouroboros
0a065d2048 fix: keep web ingress off ASGI loop and align rights closure tests
Await durable web acceptance in the existing off-loop completion seam; test responsiveness with a held ingress lock. Diagnose long edit needles honestly, and update old access-matrix assertions while preserving child secret/write refusals. Version-neutral contributor fixup.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 03:30:38 +03:00
Ouroboros
90b757712b feat: make owner ingress and startup status honest
Close tool-rights affordances and bounded edit diagnostics, retain per-message batch transport with live-process tail handback, preserve accepted web input evidence through quiz responses, and show Starting/Input saved without implying a running supervisor. This version-neutral contributor slice deliberately defers artifact custody, directory ZIP, and V10 to a separate structural PR.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 02:31:33 +03:00
Ouroboros
aafaa37138 fix(tz2): close seven exact-range review findings (D15 quiz ingress, paid interruptions below adapters, action-only stop, suppressed-origin corpus, telegram card ordering, degraded reflection/promotion, split-root ceiling)
Repair batch authored by a delegated Claude Fable run; reproduced each finding via a failing test through the real seam.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-26 00:59:44 +03:00
Ouroboros
d4b7996e02 fix: keep post-task cancellation and steering receipts truthful 2026-09-25 09:28:40 +03:00
Ouroboros
32e8ddd8fe fix: retain free task facts and refuse post-final owner mail 2026-09-25 08:59:21 +03:00
Ouroboros
0ab204da11 Merge remote-tracking branch 'origin/ouroboros' into ouroboros-agent/tz2-bc-open-question
# Conflicts:
#	docs/inventories/FACADE_INVENTORY.md
2026-09-25 08:14:56 +03:00
Ouroboros
3f9fe1b654 WIP: preserve TZ-2 open-question and factual-terminal candidate
Keep this version-neutral contribution local. B1/B2/B3 and C1/C4/C5 are partial; C2/C3, exact diff review, full suite and PR are outstanding. Do not merge this checkpoint without final-candidate checks.
2026-09-25 06:03:38 +03:00
Ouroboros
5fb287a63d Apply the review wave: slim budget line everywhere, flush before restart, honest batch time budget
Findings from the simulated triad (gpt-6-astra, grok-4.7, fable-5.1), one batch:

- The budget line's unbounded-budget branch still rendered the full
  usage_breakdown; it now reads usage_writer_snapshot like /api/state, so the
  chapter 05 sentence that both read the slim shapes is true on every install.
- A restart that lands right after a drain now flushes that turn's dirty budget
  projection before leaving the loop, so the last drained llm_usage events
  still reach state.json.
- The events batch time budget is checked between handlers; the docstring,
  the constant comment and chapter 05 now say a running handler can overrun it
  instead of promising a hard elapsed-time bound.
- The socket-open project refresh forces its read: the in-flight read predates
  the socket and was just invalidated, so joining it would bring nothing fresh.
- The update_budget_from_usage docstring distinguishes the loop's once-per-turn
  llm_usage path from direct callers; plan labels removed from test docstrings.
2026-09-25 03:56:23 +03:00
Ouroboros
6ce5df9542 Bound the supervisor events pass and write the budget projection once per turn
Every llm_usage worker event made the supervisor loop render the FULL usage
breakdown (five grouped axes plus the per-root projection over the whole
ledger) to persist a handful of totals, take STATE_LOCK and rewrite
state.json, and the loop drained the event queue to empty before it read the
owner's bridge; at one settled attempt per second the events phase never
returned to intake and the chat went silent while the header said Online.

The compatibility writer now reads a slim render-cached snapshot
(usage_writer_snapshot: totals, the ordering marker, the OpenRouter bucket its
drift check compares, the totals-only projection) computed from the same
validated rows as usage_breakdown, so the marker and the money still come from
one read; /api/state and the budget line read the same slim shapes.
_projection_from_final moves verbatim into the _usage_rows leaf and is
re-exported. An llm_usage event only marks the loop context dirty (its row and
live frame say "deferred"); the loop flushes one write per turn AFTER bridge
intake, keeps the flag on a refused or failed write and retries no more often
than BUDGET_PROJECTION_RETRY_SEC, and the OpenRouter ground-truth check fires on
crossing each multiple of 50 physical calls (a coalesced write may jump 49 to
51). One events pass drains at most SUPERVISOR_EVENT_BATCH_MAX_EVENTS or
SUPERVISOR_EVENT_BATCH_MAX_SEC, FIFO, restart requests and event-lag
observation preserved, the remainder next turn; a turn that hit its bound skips
the idle sleep. The drain and the flush live in server_liveness beside the
other loop-phase helpers, so server.py shrinks.

Tests pin the snapshot against the full breakdown on every key the writer reads
with the full render spied out, byte-identical state.json from either render,
the 49 to 51 crossing, N events to one write after intake with "deferred" rows,
the dirty-flag retry contract including a corrupt ledger, the count and time
bounds with intake reached while events remain queued, and /api/state serving
the same values for both budget branches without the full breakdown.
2026-09-25 03:21:41 +03:00
Ouroboros
42784b7604 Decide orphan reconciliation on a status-only read; stamp its cadence at pass end
The 300-s zombie reconcile called the materializing task-result loader
before it knew whether a RUNNING row was alive, so every pass copied a live
child's whole artifact tree into the canonical store, hashed it, and rewrote
the artifact manifest once per file; its cadence marker was stamped before
the pass, so a pass slower than 300 s re-armed on the very next tick and the
supervisor loop stayed in maintenance for minutes (issue #1230). The
reconciler now decides on the status-only projection and materializes only a
row it is about to settle (the healed row keeps full artifact custody, quiz
and owner-wait settlement), the cadence is stamped when the pass ends, and a
manifest registration whose merged document equals the current one skips
the rewrite for every materializer. Tests pin a live child (one projection,
zero copies, untouched manifest), a genuine orphan (decide then heal, file
promoted, quiz expired, persisted fields equal a direct read), the cadence
stamp including a failing pass, and the no-op manifest write by bytes and
mtime; the loop chapter, invariant 24 and the behavioural docstrings say so.
2026-09-25 03:03:09 +03:00
Ouroboros
b50f10d4b5 Stop persisting the per-root cost map in state.json
The compatibility budget projection re-attached usage_projection's by_root
map to every state.json write, so the file grew with the number of root
tasks (899 KB on a long-lived install) and every state save, load and
STATE_LOCK hold paid for it, while nothing reads the persisted map back:
every per-root reader renders from the usage ledger. The writer now keeps
totals only and the fallback branch asks the ledger for the slim
projection, which makes state.json size independent of root count
(issue #1002). Tests pin the scale invariance, the untouched ledger-side
per-root capability and the slim fallback branch; the chapter and the
persistence inventory say what the persisted projection carries.
2026-09-25 02:54:36 +03:00
Anton
3fb5ddd518 Merge origin/ouroboros c9180ef8 (#1252 cards, #1244 safe-check, #1254 shutdown, v7.4.10) into feature/1196-continuity
Conflicts: generated inventories/DOMAIN_MAP taken from upstream and to be regenerated; chapter 06 money paragraph taken from upstream (ours was a compression of the same text). contracts.py/chat.js trimmed back under their ratchet limits. KNOWN RED: size_ratchet_manifest.py stale — runtime function count 10036 > 10000 after the merge (upstream grew ~42 functions since 87fd00f4); to be paid down by simplification, not a cap raise.
2026-09-24 17:21:56 +03:00
Ouroboros
23f52d0d96 refactor: move the signal-owning uvicorn shapes to server_process (#1142 CI)
quick-test/full-test failed two extraction contracts: the runtime_limits owner
inventory had no row for the new constants, and server.py crossed its 1700-line
bound. _SignalStopServer and _embedded_uvicorn_server now live beside the events
they set in ouroboros/server_process.py (server re-exports them; test seams
unchanged); the inventory maps the two constants; the server bound moves
1700 -> 1730 with the reason recorded next to it.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-24 15:39:34 +03:00
Ouroboros
78780e0b7d fix: keep an ordinary close inside the launcher budget (#1142)
The launcher's graceful stop SIGTERMed the whole server process group, so the
multiprocessing Manager and pooled workers died before the server's own
lifespan teardown began; the supervisor loop met BrokenPipe on EVENT_Q three
times and raised the false "Supervisor loop died" owner alarm, while uvicorn's
unbounded graceful drain parked the terminal-custody write past the launcher's
10 s wait, ending three window closes in SIGKILL.

Server half (self-sufficient against an old immutable launcher):
- server._SignalStopServer.handle_exit sets _supervisor_stop at the signal, so
  the loop leaves its tick and a torn-Manager error is never a crash;
- uvicorn.Config carries timeout_graceful_shutdown=SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC
  (runtime_limits, paired with LAUNCHER_STOP_GRACE_SEC as one budget).

Launcher half (lands with the next release): the POSIX graceful phase signals
only the server PID; the group SIGKILL fallback and post-exit sweep stay.

Tests: unit coverage of the handler, the constant wiring and stop_agent, plus a
serial real-process test that reproduces the old group SIGTERM with an in-flight
request open and asserts the teardown lands and no supervisor_failure is written
(it fails on the unfixed server). Docs 01/05/09 updated; version-neutral.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-24 14:22:26 +03:00
Anton
bea4049572 WIP: preserve merged no-tool pause semantics and structural cleanup
Focused verification passed; global function-count gate and full consumer verification remain unresolved. Not publication-ready.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-24 05:59:39 +03:00
Anton
612067385c WIP: preserve #1196 continuity repairs before upstream integration
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-24 05:09:55 +03:00
Ouroboros
0abab76225 ouroboros: checkpoint after task c6207d8d665040b0 — Долгая работа #1196 — продолжение до реализации 2026-09-23 23:52:17 +03:00
Ouroboros
95d50ab703 Merge remote-tracking branch 'upstream/ouroboros' into ouroboros/cross-focus-awareness
# Conflicts:
#	docs/DOMAIN_MAP.md
#	docs/inventories/DATA_LAYOUT_INVENTORY.md
#	docs/inventories/FACADE_INVENTORY.md
#	prompts/CONSCIOUSNESS.md
2026-09-22 17:56:57 +03:00
Ouroboros
334a4b2ac7 Cross-focus awareness: authored focus, live-root catalogue, per-room memory consolidation
One Ouroboros across Main and project rooms: each root may publish a short
authored focus (update_focus) that rides the durable task result and the
[INDEPENDENT_ROOTS] tail, so concurrent foci can see one another without a
shared chat, a new wake or any widened authority; explicit cross-room
journal/workpad reads are honoured instead of silently redirected.

Dialogue consolidation summarizes each source room of a logical chunk from
its own bytes (Light draft + source-grounded Light correction), assembles the
typed room sections deterministically into one shared block and carries
them through era compression; a failed room withholds its whole chunk while
earlier complete chunks stay published; legacy mixed blocks keep unknown
provenance. Room labels resolve against the canonical registry root even on
a forked task drive. BIBLE P1 states the principle in three sentences.

Version-neutral contribution: release carriers untouched.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-22 17:25:35 +03:00
Ouroboros
4ef3acb1e3 feat: consciousness Observe boundary, agency-first wake prompt, owner-governed schedule lifecycle
Wake prompt (prompts/CONSCIOUSNESS.md) opens with judgment and continuity instead of
budget and a closed menu; cash, subscription quota and time are named separately.
Observe autonomy keeps research, internal memory, task artifacts, project notes and
read-only children (argument-narrowed at dispatch) while every world-mutating verb,
patch disposition and publication route is withheld (OBSERVE_WORLD_MUTATION_TOOLS).
Schedules gain an audited owner-governed lifecycle (manage_schedules tool +
POST /api/schedules/{id}/action): disable/delete/restore with intent-then-outcome
audit rows, durable suppression of skill-manifest rows keyed by schedule id (restore
lifts it only into an extant manifest), consumed one-shots as collapsed history in
Activity, typed lock-timeout and unreadable-store refusals; the owner lifecycle lives
in supervisor/schedule_lifecycle.py beside the store in queue_schedules.py.

Version-neutral contribution; release carriers untouched.
2026-09-22 17:08:03 +03:00
Ouroboros
cf7daa3e8d Make the supervisor-loop chapter's pointer to the reader contract a sentence
'A display read (invariant 28); admission reads exactly.' was a fragment.
It now reads 'This read is display (invariant 28); admission is exact.',
one byte longer, inside the chapter's budget.
2026-09-21 17:07:26 +03:00
Ouroboros
8abd444aec Point the two chapters at the reader contract's real number
The money-lock package was written against a base where its rule was
invariant 26. On the current target 26 is the cross-process compaction guard
and 27 the predecessor door, so the reader contract is invariant 28; the
key-invariants chapter already says so, the supervisor-loop chapter and the
rules chapter still named 26.
2026-09-21 15:11:29 +03:00
Ouroboros
13588ef8e8 Document the reader contract of the usage ledger: display rides a snapshot, money never does
- ARCHITECTURE §10: invariant 26 states the contract; invariant 24 no longer
  names the heartbeat handler's ledger lock as a residual of the loop thread —
  what is left there are the exact reads that settle or refuse.
- ARCHITECTURE §5: the `state.json` projection read is named as a display
  read (byte-neutral replacement of one sentence; the chapter had 4 bytes).
- ARCHITECTURE §1: the rows memo's module line.
- DEVELOPMENT "rules by change class": the ledger-lock rule gains the reader
  contract, and the sentence about the lock's caller wait is replaced. The
  chapter had 5 bytes left, so its budget rises by 323 with the reason
  recorded beside the number; both architecture chapters fit their budgets.
- The data-layout inventory is regenerated: it pins the SHA-256 of §1.
2026-09-21 14:56:51 +03:00
Ouroboros
be203df868 Run the periodic custody block off the thread that answers workers
Every 600 s the supervisor loop thread ran skill-payload hashing, the
orphaned-process reaper, delegated-run reconciliation (gateway handshake, several
full custody replays, registration retirement) and the settled-terminal cursor
INLINE: 11.6-17.6 s on a healthy install, far longer on a sick one, ahead of
assignment and behind every worker ack. It now rides a daemon thread under a
non-blocking module lock exactly like the 20 s sweep — the cadence marker still
advances on the loop thread, a busy pass makes the tick SKIP rather than queue a
second one, and the latch is released in finally.

Running off-thread removes the hidden invariant that no assignment could interleave
with the sweep, so the read order becomes load-bearing: reap_orphaned_processes and
reconcile_orphaned_runs now accept the live set as a set OR a zero-arg callable and
evaluate it AFTER they have read their candidates (ledger, open_runs, pending
invocations). A candidate exists implies its owner was registered earlier — admission
takes _queue_lock before any spawn — so an owner absent from the LATER snapshot is
really gone, while a snapshot taken first reaps a task admitted during the read.
None still means unknown and touches nothing, on both surfaces.

One live source now serves both custody surfaces and the settled-terminal cursor:
_live_task_ids over _startup_live_task_ids — RUNNING and busy worker slots under
_queue_lock, the direct-activity registry, in-flight post-task synthesis, memory
only, no durable result load. The periodic set used to be narrower than the startup
one (no busy workers, no post-task synthesis, RUNNING read without the lock), so the
two surfaces could disagree about the same owner. The cursor pass defers a task whose
owner is still billing instead of healing its disclosure under a live writer.

The pass re-reads the loop's per-generation _watchdog_stop token before each mutation
and stops, and its reconciliation gateway is ATTACH-ONLY once a stop, restart or panic
is in flight: an ensure there would start the engine the teardown is about to end
(ARCHITECTURE section 9, DEVELOPMENT Process Custody Rule).

server_maintenance.py is at its 1000-line ceiling, so the block was paid for in
place, not extracted: the eight startup prune/GC steps wrote the same
events.jsonl row through eight copies of the same five lines and now share one
writer, and the runtime-artifact condition loses three clauses that named keys
neither of its reports has ever carried.

Docs: ARCHITECTURE 05 and 09 replace the sentences describing the node that moved
(05 is byte-negative), and key invariant 21 states the rule with its owners and
names the residuals it does NOT close — the 300 s zombie reconcile and the
usage-ledger lock in the heartbeat handler still run on the loop thread.

Consciously updated tests: test_delegated_reconciliation.py's live-set expectations
(both surfaces now receive the one callable, so the assertions read what it
produces), test_claudexor_startup_failure.py's sweep tests (the block owns a
"custody-maintenance" thread beside the latch-retry one), test_claudexor_custody_lifetime.py
and test_supervisor_loop_measurements.py (join that thread before reading its effects).
2026-09-21 04:58:32 +03:00
Ouroboros
269ad6ced5 Measure a stalled supervisor loop by phase, CPU and event lag
On 2026-09-19 the loop stalled 64 times (95-1091 s) and the journal could
say nothing but how long: no phase, no end, no CPU-vs-wall split. Every
candidate cause - the inline 600 s custody block, the usage-ledger lock in
the heartbeat handler, GIL starvation by heavy direct-chat threads of the
same process, a daemon/pin mismatch generation - stays unproven because the
onset row carries no fact that separates them.

The loop now stamps liveness once per coarse tick phase (events |
maintenance | assign) and publishes with each stamp the facts only the loop
thread can take honestly: its own time.thread_time() delta over the interval
that just ended (a thread that BURNED the wall gap reads differently from
one blocked or starved), the worst worker-stamped lag of the last drain
(events without a worker ts are skipped, never invented), and the in-memory
daemon-pin match. supervisor_loop_stall carries them beside the wall gap and
one supervisor_loop_stall_end closes the episode when the loop ticks again -
an onset without an end is a generation that never recovered.

The watchdog thread only reads that list: no lock, no daemon call, no disk,
because a watchdog that waits on the thread it watches reports nothing. Its
monotonic contract is unchanged, and so are the owner chat notice and toast.
No registered-project count rides along: that set exists only as a full
custody-log replay, and a measurement may not pay disk on the thread it
measures.

server.py sits at its pinned 1700-line bound, so the four added lines are
paid inside it: a duplicated message_bus import, a single-use state
temporary and two intermediate locals in _get_owner_chat_id.

tests/test_server_shutdown.py drives the real loop against a stubbed clock
namespace; it gains thread_time, which the loop now samples per phase.
2026-09-21 04:26:45 +03:00
Ouroboros
c55876238b Leave no orphan behind an interrupted root on restore
A PLANNED shutdown already refuses to start the PENDING children of an
interrupted root: kill_workers(preserve_pending=True) settles each with the
trigger `pending_parent_interrupted`. An UNPLANNED stop never reached that
code — boot-time kill_workers sees an empty RUNNING — so the snapshot fenced
the surviving RUNNING roots while their queued children were restored and
started on the first tick as orphans under a parent that no longer exists
(#1104). Direct-chat roots missed the fence entirely: they never enter the
queue, and the only record of them, `state/direct_roots.json`, was cleared by
queue init BEFORE restore could read it, so a window closed on a live turn
projected to failed / orphaned_running_after_worker_restart instead of the
honest cancel the owner decided on 2026-09-12, and its children were orphaned
with it.

Restore now builds the set of interrupted ancestors — the rows it just fenced,
plus the ancestors an earlier boot already handed to custody (an active intent,
or a stored cancelled result whose cancel_origin names `server_shutdown`) — and
marks every otherwise-revivable PENDING child below one with the SAME
shutdown-custody marker the planned path writes. The boot's own kill step then
settles it through the existing path, with a ledger-reconstructed cost and a
published task_done. The marker is written only after the existing gates have
proved the row unowned and revivable, so a child custody already owns, a row
that arrived carrying its own marker, and a row with no parent (roots,
schedules, evolution) are all untouched, as are the children of a root the
restore is reviving — an owner-wait handoff is a continuation, not an
interruption. The acceptance-fence ancestry walk and this one are now the same
helper instead of two copies of the same loop.

Queue init keeps the clear that stops a stale process's turns from outliving
it, but takes the roster over in the same step and hands the ids to restore,
which fences them like any other surviving RUNNING row; an `incomplete` roster
names no complete set of live turns, so it hands over nothing and its roots
keep today's projection, disclosed on the durable restore row. No second store:
the handover is the same fragment, in memory, read once.
2026-09-21 04:17:27 +03:00
Ouroboros
84d449139b docs: keep recovery disclosure within chapter budget 2026-09-20 15:03:29 +03:00
Ouroboros
b294fa48ce docs: disclose compatibility projection recovery gaps 2026-09-20 14:57:09 +03:00
Ouroboros
03fac785c7 Merge current ouroboros and repair Windows history-ACK test
Normalize checkout line endings before extracting the tested function; align the epoch documentation with the monotonic writer. Version-neutral contribution continuation.
2026-09-20 13:45:30 +03:00
Ouroboros
2ed63f8aa9 fix(state): read the usage ledger outside STATE_LOCK
update_budget_from_usage held STATE_LOCK (4.0s timeout) across
ensure_legacy_imported, usage_breakdown and usage_projection, which
contend for the monetary lock (45.0s timeout). A short-timeout lock held
across a long-timeout one starved every state.json reader, /api/state
among them, and stalled the supervisor loop. The ledger read now happens
outside the critical section; only load, compare, mutate and save remain
inside.

The lock was held deliberately: reading the ledger outside it lets an
older concurrent snapshot acquire STATE_LOCK later and regress the
monetary fields. That invariant is preserved by an ordering marker
instead of by lock duration. The marker is the lexicographic pair
(compaction_epoch, max live row seq), derived by _usage_rows from the
same validated rows the buckets come from and exposed as the private
_ledger_high_water_seq. Neither component works alone: compaction starts
a fresh dense epoch, so the live seq restarts at 1, header seq is always
1, pre_compaction_seq is per-row provenance into the archived range, and
source_last_seq is the just-folded file's length. Only compaction_epoch
is globally monotonic, and only the pair orders snapshots both inside an
epoch and across a fold.

Any marker strictly lower than the saved one is refused, whether the
epoch or only the sequence differs. An earlier revision of this change
wrote a lower epoch on the theory that it could only mean a restore.
That is wrong, and no restore is needed to break it: a writer can read
at epoch N, a reserve_attempt can compact to N+1, a second writer can
save the post-compaction projection, and the delayed writer would then
overwrite state.json with stale money. Equal and higher markers take the
normal write path. A genuine restore can therefore freeze the
compatibility projection until usage_ledger_high_water_seq is cleared;
that belongs to the restore path, not to this writer, and the refusal is
logged rather than silent. Each decision carries a greppable substring
(FRESHNESS MARKER UNKNOWN, STALE SNAPSHOT REJECTED).

The marker and every persisted accounting field now come from one
validated snapshot: usage_breakdown renders the projection from the same
final rows, so two writers cannot share a marker while saving
projections read at different sequences. An unreadable ledger is
unknown, never zero: the previous projection is preserved,
update_budget_from_usage returns a typed false, and _handle_llm_usage
records projection_update_status=unavailable instead of claiming an
availability it does not have.

usage_breakdown's private provenance is filtered at the gateway
boundary, where the unbounded-budget branch reuses the whole mapping as
its /api/state accounting projection. The filter drops every
leading-underscore key, so the class is closed rather than this one key.

Measured on a copy of a real 33 MB ledger in a throwaway data root, one
writer against six readers for 25s: reader lock-timeout fallbacks 21 to
0, reader max latency 4.05s to 1.07s (the 4.0s timeout ceiling was being
hit), reader throughput 5019 to 20503 calls. With six writers, 18 to 0
across three runs. Writers alone are faster, not slower: 431 to 855
calls, p50 29.3ms to 16.6ms.
2026-09-20 04:14:09 +03:00
Ouroboros
782412a1a7
Merge PR #1064: preserve outcomes and cancellation facts with cost custody
Preserve terminal outcomes, usage accounting and cancellation provenance
2026-09-18 15:26:34 +03:00
Anton Razzhigaev
7e4a7796b7 Keep live turns visible and honor worker startup progress
Reserve chat composer space with a flex item for native WebKit. Combine queue and live activity identities through the existing task controls. Give a worker with its own early progress one readiness extension bounded to 300 seconds from spawn. Wait for the actual chat socket before the large-attachment browser test sends its files.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 12:12:28 +03:00
Ouroboros
6e67c353a6 Preserve terminal task evidence and abandoned usage custody
Phase A candidate for issues #641, #639 and #738. Focused checks passed; external phase review is pending.
2026-09-18 04:06:08 +03:00
Ouroboros
439b9e65ca Merge landed runtime fixes into the consolidated reference books
Preserve current request accounting, recovered tool evidence, execution-start binding and shared owner restart descriptions while retaining the compressed chapter structure. Regenerate inventory provenance.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 03:06:04 +03:00
Ouroboros
9a20a1bf59 Consolidate reference books and add documentation checks
Rewrite repeated mechanism prose while preserving inventories, rationale and mode distinctions. Correct stale documentation and move the live E2E operator manual beside its implementation. Add official-CI chapter byte budgets and narrower residue checks.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-18 01:54:01 +03:00
Ouroboros
5152832d28 Record a running task in the canonical root when the execution drive is split
A task whose durable record lives in the install's data root while its work runs on a
separate headless or project drive only ever wrote its "running" row to that execution
drive, so the canonical result stayed at "scheduled" while the child said "running".
When the worker or the machine died, the queue snapshot was gone and both orphan healers
start from rows already stored as running, so a scheduled row was invisible to them
forever and the task became a permanent ghost that never finished and never booked its
cost. The agent now mirrors the same start timestamp and the child drive address into the
canonical root through the existing locked writer, which still refuses a late start over a
row that is already terminal. Startup terminal-file recovery additionally rebinds the
canonical row for ghosts older installs already carry, but only from a known non-direct
child's positive running record and only when the existing fresh-queue, later-worker-boot
and no-active-cancel checks already prove that child orphaned; it repairs the record and
never resumes execution. Carved out of PR #873, where this core fix arrived mixed into
the Android host port.
2026-09-17 23:54:19 +03:00
Ouroboros
4408d25b40 Merge managed/ouroboros (PR #964, promote refusals) into the consciousness redesign
Conflicts: ouroboros/tools/control_routing.py (both kept: the presence note inside
its branch, the consciousness origin stamp after it), docs/architecture/03, 06, 12
(their new sentences kept, the consciousness wording re-applied), and the generated
v7next inventories (regenerated).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-16 11:33:26 +03:00
Ouroboros
e3b4981e1b Rewrite the books for a consciousness that takes an ordinary turn
The books still described a private background loop with its own prompt, its
own context, its own card and an observation inbox. They now describe what the
code does.

Architecture: 01 gives `consciousness.py` its alarm-clock description and adds
the missing `consciousness_wake.py` row beside the two P3 modules, drops the
retired observation inbox from the data layout and the deferred-frames clause
from owner_delivery, and names CONSCIOUSNESS.md as the wake template rather
than a second system prompt. 03 says the alarm does not start a wake while an
owner direct turn is live (a running wake is not interrupted by one), drops the
always-shown kind from the block predicate, and describes what Activity and
Evolution actually render. 05 puts the alarm tick in the supervisor pass and
stops claiming a separate consciousness event producer. 06 replaces the
"Background consciousness and Evolution" opening with the new design — an
ordinary Main turn, its origin, its Observe/Act/Full authority with
dispatch-only `disabled_tools` and the per-task mode cap, and one rolling-24h
allowance through the single admission door — and keeps the evolution
paragraphs untouched. 07 loses the retired max_tokens row; 10 loses the inbox
invariant; 11 records the accepted late quiz answer, `max_wait_minutes`, and
that no third activity `kind` was introduced.

Development: 02 stops saying consciousness carries its own prompt; 03 drops the
retired observation growth row; 04 says its context IS the Main context; 06
drops the wake-scoped `ModelTurnState` and states the cache consequence of the
shared prefix. DESIGN.md loses the always-shown kind and the reusable
background card, and states the quiz card's new "you can still answer" family
and the bounded-wait notice. PERSISTENCE.md marks the inbox retired with its
one-time archive move.

The three new modules were missing from `ouroboros/domains.toml`, which left
the domain manifest red; they are D15 and the generated sections and
DOMAIN_MAP.md are regenerated (D15 shed two strict edges, D07 gained the
admission door's edge). The hot-store count is eight everywhere, and the
generated v7next inventories follow their sources.

CHECKLISTS.md (a protected file, owner-sanctioned В30=A) gets three minimal
wording fixes: item 13(b) drops "(or the background whitelist)" because no
whitelist exists, 13(d) points at the wake template and at wakes rather than
"background loop behavior", and the Tool schema critical-surface row names the
wake template.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-16 06:32:26 +03:00
Anton Razzhigaev
2f8aa7d4e2 Promote refusals: cause to the model, receipt-only for live turns, one typed System row for host-issued acts
A supervisor refusal of a chat-issued promote used to reach the owner as a
standalone Ouroboros bubble (WORKSPACE_UNUSABLE ...), reach the model without
its cause, and label the owner's message "Choose a target" with no options.

- Every workspace refusal returns `detail` (cause + repair) composed by
  `workspace_admission.workspace_repair_hint` from the typed source of the
  refused folder; `_persist_promote_rejection` is the single durable writer.
  `_fail_promoted_task_loudly` and `_explicit_workspace_remedy` are gone.
- Placement follows who can narrate: a tool-issued act gets its receipt, the
  failed-call error row and `detail`; a host-issued act (skill card, Swarm,
  picker click, stamped `host_initiated`) gets ONE typed System row
  (`task_not_started` / `task_start_unconfirmed`) in the chat the owner wrote
  in, from the one publication boundary `_handle_promote_chat_to_task` wraps
  around every promote outcome. The skill-repair untyped bubble is removed;
  steer cancel-pending notices obey the same owner-labelled rule.
- The host owns the owner-facing sentence: `project_dialogue.routing_refusal_cause`
  (action + status + reason, e.g. "Not started: the working folder can't be
  used") rides the annotation, the live `message_annotation` frame, history
  replay, `MessageAnnotationOutbound`/`DecisionResponse` and the picker's 409
  body; the browser renders `cause` verbatim and keeps no client table.
- Admission-notice rows stay in the chat they were sent to on replay
  (`room_membership` ignores the never-started task's project binding); the
  "Project · Started" row is announced only after the task is really queued.
- `workspace_root` naming the Ouroboros repository itself maps to the existing
  `workspace="none"` sentinel at the promote tool with a disclosure; subfolders,
  the data drive and every other caller keep the typed refusal.
- Docs (DESIGN, architecture 01/03/04/05/06/12, DEVELOPMENT naming rule),
  generated inventories, python and web tests.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-16 03:56:03 +03:00
Ouroboros
35b8151819 Split the two reference books into verbatim chapters
The books were one physical file each: ARCHITECTURE 2,386 lines and
DEVELOPMENT 3,886, with 13 and 14 `##` sections. `reference_books.py` had
shipped the chaptered reader — membership, authored introductions, exact
physical source views — with the migration still at zero, so every reader
took the legacy monolith branch and the validator had no production caller.

Each `##` section is now one chapter file under `docs/architecture/` or
`docs/development/`. The only new bytes per chapter are its prologue: the old
section title at H1 (numbering text kept, so every `ARCHITECTURE "8. Git
Branching, CI, and Build"` cross-reference still reads) and one authored
introductory paragraph saying what the chapter owns and why it exists.
Everything after that prologue is the old section body byte for byte, with
`###`/`####` levels untouched — so the residue rules keep reading the exact
subsection headings they exempt, and no section title was renamed.

The move is therefore INVERTIBLE, and `tests/test_reference_book_migration.py`
inverts it: drop each chapter's H1 line and its one introduction, re-prefix
`## `, concatenate in membership order, and require the recorded SHA-256 of
the old body — plus, whenever the base commit is reachable, byte equality with
`git show <base>:<path>`. `docs/reference-books-migration.md` is the operator
transfer table: every row a verbatim move, with its line range at the base, its
destination and an empty rename column. It lives directly under `docs/`, so it
is reviewable without becoming a book member.

`_preamble` now takes the FIRST paragraph under the H1 instead of demanding the
only one before the first H2. Most sections open with prose at the level they
already had, so the old rule could only be satisfied by promoting `###` to `##`
or inventing a sub-heading — either of which would rewrite what this migration
relocates verbatim. A source whose H1 is followed straight by a subsection
still has no introduction and is still refused, which is the property that
keeps an overview from quoting body prose as authored orientation.

`.gitattributes` pins `docs/**/*.md` to LF: chapter line ranges, byte spans and
SHA-256s are physical facts that a Windows checkout must not rewrite, and
`full-test` runs on windows-latest for every PR. The two ARCHITECTURE-derived
generated inventories are regenerated, because a chaptered section now carries
its physical provenance note.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-15 19:20:44 +03:00