Review round 1: recovered splits are not failed runs, Light authors summaries, index fallback, explicit requests stay explicit

A refused-then-split chunk that was fully summarized carried its attempt errors
in usage and marked the run failed, so the stale error survived the advance and
Health said FAILED after a success; only a chunk that produced no content fails
now. A chronicle-only consolidation pass reports zero written blocks. The Light
maintenance prompt — its only carrier, since Light never sees the knowledge_write
schema — now asks for an authored YAML summary on new and meaningfully revised
notes and keeps an explicit standing request explicit; SYSTEM.md says the same
about corrections and holds interpretations as testable. A missing generated
index falls back to the fresh inventory so summaries stay resident. Docs state
the batch receipt, the local compactor's preserved sections, memory_mode's
canonical writes, and the read side of the Presence baseline; the Presence
end-to-end test also executes update_scratchpad and update_identity.

Prompt bytes: SYSTEM.md 24071 -> 24131.
This commit is contained in:
Ouroboros 2026-09-15 00:22:41 +03:00
parent 9ac697eed1
commit 17dbc3687c
11 changed files with 80 additions and 17 deletions

View file

@ -173,7 +173,7 @@ server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (de
├── mcp_client.py ← MCP client: parses MCP_SERVERS, validates transports, masks tokens, prefixes tools `mcp_<server>__<tool>`, guarded SDK import; MCP descriptions/results stay untrusted data (§6)
├── safety.py ← Safety supervisor call with a bounded newest-first context budget (the omission marker is reserved INSIDE the budget); a 429 on the safety check is an infrastructure fact about the supervisor, not a verdict about the tool call — one deadline-capped retry, then the typed `⚠️ SAFETY_UNAVAILABLE` non-verdict telling the agent to retry the same call, not reword it; process-local storm latch answers in-window checks without provider calls; durable `safety_check_rate_limited` audit event; structured insufficient-quota keeps its PERMANENT classification and still blocks as a verdict; the serialized SUBJECT has its own 250k-char budget (`_SAFETY_SUBJECT_CHAR_BUDGET`, rendered `ensure_ascii=False`) and an over-budget subject is refused fail-closed with the typed `⚠️ SAFETY_SUBJECT_TOO_LARGE_BLOCKED` denial plus a durable `safety_subject_too_large` event — never truncated, because anything past a cut would run unreviewed; the fail-open cases are owned by prompts/SYSTEM.md
├── consciousness.py ← Background thinking loop with live progress emission (§6)
├── consolidator.py ← Dialogue consolidation with a generation-aware cursor over the ordered archive chain; an unfindable generation appends a loud durable `[MEMORY GAP]` block, never a silent offset reset; a run that advances without a new failure clears `last_consolidation_error`, and each knowledge-nomination batch leaves its publication receipt in `dialogue_meta.json`
├── consolidator.py ← Dialogue consolidation with a generation-aware cursor over the ordered archive chain; an unfindable generation appends a loud durable `[MEMORY GAP]` block, never a silent offset reset; a run that advances without a new failure clears `last_consolidation_error`, and a knowledge-nomination batch that is not fully published leaves `last_unpublished_nominations` in `dialogue_meta.json` (a fully published batch clears it)
├── memory.py ← Scratchpad, identity, chat history
├── knowledge.py ← `ouroboros/knowledge.py`: shared linked-Markdown note addressing, exact source reads, revision-checked writes and generated shelf indexes for global and project knowledge; one writer preserves authored understanding, unknown metadata and provenance so concurrent cognition cannot silently overwrite a newer note or invent understanding from a truncated prefix
├── memory_journal_compaction.py ← Digest-only compaction of old memory-journal snapshots: the digest replaces the snapshots it summarizes, never a silent drop
@ -493,7 +493,7 @@ Frontend calls go through `web/modules/api_client.js` with the JSDoc mirror `web
`POST /api/tasks` creates an ordinary managed root; `GET /api/tasks` is a non-materializing list; `GET /api/tasks/<id>` returns the effective durable result; `/events` is the archive-aware SSE stream (§3 Chat); `/artifacts/<name>` serves simple filenames confined to `data/task_results/artifacts/<task_id>/` — a stored arbitrary path is not a download capability. The CLI refuses any `delegation_role` other than `root`, the gateway rejects caller lineage/subagent labels, and only `schedule_subagent` creates children. Reserved service metadata is written after caller metadata. Admission reserves the task id plus a worker-pool slot under one queue lock; a failure rolls back only the token-owned row with a loud typed refusal. The parsed request's complete admission and attachment staging run off the HTTP event loop through `gateway._helpers.run_sync_to_completion`; a cancelled HTTP waiter retains that worker until durable admission or rollback settles, without cancelling the admitted task. Detail reads and SSE terminal materialization use the same settled wait, and v2 closes its row iterator only after the outstanding read finishes. Attachments are copied into the effective task drive before enqueue; artifact-store references are not host-path authority.
Workspace tasks default `memory_mode=forked`; `shared` is rejected for an external workspace and materialized on a forked child drive for project scope — the stored `memory_mode` reports what was requested while `drive_root` reports where the task executes, so isolation does not depend on relabelling the request.
Workspace tasks default `memory_mode=forked`; `shared` is rejected for an external workspace and materialized on a forked child drive for project scope — the stored `memory_mode` reports what was requested while `drive_root` reports where the task executes, so isolation does not depend on relabelling the request; the mode isolates the execution drive and the knowledge seed, while identity and scratchpad writes still land on the canonical root the next context reads.
`--detach` returns only after durable admission; `--no-stream` polls to completion; waiters treat a result terminal only after the artifact state leaves `pending`/`finalizing`, and an explicitly-partial cost gets a bounded 60-second finality wait before partial flags stay visible. `ouroboros run` exits 0 only for a completed lifecycle, a clean execution axis, no failed/degraded objective, and a finished artifact bundle — strict exit semantics keep shell automation from interpreting "the model answered" as "the requested workspace deliverable exists"; `--patch`/`--patch-out` are stricter still (failed/missing patch, no-change, empty payload, or unfinished finalization is an error).
@ -1191,7 +1191,7 @@ For ordinary delegated profiles, the URL, route, private-range, and control-plan
#### Context fitting, retry, and compaction
`config.get_context_mode()` is the effective Main sizing/rendering source, while `config.get_owner_context_mode()` is the persistent owner-intent/P3 source; they differ only during the auto-Low compatibility window. In Max, `ARCHITECTURE.md` is full-resident for every task class because it is Ouroboros's capability/tools/access map; in Low it is replaced by its lossless navigation map. `DEVELOPMENT.md` is mode-independent: full when the active repository binding says the work targets Ouroboros's own body; a bound external workspace, any subagent, or an API/CLI/scheduled external surface receives the visible on-demand pointer (`context_requires_development`/`context_requires_self_body_docs` override) — context economy comes from dropping the self-engineering handbook for external work, never from hiding the capability map in Max. `prompts/SYSTEM.md` and `BIBLE.md` are tier-0 and full in both projections (`context_layout.TIER0_ALWAYS_FULL`); the one path that omits SYSTEM.md sections is the local-model overflow compactor (`llm.py::_compact_local_text`), hence the prompt's load-bearing floor lives in its preamble; its semi-stable pass keeps Identity and Shared understanding by heading prefix and leaves the knowledge index as an omission notice — the one disclosed exception to the always-loaded index. A tool's contract lives in its `get_tools()` schema, sent every round; mechanism documentation lives in this file — a prompt sentence restating either is a second copy that drifts.
`config.get_context_mode()` is the effective Main sizing/rendering source, while `config.get_owner_context_mode()` is the persistent owner-intent/P3 source; they differ only during the auto-Low compatibility window. In Max, `ARCHITECTURE.md` is full-resident for every task class because it is Ouroboros's capability/tools/access map; in Low it is replaced by its lossless navigation map. `DEVELOPMENT.md` is mode-independent: full when the active repository binding says the work targets Ouroboros's own body; a bound external workspace, any subagent, or an API/CLI/scheduled external surface receives the visible on-demand pointer (`context_requires_development`/`context_requires_self_body_docs` override) — context economy comes from dropping the self-engineering handbook for external work, never from hiding the capability map in Max. `prompts/SYSTEM.md` and `BIBLE.md` are tier-0 and full in both projections (`context_layout.TIER0_ALWAYS_FULL`); the one path that omits SYSTEM.md sections is the local-model overflow compactor (`llm.py::_compact_local_text`), hence the prompt's load-bearing floor lives in its preamble; its passes keep Identity, Shared understanding and the dynamic memory sections (Scratchpad included) by heading prefix and leave the knowledge index as an omission notice — the one disclosed exception to the always-loaded index. A tool's contract lives in its `get_tools()` schema, sent every round; mechanism documentation lives in this file — a prompt sentence restating either is a second copy that drifts.
`context_fit.py` renders Max, Low and Nano projections from one immutable core and measures each ordinary Main candidate on one labelled density basis against the selected route capacity. Owner Nano uses its compact source view and bounded output target; Low adds an elastic 200K total-context economy target; crossing either target is never a synthetic failure. With unknown capacity, owner Max gets one honest Max call while owner Low or Nano may reclaim toward its known target. Predicted pressure retains the selected document projection and may request one deficit-sized mutable-history pass; only an actual provider overflow authorizes task-local Low, and one same-route semantic recovery is permitted only when the final candidate has the same route, round, and response reserve and strictly fewer context-bearing bytes. Owner mode and P3 applicability never change.
@ -2327,7 +2327,7 @@ Operation correlation (#667): a named injected message has `operation_ref=<chat_
Successful extension children also carry a bounded `ws_relay_failures` count map from the existing PluginAPI transport owner through their result envelope and process facts. Missing transport, network errors and HTTP refusals remain visible without changing `send_ws_message -> None` or the producer's success. One host warning reports the aggregate; diagnostics contain no message bodies, URLs or credentials. A child that dies before its final envelope may lose this aggregate; the existing measured death/timeout facts remain authoritative. The separate streaming route implementation carries the same aggregate in completion X after body/background work.
Presence flow: `POST /presence/turn` requires the content-hash-bound `presence` permission, one `binding_id`, one exact transport event, and optionally staged files confined to the skill's state root; the binding resolves from `state/presence_bindings.json` with exact provider/account/conversation/thread origin verification; cross-process locks enforce the installation-wide cap and serialize one `conversation_key`; a stable event-derived task id makes transport retries idempotent; input and output join ordinary dialogue history with full transport/actor provenance. The one typed outcome is message/silent/tool_delivered/deferred — `deferred` only with a correlated `work_ref`, because an unanchored "deferred" would be an unanchored promise — and `GET /presence/work/{work_ref}` polls the bound late result without exposing the general task API. Promotion of Presence work into a managed task clears requested Project/workspace/source widening: a public conversation may promote long work but cannot choose new authority, and the cost ceiling plus return destination follow the promoted root by value. `presence_cancel_work` acts only on a `work_ref` whose stored binding and conversation match the current turn; owner chat and Background Consciousness may `initiate_presence` on an existing enabled binding. Every admitted turn's ceiling also carries the cognitive-memory baseline (knowledge read/write/list, scratchpad, identity, chat history): one mind keeps one memory in every channel, so an external message can prompt an in-turn memory or identity revision on the profile's model slot rather than only a post-task knowledge write; the correspondent gains no tool, and public text still carries no owner-command authority. Ceilings frozen before the baseline existed verify their own digest and carry no baseline until their profile is recompiled.
Presence flow: `POST /presence/turn` requires the content-hash-bound `presence` permission, one `binding_id`, one exact transport event, and optionally staged files confined to the skill's state root; the binding resolves from `state/presence_bindings.json` with exact provider/account/conversation/thread origin verification; cross-process locks enforce the installation-wide cap and serialize one `conversation_key`; a stable event-derived task id makes transport retries idempotent; input and output join ordinary dialogue history with full transport/actor provenance. The one typed outcome is message/silent/tool_delivered/deferred — `deferred` only with a correlated `work_ref`, because an unanchored "deferred" would be an unanchored promise — and `GET /presence/work/{work_ref}` polls the bound late result without exposing the general task API. Promotion of Presence work into a managed task clears requested Project/workspace/source widening: a public conversation may promote long work but cannot choose new authority, and the cost ceiling plus return destination follow the promoted root by value. `presence_cancel_work` acts only on a `work_ref` whose stored binding and conversation match the current turn; owner chat and Background Consciousness may `initiate_presence` on an existing enabled binding. Every admitted turn's ceiling also carries the cognitive-memory baseline (knowledge read/write/list, scratchpad, identity, chat history): one mind keeps one memory in every channel, so an external message can prompt an in-turn memory or identity revision on the profile's model slot rather than only a post-task knowledge write; the correspondent gains no tool, and public text still carries no owner-command authority; on the read side, an unselected baseline `chat_history` grant is unbound, so — like `knowledge_read` over the notes kept about people — a reply can surface the owner's own words to the correspondent through the model's judgment. Ceilings frozen before the baseline existed verify their own digest and carry no baseline until their profile is recompiled.
Companion processes are host-supervised: reviewed manifest-declared descriptors enter durable custody, reconcile after lifecycle changes and restart, and stop on disable/unload/panic; `state/extension_generation.json` carries the opposite direction — the server's published live set, which a running task worker adopts at a task's start or at a dispatch miss, so an enable after boot is not invisible until the pool respawns. Worker-side changes write durable reconcile requests (`state/extension_reconcile/`) rather than spawning server-owned children; every reconcile state names the process that answered and whether that marker request was written, and the tool and review receipts pass both facts through. Health observations are process-qualified: aggregate and Skills UI health use the server observation as authority and expose the worker observation with its handoff outcome only as a qualifier, so failed handoff success cannot advance `last_known_good`; restart-budget exhaustion persists a terminal reason in that health state, cleared only by a later successful start. A companion's cwd is the reviewed payload directory, so a payload edit stales review before reload instead of silently mutating a live process. The live projection is `state/extension_companions.json`.

View file

@ -158,8 +158,8 @@ def consolidate(
and not (pressure_fits is not None and pressure_fits())):
reduced = _compact_chronicle(blocks_path, llm_client, identity_text, knowledge_context)
merged = _merge_consolidation_usage(*([usage] if usage else []), reduced)
if usage and "_blocks_written" in usage: # a fixed-key merge would drop this receipt
merged["_blocks_written"] = usage["_blocks_written"]
# A fixed-key merge would drop this receipt; a chronicle-only pass wrote no block.
merged["_blocks_written"] = (usage or {}).get("_blocks_written", 0)
usage = merged
return usage
finally:
@ -322,7 +322,9 @@ def _run_block_consolidation(
meta.pop("consolidation_retry", None)
if not content and usage.get("_consolidation_retry"):
meta["consolidation_retry"] = {"source_sha256": source_hash, "input_limit": usage["_consolidation_retry"]}
if usage.get("_consolidation_errors"):
# A refused part that was split and then fully summarized still carries its
# attempt errors in usage; only a chunk that produced NO content failed.
if usage.get("_consolidation_errors") and not (content and content.strip()):
run_failed = True
meta["last_consolidation_error"] = dict(
usage["_consolidation_errors"][-1], cursor_offset=last_offset + processed,
@ -654,7 +656,9 @@ unknown metadata and useful links. A new observation may correct an old interpre
do not merely repeat fragments. New topics may be created without a prior read.
Understanding of the people involved — preferences, recurring reactions, shared history,
tentative interpretations with their source — is ordinary knowledge to nominate in global scope;
a pattern across several moments is worth more than one; revise the existing note rather than minting a rule.
a pattern across several moments is worth more than one; revise the existing note rather than minting a rule,
and an explicit standing request stays explicit. Author a YAML summary for a new or meaningfully revised note —
the summary is what stays resident in the index — and revise it when the note's meaning changes.
The note overview (scope global) is the shared orientation loaded into every future context; keep it
current, and when none exists and this episode gives real understanding, create it after reading the index.
Scope is a separate field, never a topic prefix.

View file

@ -651,8 +651,9 @@ def build_knowledge_sections(
# The authored summary is the resident face of a note, so the index carries it
# whether or not a common orientation exists; the fresh inventory render stays
# for the case where a note was written but its index rebuild did not land.
is_global_index = path == global_address.shelf / INDEX_FILE
text = (render_knowledge_index(inventory_knowledge(global_address), include_summaries=True)
if authored_overview and path == global_address.shelf / INDEX_FILE else safe_read(path))
if is_global_index and (authored_overview or not path.exists()) else safe_read(path))
if not text.strip():
continue
if warn_large and len(text) > _LARGE_CONTEXT_SECTION_CHARS:

View file

@ -13,7 +13,7 @@ A project-scoped task (an external/workspace task, or one given an explicit
This is a thin SSOT helper, NOT a parallel memory subsystem (P7): the existing
knowledge tool + context loader simply redirect their base dir when a task is
project-scoped, and the post-task canonical dual-run is suppressed for such tasks
so project facts cannot contaminate the global memory.
so the automatic post-task writers cannot contaminate the global memory.
"""
from __future__ import annotations

View file

@ -343,13 +343,15 @@ artifact. I distinguish known, stale, missing, and inferred, preserving source
and timestamp where it affects decisions. Knowledge holds understanding of
every kind: verified operational facts, recipes and gotchas, and the people I
work with — who they are, what matters to them, how we work well together, what
we have been through, and what I make of it. I keep apart what someone told me,
we have been through, and what I make of it, held as interpretations I can test.
I keep apart what someone told me,
what I observed, and what I infer, dated and sourced, revised in place rather
than piled up. A correction is evidence about that person in that moment, not a
standing rule — later moments refine it — and one interpretation restated across
standing rule unless they make it one — later moments refine it — and one
interpretation restated across
several notes is still one interpretation. The authored summary of a note is
what stays in front of me through the index, so I write it myself whenever I
create or meaningfully revise one, and the global note overview is the shared
create or meaningfully revise one, and the global overview note is the shared
orientation loaded into every context. Understanding of people is global
knowledge, whatever room I am working in. `knowledge_list` shows the topics;
`knowledge/index-full.md` is a reserved internal name — Do NOT call it

View file

@ -31,6 +31,42 @@ def _health_env(tmp_path):
# --- stale error is retired only by a run that recorded none of its own -----------
def test_a_chronicle_only_pass_still_reports_zero_written_blocks(tmp_path, fit, monkeypatch):
chat, blocks, meta = _paths(tmp_path)
_write_chat(chat, count=5, text_size=0) # below the block size: nothing to consolidate
monkeypatch.setattr(c, "_compact_chronicle", lambda *a, **k: {"prompt_tokens": 1, "completion_tokens": 1,
"total_tokens": 2, "cost": 0.0})
usage = c.consolidate(chat, blocks, meta, _LLM(), compact_chronicle=True, pressure_fits=lambda: False)
assert usage["_blocks_written"] == 0
def test_a_chronicle_pass_after_a_real_run_keeps_the_written_count(tmp_path, fit, monkeypatch):
chat, blocks, meta = _paths(tmp_path)
_write_chat(chat, text_size=0)
monkeypatch.setattr(c, "_compact_chronicle", lambda *a, **k: {"prompt_tokens": 1, "completion_tokens": 1,
"total_tokens": 2, "cost": 0.0})
usage = c.consolidate(chat, blocks, meta, _LLM(), compact_chronicle=True, pressure_fits=lambda: False)
assert usage["_blocks_written"] == 1
def test_light_is_told_to_author_summaries_and_keep_explicit_requests_explicit():
# Light never sees the knowledge_write schema, so the maintenance prompt is its only carrier.
assert "YAML summary" in c.KNOWLEDGE_MAINTENANCE_PROMPT
assert "resident in the index" in c.KNOWLEDGE_MAINTENANCE_PROMPT
assert "explicit standing request stays explicit" in c.KNOWLEDGE_MAINTENANCE_PROMPT
def test_a_nominated_note_with_a_summary_becomes_resident_in_the_index(tmp_path):
from ouroboros.knowledge import inventory_knowledge, render_knowledge_index, resolve_knowledge_address
ctx = ToolContext(repo_dir=tmp_path, drive_root=tmp_path, budget_drive_root=str(tmp_path), task_id="t")
entry = {"topic": "people/alex", "scope": "global", "expected_revision": None,
"content": "---\ntype: understanding\nsummary: Alex asks for brevity; an interpretation to test.\n---\nEvidence."}
outcomes = c._write_knowledge_entries(tmp_path / "memory" / "knowledge", [entry], context=ctx)
assert outcomes and outcomes[0]["ok"]
rendered = render_knowledge_index(inventory_knowledge(resolve_knowledge_address(tmp_path, "overview", "global")))
assert "Alex asks for brevity" in rendered
def test_clean_run_clears_a_stale_error_from_an_earlier_run(tmp_path, fit):
chat, blocks, meta = _paths(tmp_path)
_write_chat(chat, text_size=0)

View file

@ -110,6 +110,9 @@ def test_oversized_logical_block_splits_complete_source_and_advances_once(tmp_pa
assert not c.should_consolidate(meta, chat)
refused = usage["_consolidation_errors"][0]
assert refused["kind"] == "context_overflow" and not refused["preflight_only"]
# A refusal that was split and then fully summarized is a recovered attempt, not a
# failed run: the block was written, so no stale error may outlive the advance.
assert "last_consolidation_error" not in json.loads(meta.read_text())
assert refused["capacity_tokens"] is None and refused["input_limit"] is None
sizes = [len(call["messages"][0]["content"].encode("utf-8")) for call in llm.calls]
for index, size in enumerate(sizes[:-1]):

View file

@ -7,9 +7,9 @@ or a pattern, and reflection nominated "self-knowledge". The observable result
on a live install was that everything ever written about the human had been
dictated word for word, never noticed and never inferred.
These assertions are about what the prompts MEAN — which memory each kind of
understanding has, and where it lives — not about a particular turn of phrase.
Rewording is free; dropping one of these contracts is not.
These are vocabulary pins for the memory contract — which memory each kind of
understanding has, and where it lives. They pin the words that carry each
contract, not a model's behaviour; dropping one of these contracts is the defect.
"""
from __future__ import annotations

View file

@ -77,6 +77,15 @@ def test_note_summaries_are_resident_without_any_authored_orientation(tmp_path):
assert "Details belong to their source." in sections
def test_note_summaries_are_resident_when_the_generated_index_is_absent(tmp_path):
env = _env(tmp_path)
_note(env)
index = resolve_knowledge_address(env.drive_root, "overview", "global").shelf / "index-full.md"
index.unlink() # a note landed but its index rebuild did not
sections = "\n".join(context.build_knowledge_sections(env))
assert "Details belong to their source." in sections
@pytest.mark.parametrize("project_id", ["", "proj-1"])
def test_the_gap_travels_into_every_room(tmp_path, project_id):
env = _env(tmp_path)

View file

@ -166,6 +166,10 @@ def test_admitted_external_turn_writes_global_knowledge_and_nothing_else(tmp_pat
"content": "Alex asked for short answers today; I read it as a preference to test.",
},
)
seen["scratchpad"] = registry.execute(
"update_scratchpad", {"content": "Alex's room: brevity today, worth testing next time."})
seen["identity"] = registry.execute(
"update_identity", {"content": "I am Ouroboros. " + "I keep one memory in every room. " * 4})
seen["refused"] = {
name: registry.execute(name, args)
for name, args in (
@ -189,6 +193,10 @@ def test_admitted_external_turn_writes_global_knowledge_and_nothing_else(tmp_pat
note = (data / "memory" / "knowledge" / "people" / "alex.md").read_text(encoding="utf-8")
assert "short answers" in note
assert "PRESENCE_CAPABILITY_BLOCKED" not in str(seen["write"])
assert str(seen["scratchpad"]).startswith("OK")
assert str(seen["identity"]).startswith("OK")
assert "brevity today" in (data / "memory" / "scratchpad_blocks.json").read_text(encoding="utf-8")
assert "one memory in every room" in (data / "memory" / "identity.md").read_text(encoding="utf-8")
assert COGNITIVE_MEMORY_TOOL_NAMES <= seen["schemas"]
for name, refusal in seen["refused"].items():
assert name not in seen["schemas"]

View file

@ -271,7 +271,7 @@ def test_forked_project_child_skips_global_knowledge_seed(tmp_path):
def test_d5_project_scoped_shared_preserves_mode_but_isolates_drive(tmp_path, monkeypatch):
"""D5 (Option A) gateway contract: a project-scoped `shared` task keeps its RECORDED
memory_mode 'shared', yet the drive is MATERIALIZED isolated (effective mode 'forked')
so project facts can never leak into global/shared memory. Guards against a future
so the automatic writers can never leak project facts into global/shared memory. Guards against a future
'simplification' back to passing memory_mode straight through to prepare_task_drive."""
import asyncio
import json