Commit graph

23 commits

Author SHA1 Message Date
Ouroboros
95078b23bc fix: recover callable tool names on exact catalog misses
Provide scoped MCP raw-to-wire discovery, pure pre-safety name resolution, and a shared policy-filtered refusal path for tool namespaces. Preserve exact dispatch and extension adoption; cover real registry consumers and classification.
2026-09-25 03:46:59 +03:00
Ouroboros
4ff299846e feat: support addressed continuing participant turns 2026-09-24 05:39:44 +03:00
Ouroboros
d643dfd44f fix: preserve release diagnostic classification and consumer fixtures 2026-09-23 03:12:20 +03:00
Ouroboros
c693c0fad3 Retain the focus source as a native task_source ref readable through get_task_result; bound it, refuse unsettled task sources 2026-09-22 19:12:51 +03:00
Ouroboros
334a4b2ac7 Cross-focus awareness: authored focus, live-root catalogue, per-room memory consolidation
One Ouroboros across Main and project rooms: each root may publish a short
authored focus (update_focus) that rides the durable task result and the
[INDEPENDENT_ROOTS] tail, so concurrent foci can see one another without a
shared chat, a new wake or any widened authority; explicit cross-room
journal/workpad reads are honoured instead of silently redirected.

Dialogue consolidation summarizes each source room of a logical chunk from
its own bytes (Light draft + source-grounded Light correction), assembles the
typed room sections deterministically into one shared block and carries
them through era compression; a failed room withholds its whole chunk while
earlier complete chunks stay published; legacy mixed blocks keep unknown
provenance. Room labels resolve against the canonical registry root even on
a forked task drive. BIBLE P1 states the principle in three sentences.

Version-neutral contribution: release carriers untouched.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-22 17:25:35 +03:00
Ouroboros
ff06711492 Pin the approved tool schema and environment result contracts 2026-09-18 04:33:05 +03:00
Ouroboros
ce6e2a137e Consciousness: register the capped-tree integration refusal; set_next_wakeup says when its interval applies
- INTEGRATE_CAPPED_TREE joins the classifier's current-producer contracts as an
  integration_blocked outcome.
- set_next_wakeup's receipt and schema say the interval is finish-relative: it applies
  from the end of the next wake-up, a pending wake-up keeps its time (astra scope,
  round 7).

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-16 15:37:05 +03:00
Ouroboros
a2ebc17a3b Pin the live classification of the scope-bind receipts
ensure_project_scope now answers SCOPE_REJECTED / SCOPE_UNCONFIRMED through the
same refusal register as the routing family, and the differential test refuses a
harvested identifier that has neither a golden answer nor a present contract.
Both identifiers are new to the tree, so their live answer is asserted here (a
recorded refusal, never ok), exactly as the promotion receipts were pinned.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-15 18:05:14 +03:00
Claude
94a9a67ce5 Report what the routing rail actually did
A Project root steered the Main root on 14.09; the supervisor wrote a refusal
receipt and the model was told STEER_UNCONFIRMED, so it kept retrying a steer
the host had already declined in writing. Four independent defects produced
that one sentence, and each also mis-reports other routing outcomes.

A task-authored steer carries no chat ingress id, so it went out with an empty
client_message_id: the supervisor wrote no annotation and the wait answered
`client_message_id_missing` before the handler had even run. Such a steer now
keys the synthetic `agent-steer:<token>` receipt the drained-delivery path
already mints, so refusal AND delivery are confirmable. Nothing in any chat
carries that id, so no owner message is labelled with the agent's act, and
`_relayed_owner_message` keeps relaying nothing under the prefix.

`wait_for_routing_annotation` accepted delivered/needs_manual_target/
unconfirmed but not `rejected`, which is exactly what the cancel-pending
refusal writes — so a settled refusal was polled to the deadline and returned
as `confirmation_timeout`. It is accepted as the settled outcome it is, and
UNCONFIRMED keeps its one meaning: no matching receipt exists.

The promotion receipts opened line 1 with no warning marker at all, and
PROMOTE_*/ROUTE_*/ROUTING_UNCONFIRMED/NEEDS_MANUAL_TARGET were absent from
`_EXACT_IDENTIFIER_CODES`, so a routing act that scheduled nothing counted as a
SUCCESSFUL tool call in the outcome classifier and the acceptance packet. The
family is registered beside the steer receipts, in the same recorded-but-not-
degrading bucket: the agent sees the refusal, and the host's "no" is not the
agent's execution failure.

An admitted promote named the project it ASKED for. An implicit scope is
re-resolved by the admission handler under the origin claim lock, so a sibling
card converted in the emit -> admission window moves the root somewhere the
tool never named; the wait now carries the effective destination back and the
result reports that, while an unconfirmed admission says what was requested and
that the outcome is unknown.

The differential could not have caught any of this: the four control_routing
corpus rows declared codes their plain-string producers never publish, so each
row asserted its own input. They carry no code now, which puts the producer's
own sentence in front of the one classifier.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-15 17:06:57 +03:00
Ouroboros
00fb549012 Align composite contracts and generated fixtures
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-13 11:48:49 +03:00
Ouroboros
2eb32f1689 Preserve direct resource contracts and single execution after agency review 2026-09-13 01:52:00 +03:00
Ouroboros
2e5f066bb4 P5-fix5b: give the listing miss the corpus case P5.2 asked for
The differential derives its expected key set from the corpus, and the
corpus derives every LIST_FILES_NOT_FOUND key from the tree itself:
`harvested_identifiers` sees the literal prefix of the producer's
f-string and yields `ident:LIST_FILES_NOT_FOUND:{plain,named}`, while
`harvested_native_pairs` sees the same `ToolResult(...)` call carrying a
literal `code="LEGACY_WARNING"` and a leading literal text, and yields
`native:LEGACY_WARNING:LIST_FILES_NOT_FOUND`. All three golden rows
therefore stay: removing any one of them fails
`test_single_classifier_matches_the_retired_pair_except_approved_deltas`
on its own `case.key in golden` assertion, the plain row additionally
fails `test_golden_covers_every_harvested_producer`, and the native row
additionally fails `test_native_golden_answers_have_identical_retired_text_inputs`.
The third row is not a hand-picked extra: the producer's native code
path publishes it and the derivation demands it. The identifier rows
were what P5.2 called "a corpus row for the new identifier", and the
harvest produces them without a hand-written entry.

What was genuinely missing is the producer's own shape. The miss
interpolates the exception into its text and publishes it under
`list_files`, and no case exercised that: the identifier rows carry a
synthetic detail under `read_file`, and the native row reuses that exact
input by construction. The shape row adds it, and the golden gains the
one answer its key forces.

That answer is derived, not invented. The retired pair no longer exists
in this repository and its source tree is unreachable here, so a new
shape answer could only be reused from a recorded one. The new test
`test_shape_golden_answers_match_their_own_identifier_line` states the
rule that permits the reuse, the sibling of the native-key rule already
asserted beside it: every shape row whose text opens with a marker holds
exactly its `ident:IDENT:plain` answer, over the rows captured from the
golden's own source tree, with no counterexample. GOLDEN_SOURCE_SHA and
the fixture's recorded `source_sha` are unchanged, and nothing was
regenerated.
2026-09-12 18:13:27 +03:00
Ouroboros
f13e5c2984 P5.1b: record a refused steer as a refusal the agent must see
Completes owner item I23 with the two rows the earlier P5.1 commit left on hold
(5ac8fbe44). The owner answered batch #3, 6g = A: the refusal is visible as a
refusal (is_error, policy-denial telemetry) and the task's execution health axis
is NOT degraded, because a host refusal is not the agent's failure.

Both steer refusals are plain sentences - control_routing publishes no typed
result for them - so the identifier table is the only reader that can classify
them, and it recorded them status=ok. The agent's own trace then said the
message had been delivered while the host had refused it.

STEER_REJECTED and STEER_UNCONFIRMED now map to the EXISTING
TOOL_REPORTED_FAILURE code, whose homing already matches the owner's answer
exactly: is_error=True for the counters and the anti-loop scan, routed to
policy_denials rather than unresolved, and declared a non-failure by
outcomes._LEDGER_NON_FAILURE_STATUSES since v6.83.0. No new spec, no new bucket,
no producer rewrite, and the register is BY CODE so stored traces of past
refusals are re-read too.

ouroboros/tools/tool_result.py is SAFETY_CRITICAL (runtime_mode_policy.py:229).
The diff in that file is exactly the two table rows the phase names.

Tests: the two APPROVED_DELTAS rows the frozen differential requires (it fails
both on an unapproved divergence and on a row that no longer fires), and a case
in tests/test_control_native_results.py that runs the real producer sentences
through the classifier and the bucket collector: is_error true, policy_denials,
unresolved empty, and a confirmed steer still classified OK.
2026-09-12 18:13:26 +03:00
Anton
323eb77400 P5.3: admit the Pattern Register on typed evidence, not a marker scan
Owner item I24. The register's gate was a substring scan: _detect_markers looked
for eight hand-listed words inside a bounded head+tail view of every result body,
and append_reflection/append_reflection_routed opened the register only when that
scan found one. It was blind twice over. A typed failure whose word nobody had
listed never opened it, and a root whose own calls all succeeded while its
CHILDREN failed never did either - 15 of the 16 failed tasks of the incident day
were children, and children do not reflect (ARCHITECTURE, Post-task reflection).

(a) _detect_markers now returns the sorted typed codes of the calls that went
wrong, read from the tool_result_code loop_tool_execution already stamps beside
every trace row, with the recorded status as the fallback for a legacy row that
predates the code (kept verbatim rather than dressed up as a code it never had).
_ERROR_MARKERS, _marker_scan_view and _has_error_evidence are deleted with their
last consumers: the scan inside should_generate_reflection's loop, the one in
_collect_error_details, and the prompt choice, which now reads the error_count
the same function already computes. reflection.py is net-negative (764 -> 748).

(b) One typed helper, _admits_pattern_register, replaces the duplicated
key_markers gate at both writers: an errored call, a typed code, or a failed
child. Deliberately NOT "reason_code is non-empty", which would open on every
terminal.

(c) _child_task_evidence returns (text, rows) from its single existing walk and
_run_reflection derives the typed child classes from rows[*].outcome_axes
.execution, stamped beside error_count as child_failure_classes. No new
collector, no second walk, no second register writer (the knowledge write lock
and its CAS stay the sole one), and children still do not reflect.

Durable-field note: key_markers now stores typed codes; older reflections keep
their marker strings and stay readable by _update_patterns and context.py, which
both only join the strings.

Tests: the substring pins in tests/test_process_memory.py are rewritten as typed
tests keeping their contracts (a doc that MENTIONS a marker is still not error
evidence; a real failure is still detected wherever its text sits), plus the
admission cases the step names - zero markers with error_count>0 admits, a root
whose only failures are children admits, a clean non-trivial task does not, and
an unknown error_count is not evidence. reflection.py's row in the differential's
_RESIDUAL_TEXT_INSPECTIONS is removed: all six hits were _ERROR_MARKERS, so the
module now holds none.
2026-09-12 18:13:26 +03:00
Anton
2e525cb138 P5.1: record a refused child-result disposition as an argument error
Owner item I23: a typed refusal was being recorded as status=ok. Add one row to
_EXACT_IDENTIFIER_CODES mapping CHILD_RESULT_DISPOSITION_INVALID to the EXISTING
TOOL_ARG_ERROR code, matching the typed arms join_ledger.py already publishes for
the single form. The register is BY CODE, so this one row also covers the three
plain-string producers that never publish a typed result (the batch envelope and
per-entry rejections in join_ledger.py, the ledger-append path in
task_tree_ledger.py, and the task_tree.py schema text) and re-reads the stored
traces of every past refusal, which a producer-only cutover cannot.

CHILD_RESULT_DISPOSITION_PARTIAL deliberately stays LEGACY_WARNING: those entries
did record.

ouroboros/tools/tool_result.py is SAFETY_CRITICAL (runtime_mode_policy.py:229).
The diff in that file is exactly the one table row named by the phase; no spec,
no bucket, no producer and no classifier branch changed.

HELD, not implemented here: the STEER_REJECTED and STEER_UNCONFIRMED rows of the
same owner item wait for an owner answer on whether a refused steer touches the
task's health axis (batch #3, R3-32).

Tests: one APPROVED_DELTAS row in tests/test_tool_classification_differential.py
(the table fails both on unapproved divergence and on rows that no longer fire),
and the four text-only pins in tests/test_child_result_disposition.py now also
assert the typed code, including the partial-batch carve-out. The frozen golden
fixture is untouched: ident:CHILD_RESULT_DISPOSITION_INVALID already exists in it.
2026-09-12 18:13:26 +03:00
Anton
13c5dcd223 Preserve transport work and reconcile exact review outcomes
Recover complete reviewer results from their original operation CAS, preserve terminal custody separately from model narrative, and consume remote streams within physical-attempt accounting. Apply caller deadlines before every recovery send, publish known tool refusals accurately, and retain raw reviewer PASS separately from completion eligibility.

Managed work waits through unknown provider outcomes and continues with a marked new attempt after route-bound upstream observation, retaining old monetary custody. Metadata HEAD observations reuse the connection bound; ordinary no-deadline clients retain their existing timeout defaults.

Preserves the complete interrupted implementation and source-only version policy. Focused consumer, golden, lifecycle, domain, inventory, lint and size checks pass. The initial full preflight failed on stale consumers and a missing manifest compatibility field; the final full preflight is still required on this committed candidate.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-11 03:45:18 +03:00
Ouroboros
cfc3b00849 test: refresh the retired classifier oracle after Repair cleanup 2026-09-07 04:43:05 +00:00
Ouroboros
b7b9311674 fix(github): report repository target refusals as argument errors 2026-09-06 16:24:46 +00:00
Ouroboros
a60307d35a test: align integration fixtures with current tool contracts 2026-09-06 15:18:11 +00:00
Ouroboros
269cd303ef F2 review wave: one fix batch over the upstream absorption
Review wave on the merge commit 9698e2e0 (four Claudexor lanes: codex sol
triad/scope, fable, grok; four adversarial lenses: hand-merge reconstruction,
lost upstream changes, relocation mechanics, docs and test honesty). Every
output was read to EOF and dispositioned; this is the one batch.

Runtime:
- context_fit imported the retired `ouroboros.delivery_protocol` (folded into
  loop_delivery on the v7 line): `import ouroboros.loop` failed on a clean
  checkout. The plain-text extractor is now a call-time wrapper over
  loop_messages (loop_messages imports ouroboros.llm at module top, so a
  top-level import would be a cycle). The venv's editable install of the live
  repo had masked this; the battery now strips that finder.
- masked-green exit disclosure gates on the process fact exit_code == 0, not
  the typed status, so the undeclared-output and artifact-error publications
  keep upstream's disclosure; one test per site.
- loop_nudges reaches the relocated skill-trace lookup through the loop handle
  (facade patches bite); the declared handle set follows.
- delegate_custody_usage: unreachable trailing return removed.
- ChatOutbound / api_types declare `cancel_physical_task_id` (the incident meta
  emits it).

Tests: patch targets follow the v7 leaves (wire-recovery preparer in
llm_attempt, every lane that binds the physical-candidate executor, the
process-role owner for reconcile receipts, the four multi-line reviewer-slot
patches that stayed on the dead name); the Windows skip marker returns to the
symlink widget test; adapted upstream tests pin exact messages and exported
values instead of loosened predicates; vacuous scans (chat-id records,
progress_meta keys, benchmark slot counts) are made to bite; residual text-
inspection inventory follows the core_secret_paths split; a keyless system-E2E
scenario (S26) proves the owner stop of an in-flight direct chat turn.

Docs: ARCHITECTURE component-map pointers follow the merged owners
(outcomes.extract_final_answer, review_verdict.aggregate_dialogue_status,
FORBIDDEN_SKILL_SETTINGS, worker_chat_lane namer, tools/tool_context.py,
process_custody.py row, nine hot stores, cost-breakdown leaf); cost_projection
row states the 7.0 contract (retired cost_usd spellings are read-only
tolerance, never emitted); contracts/api_v1.py and delivery_protocol are no
longer named as live modules; relocation-ledger corrections appended.
2026-09-04 20:17:45 +00:00
Anton Razzhigaev
8ac736a82e tests: close the sync's red battery — 29 failures, each at its real cause
The full non-serial run after the merge surfaced 29 reds. None was fixed by
weakening an assertion; each is either a pin retargeted onto the campaign owner
or a real gap the sync opened.

Product gaps closed:

- the skill owner-state read carve never reached its call site: the guard was
  taught `writeish` but still called without it, so `rg review.json` stayed
  refused with a WRITE-named marker;
- a post-exec tripwire lost its typed fact whenever the producer returned plain
  text. With the notes now TRAILING the payload the marker no longer owns line
  1, so a text-only reader could not re-derive the classification — the guard
  adapts once through the one legacy adapter instead;
- the reasoning-pin ContextVar sat in llm_messages, which imports llm_attempt,
  closing an import cycle the leaf-graph test caught. It moves to
  reasoning_artifacts, beside the fact it carries, where neither producer nor
  reader imports the other.

Pins retargeted with their reason named: the composer's retired
ambiguous_safety_wrapper, the LLMClient member inventory (second documented
move), the declared handle sets of three leaves, the registry facade count,
the gh-auth denial text, and two observation seams the campaign split moved
(events_budget's append, registry_guard_process's guard entry).

Goldens re-recorded by their own documented recipes, every diff explained: the
llm route goldens now carry the reasoning_pin fact and pin openai/* on a sealed
artifact (#468 absorbs the family carve-out); the classification golden gains
five new warning identifiers and loses two retired ones. One golden case was
repaired rather than re-baselined — `or_provider_never_unpins_reasoning` held a
READABLE artifact, which the shape-first classifier never pins, so the case had
silently stopped testing its own contract; it now carries a sealed one. The one
real behavior delta is recorded as A.24: a structured `{"ok": false}` answer
behind an appended host note is a failure again.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-01 20:54:15 +00:00
Ouroboros
fadb7e33dc fix(f6-sync): battery round 1 — retarget pins, stamp fixtures, regen goldens
- module-handle declared sets updated to the folded reality (hold controls,
  body parser, quiz drain, cached ledger read)
- node-resolver tests retarget the campaign guard/attestation seams; frozen
  ToolEntry replaced via dataclasses.replace
- sweep-refresh test targets ouroboros.server_maintenance (campaign owner)
- process-signal tests drive the typed dispatcher; the upstream regex-fallback
  pin replaced by the D02 contract pin (prose forges no process facts) and the
  stale regex-fallback comments corrected
- routing-decision/find-child fixtures stamp _schema_version (ABI 7.0 readers
  quarantine unstamped rows); emit-and-wait patch retargets control_routing
- terminal-writers manifest gains _rewrite_execution_evidence (cursor/backfill)
- core catalog pins updated for send_links/escalate (owners: core_artifacts);
  node argv policy moved beside its PATH-prepend half in extension_child_catalog
  to keep extension_plugin_api at the 1000-line bound
- classification golden regenerated per the corpus recipe at 0f715831; three
  A.23 approved deltas record the F6-sync classification changes
2026-09-01 11:35:50 +00:00
Ouroboros
ccbb933a95 v7next F3.1 lane A: loop cutover to published ToolResult codes, typed extension/MCP dispatch, T1 outcome partition (D02/D04/D14/D15)
Re-derived on tip bytes; oracle v7_wip @ 9f691656 is the structural contract.

- loop_tool_execution (rows 157-164, 826-828): the retired result-text
  classifiers (_FAILURE_PREFIXES/_FAILURE_MARKERS/_EXIT_CODE_RE/_SIGNAL_RE and
  the elif ladder) are gone; status/is_error and the process/plan facts are
  READ from the dispatcher's published ToolResult (ABI-6(b): the unreachable
  _typed_or_adapted branch is NOT reproduced — the loop holds the typed result,
  text-only callers get the ONE adapter through compatibility wrappers).
  trace rows carry tool_result_status/code/meta alongside the legacy fields.
- extension_dispatch (D04 entry 5, rows 187/188) ADOPTED WHOLE from the
  reference WITH BYTE PROOF (tip == merge-base == v6.64.0 for this file, md5
  4e9ad3ba, so reference == tip+delta exactly): _dispatch_extension_tool_result
  / _dispatch_mcp_tool_result / _extension_dispatch_candidate produce native
  EXTENSION_*/MCP outcome facts; dispatch_extension_tool stays the text
  projection. registry_core retires the ToolRegistry dispatch methods and the
  hoisted candidate; the dead-extension unknown-name answer is typed
  EXTENSION_UNAVAILABLE (helper extracted per the function-size law).
- extension_process_runner (D14 entry 10): ExtensionProcessError gains
  failure_kind ('timeout' at the deadline kill) consumed by the typed
  dispatcher's EXTENSION_TIMEOUT arm.
- _outcome_tool_errors T1-partition (D15 entries 3-4, re-derived against the
  upstream status handling): every produced status is homed by its nearest
  analogue; tool_reported_failure and unavailable are policy-denial-partitioned
  (spec 1.15) while argument_error stays a real failure; retired codes'
  status names survive for stored traces; _UNPARTITIONED_BUCKETS makes the
  deliberate vlm_error hole explicit; untyped joins _OK_TOOL_STATUSES.
- reflection (row 166): the 4 CLAUDE_CODE markers retire (0 emitters, pinned
  by a repo scan test); _trace_call_errored reads the ok-status SSOT so
  untyped/ok_autocorrected successes stop triggering error reflections.
- plan-review typed control: _parse_plan_review_control moves to plan_render
  (its render-side home); publish_plan_review_projection/publish_rendered_wave
  emit the typed plan result whose meta the loop trusts (wave_control_state is
  the same projection the rendered footer reads); the plan handler pool-hops
  through contextvars.copy_context so the sidecar publication reaches the
  dispatching thread (reference delta, tip timeout design kept).
- tools/git facade re-exports _publish_git_error/_publish_review_blocked.
- Carried suites adapted to this tree (docstring/comment disclosures at each
  non-verbatim spot): test_tool_result{,_meta_boundaries,_t46}, test_registry_core,
  test_registry_guard_process, test_process_guard_codes, test_tool_catalog,
  test_tool_classification_differential + corpus + legacy fixture; the two
  loop_misc structured-failure tests re-home into the classification suite.
  Notable adaptations: facade keeps the broad historical surface (AST 'defines
  nothing' pin replaces the reference's exact-32 vars equality); guard patch
  points follow the _registry() call-time handle; the strict managed-update
  resolver is pinned with its corrupt-marker A4 channel; SCOPE_REVIEW_FLOOR
  rows dropped (ABI-5, Q10=A).
- ARCHITECTURE.md same-commit delta (extension_dispatch + loop rows); size
  ratchet regenerated officially (test_tool_result.py enters the band with
  rationale); ruff F clean; -m size_ratchet 5 passed.

(cherry picked from commit c12800b3cf320490f59f1d2532fa45ce62e9486a)
2026-08-31 18:05:03 +00:00