Commit graph

31 commits

Author SHA1 Message Date
Ouroboros
bbd0aa0be9 Improve autonomous review and preserve acceptance evidence
Implement the owner-approved P3 package: optional planning advice, substantive auto acceptance, separate host notices, durable source citations and promotion, HEAD-bound hermetic proof reuse, and focused documentation and CI coverage.

This version-neutral contributor checkpoint is pending exact-SHA final review. Publish browser coverage depends on the separately owned P1 Main-routing change; no P1 implementation is included.
2026-09-08 16:05:22 +03:00
Ouroboros
7b564c5cd8 fix(review): reconcile skipped delegated preflight without stranding new work 2026-09-06 18:25:55 +00:00
Ouroboros
24fbd48bc1 fix(review): prepare and recover authoritative review candidates 2026-09-06 12:31:01 +00:00
Anton Razzhigaev
b9ceed6ed7 merge: absorb upstream f3fbfdbb into the v7next campaign tree
Rolling-upstream sync #2 (8d13373b..f3fbfdbb, 101 commits, 180 files) under
the standing principle: upstream is the SEMANTIC truth, the campaign is the
STRUCTURAL truth. Every upstream semantic delta is re-seated in the campaign
leaf that owns its span; the campaign's split shape, typed organs and ratified
contracts are preserved.

Notable seats: registry post-exec organ (owner-state restore DELETED, replaced
by the settings tripwire) -> registry_guard_process/registry_core; the #447 H1
notes-trail-the-payload contract -> tool_result composition; the В23=A owner-home
read carve + egress masking -> tool_access_user_files/core_file_tools; the #468
shape-first reasoning pin -> llm_messages/llm_fallback/llm_attempt/
llm_openai_compatible; delivery-control provenance and the forced-control body
resolver -> loop_delivery/loop_forced_finalization; D4 export policy ->
shell_outputs; A5 literal-argv disclosure -> tools/shell.

Upstream twins of campaign organs do not live: tools/read_inspection.py folds
into registry_guard_process, tools/result_envelope.py into tools/tool_result,
delivery_protocol.py stays folded in loop_delivery.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-01 20:26:20 +00:00
Anton Razzhigaev
bcda779ce0 stage 3, governance docs: item-21 executable triggers + coaching posture (#447)
WS3-a — docs/CHECKLISTS.md item 21 (capability_regression) becomes an
EXECUTABLE contract (В7=A with В1's emphasis): a guard/filter/allowlist
change fires the item and the reviewer must see a STAGED POSITIVE test
proving the surviving capability path and a NAMED surviving flow — a guard
change without one is a capability-regression finding, not a safety
improvement. Owner acceptance of a narrowing = a GREEN plan review that
explicitly names it (В4=C/В22=B); disclosure alone is not acceptance. The
Critical threshold gains question 4 ("name the surviving positive path",
auto-YES for non-restriction findings) and extends to proposed REMEDIES —
a capability-deleting remedy meets its own concreteness bar or stays
advisory; repetition confers no additional authority. The four standing
disclosures move byte-verbatim to docs/CHECKLISTS_ARCHIVE.md (new,
reviewer-binding, append-only) with a pointer; ARCHITECTURE.md maps the new
file (docs-sync).

WS3-b — the REVIEW_BLOCKED retry coaching (review.py) opens the
proportionality channel (В8=A): rebuttal is legitimate for a factual error,
unsupported severity, or a disproportionate remedy that would remove a
working capability the accepted plan did not narrow — "change what you can
argue for, not what you can override"; a rebuttal never overrides
owner-chosen enforcement. In the Skill Review Checklist (В21): the
".env.example rename" advice is deleted and "permanent text-only" becomes
current owner posture; delegate_integration's refusal message now reads the
same truth (one author for one fact).

Test disposition: test_review_blocked_message_prefers_fix_over_rebuttal
pinned the exact restriction-posture text В8=A replaces — rewritten to pin
the new contract with strictly more asserts (proportionality channel +
non-override clause).

The archive extraction moved the standing 2026-08-15 delegate_start recipe
with its block; the recipe-schema doc-rot test now scans the archive (its
recipes are binding reviewer material and must stay schema-valid too).
2026-09-01 16:00:32 +00:00
Ouroboros
71e1f13fab v7next F3.3: comma-list remnant sweep gate + phase-5 residual cleanup
ABI-10 tail (base 1bd342b1): tests/test_comma_list_remnant_sweep.py is the
phase CI gate - retired-key mentions snapped dynamically from
RETIRED_COMMA_LIST_SETTING_KEYS with a count-anchored per-site allowlist,
comma-split parsing anchored in model/review modules, phase-5 plumbing
absence, and a retired-envs-are-ignored runtime pin.

Residual cleanup per the F3.1 lane D disclosure (LEDGER_CORRECTIONS 10):
- remove the dead phase-5 per-row route env plumbing
  (configured_review_routes, TRIAD/SCOPE_REVIEW_ROUTES_ENV, route_env_key):
  rows built from plain model lists are pinned api_chat; per-row delegated
  delivery is a structured OUROBOROS_REVIEWER_SLOTS fact only;
- remove the dead OUROBOROS_ADVISORY_REVIEW_ROUTE constant/prose (nothing
  reads the env; guidance now names the reviewer-slots advisory row);
- web: delete the 8 HOT-DEFERRED stale JSDoc alias lines in api_types.js
  and the _abi3_deferred_js_extras excuse set in test_gateway_parity (the
  browser mirror is exact again); the GATEWAY_CONTRACT_VERSION carrier
  switch stays DEFERRED to the release tact (version-carrier mechanics).

Behavior is byte-identical per configuration class: the removed reads fired
only for retired spellings. Docs same-commit: ARCHITECTURE review-substrate
paragraph, ADOPTION ABI-10 hook column, ledger F3.3 section.
2026-09-01 06:52:14 +00:00
Anton Razzhigaev
5a190512fa WIP: rename the preflight organ (Q1=A): preflight_review with a callable advisory_review alias
The advertised tool is now preflight_review; advisory_review stays a
CALLABLE compatibility alias (ToolEntry.alias_for — dispatched like any
entry, never advertised by schemas()/available_tools(), same parameter
schema so old calls keep their args). Contracts that withhold either
spelling silence both (D10-style mapping in _disabled_tools). Synced:
safety.py POLICY_SKIP rows (both spellings), tool_capabilities core and
untruncated sets, prompts/SYSTEM.md canonical tool list, docs/CHECKLISTS.md
invocation lines, update_merge_policy prompt, and every agent-facing
guidance string that teaches the tool name. Enforcement/severity
vocabularies, durable state paths (state/advisory_review.json), recorded
tool_name history, and the skip_advisory_review parameter are deliberately
untouched.
2026-08-29 22:46:59 +00:00
Anton Razzhigaev
c510de72d1 WIP: Claude-runtime retirement (Q4) + shared reviewer-row UI vocabulary (5A)
Engine: delete gateways/claude_code.py and the /api/claude-code/* endpoints,
launcher verify block, platform-layer SDK/CLI probes, CLAUDE_CODE_* dispatch
branches, CLAUDE_CODE_MODEL setting, claude-agent-sdk dependency; advisory
module loses the dead SDK budget/paid-suspect blocks (event renamed
advisory_suspect_result), error formatter drops retired probe fields.

Web: Claude-runtime status/repair panel and onboarding card removed; reviewer
rows get the Direct model / Configured subagent source picker with read-only
derived disclosure; advisory editor on the shared api_chat vocabulary (legacy
kind 'api' is parse-only); GET /api/reviewer-slots now round-trips
subagent_id references (resolved route as disclosure only).

Bench devtools: CLAUDE_CODE_MODEL dropped from forward lists, templates,
manifest active projection (read vocabulary keeps it), cybergym snapshot,
TB metadata roles; advisory bench rationale is comparability, not
impossibility.

Tests: suites repinned to the successor contracts (native episode caps,
in-process usage scope, credentials-based availability, endpoint round-trip);
size-ratchet manifest regenerated with band rationales.
2026-08-29 22:29:46 +00:00
Anton Razzhigaev
f97791afe0 WIP: commit-gate bypass test repinned to credentials-based availability 2026-08-29 21:56:26 +00:00
Anton Razzhigaev
483cb2fbd8 review subject: managed resolutions review the M0->S delta; pre-dispatch packet admission; honest advisory guidance
Lane L-review of the update-flow redesign (Delta-4, Q25=A, Q28=A, O1):

- ouroboros/tools/review_subject.py (new): ManagedReviewSubject — the
  resolution delta between the pinned mechanical-merge baseline M0 (from
  the durable update tx) and the candidate tree S, whose definition
  follows the surface (M1 contract): the COMMIT GATE serializes the real
  index write-tree - the exact tree the review-binding fingerprint pins
  and the commit writes - while the ADVISORY pre-review serializes the
  private-index worktree snapshot (work-in-progress by contract). The
  disclosure header carries M0/S identities, both real merge parents,
  conflict anchors and dual counters (full candidate paths vs reviewed
  resolution paths), and the M0-missing fallback renders the FULL
  candidate from the same pinned S (HEAD..S) behind a loud disclosure
  line. capture_review_diff(ctx, repo, unified) stays
  byte-identical to capture_staged_diff for every non-managed caller.
- Every managed review consumer reads the artifact: triad diff, changed
  list, -U0 rung and session task; scope diff, touched list (union with
  conflict anchors), -U0 and session task; advisory diff, changed-file
  context and the displayed counts. Review binding fingerprint, advisory
  freshness, preflight staged lists, doc-only classification and the
  scope snapshot key stay on the FULL candidate (I2).
- review_binary_context: managed binary rows render the mechanical-merge
  M0 baseline blob beside both real merge parents.
- review_admission.py (new) + parallel_review: BOTH gate packets (triad
  api pack, every scope row's pack) are assembled and fit-checked BEFORE
  any reviewer dispatch; a deterministic assembly block anywhere spends
  $0 everywhere (typed not_dispatched placeholders). the fit ladder
  moved whole (review.py stays under the module cap; review's
  _fit_triad_prompt remains the importable, patchable seam); _run_unified_review and
  run_scope_review are split into prepare/dispatch halves with unchanged
  verdict logic.
- Q28-A oversized outcomes: a panel whose agent-session rows alone
  satisfy the quorum drops (triad) or yields (scope) its fit-blocked api
  rows — recorded loudly — and proceeds; a panel that cannot reach quorum
  without them gets a typed zero-spend terminal, and the managed
  resolver's refusal carries the Settings -> Agents -> Review lanes
  guidance. The managed advisory oversize (>500k chars) becomes an
  audited non-blocking skip instead of the structurally impossible
  "split the commit" error; the exported constant and every non-managed
  path keep today's behavior byte-for-byte.
- Advisory honesty (O1): _next_step_guidance and the fresh-run message
  take the enforcement mode — blocking keeps the historical wording;
  advisory truthfully states that findings are recorded durably, the
  agent decides which to apply, and commit_reviewed is available.
- prompts/SYSTEM.md commit-review section (owner-sanctioned): the inlined
  managed artifact is authoritative; reviewers must not substitute their
  own `git diff --cached`.
- Superseded pins are updated in-test with rationale: the parallel-review
  wiring/ordering pins (Q25-A two-phase contract), the scope
  represent_binary pin (now driven by the managed subject), and the triad
  session-task builder (moved to review_subject behind a compat shim).
- tests/test_managed_review_subject.py (new): delta-only capture, -U0
  rung, non-managed byte-compat, M0-missing disclosure, binary M0 rows,
  dual counters, session-task inlining, $0 admission on deterministic
  no-fit, Q28-A drop/yield/terminal outcomes, oversize skip wording, and
  the enforcement-branch guidance texts.

Adversarial fix round (verified findings M1-M5, m1-m11; amended in place):

- M1: the reviewed gate subject S is bound to the tree that commits. The
  commit-gate consumers (triad capture, scope prepare, the -U0 fit rung)
  now serialize the REAL index (git write-tree - the same tree
  _fingerprint_staged_diff pins), never the live-worktree snapshot, so a
  pre-staged payload can no longer commit while reviewers read a restored
  worktree. The advisory pre-review keeps the worktree snapshot by
  contract (surface="advisory": advisory reviews work-in-progress;
  freshness handles staleness). Defense-in-depth: every gate subject
  records its S tree on the attempt and the commit gate asserts - a typed
  review_subject_binding_mismatch block - that the reviewed tree equals
  the binding fingerprint's tree_sha. Side benefit: gate-side S is stable
  and cheap across every consumer within one attempt, closing the
  disclosed multi-serialization TOCTOU consequence for the gate. A
  regression test reproduces the EVIL-index/GOOD-worktree divergence.
- M2: the M0-missing fallback populates name_status from the FULL
  candidate (HEAD..S), so touched packs, scope touched context and the
  advisory context/syntax preflight cover exactly what the full diff and
  the commit contain - never just the conflict anchors.
- M3: managed oversize terminals REPLACE the structurally impossible
  "split the commit" imperative instead of appending guidance under it:
  the triad fit ladder's block message, the scope ladder terminal remedy
  (ladder_terminal_cause managed=True) and the admission-side guidance
  all state that a managed resolution stages the whole two-parent tree
  and point at agent-route rows / larger-window models.
- M4: the SESSION variant of the M0-missing packet is honest - the header
  instructs "retrieve the FULL staged candidate diff yourself
  (git diff --cached)" when no body is rendered, and the counters line
  reports "reviewed resolution paths: n/a (M0 missing - full candidate
  under review)".
- M5: the managed advisory oversize skip's trailing sentence branches on
  enforcement (blocking: "still gate the commit"; advisory: findings are
  recorded rather than blocking).
- m1: split advice is managed-aware in the post-skip next-step guidance
  and in the 1.6M prompt-size gate (both pinned by tests).
- m2: enforcement wiring is covered end-to-end (_handle_review_status
  next_step and the pre-review completion message under both modes).
- m3: three stub-heavy tests replaced with real-flow drives - the healthy
  admission runs the REAL _prepare_unified_review and asserts the
  dispatched packet carries the actual staged hunk; the Q28-A
  drop/terminal tests drive the REAL fit ladder over the managed repo
  (tiny calibrated limit through the documented seam, no injected error
  strings); the advisory oversize test uses a genuinely >500k resolution
  delta with real routing and a persisted skip record.
- m4: the dead managed seam in _build_advisory_prompt is deleted (managed
  routing lives in _advisory_review_diff; no ADVISORY_ERROR string can
  become a review subject).
- m5/m6: docs/ARCHITECTURE.md states the two-phase admission contract;
  prompts/SYSTEM.md qualifies the inlined-artifact claim by M0
  availability (owner-sanctioned edit).
- m7: ADVISORY_REVIEW_CHOICE_GUIDANCE reworded enforcement-neutrally
  ("blocking where enforcement makes them binding").
- m8: only LIVE agent-session rows (still to be dispatched) count toward
  the Q28-A yield quorum; a session row that terminated at assembly is a
  dead seat.
- m9: an all-not_dispatched scope panel records the distinct typed
  scope_not_dispatched_assembly_block reason instead of a false
  "diversity was not achieved" quorum advisory.
- m10: a yielded api row's pre-yield refusal advisories are superseded in
  place; the never-produced budget_exceeded status left
  SCOPE_FIT_BLOCK_STATUSES (measurement documented in the comment).
- m11: the deliberate admission-time-only scope of the Q28-A yield is
  documented beside SCOPE_FIT_BLOCK_STATUSES (dispatch-time
  provider-tokenizer oversize stays fail-closed; the owner-level
  consistency question is escalated separately).

Panel fix round (5-seat review convergence; accepted R1-R10, amended in
place):

- R1: ARCHITECTURE/DEVELOPMENT truth - the structural tools tree gains
  review_subject.py and review_admission.py rows; the git/commit-review
  paragraph, the Review-stack rationale and DEVELOPMENT's "Triad slots
  review the staged diff" all carry the managed exception (managed
  resolution commits review the declared M0->S subject; the commit gate
  binds S to the index write-tree the fingerprint pins). The full
  handbook pass stays scheduled for synthesis S.4.
- R2: the M0-missing fallback body (every width, both surfaces) renders
  git diff HEAD..S from the PINNED subject tree, never a second --cached
  capture - on the advisory surface the old body omitted the unstaged
  work the counters/name-status described; on the gate the second
  capture weakened the binding to the pinned S. Regression tests combine
  M0-missing with index/worktree divergence on both surfaces.
- R3: authorized-managed fail-open closed - once the authority predicate
  says MANAGED, an unreadable or missing tx builds the LOUD M0-missing
  fallback subject (reason tx_unreadable/tx_missing), never a silent
  None that would review official code as resolver work; None remains
  only for genuinely non-managed contexts.
- R4: managed binary deletion is visible - the deletion renderer accepts
  the M0->S deletion (path present in M0, absent from stage-0, absent
  from the HEAD->index deletion set) and renders the row against
  M0/parent evidence; scope's deleted-path classifier takes the subject
  trees (branching only on subject presence - the non-managed call stays
  byte-identical) so an official-added, resolver-deleted binary can no
  longer be classified as text. Test reproduces the exact topology.
- R5: seat identity survives $0 paths - a prepared-but-withheld triad
  panel (deterministic scope assembly block) and Q28-dropped api rows
  leave typed $0 not_dispatched actor records (model, slot, slot_id,
  reason) in the durable raw results, matching the scope rows'
  placeholders; not_dispatched seats are excluded from the failed-actor
  diagnostics. Pinned by tests.
- R6: the blocking-mode fresh-critical advisory guidance drops the false
  "fix all and re-run, or audited skip" dichotomy: a fresh advisory
  already satisfies the freshness requirement, findings are recorded as
  debt the commit gate acknowledges, commit_reviewed is available, and
  the audited skip bypasses only the freshness/debt checks. Both
  enforcement branches pinned (superseded pins updated with rationale).
- R7: _handle_advisory_pre_review no longer builds a second
  advisory-surface subject for counters_line() - the counters thread out
  of the ONE subject _advisory_review_diff builds (test pins exactly one
  subject construction per pre-review).
- R8: a failed full-candidate path count renders "n/a (count
  unavailable)" instead of a fake 0.
- R9: review_status renders a friendly message for
  review_subject_binding_mismatch.
- R10: the M0-missing header lead is branched per mode (it no longer
  claims the official delta "is not re-rendered" above a fallback body
  that includes it); the managed packet header disclosed the accepted
  Delta-2 residual (M0 is a pinned forensic baseline from the update tx
  and is not re-verified at review time - disclosure only, no re-merge
  check); this commit message states the M1 surface-following S contract
  and the real _fit_triad_prompt seam name.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

[synthesis] conflict resolution: ouroboros/tools/git.py — function-boundary
collision between lane-cycles' extracted _advisory_and_tests_gate and
lane-review's inserted _subject_binding_mismatch_outcome: kept BOTH functions
(subject binding check first, then the advisory/tests gate helper); the binding
call auto-merged into the cycles version of _run_reviewed_stage_cycle after
post-fingerprint revalidation and before block classification, preserving both
lanes' gate-order contracts (free-cycle -> advisory -> paid stamp -> dispatch
-> binding check). prompts/SYSTEM.md — merged paragraphs: kept lane-cycles'
paid-cycles/identical-refusal wording (supersedes attempt_cap_reached) and
appended lane-review's M0 resolution-delta packet contract.
tests/test_commit_gate.py — kept cycles' _advisory_and_tests_gate assertions
plus review's prepare/dispatch-phase assertions (Q25-A).
ouroboros/review_evidence.py — union of both lanes' reason_map entries.
2026-08-21 04:46:25 +00:00
Anton Razzhigaev
386e9417b1 Max Review Cycles counts PAID cycles on the commit gate and skill review; identical bytes are never re-reviewed for pay
Owner decisions Q12/Q16-A/Q17-A/Q22-A/Q23-A (update-flow redesign sprint,
lane L-cycles, deltas 5-6). The one shared knob (OUROBOROS_REVIEW_MAX_CYCLES)
now means the same thing on every review gate: paid cycles per task.

Commit gate (commit_gate.py, git.py, review_state.py):
- Typed block-row classification at record time: verdict-blocks (triad/scope
  reviewer FINDINGS) vs infra-blocks (fit fixed_overflow, scope sub_floor,
  revalidation, transport/no-quorum). New durable CommitAttemptRecord fields
  block_class, rebuttal_sha256, paid, review_contract_fingerprint,
  root_task_id (merge- and roundtrip-safe; legacy rows classify by reason).
- A byte-identical resubmission of a verdict-blocked staged diff is refused
  FREE from the FIRST block (identical_diff_refused), before the advisory
  freshness gate and before any paid dispatch, quoting the recorded verdict;
  the refusal is cross-task by fingerprint (anti-laundering). Infra-blocks,
  infra/expired failures and preflight facts neither build nor reset the
  streak; a post-review failure or a success ends it.
- Rebuttals are identified by CONTENT (sha256): a hash new to the streak buys
  exactly one paid re-review; a repeated hash is refused free. A rebuttal is
  "spent" only when it bought a dispatched, verdict-answered wave — one
  refused undispatched (e.g. by the ceiling) stays fresh.
- The knob bounds PAID triad+scope cycles per ROOT task (the task tree shares
  one ceiling; a manual session is its own task; a follow-up task is a fresh
  root — origin_root_task_id is deliberately not honored, the cross-task
  identical-fingerprint refusal stays the anti-laundering backstop). Paid is
  recorded at dispatch on the attempt ledger (plan-review precedent; P7 - no
  counter file) and the ceiling counts MONEY: every dispatched wave counts
  whatever its terminal; only undispatched attempts stay outside the count.
  Exhaustion is a free typed refusal plus
  emit_review_cycles_exhausted(surface="commit_gate"); unlimited disables it.
- Free refusal/replay honors a review-contract fingerprint (triad roster and
  routes, scope rows, enforcement, shipped prompt contract); a changed
  contract lapses the streak and the refusal never quotes across the change.
- Under ADVISORY enforcement neither state hard-blocks: the commit proceeds
  with a loud typed disclosure (commit_review_free_replay event + result
  warning; the wording distinguishes a replayed verdict from a
  ceiling-exhausted commit that has no verdict to reuse) and no further paid
  dispatch. Managed resolvers keep MERGE_HEAD repaired on the free refusal
  path.
- _run_reviewed_stage_cycle order: fingerprint (hoisted only because the free
  gate needs it) -> free cycle gate -> advisory/tests gate -> binding
  precondition (kept at its original position) -> paid dispatch; the
  advisory+tests section moved whole to _advisory_and_tests_gate at the
  function-size gate.

Skill review (skill_review.py, new skill_review_cycles.py,
skill_review_history.py, skill_review_runner.py):
- review_max_cycles() wired in for the first time: paid panel dispatches are
  counted per ceiling key - the root task for task-driven review groups
  (across every skill that task reviews; follow-ups start fresh), and, for
  the manual lane, the CURRENT content snapshot (revised content restarts
  the manual count; marketplace-install rows share the same per-snapshot
  scoping). One chunked wave = ONE cycle; counts derive from the append-only
  review history; exhaustion is a typed refusal plus
  emit_review_cycles_exhausted(surface="skill_review").
- Q17-A free replay: an identical (group, content_hash, panel-contract
  fingerprint) snapshot with a recorded SUBSTANTIVE verdict (clean/warnings/
  blockers) replays at USD 0, quoting it - but only while the persisted
  review state still covers that exact snapshot+verdict; a diverged state
  falls through to a PAID rerun that re-persists (never a live-lock).
  pending/interrupted/timeout/failed/cancelled are infra facts that never
  replay; a rebuttal recorded on such a row stays fresh and buys the paid
  rerun it is owed; a content-new rebuttal buys exactly one paid rerun.
- The panel-contract fingerprint pins the roster, the required items, the
  shipped prompt/aggregation contract text and the skill's RESOLVED review
  profile (computed once and threaded through aggregation).
- Terminal history rows carry the paid fact, the panel-contract fingerprint,
  the rebuttal hash and replay provenance.
- Module splits at the size gates (update_candidate.py precedent): the
  accepted-rebuttal ledger and the review-wave budget admission moved whole
  to skill_review_cycles.py with same-name re-exports; review_skill's persist
  tail extracted to _persist_reviewed_outcome and the cycles gate to
  _skill_cycles_gate.

Disclosed residual: the ceiling checks are read-at-gate-time with no
reservation (TOCTOU) - concurrent dispatches sharing one ceiling key can
overshoot the cap by the concurrency width.

Pinned wording synced in this same commit (CHECKLISTS row 10):
review_cycles.py per-gate docstring (now four gates), commit_gate.py header
and module docstring, the advisory _identical_diff_cap_note, commit_reviewed
rebuttal schema texts, prompts/SYSTEM.md commit-gate wording,
docs/DEVELOPMENT.md cycles section plus the new anti-pattern instruction
(never pay for byte-identical review material; the limit counts PAID cycles;
exhaustion is a typed event), docs/ARCHITECTURE.md (review_cycles and
commit_gate module rows, skill review module map rows, evolution cleanup
paragraph, settings table), ouroboros/config.py settings comment, web
settings_ui.js copy, and review_evidence.py block-reason guidance - all
stating the per-ROOT-task scope. Pinned tests rewritten as contract tests
(test_commit_gate.py, test_review_cycles.py, plus the structural pins in
test_scope_review.py and test_advisory_delegated_route.py) and a new
tests/test_review_cycles_gates.py suite added, including behavior tests that
drive the real recorder to the ledger, the real MERGE_HEAD repair against a
temp git dir, and the real review_skill pipeline through replay, refusal and
paid dispatch.

Retired API: check_blocked_attempt_cap / blocked_attempt_fingerprint_cap
(pay-for-identical-re-review-up-to-a-cap semantics). Legacy
attempt_cap_reached ledger rows remain recognized as refusal records.

Fix round (5-seat review panel convergence; accepted findings F1-F8, F10,
fable P3-2/P3-3):
- F1 (CRITICAL) review_state.py: the 50-row attempt-history trim is now
  authority-preserving - it evicts only non-authoritative noise (free
  refusals, preflight facts, unpaid infra rows), oldest first, and never
  paid rows (the money ledger the per-root-task ceiling derives from),
  review-verdict anchors (the evidence an identical-diff refusal quotes) or
  in-flight rows. The real bound (delta-review follow-up): the ACCOUNTING
  FACTS of preserved rows are immortal, their heavy forensic payloads are
  not - a preserved row that falls outside the newest-50 window is compacted
  (raw_stripped=True): triad_raw_results/scope_raw_result dropped
  (block_class materialized first, so a legacy scope verdict keeps its
  refusal-anchor authority after the classifying evidence is gone),
  block_details bounded to the 600 chars the refusal quote renders,
  commit_message bounded to 300; paid/root/task ids, fingerprints, rebuttal
  hash, timestamps and the quoted critical_findings all survive. The
  serialized ledger therefore grows ~O(preserved rows x small record) - per
  root task bounded by the money ceiling - instead of by full reviewer raw
  output per reviewed commit forever. Regressions: 60+ free refusals after
  the cap change neither the paid count nor the quoted verdict, and a
  120-row heavy flood keeps full raw payloads only inside the newest-50
  window (every over-window row serializes under 4KB).
- F1-adjacent recorder bug (exposed by the eviction test): a NEW attempt
  recorded with no in-flight marker inherited the PREVIOUS terminal
  attempt's fields - paid=True, block_class, rebuttal/contract/fingerprints
  - so once any paid row existed, every free refusal was silently counted as
  a paid cycle. The terminal-latest branch of _record_commit_attempt now
  nulls the inherited row for a fresh attempt number, mirroring the
  contract the reviewing branch already enforced.
- F2 (MAJOR) the paid stamp moved from gate passage to a WRITE-AHEAD record
  at the first PHYSICAL reviewer dispatch: new module
  ouroboros/review_dispatch.py (which also absorbs the slot-identity mint
  moved whole from review_substrate.py at the module-size gate, same-name
  re-exports kept) provides the idempotent thread-safe ReviewPaidStamp;
  the commit gate installs it on ctx (_install_paid_dispatch_stamp) and the
  shared transport entry run_review_request invokes it immediately before
  the first reviewer call on EITHER side - triad and scope dispatch in
  parallel, so any side dispatching makes the cycle paid. Assembly-only
  exits (triad fit ladder, scope pack signals) stay $0 and outside the
  ceiling; a crash after dispatch keeps the durable paid fact. The three
  docstring enumerations ("assembly failures stay outside the count") are
  now literally true. Design note: this seam is where the L-review lane's
  two-phase admission slots in at synthesis.
- F3 (MAJOR) skill review shares ONE durable write-ahead dispatch marker
  (state/skills/<name>/review_dispatch.json) between the lifecycle runner
  and direct review_skill callers, written at physical panel dispatch with
  paid=True + contract fingerprint + rebuttal hash + wave id. Terminal
  history rows merge it idempotently (a lifecycle timeout with no result
  still lands the paid facts) and clear it; an orphaned predecessor is
  flushed into the history as an infra terminal; the derived count includes
  an unmerged marker, so a crashed or swallowed wave never un-counts spent
  money. The quorum-failure append now receives the paid facts (the direct
  marketplace-install path no longer lands unpaid), and history append
  failures surface as loud typed skill_review_history_append_failed events
  instead of a silent pass. Four-outcome behavior tests (substantive
  verdict / quorum-failed / transport-failed / crash-after-dispatch) each
  pin exactly one paid unit on a real jsonl.
- F4 (wording + scoping, no semantics flip): spent-rebuttal memory on the
  skill side is scoped to the CURRENT panel-contract fingerprint (a
  contract lapse clears it along with replay); the contradictory wordings
  (skill_review_cycles.py header, docs/DEVELOPMENT.md) now state the exact
  rule - refusal-streak eligibility (substantive verdicts only) is distinct
  from money accounting (every dispatched wave counts); a rebuttal is spent
  by the substantive verdict it bought; a dispatched infra terminal
  consumes money but not the rebuttal.
- F5 advisory-honesty sync: prompts/SYSTEM.md, web settings copy, the
  advisory-review schema note (claude_advisory_review.py) each carry the
  one enforcement caveat - under blocking the commit is refused for free;
  under advisory it proceeds with a loud durable disclosure and no new
  review spend.
- F6 the blocked-classification comment in git.py no longer claims a
  dispatched infra block costs nothing - money is stamped at dispatch.
- F7 the skill_review tool schema documents the replay key (group,
  content-hash, contract fingerprint), the paid-cycle ceiling, exhaustion
  leaving the skill pending (never executable), and the rebuttal as content
  whose sha256 buys exactly one paid wave when new to the streak.
- F8 auditability: review_evidence attempt projections expose block_class,
  paid, rebuttal_sha256, review_contract_fingerprint and root_task_id; the
  skill review-history UI detail includes the accounting facts inside the
  existing free-form markdown detail string - the
  SkillReviewHistoryDetailResponse contract is typed to exactly four
  fields, so no api_types.js change and no version bump.
- F10 dead _POST_REVIEW_FAILURE_PHASES deleted (the streak walker's generic
  terminal break already covers post-review failure phases).
- fable P3-2: the advisory identical-replay branch is marked as deliberate
  defense-in-depth (verdict rows are only minted under blocking and the
  contract fingerprint includes enforcement, so today it degrades to an
  ordinary paid review).
- fable P3-3: end-to-end tests drive _run_reviewed_stage_cycle through both
  advisory replay reasons, pinning the distinct progress notes and the
  disclosure landing in the commit result formatting.

Residual disclosures:
- Bench settings_base templates set blocking enforcement and do not pin
  OUROBOROS_REVIEW_MAX_CYCLES; the trace audit and the bench-cap option are
  escalated to the owner before landing (per the sprint plan's I1
  verification duty). Only the SWE-Pro evolution lane executes the commit
  gate, and the E1v2 METHODOLOGY already holds advisory.
- Successful reviewed commits also consume paid cycles, so a root task's
  third reviewed commit refuses under blocking at the default cap of 2
  (settled Q16-A consequence).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 04:46:25 +00:00
Ouroboros
7de26338b7 test: stabilize cross-platform release gates
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-21 04:13:58 +03:00
Anton Razzhigaev
d6210b1a05 plan review becomes a domain-neutral spec gate
The pre-implementation gate reviewed an unbounded prose plan with a ~100k-token
governance pack, planning-scout subagents and an Atlas, and every rewritten plan
minted a fresh fingerprint that bought a whole new paid wave from cold reviewers.
Convergence was structurally impossible: the reviewer was asked to author a
competing plan every wave, REVISE_PLAN could not be dispositioned, and nothing
capped the cycles.

plan_task now reviews an INTENTION — the same organ whether the work is code,
research, a deliverable or an action in the world:

- a typed domain-neutral SPEC (goal, in_scope, non_goals, acceptance_claims,
  invariants, decisions with rejected alternatives, deferred, affected_resources,
  evidence) with host-minted ids that are the only valid `breaks` targets;
- ONE structural fact tiers the governance pack: `constitutional` iff a declared
  target resolves under the system repo (never prose, never a plan-kind taxonomy);
- agent-declared evidence, bounded, with EVERY absence named, the runtime data
  plane denied outright, and the exploration log redacted through the same SSOT
  task acceptance uses;
- typed findings (blocking with a `breaks` id | note | need_evidence) with the
  HOST computing the aggregate through adaptive_quorum — no reviewer emits GREEN
  as authority, none writes a competing plan;
- ONE owner setting OUROBOROS_REVIEW_MAX_CYCLES (default 2, unlimited available)
  bounds paid cycles for plan review, task acceptance (passes = cycles - 1) and
  the commit gate's identical-diff attempt cap; an identical envelope replays for
  free and a DEGRADED wave costs nothing;
- under blocking, an open plan holds implementation and a spent cap escalates
  with a typed review_cycles_exhausted reason and an honest blocked terminal;
  under advisory the agent may proceed with the wave open and a loud disclosure.

Deleted: planning scouts, the plan Atlas, plan_class, context_level, the
governance mega-pack, the generative reviewer stance, the hidden 32-wave limit
and the api_chat-only pin. Net effect on the tree is negative.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-17 11:37:32 +00:00
Ouroboros
fe11e54fc1 Disclose unavailable advisory pre-review routes 2026-08-14 04:23:01 +03:00
Anton Razzhigaev
bd9aaf3d40 fix: close evolution authority race windows
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-04 00:55:14 +03:00
Anton Razzhigaev
cacafa7d6c fix: bind evolution to isolated campaign state
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-08-03 22:04:42 +03:00
ndrew1337
45d77ff785
tests+ci: green & de-flake the pre-push suite (P3/P6/xdist-safety) + parallelize CI ~9x (gate stays serial) (#55)
* test(improvement_backlog): stub semantic-dedup LLM in groom tests (P6)

_seed_many() seeds items via append_backlog_items(), whose C9.2 semantic-redirect
pre-pass calls semantic_dedup.find_semantic_duplicate_id() once per fingerprint-MISS
with candidates — a real light-model NETWORK call. The seeding runs BEFORE
_patch_groom_llm installs its mock, and that mock only covers chat_observed, not the
detector's own client path. With no API key the call retry-storms for minutes before
failing open to None, making test_groom_backlog_rejects_invented_items ~129s alone
(~40% of the whole suite) and non-deterministic.

Add a module-level autouse fixture that stubs find_semantic_duplicate_id to its own
fail-open default (None = no duplicate — exactly what the doomed call eventually
returns for the distinct seeded items), so the module is network-free and
deterministic. No production change; the dedup contract stays covered by
test_semantic_dedup_v6370.

Effect: test_groom_backlog_rejects_invented_items 129.12s -> 0.02s; whole file
~165s -> 2.37s, all 14 tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(core): validate regex up front in _code_search so invalid-regex contract holds for both backends (P3)

_code_search ran the ripgrep path first and returned its formatted result before
ever reaching the Python-fallback re.compile guard. ripgrep accepts some malformed
patterns permissively — an unterminated '[' yields "no matches" instead of erroring
— so an invalid regex like "[invalid" silently returned no-match on the rg path while
only the fallback (rg absent/failed) emitted "⚠️ SEARCH_ERROR: invalid regex". This
made test_code_search_invalid_regex fail whenever ripgrep is present (i.e. always, in
CI and the pre-push preflight), so the suite exited non-zero on every candidate diff
regardless of the change under test.

Compile the regex once up front (regex queries only; literal queries need no check)
and return SEARCH_ERROR on re.error before dispatching to either backend, so both the
rg path and the Python fallback share the same invalid-regex contract.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: guard the conftest repo-root pollution sweep to the xdist controller (xdist-safety)

Under pytest-xdist, pytest_sessionfinish fires on the controller AND every worker against
the SHARED repo root, so the mock-pollution sweep (shutil.rmtree of leaked <MagicMock>
paths + session.exitstatus=1) had workers racing the same rmtree and each independently
failing the run — a non-deterministic, failed-shaped result. Guard the repo-root sweep +
exitstatus mutation behind `if not hasattr(session.config, "workerinput")` so it runs only
on the controller (the single authority); the per-process _PYTEST_DATA_DIR cleanup stays
outside the guard and runs on every process. Serial runs are unaffected (the guard is
always True without -n).

Adds tests/test_conftest_xdist_guard.py (sweep runs on the controller, skipped on a worker).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci: parallelize the test suite with pytest-xdist (~9x faster), keep the gate serial

The full-suite CI jobs (quick-test, full-test) now run a PARALLEL pass plus a short SERIAL
pass. Empirically ~270s serial -> ~30s parallel (~9x). The per-commit preflight GATE stays
serial by design (a flaky parallel fail-closed gate manufactures non-deterministic
TESTS_FAILED indistinguishable from a real immune rejection).

- requirements: pytest-xdist + pytest-timeout (a hang-guard).
- pyproject: register the `serial` marker (addopts unchanged).
- ci.yml (quick-test, full-test): two-step. Parallel
  `-m "not serial and <default lane exclusions>" -n auto --dist loadscope
  --max-worker-restart=0 --timeout=300 --timeout-method=thread`, then serial
  `-m "serial and <exclusions>"`. A command-line -m REPLACES the pyproject addopts markexpr,
  so the default lane exclusions are repeated and ANDed with the serial split — the union
  exactly reproduces the old default suite (4594 parallel + 122 serial = 4716, disjoint).
- conftest: a tryfirst pytest_collection_modifyitems hook marks the real-process/port files
  (workspace_executor[+cleanup], process_custody, kill_process_tree_orphans, zombie_prevention,
  worker_crash_retry, process_resource_leaks, restart_reconnect, preflight_runner,
  services_tool_v2) `serial`; plus an autouse fixture isolating workspace_executor._SERVICES/
  _FOREGROUND between tests (a latent ordering bug -n redistribution exposes).

Test-isolation fixes surfaced by running the suite under -n:
- test_task_constraint_tools: 3 bare `sys.modules["...claude_code"] = mock` (no restore ->
  polluted the worker's sys.modules -> later SDK-dependent tests failed) -> monkeypatch.setitem.
- test_workspace_executor: poll until the spawned process's command-sha is readable before
  registering (the PID-reuse safety check compared the sha recorded at registration vs
  recomputed at kill; right after fork+exec the command line is unreadable -> shas diverge ->
  kill silently skipped -> flaky), + 5->15s kill-confirmation deadlines.

docs/DEVELOPMENT.md + docs/CHECKLISTS.md: guidance so future tests are parallel-safe or
marked `serial`. Reviewed by adversarial subagents (ship); verified parallel 8/8 + serial 3/3
green, partition exact + disjoint, gate byte-identical.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* tests+ci: parallel-safety follow-ups (services-global isolation, serial-lane guard, monkeypatch conversions)

Three non-blocking follow-ups from the parallelization review, hardening the
xdist split shipped earlier in this PR:

F1 (tests/conftest.py): extend the autouse _isolate_workspace_executor_globals
fixture to ALSO snapshot/clear/restore ouroboros.tools.services._SERVICES (the
legacy services registry, a separate module-global from workspace_executor's)
under its plain threading.Lock _LOCK. Closes the legacy-services-path
global-leak class generally; raw dict ops only under the lock (no re-entrant
deadlock on the plain Lock), registry-only (never reaps the live Popen
handles), each module lazy-imported under its own guard.

F2 (.github/workflows/ci.yml): add a "Guard non-empty serial marker lane" step
to marker-guards. `pytest --collect-only -m serial` exits 5 on an emptied
_SERIAL_TEST_FILES; set -euo pipefail + tee surfaces it, and a positive anchor
grep on test_workspace_executor.py is the working assertion (the existing
browser-guard's `! grep "no tests collected"` is a dead no-op under -q).

F3 (test_skill_loader / test_iteration2_fixes / test_marketplace_clawhub /
test_devtools_benchmarks): convert remaining bare os.environ / sys.modules
mutations (no save-restore) to auto-reverting monkeypatch.setenv/delenv/setitem,
so the parallel suite has no cross-test env/module leaks. tests/_shared.py's
intentional process-wide SDK mock left untouched.

Verified: parallel lane 4574 pass (-n auto --dist loadscope), serial lane 140
pass, the 4 converted files 159 pass. Two subagent reviews: no regressions.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: prune 26 redundant/tautological/obsolete tests (immune coverage retained)

Three-pass audit (finder+skeptic -> 31 empirical verifiers with rename+mutation
experiments -> delete-and-run dry-run) of the 275-file / 4067-test suite found 26
tests whose value is fully backstopped by a stronger survivor or is near-zero:

- tautological / zero-production pins (assert hasattr/callable; inline-simulated
  branches; closed-loop regex; TypedDict dict-literal; CPython-only checks);
- proven duplicates with strict-superset twins (test_consolidator, test_commit_gate,
  test_cache_optimization, test_review_intent_split, TestGrepRegexHint, ...);
- registry-registration one-liners, all backstopped by test_smoke::test_tool_set_matches
  (set-equality vs EXPECTED_TOOLS -- mutation-proven to FAIL on a lost registration);
- dead-feature tests (retired skill_migrations module) + obsolete version regressions
  (v636 import now unconditional; README "(N tests)" convention abandoned).

Removes ~47 test functions across 20 files (incl. the whole test_shell_regex_hint.py,
mirrored 1:1 in test_shell_run_shell.py::TestGrepRegexHint, and the whole
TestGoalScopePrecedence class). Also drops 2 now-unused imports + 1 orphaned helper.

KEPT (not redundant): test_smoke::test_git_commit_with_tests_exists -- sole guard of the
post-commit gate seam (a rename experiment proved nothing else catches its loss).

Verified: ruff F clean; collection exit 0 (4703 collected, no emptied class); full
parallel + serial suites green; two subagent reviews confirm exact set + no broken refs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Andrew <andgri200@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 02:56:34 +03:00
Ouroboros
f328b72821 feat(evolution): window-fit context, guaranteed-fit scope assembly, plan_task degraded fallback, identical-diff review-attempt cap, solve-capability ledger, lifecycle extraction with deterministic cycle cleanup (v6.30.0)
Phase 1 Block 5 of the combined release plan (A+B+C+D) plus the owner-directed
guaranteed-fit scope assembly:

- A window-fit context: evolution reads ARCHITECTURE as an on-demand nav map;
  emergency compaction threshold derives from the active model's real window
  (profile ceiling for 1M-window models); small-window remote models get
  routine compaction in max mode.
- B plan_task capacity-class failures degrade to one labeled inline critique
  pass (stalled/infra stay fail-closed); window-aware scope input cap with
  Claude tokenizer calibration and window-scaled reserves (gigachat 131K maps
  to a workable input limit; >=1M reserves byte-identical); provider oversize
  400 downgrades to the disclosed non-blocking budget_exceeded skip
  (owner-approved safety net); identical-diff commit-review attempt cap
  (3 genuine verdicts, diff-scoped, preflights stamped phase=preflight,
  review_rebuttal lifts, fail-open cost guard).
- Guaranteed-fit scope assembly (owner directive: scope review must actually
  run): deterministic ladder full atlas -> compact atlas -> atlas required
  files to explicit budget_omitted manifest entries -> largest touched files
  to diff-only with disclosed notes; the >=1M blocking reviewer's irreducible
  overflow fails CLOSED (fixed_overflow); known sub-floor reviewers are
  advisory-only (P3 floor) with the documented skip terminal.
- C solve-capability ledger: kind=cycle_outcome rows joined at absorb/abandon
  resolution; digest with explicit omission notes feeds the promotion chooser,
  upgraded to the main-model slot with a small-targeted-solve-helper bias.
- D campaign/transaction lifecycle extracted from supervisor/queue.py into
  supervisor/evolution_lifecycle.py (protected git_ops untouched);
  deterministic no_op/abandoned cycle cleanup (stash + evolution-leftover
  branch + recorded skip gates); boot dangling-tx reconcile; hard-kill
  owner-mailbox cleanup.

Note: ouroboros/tools/git.py remains grandfathered-oversized (review.py:41);
the new cap logic lives in commit_gate.py.

Gauntlet: adversarial x2 (GPT+Gemini pass 1; Fable pass 2 SAFE TO COMMIT) +
triad+scope rounds 1-6 with the guaranteed-fit assembly reviewing itself from
round 3 (scope responded every round since); all verified criticals fixed,
oscillating/owner-decided findings rejected with recorded evidence.
2026-06-13 01:01:40 +03:00
Ouroboros
ae53779ccd release: Ouroboros v6.26.0-rc.1 — systemic hardening: immune system, custody, memory integrity, native multimodal chat
Owner-approved big release (plan: большой_релиз_v6.26.0). Workstreams:

- WS-A lint gate: ruff F-rules step in CI quick-test; fixed the real F823
  NameError class (supervisor/events.py utc_now_iso shadowing); make lint/health.
- WS-B memory integrity: atomic dialogue_blocks with corrupt-quarantine
  (.corrupt-<ts>.bak + memory_store_corrupt event); honest scratchpad journal
  (block_append_failed, corrupt storage renders as corruption not "(empty)");
  full Pattern Register window (16K cap vs 3K cut); merge-aware scratchpad
  consolidation under the sidecar lock; backlog writes through locked helpers;
  chat omission notes ("[N older unconsolidated messages omitted]").
- WS-C provider SSOT: provider registry (prefixes/credentials/resolution) in
  provider_models.py replacing 5 duplicated knowledge sites; credential-aware
  consolidation model; pricing empty-fetch retry; direct-route correctness
  (o-series max_completion_tokens, reasoning_effort, Anthropic error bodies,
  per-request timeouts on cached clients); deep/plan review budgets reserve
  output headroom inside the 1M window (min(SSOT, window − output − margin)).
- WS-D races/state: utils.update_json_locked (locked RMW, loud TimeoutError);
  task_results merge under per-file lock (cancel-latch safe); update_state
  migration for owner binding, evolution counters, budget_line, post-task
  activation; queue-lock coverage (enqueue, timeouts, worker health, task_done,
  snapshot); visible supervisor death (supervisor_error + owner notice + ghost
  consciousness/chat-agent cleanup on re-init); chat.jsonl rotation via
  os.replace under the append lock; unbounded _outbox removed; host-service
  port probe + actual-port panic sweep; skill lifecycle lane deadline
  (OUROBOROS_SKILL_LIFECYCLE_TIMEOUT_SEC, dedupe-leak-safe); consilium
  force-plan structured flag (plan_review_aggregate from the FULL result);
  atomic writers (update-intent, repo-manifest, metrics cache, post-task pair);
  review_state.save_state raises on lock timeout (advisory ledger honesty);
  apply_pending_request no-ops while a campaign is active;
  api_update_apply kills workers only AFTER update validation, with respawn
  on aborted checkout and "interrupted" terminal status.
- WS-E immune hardening: triad anti-refusal coverage contract (empty array
  needs NO_FINDINGS sentinel or bare-[] body; refusal prose with [] is
  parse_failure and never enters quorum); BIBLE P3 "Owner-chosen enforcement,
  loud advisory" bound + CHECKLISTS sync; every advisory pass-through of a
  blocking signal writes review_advisory_override + persistent
  advisory_overrides (surfaced by review_status); ToolEntry.mutates_worktree
  with dispatcher-level before/after worktree diff invalidation (covers error
  paths; read-only runs no longer invalidate; redundant manual calls removed);
  scope review fails closed without its checklist; advisory final write is a
  locked re-read merge; synthesis severity defaults to critical with WARNING
  fallbacks; find removed from SAFE_SHELL_COMMANDS; anti-thrashing state
  survives advisory criticals; preflight env scrubs secret-class variables;
  .git/index.lock age gate in worker startup checks.
- WS-F security: file-browser symlink containment on the RESOLVED path across
  all endpoints (out-of-root symlink targets are listed but inert; tests
  rewritten to the new contract deliberately); HMAC-signed session cookies
  (server-side persisted key, 30-day TTL, Secure on TLS) replacing the
  permanent password-derived cookie; password-class settings mask to a
  constant placeholder; conservative SSRF guard for the MAIN agent (link-local
  /cloud-metadata only, LAN stays reachable, per-request route re-validation);
  ClawHub zip-slip hardening (":"/backslash segments rejected + post-join
  containment) and lazy no-proxy OuroborosHub opener; single-execution
  signature dispatch for extension handlers (no TypeError re-run after side
  effects); onboarding postMessage origin checks; SHA256-pinned
  python-standalone download with pipefail; payload-resident dependency
  fingerprints only corroborate durable deps.json; skill payload re-hash
  immediately before spawn (TOCTOU narrowing).
- WS-G process custody: ouroboros/process_custody.py — spawn_supervised
  chokepoint + durable data/state/process_ledger.jsonl (pid, pgid,
  fingerprint{start_time, cmd_sha256}, purpose, scope task|session|daemon,
  owner_task, session_id); platform_layer.process_start_time primitive;
  startup + periodic reaper killing ONLY strict-fingerprint matches from dead
  generations/tasks (never by command-line class); migrations: services
  (the orphan hole), workspace executor + local model + extension companions
  (ledger write-through), worker_pids (write-through; legacy path retained);
  parent lifelines (ppid watchdog, group-suicide only as group leader) in
  worker_main, extension runner, Claude readonly child; conformance test
  pinning the Popen allowlist; ARCHITECTURE/DEVELOPMENT/CHECKLISTS entries.
- WS-H native multimodal chat: supports_vision capability map (static
  prefixes + OpenRouter /models input_modalities overlay); web chat uploads
  ride the WS frame as structured attachments (additive ChatInbound field)
  and image uploads become NATIVE image blocks via the existing Path B;
  browser screenshots inject natively for vision models via the multipart
  user-merge (tool result stays a string; file persisted under
  data/uploads/screenshots for re-view); K=3 newest-image eviction with
  caption placeholders carrying the vlm_query re-view path; image-aware token
  estimates (fixed ~1.1K-token equivalent instead of base64 length, fixing
  permanent emergency-compaction wedges); compaction renders images as
  captions (no base64 into the summarizer); GigaChat/local lanes emit explicit
  "[image omitted: model has no vision]"; internal _caption/_source_path
  metadata stripped from provider payloads.
- WS-I housekeeping (partial): SETTLED_STATUSES SSOT (+ cycle-safe mirror pin);
  owner_inject.py renamed to owner_mailbox.py; version-neutral envelope
  wording; files.py import-block cleanup. Remaining WS-I/WS-J/WS-K items are
  deferred with the owner's context-budget priority on review+release.

Review notes: triad+scope ran via scripts/run_external_review.py on the core
pack across 3 rounds to convergence (scope responded=PASS each round; round-2
criticals fixed: update_apply kill-order + respawn, lifecycle dedupe leak on
lane timeout, OUROBOROS_MAX_ROUNDS hot-reload + docs, toggle_evolution
NameError, async Anthropic timeout forwarding, supports_vision local check,
budget-update lock visibility, ChatInbound additive attachment contract,
file-browser doc sync). The FULL combined diff exceeds every triad model's
context window (~1.59M tokens > 1.05M) — reviewed in packs; remaining
cross-pack findings were verified as slicing artifacts. Adversarial critics
(GPT; Gemini/Opus rounds) ran on the working tree. Deliberate tradeoffs:
metadata-based eviction captions (no light-LLM call in the hot path);
worker ledger records use live-cmdline fingerprints with a synthetic-arg
fallback only where the OS offers no cmdline.
2026-06-10 16:08:53 +03:00
Ouroboros
8ab6c1ce25 feat(review): add review substrate and code inventory 2026-05-27 05:42:56 +03:00
Ouroboros
a1f4341884 v5.24.0-rc.1: shrink scope-review pack without weakening gates
Reduce the scope-review pack by retiring legacy tool surfaces, consolidating review helpers, folding advisory compatibility views into canonical ledgers, and compressing operational prompts/docs while preserving BIBLE P3 review integrity and rationale.
2026-05-16 20:13:28 +03:00
Ouroboros
0f37e6ac9f Fix advisory acknowledgment audit completeness. 2026-05-13 23:44:52 +03:00
Ouroboros
a5993dc30e v5.20.1-rc.2: free advisory and surface skill auto-grant 2026-05-13 23:32:33 +03:00
Ouroboros
56607bbdde v5.17.0-rc.3: retire legacy migrations and centralize JSON state IO 2026-05-12 14:56:16 +03:00
Ouroboros
a8cac05981 v5.15.0-rc.9: test-suite structural consolidation pass
Refactor pass on the test suite: 24 file deletes, 7 cross-file merges,
~-1.7k LOC, 3399 tests passing (up from 3396 after adversarial review
caught 3 wrongly-removed tests during merge B.11 and restored them).

Outright deletes (7 files, -353 LOC):
- test_constitution.py: self-contained spec/DSL, no production code
  exercised; BIBLE.md numbering spine guard remains in
  test_smoke.py::test_bible_exists_and_has_principles
- test_module_size_gate.py / test_no_port_collision_8767.py: literal
  constant pins; real gates live in test_smoke.test_no_oversized_modules
  and DEFAULT_HOST_SERVICE_PORT use sites
- test_git_imports.py: redundant with test_smoke.test_import parametrize
- test_lmstudio_cached_tokens.py: comment archaeology in llm.py
- test_fixtures_mock_clawhub.py / test_fixtures_mock_llm.py: fixture
  self-tests

Cross-file merges (13 groups, 17 source files → 7 target/new files):
- test_advisory_workflow.py absorbs test_advisory_workflow_ext.py
- test_plan_review.py absorbs test_plan_review_quorum.py
- test_shell_run_shell.py NEW: test_shell_recovery + test_shell_regex_hint
  + test_shell_no_match_semantics merged
- test_web_search.py NEW: test_search_tool + test_web_search_streaming
- test_loop_misc.py NEW: test_loop_incoming_messages + test_loop_skill_finalization
- test_runtime_mode_core.py NEW: test_runtime_mode + test_runtime_mode_gating
  (test_runtime_mode_elevation.py kept separate per its 50+ attack vectors)
- test_build_scripts.py absorbs test_packaging_assets + test_release_workflow
  CI parts
- test_skills_marketplace_ui.py NEW: test_skills_ui_static
  + test_marketplace_ui_static + test_skill_toggle_smart_ui
- test_page_chrome_static.py NEW: test_page_header_ui_static
  + test_settings_and_page_layout_static + test_evolution_ui_guards
- test_skill_dependencies.py absorbs test_skill_token
  + test_skill_requested_secret_keys + test_skill_dependency_specs
- test_context.py absorbs test_context_memory_overhaul.py
- test_launcher_sync.py absorbs test_launcher_host_service_cleanup.py
- test_chat_logs_ui.py / test_chat_js_contracts.py: dedup
  TestVisualViewportListener + remove vacuous L317 assert

Rename: test_phase7_pipeline.py -> test_git_review_pipeline.py.
Production-code surface (2 lines):
- ouroboros/tools/review.py:110: comment now references the renamed file
- docs/ARCHITECTURE.md:1222: reference now points at test_plan_review.py
  (test_plan_review_quorum.py was merged into it)

In-place trims across ~28 files:
- Inspect-only source-string pin tests dropped under delete_trust_behavioral
  (test_review_v4_33 TestBuildReviewContextRelaxed/CircuitBreakerHintThreshold,
  test_review_synthesis 3 import tests, test_block1_review_pipeline
  TestRepoWriteCommitScopeReview, test_bughunt_fixes test_chat_id_zero,
  test_review_observability dataclass-only + manual mirror)
- Parametrize wins: test_browser_isolation 13 tests -> 2 tables;
  test_provider_integration 8 tests -> 2 tables; test_advisory_observability
  SDK-break readonly+edit collapsed; test_budget_tracking
  TestProviderAttributionHelper 7 tests -> 1 table; test_block1_review_pipeline
  TestIsProbablyBinary 5 tests -> 1 table; test_marketplace_adapter OS-field
  pair -> 1 table; test_skill_loader fail-closed pair -> 1 table
- Cosmetic CSS literal pins removed from test_chat_logs_ui
- Declarative-schema enumeration in test_widgets_ui_static collapsed to
  4 sentinels (was 15+ markers)
- test_contracts.py: 5 paranoid manifest YAML edge cases removed
- test_commit_gate test_auto_tag_function_exists / test_credential_helper_exists
  / test_auto_push_function_exists / triad-prompt substring pins dropped
- Settings explainer copy pins removed from test_settings_ui_guards
- test_settings_updates_ui::test_update_panel_contract_exists shortened

Preserved per P3 immune system + user constraints:
- All AST smoke walkers (test_no_oversized_modules, test_no_bare_except_pass,
  test_no_extremely_oversized_functions, test_function_count_reasonable,
  test_no_env_dumping)
- test_runtime_mode_elevation.py 50+ self-elevation attack vectors
- test_safety_policy.py::test_tool_policy_covers_all_builtin_tools
- test_platform_guard.py AST scan
- tests/fixtures/chat_logs_ui_static_checks.json (167 rows untouched)
- test_skill_exec.py and test_extensions_api.py kept as separate layers

Adversarial multi-model review (gemini-critic + gpt-critic + opus-critic,
round 1) caught two blockers: (1) the agent had wrongly removed 3 tests
from test_context.py during merge B.11 due to bad copy-paste and
misdiagnosed the failures as "consolidator API drift"; round 2 restored
them verbatim from git show HEAD:tests/test_context_memory_overhaul.py.
(2) docs/ARCHITECTURE.md:1222 still referenced a deleted regression-guard
test; fixed in same commit per P6.

Plus 5 stale audit-comment cleanups across test_contracts, test_tool_capabilities,
test_skill_exec, test_bughunt_fixes, test_chat_logs_ui, test_browser_isolation,
test_git_review_pipeline pointing at the new file layout.
2026-05-11 18:36:08 +03:00
Ouroboros
317150f3c8 v5.15.0-rc.1: aggressive test consolidation pass — fixtures + parametrize
First catch-up RC of the v5.15.0 reduction line. The v5.14.0 release missed its
-8000 LOC target by an order of magnitude (delivered -539 LOC vs goal); this RC
starts the honest catch-up by attacking the test layer where audit confirmed
real consolidation room exists.

Shared helpers (tests/_shared.py + tests/conftest.py):
- New `clean_extension_runtime_state()` superset cleanup helper. Three
  duplicated copies in test_skill_exec / test_extensions_api / test_extension_loader
  collapse to one shared helper with three thin autouse fixtures.
- New `ensure_claude_agent_sdk_mock()` SDK-mock installer. Three duplicated
  copies in test_commit_gate / test_advisory_observability / test_claude_code_gateway
  share one implementation.
- `tests/conftest.py` drops four unused fixtures (`make_git_repo`, `tool_context`,
  `make_chat_mock`, `make_extension_skill`) — verified by grep that no test
  parameter ever requested them. -73 LOC.

Per-file parametrization:
- test_extension_loader.py: 13 UI-tab rejection tests + 3 route-rejection tests
  collapse into two parametrize tables.
- test_phase7_pipeline.py: 11-case preflight cluster + 4-case JSON parser cluster
  parametrized; new `git_ctx` and `review_ctx` fixtures replace 12 paired
  `_get_git_module() + _make_ctx(tmp_path)` setups.
- test_review_fidelity.py: 3-case truncation suite parametrized.
- test_advisory_observability.py: resolve-claude-code-model trio parametrized;
  budget-gate trio (`_advisory_budget_gate_returns_skipped`,
  `_handle_advisory_pre_review_returns_skipped_status`, `_budget_gate_skip_persists`)
  share one new `budget_gate_env` fixture + `_budget_ctx` helper.
- test_commit_gate.py: registered-tools triplet parametrized.
- test_pr_tools.py: 37 paired `_make_temp_git_repo + _make_ctx` setups
  collapse to a single `_setup` helper call.

Net: -317 LOC across 13 files, 3376 tests still green (full suite).

Note on plan-vs-actual: RC.1 budget was -3500 LOC. Audit-confirmed safely-
achievable ceiling for tests-only consolidation is ~1200-1350 LOC across the
12 large test files; the remaining ~700-1000 LOC will come from continued
parametrize work in test_advisory_workflow / test_skill_loader / test_runtime_mode_elevation
during stretch RC.5 if the diff-size watchdog flags a cumulative shortfall.
The bulk of the missing -8000 LOC will come from RC.2 (production dedup) and
RC.3 (legacy/compat removal) which have more deterministic surface.
2026-05-10 21:49:06 +03:00
Ouroboros
5acde5c398 v5.8.3-rc.1: remove dead compatibility paths
Trim verified dead code and retire the hidden Evolution Versions surface while moving web rendering toward explicit text-vs-attribute escaping.
2026-05-08 01:41:06 +03:00
Ouroboros
39b732df78 v5.5.0: Add skills hub and lifecycle installer UX 2026-05-01 20:55:17 +03:00
Ouroboros
1cc3046095 fix(runtime-mode): make pro protected commits use normal review and close follow-up findings
Follow up on the post-CI review findings for the rc.9 runtime-mode work.

Key changes:
- Remove the separate pro-specific core_patch_gate review path. Pro protected
  commits now flow through the normal repo_commit triad + scope review exactly
  once, while advanced still blocks protected paths before review.
- Broaden restore_to_head and revert_commit protections from the old
  SAFETY_CRITICAL_PATHS set to the shared runtime-mode protected-path policy.
  Direct revert_commit now blocks commits that would modify protected surfaces
  and instructs callers to use staged edits + repo_commit so the normal review
  gate covers the change.
- Normalize legacy scope-review defaults for official OpenAI-only installs so
  stale Anthropic/OpenRouter scope defaults migrate to openai::gpt-5.5.
- Preserve full CHECKLISTS.md in plan-review canonical context exactly once.
- Add runtime mode to the safety LLM prompt so pro-mode protected edits are
  evaluated with the correct policy context.
- Distinguish incomplete rescue metadata in git_ops logs and add tests for
  truncated untracked-file rescue.

Validation:
- python3 -m py_compile on changed runtime/review modules
- focused pytest: runtime_mode_gating, server_runtime, plan_review,
  git_ops_recovery, pricing, packaging_sync, safety_policy
- full pytest tests/ -q (green; one expected skip; only pre-existing warnings)
2026-04-25 14:24:53 +03:00
Ouroboros
c16e13c134 feat(runtime-mode): implement direct pro core-patch gate and bump VERSION to 4.50.0-rc.9
Complete the three-layer runtime-mode boundary:
- light keeps the existing repo-mutation / skill_exec / extension dispatch block.
- advanced can still evolve the application layer, but now blocks protected core,
  frozen-contract, and release/managed-repo invariant paths.
- pro can edit protected paths on disk, but protected commits must pass the extra
  core-patch review gate before they land.

Adds a shared runtime-mode policy module and wires it through registry, git commit
paths, and the Claude gateway. Adds the pro core-patch gate, extracted review
revalidation helpers, protected-path rename/backslash coverage, pro CORE_PATCH_NOTICE
coverage, and fail-closed handling when core-patch scope review does not respond.

Moves shipped review defaults to the GPT-5.5 family: commit triad starts with
openai/gpt-5.5, scope review defaults to openai/gpt-5.5, and deep self-review
uses openai/gpt-5.5-pro. Direct-provider normalization now also maps scope review
models for OpenAI-only and Anthropic-only installs.

Fixes audit debts found during the re-review: full-repo prompt secret redaction,
canonical-doc de-duplication for scope/plan review, no accidental .review-drive
telemetry, managed-repo rescue reset fail-close on incomplete rescue, CI quick-test
coverage for ouroboros-three-layer, and updated SYSTEM/SAFETY/ARCHITECTURE docs.

Validation:
- python3 -m py_compile on changed runtime/review modules
- pytest tests/ -q (full suite green; 1 skipped; warnings are pre-existing local env/deprecation warnings)
- gpt-5.5 scope + triad review cycles until no unrebutted critical findings remained within the $100 review budget

No GitHub prerelease is published in this commit by user request; no tag is created here.
2026-04-25 03:01:05 +03:00
Ouroboros
6700358a41 Initial commit from app bundle 2026-04-22 01:04:48 +03:00