mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 20:27:56 +00:00
31 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bbd0aa0be9 |
Improve autonomous review and preserve acceptance evidence
Implement the owner-approved P3 package: optional planning advice, substantive auto acceptance, separate host notices, durable source citations and promotion, HEAD-bound hermetic proof reuse, and focused documentation and CI coverage. This version-neutral contributor checkpoint is pending exact-SHA final review. Publish browser coverage depends on the separately owned P1 Main-routing change; no P1 implementation is included. |
||
|
|
7b564c5cd8 | fix(review): reconcile skipped delegated preflight without stranding new work | ||
|
|
24fbd48bc1 | fix(review): prepare and recover authoritative review candidates | ||
|
|
b9ceed6ed7 |
merge: absorb upstream f3fbfdbb into the v7next campaign tree
Rolling-upstream sync #2 (8d13373b..f3fbfdbb, 101 commits, 180 files) under the standing principle: upstream is the SEMANTIC truth, the campaign is the STRUCTURAL truth. Every upstream semantic delta is re-seated in the campaign leaf that owns its span; the campaign's split shape, typed organs and ratified contracts are preserved. Notable seats: registry post-exec organ (owner-state restore DELETED, replaced by the settings tripwire) -> registry_guard_process/registry_core; the #447 H1 notes-trail-the-payload contract -> tool_result composition; the В23=A owner-home read carve + egress masking -> tool_access_user_files/core_file_tools; the #468 shape-first reasoning pin -> llm_messages/llm_fallback/llm_attempt/ llm_openai_compatible; delivery-control provenance and the forced-control body resolver -> loop_delivery/loop_forced_finalization; D4 export policy -> shell_outputs; A5 literal-argv disclosure -> tools/shell. Upstream twins of campaign organs do not live: tools/read_inspection.py folds into registry_guard_process, tools/result_envelope.py into tools/tool_result, delivery_protocol.py stays folded in loop_delivery. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com> |
||
|
|
bcda779ce0 |
stage 3, governance docs: item-21 executable triggers + coaching posture (#447)
WS3-a — docs/CHECKLISTS.md item 21 (capability_regression) becomes an
EXECUTABLE contract (В7=A with В1's emphasis): a guard/filter/allowlist
change fires the item and the reviewer must see a STAGED POSITIVE test
proving the surviving capability path and a NAMED surviving flow — a guard
change without one is a capability-regression finding, not a safety
improvement. Owner acceptance of a narrowing = a GREEN plan review that
explicitly names it (В4=C/В22=B); disclosure alone is not acceptance. The
Critical threshold gains question 4 ("name the surviving positive path",
auto-YES for non-restriction findings) and extends to proposed REMEDIES —
a capability-deleting remedy meets its own concreteness bar or stays
advisory; repetition confers no additional authority. The four standing
disclosures move byte-verbatim to docs/CHECKLISTS_ARCHIVE.md (new,
reviewer-binding, append-only) with a pointer; ARCHITECTURE.md maps the new
file (docs-sync).
WS3-b — the REVIEW_BLOCKED retry coaching (review.py) opens the
proportionality channel (В8=A): rebuttal is legitimate for a factual error,
unsupported severity, or a disproportionate remedy that would remove a
working capability the accepted plan did not narrow — "change what you can
argue for, not what you can override"; a rebuttal never overrides
owner-chosen enforcement. In the Skill Review Checklist (В21): the
".env.example rename" advice is deleted and "permanent text-only" becomes
current owner posture; delegate_integration's refusal message now reads the
same truth (one author for one fact).
Test disposition: test_review_blocked_message_prefers_fix_over_rebuttal
pinned the exact restriction-posture text В8=A replaces — rewritten to pin
the new contract with strictly more asserts (proportionality channel +
non-override clause).
The archive extraction moved the standing 2026-08-15 delegate_start recipe
with its block; the recipe-schema doc-rot test now scans the archive (its
recipes are binding reviewer material and must stay schema-valid too).
|
||
|
|
71e1f13fab |
v7next F3.3: comma-list remnant sweep gate + phase-5 residual cleanup
ABI-10 tail (base
|
||
|
|
5a190512fa |
WIP: rename the preflight organ (Q1=A): preflight_review with a callable advisory_review alias
The advertised tool is now preflight_review; advisory_review stays a CALLABLE compatibility alias (ToolEntry.alias_for — dispatched like any entry, never advertised by schemas()/available_tools(), same parameter schema so old calls keep their args). Contracts that withhold either spelling silence both (D10-style mapping in _disabled_tools). Synced: safety.py POLICY_SKIP rows (both spellings), tool_capabilities core and untruncated sets, prompts/SYSTEM.md canonical tool list, docs/CHECKLISTS.md invocation lines, update_merge_policy prompt, and every agent-facing guidance string that teaches the tool name. Enforcement/severity vocabularies, durable state paths (state/advisory_review.json), recorded tool_name history, and the skip_advisory_review parameter are deliberately untouched. |
||
|
|
c510de72d1 |
WIP: Claude-runtime retirement (Q4) + shared reviewer-row UI vocabulary (5A)
Engine: delete gateways/claude_code.py and the /api/claude-code/* endpoints, launcher verify block, platform-layer SDK/CLI probes, CLAUDE_CODE_* dispatch branches, CLAUDE_CODE_MODEL setting, claude-agent-sdk dependency; advisory module loses the dead SDK budget/paid-suspect blocks (event renamed advisory_suspect_result), error formatter drops retired probe fields. Web: Claude-runtime status/repair panel and onboarding card removed; reviewer rows get the Direct model / Configured subagent source picker with read-only derived disclosure; advisory editor on the shared api_chat vocabulary (legacy kind 'api' is parse-only); GET /api/reviewer-slots now round-trips subagent_id references (resolved route as disclosure only). Bench devtools: CLAUDE_CODE_MODEL dropped from forward lists, templates, manifest active projection (read vocabulary keeps it), cybergym snapshot, TB metadata roles; advisory bench rationale is comparability, not impossibility. Tests: suites repinned to the successor contracts (native episode caps, in-process usage scope, credentials-based availability, endpoint round-trip); size-ratchet manifest regenerated with band rationales. |
||
|
|
f97791afe0 | WIP: commit-gate bypass test repinned to credentials-based availability | ||
|
|
483cb2fbd8 |
review subject: managed resolutions review the M0->S delta; pre-dispatch packet admission; honest advisory guidance
Lane L-review of the update-flow redesign (Delta-4, Q25=A, Q28=A, O1):
- ouroboros/tools/review_subject.py (new): ManagedReviewSubject — the
resolution delta between the pinned mechanical-merge baseline M0 (from
the durable update tx) and the candidate tree S, whose definition
follows the surface (M1 contract): the COMMIT GATE serializes the real
index write-tree - the exact tree the review-binding fingerprint pins
and the commit writes - while the ADVISORY pre-review serializes the
private-index worktree snapshot (work-in-progress by contract). The
disclosure header carries M0/S identities, both real merge parents,
conflict anchors and dual counters (full candidate paths vs reviewed
resolution paths), and the M0-missing fallback renders the FULL
candidate from the same pinned S (HEAD..S) behind a loud disclosure
line. capture_review_diff(ctx, repo, unified) stays
byte-identical to capture_staged_diff for every non-managed caller.
- Every managed review consumer reads the artifact: triad diff, changed
list, -U0 rung and session task; scope diff, touched list (union with
conflict anchors), -U0 and session task; advisory diff, changed-file
context and the displayed counts. Review binding fingerprint, advisory
freshness, preflight staged lists, doc-only classification and the
scope snapshot key stay on the FULL candidate (I2).
- review_binary_context: managed binary rows render the mechanical-merge
M0 baseline blob beside both real merge parents.
- review_admission.py (new) + parallel_review: BOTH gate packets (triad
api pack, every scope row's pack) are assembled and fit-checked BEFORE
any reviewer dispatch; a deterministic assembly block anywhere spends
$0 everywhere (typed not_dispatched placeholders). the fit ladder
moved whole (review.py stays under the module cap; review's
_fit_triad_prompt remains the importable, patchable seam); _run_unified_review and
run_scope_review are split into prepare/dispatch halves with unchanged
verdict logic.
- Q28-A oversized outcomes: a panel whose agent-session rows alone
satisfy the quorum drops (triad) or yields (scope) its fit-blocked api
rows — recorded loudly — and proceeds; a panel that cannot reach quorum
without them gets a typed zero-spend terminal, and the managed
resolver's refusal carries the Settings -> Agents -> Review lanes
guidance. The managed advisory oversize (>500k chars) becomes an
audited non-blocking skip instead of the structurally impossible
"split the commit" error; the exported constant and every non-managed
path keep today's behavior byte-for-byte.
- Advisory honesty (O1): _next_step_guidance and the fresh-run message
take the enforcement mode — blocking keeps the historical wording;
advisory truthfully states that findings are recorded durably, the
agent decides which to apply, and commit_reviewed is available.
- prompts/SYSTEM.md commit-review section (owner-sanctioned): the inlined
managed artifact is authoritative; reviewers must not substitute their
own `git diff --cached`.
- Superseded pins are updated in-test with rationale: the parallel-review
wiring/ordering pins (Q25-A two-phase contract), the scope
represent_binary pin (now driven by the managed subject), and the triad
session-task builder (moved to review_subject behind a compat shim).
- tests/test_managed_review_subject.py (new): delta-only capture, -U0
rung, non-managed byte-compat, M0-missing disclosure, binary M0 rows,
dual counters, session-task inlining, $0 admission on deterministic
no-fit, Q28-A drop/yield/terminal outcomes, oversize skip wording, and
the enforcement-branch guidance texts.
Adversarial fix round (verified findings M1-M5, m1-m11; amended in place):
- M1: the reviewed gate subject S is bound to the tree that commits. The
commit-gate consumers (triad capture, scope prepare, the -U0 fit rung)
now serialize the REAL index (git write-tree - the same tree
_fingerprint_staged_diff pins), never the live-worktree snapshot, so a
pre-staged payload can no longer commit while reviewers read a restored
worktree. The advisory pre-review keeps the worktree snapshot by
contract (surface="advisory": advisory reviews work-in-progress;
freshness handles staleness). Defense-in-depth: every gate subject
records its S tree on the attempt and the commit gate asserts - a typed
review_subject_binding_mismatch block - that the reviewed tree equals
the binding fingerprint's tree_sha. Side benefit: gate-side S is stable
and cheap across every consumer within one attempt, closing the
disclosed multi-serialization TOCTOU consequence for the gate. A
regression test reproduces the EVIL-index/GOOD-worktree divergence.
- M2: the M0-missing fallback populates name_status from the FULL
candidate (HEAD..S), so touched packs, scope touched context and the
advisory context/syntax preflight cover exactly what the full diff and
the commit contain - never just the conflict anchors.
- M3: managed oversize terminals REPLACE the structurally impossible
"split the commit" imperative instead of appending guidance under it:
the triad fit ladder's block message, the scope ladder terminal remedy
(ladder_terminal_cause managed=True) and the admission-side guidance
all state that a managed resolution stages the whole two-parent tree
and point at agent-route rows / larger-window models.
- M4: the SESSION variant of the M0-missing packet is honest - the header
instructs "retrieve the FULL staged candidate diff yourself
(git diff --cached)" when no body is rendered, and the counters line
reports "reviewed resolution paths: n/a (M0 missing - full candidate
under review)".
- M5: the managed advisory oversize skip's trailing sentence branches on
enforcement (blocking: "still gate the commit"; advisory: findings are
recorded rather than blocking).
- m1: split advice is managed-aware in the post-skip next-step guidance
and in the 1.6M prompt-size gate (both pinned by tests).
- m2: enforcement wiring is covered end-to-end (_handle_review_status
next_step and the pre-review completion message under both modes).
- m3: three stub-heavy tests replaced with real-flow drives - the healthy
admission runs the REAL _prepare_unified_review and asserts the
dispatched packet carries the actual staged hunk; the Q28-A
drop/terminal tests drive the REAL fit ladder over the managed repo
(tiny calibrated limit through the documented seam, no injected error
strings); the advisory oversize test uses a genuinely >500k resolution
delta with real routing and a persisted skip record.
- m4: the dead managed seam in _build_advisory_prompt is deleted (managed
routing lives in _advisory_review_diff; no ADVISORY_ERROR string can
become a review subject).
- m5/m6: docs/ARCHITECTURE.md states the two-phase admission contract;
prompts/SYSTEM.md qualifies the inlined-artifact claim by M0
availability (owner-sanctioned edit).
- m7: ADVISORY_REVIEW_CHOICE_GUIDANCE reworded enforcement-neutrally
("blocking where enforcement makes them binding").
- m8: only LIVE agent-session rows (still to be dispatched) count toward
the Q28-A yield quorum; a session row that terminated at assembly is a
dead seat.
- m9: an all-not_dispatched scope panel records the distinct typed
scope_not_dispatched_assembly_block reason instead of a false
"diversity was not achieved" quorum advisory.
- m10: a yielded api row's pre-yield refusal advisories are superseded in
place; the never-produced budget_exceeded status left
SCOPE_FIT_BLOCK_STATUSES (measurement documented in the comment).
- m11: the deliberate admission-time-only scope of the Q28-A yield is
documented beside SCOPE_FIT_BLOCK_STATUSES (dispatch-time
provider-tokenizer oversize stays fail-closed; the owner-level
consistency question is escalated separately).
Panel fix round (5-seat review convergence; accepted R1-R10, amended in
place):
- R1: ARCHITECTURE/DEVELOPMENT truth - the structural tools tree gains
review_subject.py and review_admission.py rows; the git/commit-review
paragraph, the Review-stack rationale and DEVELOPMENT's "Triad slots
review the staged diff" all carry the managed exception (managed
resolution commits review the declared M0->S subject; the commit gate
binds S to the index write-tree the fingerprint pins). The full
handbook pass stays scheduled for synthesis S.4.
- R2: the M0-missing fallback body (every width, both surfaces) renders
git diff HEAD..S from the PINNED subject tree, never a second --cached
capture - on the advisory surface the old body omitted the unstaged
work the counters/name-status described; on the gate the second
capture weakened the binding to the pinned S. Regression tests combine
M0-missing with index/worktree divergence on both surfaces.
- R3: authorized-managed fail-open closed - once the authority predicate
says MANAGED, an unreadable or missing tx builds the LOUD M0-missing
fallback subject (reason tx_unreadable/tx_missing), never a silent
None that would review official code as resolver work; None remains
only for genuinely non-managed contexts.
- R4: managed binary deletion is visible - the deletion renderer accepts
the M0->S deletion (path present in M0, absent from stage-0, absent
from the HEAD->index deletion set) and renders the row against
M0/parent evidence; scope's deleted-path classifier takes the subject
trees (branching only on subject presence - the non-managed call stays
byte-identical) so an official-added, resolver-deleted binary can no
longer be classified as text. Test reproduces the exact topology.
- R5: seat identity survives $0 paths - a prepared-but-withheld triad
panel (deterministic scope assembly block) and Q28-dropped api rows
leave typed $0 not_dispatched actor records (model, slot, slot_id,
reason) in the durable raw results, matching the scope rows'
placeholders; not_dispatched seats are excluded from the failed-actor
diagnostics. Pinned by tests.
- R6: the blocking-mode fresh-critical advisory guidance drops the false
"fix all and re-run, or audited skip" dichotomy: a fresh advisory
already satisfies the freshness requirement, findings are recorded as
debt the commit gate acknowledges, commit_reviewed is available, and
the audited skip bypasses only the freshness/debt checks. Both
enforcement branches pinned (superseded pins updated with rationale).
- R7: _handle_advisory_pre_review no longer builds a second
advisory-surface subject for counters_line() - the counters thread out
of the ONE subject _advisory_review_diff builds (test pins exactly one
subject construction per pre-review).
- R8: a failed full-candidate path count renders "n/a (count
unavailable)" instead of a fake 0.
- R9: review_status renders a friendly message for
review_subject_binding_mismatch.
- R10: the M0-missing header lead is branched per mode (it no longer
claims the official delta "is not re-rendered" above a fallback body
that includes it); the managed packet header disclosed the accepted
Delta-2 residual (M0 is a pinned forensic baseline from the update tx
and is not re-verified at review time - disclosure only, no re-merge
check); this commit message states the M1 surface-following S contract
and the real _fit_triad_prompt seam name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
[synthesis] conflict resolution: ouroboros/tools/git.py — function-boundary
collision between lane-cycles' extracted _advisory_and_tests_gate and
lane-review's inserted _subject_binding_mismatch_outcome: kept BOTH functions
(subject binding check first, then the advisory/tests gate helper); the binding
call auto-merged into the cycles version of _run_reviewed_stage_cycle after
post-fingerprint revalidation and before block classification, preserving both
lanes' gate-order contracts (free-cycle -> advisory -> paid stamp -> dispatch
-> binding check). prompts/SYSTEM.md — merged paragraphs: kept lane-cycles'
paid-cycles/identical-refusal wording (supersedes attempt_cap_reached) and
appended lane-review's M0 resolution-delta packet contract.
tests/test_commit_gate.py — kept cycles' _advisory_and_tests_gate assertions
plus review's prepare/dispatch-phase assertions (Q25-A).
ouroboros/review_evidence.py — union of both lanes' reason_map entries.
|
||
|
|
386e9417b1 |
Max Review Cycles counts PAID cycles on the commit gate and skill review; identical bytes are never re-reviewed for pay
Owner decisions Q12/Q16-A/Q17-A/Q22-A/Q23-A (update-flow redesign sprint,
lane L-cycles, deltas 5-6). The one shared knob (OUROBOROS_REVIEW_MAX_CYCLES)
now means the same thing on every review gate: paid cycles per task.
Commit gate (commit_gate.py, git.py, review_state.py):
- Typed block-row classification at record time: verdict-blocks (triad/scope
reviewer FINDINGS) vs infra-blocks (fit fixed_overflow, scope sub_floor,
revalidation, transport/no-quorum). New durable CommitAttemptRecord fields
block_class, rebuttal_sha256, paid, review_contract_fingerprint,
root_task_id (merge- and roundtrip-safe; legacy rows classify by reason).
- A byte-identical resubmission of a verdict-blocked staged diff is refused
FREE from the FIRST block (identical_diff_refused), before the advisory
freshness gate and before any paid dispatch, quoting the recorded verdict;
the refusal is cross-task by fingerprint (anti-laundering). Infra-blocks,
infra/expired failures and preflight facts neither build nor reset the
streak; a post-review failure or a success ends it.
- Rebuttals are identified by CONTENT (sha256): a hash new to the streak buys
exactly one paid re-review; a repeated hash is refused free. A rebuttal is
"spent" only when it bought a dispatched, verdict-answered wave — one
refused undispatched (e.g. by the ceiling) stays fresh.
- The knob bounds PAID triad+scope cycles per ROOT task (the task tree shares
one ceiling; a manual session is its own task; a follow-up task is a fresh
root — origin_root_task_id is deliberately not honored, the cross-task
identical-fingerprint refusal stays the anti-laundering backstop). Paid is
recorded at dispatch on the attempt ledger (plan-review precedent; P7 - no
counter file) and the ceiling counts MONEY: every dispatched wave counts
whatever its terminal; only undispatched attempts stay outside the count.
Exhaustion is a free typed refusal plus
emit_review_cycles_exhausted(surface="commit_gate"); unlimited disables it.
- Free refusal/replay honors a review-contract fingerprint (triad roster and
routes, scope rows, enforcement, shipped prompt contract); a changed
contract lapses the streak and the refusal never quotes across the change.
- Under ADVISORY enforcement neither state hard-blocks: the commit proceeds
with a loud typed disclosure (commit_review_free_replay event + result
warning; the wording distinguishes a replayed verdict from a
ceiling-exhausted commit that has no verdict to reuse) and no further paid
dispatch. Managed resolvers keep MERGE_HEAD repaired on the free refusal
path.
- _run_reviewed_stage_cycle order: fingerprint (hoisted only because the free
gate needs it) -> free cycle gate -> advisory/tests gate -> binding
precondition (kept at its original position) -> paid dispatch; the
advisory+tests section moved whole to _advisory_and_tests_gate at the
function-size gate.
Skill review (skill_review.py, new skill_review_cycles.py,
skill_review_history.py, skill_review_runner.py):
- review_max_cycles() wired in for the first time: paid panel dispatches are
counted per ceiling key - the root task for task-driven review groups
(across every skill that task reviews; follow-ups start fresh), and, for
the manual lane, the CURRENT content snapshot (revised content restarts
the manual count; marketplace-install rows share the same per-snapshot
scoping). One chunked wave = ONE cycle; counts derive from the append-only
review history; exhaustion is a typed refusal plus
emit_review_cycles_exhausted(surface="skill_review").
- Q17-A free replay: an identical (group, content_hash, panel-contract
fingerprint) snapshot with a recorded SUBSTANTIVE verdict (clean/warnings/
blockers) replays at USD 0, quoting it - but only while the persisted
review state still covers that exact snapshot+verdict; a diverged state
falls through to a PAID rerun that re-persists (never a live-lock).
pending/interrupted/timeout/failed/cancelled are infra facts that never
replay; a rebuttal recorded on such a row stays fresh and buys the paid
rerun it is owed; a content-new rebuttal buys exactly one paid rerun.
- The panel-contract fingerprint pins the roster, the required items, the
shipped prompt/aggregation contract text and the skill's RESOLVED review
profile (computed once and threaded through aggregation).
- Terminal history rows carry the paid fact, the panel-contract fingerprint,
the rebuttal hash and replay provenance.
- Module splits at the size gates (update_candidate.py precedent): the
accepted-rebuttal ledger and the review-wave budget admission moved whole
to skill_review_cycles.py with same-name re-exports; review_skill's persist
tail extracted to _persist_reviewed_outcome and the cycles gate to
_skill_cycles_gate.
Disclosed residual: the ceiling checks are read-at-gate-time with no
reservation (TOCTOU) - concurrent dispatches sharing one ceiling key can
overshoot the cap by the concurrency width.
Pinned wording synced in this same commit (CHECKLISTS row 10):
review_cycles.py per-gate docstring (now four gates), commit_gate.py header
and module docstring, the advisory _identical_diff_cap_note, commit_reviewed
rebuttal schema texts, prompts/SYSTEM.md commit-gate wording,
docs/DEVELOPMENT.md cycles section plus the new anti-pattern instruction
(never pay for byte-identical review material; the limit counts PAID cycles;
exhaustion is a typed event), docs/ARCHITECTURE.md (review_cycles and
commit_gate module rows, skill review module map rows, evolution cleanup
paragraph, settings table), ouroboros/config.py settings comment, web
settings_ui.js copy, and review_evidence.py block-reason guidance - all
stating the per-ROOT-task scope. Pinned tests rewritten as contract tests
(test_commit_gate.py, test_review_cycles.py, plus the structural pins in
test_scope_review.py and test_advisory_delegated_route.py) and a new
tests/test_review_cycles_gates.py suite added, including behavior tests that
drive the real recorder to the ledger, the real MERGE_HEAD repair against a
temp git dir, and the real review_skill pipeline through replay, refusal and
paid dispatch.
Retired API: check_blocked_attempt_cap / blocked_attempt_fingerprint_cap
(pay-for-identical-re-review-up-to-a-cap semantics). Legacy
attempt_cap_reached ledger rows remain recognized as refusal records.
Fix round (5-seat review panel convergence; accepted findings F1-F8, F10,
fable P3-2/P3-3):
- F1 (CRITICAL) review_state.py: the 50-row attempt-history trim is now
authority-preserving - it evicts only non-authoritative noise (free
refusals, preflight facts, unpaid infra rows), oldest first, and never
paid rows (the money ledger the per-root-task ceiling derives from),
review-verdict anchors (the evidence an identical-diff refusal quotes) or
in-flight rows. The real bound (delta-review follow-up): the ACCOUNTING
FACTS of preserved rows are immortal, their heavy forensic payloads are
not - a preserved row that falls outside the newest-50 window is compacted
(raw_stripped=True): triad_raw_results/scope_raw_result dropped
(block_class materialized first, so a legacy scope verdict keeps its
refusal-anchor authority after the classifying evidence is gone),
block_details bounded to the 600 chars the refusal quote renders,
commit_message bounded to 300; paid/root/task ids, fingerprints, rebuttal
hash, timestamps and the quoted critical_findings all survive. The
serialized ledger therefore grows ~O(preserved rows x small record) - per
root task bounded by the money ceiling - instead of by full reviewer raw
output per reviewed commit forever. Regressions: 60+ free refusals after
the cap change neither the paid count nor the quoted verdict, and a
120-row heavy flood keeps full raw payloads only inside the newest-50
window (every over-window row serializes under 4KB).
- F1-adjacent recorder bug (exposed by the eviction test): a NEW attempt
recorded with no in-flight marker inherited the PREVIOUS terminal
attempt's fields - paid=True, block_class, rebuttal/contract/fingerprints
- so once any paid row existed, every free refusal was silently counted as
a paid cycle. The terminal-latest branch of _record_commit_attempt now
nulls the inherited row for a fresh attempt number, mirroring the
contract the reviewing branch already enforced.
- F2 (MAJOR) the paid stamp moved from gate passage to a WRITE-AHEAD record
at the first PHYSICAL reviewer dispatch: new module
ouroboros/review_dispatch.py (which also absorbs the slot-identity mint
moved whole from review_substrate.py at the module-size gate, same-name
re-exports kept) provides the idempotent thread-safe ReviewPaidStamp;
the commit gate installs it on ctx (_install_paid_dispatch_stamp) and the
shared transport entry run_review_request invokes it immediately before
the first reviewer call on EITHER side - triad and scope dispatch in
parallel, so any side dispatching makes the cycle paid. Assembly-only
exits (triad fit ladder, scope pack signals) stay $0 and outside the
ceiling; a crash after dispatch keeps the durable paid fact. The three
docstring enumerations ("assembly failures stay outside the count") are
now literally true. Design note: this seam is where the L-review lane's
two-phase admission slots in at synthesis.
- F3 (MAJOR) skill review shares ONE durable write-ahead dispatch marker
(state/skills/<name>/review_dispatch.json) between the lifecycle runner
and direct review_skill callers, written at physical panel dispatch with
paid=True + contract fingerprint + rebuttal hash + wave id. Terminal
history rows merge it idempotently (a lifecycle timeout with no result
still lands the paid facts) and clear it; an orphaned predecessor is
flushed into the history as an infra terminal; the derived count includes
an unmerged marker, so a crashed or swallowed wave never un-counts spent
money. The quorum-failure append now receives the paid facts (the direct
marketplace-install path no longer lands unpaid), and history append
failures surface as loud typed skill_review_history_append_failed events
instead of a silent pass. Four-outcome behavior tests (substantive
verdict / quorum-failed / transport-failed / crash-after-dispatch) each
pin exactly one paid unit on a real jsonl.
- F4 (wording + scoping, no semantics flip): spent-rebuttal memory on the
skill side is scoped to the CURRENT panel-contract fingerprint (a
contract lapse clears it along with replay); the contradictory wordings
(skill_review_cycles.py header, docs/DEVELOPMENT.md) now state the exact
rule - refusal-streak eligibility (substantive verdicts only) is distinct
from money accounting (every dispatched wave counts); a rebuttal is spent
by the substantive verdict it bought; a dispatched infra terminal
consumes money but not the rebuttal.
- F5 advisory-honesty sync: prompts/SYSTEM.md, web settings copy, the
advisory-review schema note (claude_advisory_review.py) each carry the
one enforcement caveat - under blocking the commit is refused for free;
under advisory it proceeds with a loud durable disclosure and no new
review spend.
- F6 the blocked-classification comment in git.py no longer claims a
dispatched infra block costs nothing - money is stamped at dispatch.
- F7 the skill_review tool schema documents the replay key (group,
content-hash, contract fingerprint), the paid-cycle ceiling, exhaustion
leaving the skill pending (never executable), and the rebuttal as content
whose sha256 buys exactly one paid wave when new to the streak.
- F8 auditability: review_evidence attempt projections expose block_class,
paid, rebuttal_sha256, review_contract_fingerprint and root_task_id; the
skill review-history UI detail includes the accounting facts inside the
existing free-form markdown detail string - the
SkillReviewHistoryDetailResponse contract is typed to exactly four
fields, so no api_types.js change and no version bump.
- F10 dead _POST_REVIEW_FAILURE_PHASES deleted (the streak walker's generic
terminal break already covers post-review failure phases).
- fable P3-2: the advisory identical-replay branch is marked as deliberate
defense-in-depth (verdict rows are only minted under blocking and the
contract fingerprint includes enforcement, so today it degrades to an
ordinary paid review).
- fable P3-3: end-to-end tests drive _run_reviewed_stage_cycle through both
advisory replay reasons, pinning the distinct progress notes and the
disclosure landing in the commit result formatting.
Residual disclosures:
- Bench settings_base templates set blocking enforcement and do not pin
OUROBOROS_REVIEW_MAX_CYCLES; the trace audit and the bench-cap option are
escalated to the owner before landing (per the sprint plan's I1
verification duty). Only the SWE-Pro evolution lane executes the commit
gate, and the E1v2 METHODOLOGY already holds advisory.
- Successful reviewed commits also consume paid cycles, so a root task's
third reviewed commit refuses under blocking at the default cap of 2
(settled Q16-A consequence).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
7de26338b7 |
test: stabilize cross-platform release gates
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com> |
||
|
|
d6210b1a05 |
plan review becomes a domain-neutral spec gate
The pre-implementation gate reviewed an unbounded prose plan with a ~100k-token governance pack, planning-scout subagents and an Atlas, and every rewritten plan minted a fresh fingerprint that bought a whole new paid wave from cold reviewers. Convergence was structurally impossible: the reviewer was asked to author a competing plan every wave, REVISE_PLAN could not be dispositioned, and nothing capped the cycles. plan_task now reviews an INTENTION — the same organ whether the work is code, research, a deliverable or an action in the world: - a typed domain-neutral SPEC (goal, in_scope, non_goals, acceptance_claims, invariants, decisions with rejected alternatives, deferred, affected_resources, evidence) with host-minted ids that are the only valid `breaks` targets; - ONE structural fact tiers the governance pack: `constitutional` iff a declared target resolves under the system repo (never prose, never a plan-kind taxonomy); - agent-declared evidence, bounded, with EVERY absence named, the runtime data plane denied outright, and the exploration log redacted through the same SSOT task acceptance uses; - typed findings (blocking with a `breaks` id | note | need_evidence) with the HOST computing the aggregate through adaptive_quorum — no reviewer emits GREEN as authority, none writes a competing plan; - ONE owner setting OUROBOROS_REVIEW_MAX_CYCLES (default 2, unlimited available) bounds paid cycles for plan review, task acceptance (passes = cycles - 1) and the commit gate's identical-diff attempt cap; an identical envelope replays for free and a DEGRADED wave costs nothing; - under blocking, an open plan holds implementation and a spent cap escalates with a typed review_cycles_exhausted reason and an honest blocked terminal; under advisory the agent may proceed with the wave open and a loud disclosure. Deleted: planning scouts, the plan Atlas, plan_class, context_level, the governance mega-pack, the generative reviewer stance, the hidden 32-wave limit and the api_chat-only pin. Net effect on the tree is negative. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com> |
||
|
|
fe11e54fc1 | Disclose unavailable advisory pre-review routes | ||
|
|
bd9aaf3d40 |
fix: close evolution authority race windows
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com> |
||
|
|
cacafa7d6c |
fix: bind evolution to isolated campaign state
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com> |
||
|
|
45d77ff785
|
tests+ci: green & de-flake the pre-push suite (P3/P6/xdist-safety) + parallelize CI ~9x (gate stays serial) (#55)
* test(improvement_backlog): stub semantic-dedup LLM in groom tests (P6)
_seed_many() seeds items via append_backlog_items(), whose C9.2 semantic-redirect
pre-pass calls semantic_dedup.find_semantic_duplicate_id() once per fingerprint-MISS
with candidates — a real light-model NETWORK call. The seeding runs BEFORE
_patch_groom_llm installs its mock, and that mock only covers chat_observed, not the
detector's own client path. With no API key the call retry-storms for minutes before
failing open to None, making test_groom_backlog_rejects_invented_items ~129s alone
(~40% of the whole suite) and non-deterministic.
Add a module-level autouse fixture that stubs find_semantic_duplicate_id to its own
fail-open default (None = no duplicate — exactly what the doomed call eventually
returns for the distinct seeded items), so the module is network-free and
deterministic. No production change; the dedup contract stays covered by
test_semantic_dedup_v6370.
Effect: test_groom_backlog_rejects_invented_items 129.12s -> 0.02s; whole file
~165s -> 2.37s, all 14 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(core): validate regex up front in _code_search so invalid-regex contract holds for both backends (P3)
_code_search ran the ripgrep path first and returned its formatted result before
ever reaching the Python-fallback re.compile guard. ripgrep accepts some malformed
patterns permissively — an unterminated '[' yields "no matches" instead of erroring
— so an invalid regex like "[invalid" silently returned no-match on the rg path while
only the fallback (rg absent/failed) emitted "⚠️ SEARCH_ERROR: invalid regex". This
made test_code_search_invalid_regex fail whenever ripgrep is present (i.e. always, in
CI and the pre-push preflight), so the suite exited non-zero on every candidate diff
regardless of the change under test.
Compile the regex once up front (regex queries only; literal queries need no check)
and return SEARCH_ERROR on re.error before dispatching to either backend, so both the
rg path and the Python fallback share the same invalid-regex contract.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: guard the conftest repo-root pollution sweep to the xdist controller (xdist-safety)
Under pytest-xdist, pytest_sessionfinish fires on the controller AND every worker against
the SHARED repo root, so the mock-pollution sweep (shutil.rmtree of leaked <MagicMock>
paths + session.exitstatus=1) had workers racing the same rmtree and each independently
failing the run — a non-deterministic, failed-shaped result. Guard the repo-root sweep +
exitstatus mutation behind `if not hasattr(session.config, "workerinput")` so it runs only
on the controller (the single authority); the per-process _PYTEST_DATA_DIR cleanup stays
outside the guard and runs on every process. Serial runs are unaffected (the guard is
always True without -n).
Adds tests/test_conftest_xdist_guard.py (sweep runs on the controller, skipped on a worker).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: parallelize the test suite with pytest-xdist (~9x faster), keep the gate serial
The full-suite CI jobs (quick-test, full-test) now run a PARALLEL pass plus a short SERIAL
pass. Empirically ~270s serial -> ~30s parallel (~9x). The per-commit preflight GATE stays
serial by design (a flaky parallel fail-closed gate manufactures non-deterministic
TESTS_FAILED indistinguishable from a real immune rejection).
- requirements: pytest-xdist + pytest-timeout (a hang-guard).
- pyproject: register the `serial` marker (addopts unchanged).
- ci.yml (quick-test, full-test): two-step. Parallel
`-m "not serial and <default lane exclusions>" -n auto --dist loadscope
--max-worker-restart=0 --timeout=300 --timeout-method=thread`, then serial
`-m "serial and <exclusions>"`. A command-line -m REPLACES the pyproject addopts markexpr,
so the default lane exclusions are repeated and ANDed with the serial split — the union
exactly reproduces the old default suite (4594 parallel + 122 serial = 4716, disjoint).
- conftest: a tryfirst pytest_collection_modifyitems hook marks the real-process/port files
(workspace_executor[+cleanup], process_custody, kill_process_tree_orphans, zombie_prevention,
worker_crash_retry, process_resource_leaks, restart_reconnect, preflight_runner,
services_tool_v2) `serial`; plus an autouse fixture isolating workspace_executor._SERVICES/
_FOREGROUND between tests (a latent ordering bug -n redistribution exposes).
Test-isolation fixes surfaced by running the suite under -n:
- test_task_constraint_tools: 3 bare `sys.modules["...claude_code"] = mock` (no restore ->
polluted the worker's sys.modules -> later SDK-dependent tests failed) -> monkeypatch.setitem.
- test_workspace_executor: poll until the spawned process's command-sha is readable before
registering (the PID-reuse safety check compared the sha recorded at registration vs
recomputed at kill; right after fork+exec the command line is unreadable -> shas diverge ->
kill silently skipped -> flaky), + 5->15s kill-confirmation deadlines.
docs/DEVELOPMENT.md + docs/CHECKLISTS.md: guidance so future tests are parallel-safe or
marked `serial`. Reviewed by adversarial subagents (ship); verified parallel 8/8 + serial 3/3
green, partition exact + disjoint, gate byte-identical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* tests+ci: parallel-safety follow-ups (services-global isolation, serial-lane guard, monkeypatch conversions)
Three non-blocking follow-ups from the parallelization review, hardening the
xdist split shipped earlier in this PR:
F1 (tests/conftest.py): extend the autouse _isolate_workspace_executor_globals
fixture to ALSO snapshot/clear/restore ouroboros.tools.services._SERVICES (the
legacy services registry, a separate module-global from workspace_executor's)
under its plain threading.Lock _LOCK. Closes the legacy-services-path
global-leak class generally; raw dict ops only under the lock (no re-entrant
deadlock on the plain Lock), registry-only (never reaps the live Popen
handles), each module lazy-imported under its own guard.
F2 (.github/workflows/ci.yml): add a "Guard non-empty serial marker lane" step
to marker-guards. `pytest --collect-only -m serial` exits 5 on an emptied
_SERIAL_TEST_FILES; set -euo pipefail + tee surfaces it, and a positive anchor
grep on test_workspace_executor.py is the working assertion (the existing
browser-guard's `! grep "no tests collected"` is a dead no-op under -q).
F3 (test_skill_loader / test_iteration2_fixes / test_marketplace_clawhub /
test_devtools_benchmarks): convert remaining bare os.environ / sys.modules
mutations (no save-restore) to auto-reverting monkeypatch.setenv/delenv/setitem,
so the parallel suite has no cross-test env/module leaks. tests/_shared.py's
intentional process-wide SDK mock left untouched.
Verified: parallel lane 4574 pass (-n auto --dist loadscope), serial lane 140
pass, the 4 converted files 159 pass. Two subagent reviews: no regressions.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: prune 26 redundant/tautological/obsolete tests (immune coverage retained)
Three-pass audit (finder+skeptic -> 31 empirical verifiers with rename+mutation
experiments -> delete-and-run dry-run) of the 275-file / 4067-test suite found 26
tests whose value is fully backstopped by a stronger survivor or is near-zero:
- tautological / zero-production pins (assert hasattr/callable; inline-simulated
branches; closed-loop regex; TypedDict dict-literal; CPython-only checks);
- proven duplicates with strict-superset twins (test_consolidator, test_commit_gate,
test_cache_optimization, test_review_intent_split, TestGrepRegexHint, ...);
- registry-registration one-liners, all backstopped by test_smoke::test_tool_set_matches
(set-equality vs EXPECTED_TOOLS -- mutation-proven to FAIL on a lost registration);
- dead-feature tests (retired skill_migrations module) + obsolete version regressions
(v636 import now unconditional; README "(N tests)" convention abandoned).
Removes ~47 test functions across 20 files (incl. the whole test_shell_regex_hint.py,
mirrored 1:1 in test_shell_run_shell.py::TestGrepRegexHint, and the whole
TestGoalScopePrecedence class). Also drops 2 now-unused imports + 1 orphaned helper.
KEPT (not redundant): test_smoke::test_git_commit_with_tests_exists -- sole guard of the
post-commit gate seam (a rename experiment proved nothing else catches its loss).
Verified: ruff F clean; collection exit 0 (4703 collected, no emptied class); full
parallel + serial suites green; two subagent reviews confirm exact set + no broken refs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Andrew <andgri200@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
f328b72821 |
feat(evolution): window-fit context, guaranteed-fit scope assembly, plan_task degraded fallback, identical-diff review-attempt cap, solve-capability ledger, lifecycle extraction with deterministic cycle cleanup (v6.30.0)
Phase 1 Block 5 of the combined release plan (A+B+C+D) plus the owner-directed guaranteed-fit scope assembly: - A window-fit context: evolution reads ARCHITECTURE as an on-demand nav map; emergency compaction threshold derives from the active model's real window (profile ceiling for 1M-window models); small-window remote models get routine compaction in max mode. - B plan_task capacity-class failures degrade to one labeled inline critique pass (stalled/infra stay fail-closed); window-aware scope input cap with Claude tokenizer calibration and window-scaled reserves (gigachat 131K maps to a workable input limit; >=1M reserves byte-identical); provider oversize 400 downgrades to the disclosed non-blocking budget_exceeded skip (owner-approved safety net); identical-diff commit-review attempt cap (3 genuine verdicts, diff-scoped, preflights stamped phase=preflight, review_rebuttal lifts, fail-open cost guard). - Guaranteed-fit scope assembly (owner directive: scope review must actually run): deterministic ladder full atlas -> compact atlas -> atlas required files to explicit budget_omitted manifest entries -> largest touched files to diff-only with disclosed notes; the >=1M blocking reviewer's irreducible overflow fails CLOSED (fixed_overflow); known sub-floor reviewers are advisory-only (P3 floor) with the documented skip terminal. - C solve-capability ledger: kind=cycle_outcome rows joined at absorb/abandon resolution; digest with explicit omission notes feeds the promotion chooser, upgraded to the main-model slot with a small-targeted-solve-helper bias. - D campaign/transaction lifecycle extracted from supervisor/queue.py into supervisor/evolution_lifecycle.py (protected git_ops untouched); deterministic no_op/abandoned cycle cleanup (stash + evolution-leftover branch + recorded skip gates); boot dangling-tx reconcile; hard-kill owner-mailbox cleanup. Note: ouroboros/tools/git.py remains grandfathered-oversized (review.py:41); the new cap logic lives in commit_gate.py. Gauntlet: adversarial x2 (GPT+Gemini pass 1; Fable pass 2 SAFE TO COMMIT) + triad+scope rounds 1-6 with the guaranteed-fit assembly reviewing itself from round 3 (scope responded every round since); all verified criticals fixed, oscillating/owner-decided findings rejected with recorded evidence. |
||
|
|
ae53779ccd |
release: Ouroboros v6.26.0-rc.1 — systemic hardening: immune system, custody, memory integrity, native multimodal chat
Owner-approved big release (plan: большой_релиз_v6.26.0). Workstreams:
- WS-A lint gate: ruff F-rules step in CI quick-test; fixed the real F823
NameError class (supervisor/events.py utc_now_iso shadowing); make lint/health.
- WS-B memory integrity: atomic dialogue_blocks with corrupt-quarantine
(.corrupt-<ts>.bak + memory_store_corrupt event); honest scratchpad journal
(block_append_failed, corrupt storage renders as corruption not "(empty)");
full Pattern Register window (16K cap vs 3K cut); merge-aware scratchpad
consolidation under the sidecar lock; backlog writes through locked helpers;
chat omission notes ("[N older unconsolidated messages omitted]").
- WS-C provider SSOT: provider registry (prefixes/credentials/resolution) in
provider_models.py replacing 5 duplicated knowledge sites; credential-aware
consolidation model; pricing empty-fetch retry; direct-route correctness
(o-series max_completion_tokens, reasoning_effort, Anthropic error bodies,
per-request timeouts on cached clients); deep/plan review budgets reserve
output headroom inside the 1M window (min(SSOT, window − output − margin)).
- WS-D races/state: utils.update_json_locked (locked RMW, loud TimeoutError);
task_results merge under per-file lock (cancel-latch safe); update_state
migration for owner binding, evolution counters, budget_line, post-task
activation; queue-lock coverage (enqueue, timeouts, worker health, task_done,
snapshot); visible supervisor death (supervisor_error + owner notice + ghost
consciousness/chat-agent cleanup on re-init); chat.jsonl rotation via
os.replace under the append lock; unbounded _outbox removed; host-service
port probe + actual-port panic sweep; skill lifecycle lane deadline
(OUROBOROS_SKILL_LIFECYCLE_TIMEOUT_SEC, dedupe-leak-safe); consilium
force-plan structured flag (plan_review_aggregate from the FULL result);
atomic writers (update-intent, repo-manifest, metrics cache, post-task pair);
review_state.save_state raises on lock timeout (advisory ledger honesty);
apply_pending_request no-ops while a campaign is active;
api_update_apply kills workers only AFTER update validation, with respawn
on aborted checkout and "interrupted" terminal status.
- WS-E immune hardening: triad anti-refusal coverage contract (empty array
needs NO_FINDINGS sentinel or bare-[] body; refusal prose with [] is
parse_failure and never enters quorum); BIBLE P3 "Owner-chosen enforcement,
loud advisory" bound + CHECKLISTS sync; every advisory pass-through of a
blocking signal writes review_advisory_override + persistent
advisory_overrides (surfaced by review_status); ToolEntry.mutates_worktree
with dispatcher-level before/after worktree diff invalidation (covers error
paths; read-only runs no longer invalidate; redundant manual calls removed);
scope review fails closed without its checklist; advisory final write is a
locked re-read merge; synthesis severity defaults to critical with WARNING
fallbacks; find removed from SAFE_SHELL_COMMANDS; anti-thrashing state
survives advisory criticals; preflight env scrubs secret-class variables;
.git/index.lock age gate in worker startup checks.
- WS-F security: file-browser symlink containment on the RESOLVED path across
all endpoints (out-of-root symlink targets are listed but inert; tests
rewritten to the new contract deliberately); HMAC-signed session cookies
(server-side persisted key, 30-day TTL, Secure on TLS) replacing the
permanent password-derived cookie; password-class settings mask to a
constant placeholder; conservative SSRF guard for the MAIN agent (link-local
/cloud-metadata only, LAN stays reachable, per-request route re-validation);
ClawHub zip-slip hardening (":"/backslash segments rejected + post-join
containment) and lazy no-proxy OuroborosHub opener; single-execution
signature dispatch for extension handlers (no TypeError re-run after side
effects); onboarding postMessage origin checks; SHA256-pinned
python-standalone download with pipefail; payload-resident dependency
fingerprints only corroborate durable deps.json; skill payload re-hash
immediately before spawn (TOCTOU narrowing).
- WS-G process custody: ouroboros/process_custody.py — spawn_supervised
chokepoint + durable data/state/process_ledger.jsonl (pid, pgid,
fingerprint{start_time, cmd_sha256}, purpose, scope task|session|daemon,
owner_task, session_id); platform_layer.process_start_time primitive;
startup + periodic reaper killing ONLY strict-fingerprint matches from dead
generations/tasks (never by command-line class); migrations: services
(the orphan hole), workspace executor + local model + extension companions
(ledger write-through), worker_pids (write-through; legacy path retained);
parent lifelines (ppid watchdog, group-suicide only as group leader) in
worker_main, extension runner, Claude readonly child; conformance test
pinning the Popen allowlist; ARCHITECTURE/DEVELOPMENT/CHECKLISTS entries.
- WS-H native multimodal chat: supports_vision capability map (static
prefixes + OpenRouter /models input_modalities overlay); web chat uploads
ride the WS frame as structured attachments (additive ChatInbound field)
and image uploads become NATIVE image blocks via the existing Path B;
browser screenshots inject natively for vision models via the multipart
user-merge (tool result stays a string; file persisted under
data/uploads/screenshots for re-view); K=3 newest-image eviction with
caption placeholders carrying the vlm_query re-view path; image-aware token
estimates (fixed ~1.1K-token equivalent instead of base64 length, fixing
permanent emergency-compaction wedges); compaction renders images as
captions (no base64 into the summarizer); GigaChat/local lanes emit explicit
"[image omitted: model has no vision]"; internal _caption/_source_path
metadata stripped from provider payloads.
- WS-I housekeeping (partial): SETTLED_STATUSES SSOT (+ cycle-safe mirror pin);
owner_inject.py renamed to owner_mailbox.py; version-neutral envelope
wording; files.py import-block cleanup. Remaining WS-I/WS-J/WS-K items are
deferred with the owner's context-budget priority on review+release.
Review notes: triad+scope ran via scripts/run_external_review.py on the core
pack across 3 rounds to convergence (scope responded=PASS each round; round-2
criticals fixed: update_apply kill-order + respawn, lifecycle dedupe leak on
lane timeout, OUROBOROS_MAX_ROUNDS hot-reload + docs, toggle_evolution
NameError, async Anthropic timeout forwarding, supports_vision local check,
budget-update lock visibility, ChatInbound additive attachment contract,
file-browser doc sync). The FULL combined diff exceeds every triad model's
context window (~1.59M tokens > 1.05M) — reviewed in packs; remaining
cross-pack findings were verified as slicing artifacts. Adversarial critics
(GPT; Gemini/Opus rounds) ran on the working tree. Deliberate tradeoffs:
metadata-based eviction captions (no light-LLM call in the hot path);
worker ledger records use live-cmdline fingerprints with a synthetic-arg
fallback only where the OS offers no cmdline.
|
||
|
|
8ab6c1ce25 | feat(review): add review substrate and code inventory | ||
|
|
a1f4341884 |
v5.24.0-rc.1: shrink scope-review pack without weakening gates
Reduce the scope-review pack by retiring legacy tool surfaces, consolidating review helpers, folding advisory compatibility views into canonical ledgers, and compressing operational prompts/docs while preserving BIBLE P3 review integrity and rationale. |
||
|
|
0f37e6ac9f | Fix advisory acknowledgment audit completeness. | ||
|
|
a5993dc30e | v5.20.1-rc.2: free advisory and surface skill auto-grant | ||
|
|
56607bbdde | v5.17.0-rc.3: retire legacy migrations and centralize JSON state IO | ||
|
|
a8cac05981 |
v5.15.0-rc.9: test-suite structural consolidation pass
Refactor pass on the test suite: 24 file deletes, 7 cross-file merges, ~-1.7k LOC, 3399 tests passing (up from 3396 after adversarial review caught 3 wrongly-removed tests during merge B.11 and restored them). Outright deletes (7 files, -353 LOC): - test_constitution.py: self-contained spec/DSL, no production code exercised; BIBLE.md numbering spine guard remains in test_smoke.py::test_bible_exists_and_has_principles - test_module_size_gate.py / test_no_port_collision_8767.py: literal constant pins; real gates live in test_smoke.test_no_oversized_modules and DEFAULT_HOST_SERVICE_PORT use sites - test_git_imports.py: redundant with test_smoke.test_import parametrize - test_lmstudio_cached_tokens.py: comment archaeology in llm.py - test_fixtures_mock_clawhub.py / test_fixtures_mock_llm.py: fixture self-tests Cross-file merges (13 groups, 17 source files → 7 target/new files): - test_advisory_workflow.py absorbs test_advisory_workflow_ext.py - test_plan_review.py absorbs test_plan_review_quorum.py - test_shell_run_shell.py NEW: test_shell_recovery + test_shell_regex_hint + test_shell_no_match_semantics merged - test_web_search.py NEW: test_search_tool + test_web_search_streaming - test_loop_misc.py NEW: test_loop_incoming_messages + test_loop_skill_finalization - test_runtime_mode_core.py NEW: test_runtime_mode + test_runtime_mode_gating (test_runtime_mode_elevation.py kept separate per its 50+ attack vectors) - test_build_scripts.py absorbs test_packaging_assets + test_release_workflow CI parts - test_skills_marketplace_ui.py NEW: test_skills_ui_static + test_marketplace_ui_static + test_skill_toggle_smart_ui - test_page_chrome_static.py NEW: test_page_header_ui_static + test_settings_and_page_layout_static + test_evolution_ui_guards - test_skill_dependencies.py absorbs test_skill_token + test_skill_requested_secret_keys + test_skill_dependency_specs - test_context.py absorbs test_context_memory_overhaul.py - test_launcher_sync.py absorbs test_launcher_host_service_cleanup.py - test_chat_logs_ui.py / test_chat_js_contracts.py: dedup TestVisualViewportListener + remove vacuous L317 assert Rename: test_phase7_pipeline.py -> test_git_review_pipeline.py. Production-code surface (2 lines): - ouroboros/tools/review.py:110: comment now references the renamed file - docs/ARCHITECTURE.md:1222: reference now points at test_plan_review.py (test_plan_review_quorum.py was merged into it) In-place trims across ~28 files: - Inspect-only source-string pin tests dropped under delete_trust_behavioral (test_review_v4_33 TestBuildReviewContextRelaxed/CircuitBreakerHintThreshold, test_review_synthesis 3 import tests, test_block1_review_pipeline TestRepoWriteCommitScopeReview, test_bughunt_fixes test_chat_id_zero, test_review_observability dataclass-only + manual mirror) - Parametrize wins: test_browser_isolation 13 tests -> 2 tables; test_provider_integration 8 tests -> 2 tables; test_advisory_observability SDK-break readonly+edit collapsed; test_budget_tracking TestProviderAttributionHelper 7 tests -> 1 table; test_block1_review_pipeline TestIsProbablyBinary 5 tests -> 1 table; test_marketplace_adapter OS-field pair -> 1 table; test_skill_loader fail-closed pair -> 1 table - Cosmetic CSS literal pins removed from test_chat_logs_ui - Declarative-schema enumeration in test_widgets_ui_static collapsed to 4 sentinels (was 15+ markers) - test_contracts.py: 5 paranoid manifest YAML edge cases removed - test_commit_gate test_auto_tag_function_exists / test_credential_helper_exists / test_auto_push_function_exists / triad-prompt substring pins dropped - Settings explainer copy pins removed from test_settings_ui_guards - test_settings_updates_ui::test_update_panel_contract_exists shortened Preserved per P3 immune system + user constraints: - All AST smoke walkers (test_no_oversized_modules, test_no_bare_except_pass, test_no_extremely_oversized_functions, test_function_count_reasonable, test_no_env_dumping) - test_runtime_mode_elevation.py 50+ self-elevation attack vectors - test_safety_policy.py::test_tool_policy_covers_all_builtin_tools - test_platform_guard.py AST scan - tests/fixtures/chat_logs_ui_static_checks.json (167 rows untouched) - test_skill_exec.py and test_extensions_api.py kept as separate layers Adversarial multi-model review (gemini-critic + gpt-critic + opus-critic, round 1) caught two blockers: (1) the agent had wrongly removed 3 tests from test_context.py during merge B.11 due to bad copy-paste and misdiagnosed the failures as "consolidator API drift"; round 2 restored them verbatim from git show HEAD:tests/test_context_memory_overhaul.py. (2) docs/ARCHITECTURE.md:1222 still referenced a deleted regression-guard test; fixed in same commit per P6. Plus 5 stale audit-comment cleanups across test_contracts, test_tool_capabilities, test_skill_exec, test_bughunt_fixes, test_chat_logs_ui, test_browser_isolation, test_git_review_pipeline pointing at the new file layout. |
||
|
|
317150f3c8 |
v5.15.0-rc.1: aggressive test consolidation pass — fixtures + parametrize
First catch-up RC of the v5.15.0 reduction line. The v5.14.0 release missed its -8000 LOC target by an order of magnitude (delivered -539 LOC vs goal); this RC starts the honest catch-up by attacking the test layer where audit confirmed real consolidation room exists. Shared helpers (tests/_shared.py + tests/conftest.py): - New `clean_extension_runtime_state()` superset cleanup helper. Three duplicated copies in test_skill_exec / test_extensions_api / test_extension_loader collapse to one shared helper with three thin autouse fixtures. - New `ensure_claude_agent_sdk_mock()` SDK-mock installer. Three duplicated copies in test_commit_gate / test_advisory_observability / test_claude_code_gateway share one implementation. - `tests/conftest.py` drops four unused fixtures (`make_git_repo`, `tool_context`, `make_chat_mock`, `make_extension_skill`) — verified by grep that no test parameter ever requested them. -73 LOC. Per-file parametrization: - test_extension_loader.py: 13 UI-tab rejection tests + 3 route-rejection tests collapse into two parametrize tables. - test_phase7_pipeline.py: 11-case preflight cluster + 4-case JSON parser cluster parametrized; new `git_ctx` and `review_ctx` fixtures replace 12 paired `_get_git_module() + _make_ctx(tmp_path)` setups. - test_review_fidelity.py: 3-case truncation suite parametrized. - test_advisory_observability.py: resolve-claude-code-model trio parametrized; budget-gate trio (`_advisory_budget_gate_returns_skipped`, `_handle_advisory_pre_review_returns_skipped_status`, `_budget_gate_skip_persists`) share one new `budget_gate_env` fixture + `_budget_ctx` helper. - test_commit_gate.py: registered-tools triplet parametrized. - test_pr_tools.py: 37 paired `_make_temp_git_repo + _make_ctx` setups collapse to a single `_setup` helper call. Net: -317 LOC across 13 files, 3376 tests still green (full suite). Note on plan-vs-actual: RC.1 budget was -3500 LOC. Audit-confirmed safely- achievable ceiling for tests-only consolidation is ~1200-1350 LOC across the 12 large test files; the remaining ~700-1000 LOC will come from continued parametrize work in test_advisory_workflow / test_skill_loader / test_runtime_mode_elevation during stretch RC.5 if the diff-size watchdog flags a cumulative shortfall. The bulk of the missing -8000 LOC will come from RC.2 (production dedup) and RC.3 (legacy/compat removal) which have more deterministic surface. |
||
|
|
5acde5c398 |
v5.8.3-rc.1: remove dead compatibility paths
Trim verified dead code and retire the hidden Evolution Versions surface while moving web rendering toward explicit text-vs-attribute escaping. |
||
|
|
39b732df78 | v5.5.0: Add skills hub and lifecycle installer UX | ||
|
|
1cc3046095 |
fix(runtime-mode): make pro protected commits use normal review and close follow-up findings
Follow up on the post-CI review findings for the rc.9 runtime-mode work. Key changes: - Remove the separate pro-specific core_patch_gate review path. Pro protected commits now flow through the normal repo_commit triad + scope review exactly once, while advanced still blocks protected paths before review. - Broaden restore_to_head and revert_commit protections from the old SAFETY_CRITICAL_PATHS set to the shared runtime-mode protected-path policy. Direct revert_commit now blocks commits that would modify protected surfaces and instructs callers to use staged edits + repo_commit so the normal review gate covers the change. - Normalize legacy scope-review defaults for official OpenAI-only installs so stale Anthropic/OpenRouter scope defaults migrate to openai::gpt-5.5. - Preserve full CHECKLISTS.md in plan-review canonical context exactly once. - Add runtime mode to the safety LLM prompt so pro-mode protected edits are evaluated with the correct policy context. - Distinguish incomplete rescue metadata in git_ops logs and add tests for truncated untracked-file rescue. Validation: - python3 -m py_compile on changed runtime/review modules - focused pytest: runtime_mode_gating, server_runtime, plan_review, git_ops_recovery, pricing, packaging_sync, safety_policy - full pytest tests/ -q (green; one expected skip; only pre-existing warnings) |
||
|
|
c16e13c134 |
feat(runtime-mode): implement direct pro core-patch gate and bump VERSION to 4.50.0-rc.9
Complete the three-layer runtime-mode boundary: - light keeps the existing repo-mutation / skill_exec / extension dispatch block. - advanced can still evolve the application layer, but now blocks protected core, frozen-contract, and release/managed-repo invariant paths. - pro can edit protected paths on disk, but protected commits must pass the extra core-patch review gate before they land. Adds a shared runtime-mode policy module and wires it through registry, git commit paths, and the Claude gateway. Adds the pro core-patch gate, extracted review revalidation helpers, protected-path rename/backslash coverage, pro CORE_PATCH_NOTICE coverage, and fail-closed handling when core-patch scope review does not respond. Moves shipped review defaults to the GPT-5.5 family: commit triad starts with openai/gpt-5.5, scope review defaults to openai/gpt-5.5, and deep self-review uses openai/gpt-5.5-pro. Direct-provider normalization now also maps scope review models for OpenAI-only and Anthropic-only installs. Fixes audit debts found during the re-review: full-repo prompt secret redaction, canonical-doc de-duplication for scope/plan review, no accidental .review-drive telemetry, managed-repo rescue reset fail-close on incomplete rescue, CI quick-test coverage for ouroboros-three-layer, and updated SYSTEM/SAFETY/ARCHITECTURE docs. Validation: - python3 -m py_compile on changed runtime/review modules - pytest tests/ -q (full suite green; 1 skipped; warnings are pre-existing local env/deprecation warnings) - gpt-5.5 scope + triad review cycles until no unrebutted critical findings remained within the $100 review budget No GitHub prerelease is published in this commit by user request; no tag is created here. |
||
|
|
6700358a41 | Initial commit from app bundle |