mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
P1-5: Let the plan envelope declare its reviewer panel's strength
Owner answers batch 1/Q6 and batch 2/Q3=A. The plan-local constant PLAN_REVIEW_EFFORT that overrode the owner's OUROBOROS_EFFORT_REVIEW setting is removed: plan_review_slots now resolves effort exactly like the commit triad, scope and acceptance surfaces (the setting is every row's default rung). plan_task gains ONE optional scalar, reviewer_effort (enum EFFORT_SCALE), placed on the default rung of row_effort's ladder, so an explicit per-row effort and a compound Cursor/Agy route slug still win; it travels only as an argument of plan_review_slots (no contextvar), and the engine wraps the builder only when the envelope declares a value, so the zero-arg stubs in the existing tests stay valid. Effort is roster identity: a different declaration re-dispatches a paid panel within OUROBOROS_REVIEW_MAX_CYCLES, the same one replays free, and the wave records the declaration so a collection rebuilds the roster it was dispatched with. The declaration is disclosed where a blocking install sees it: the wave's Reviewer slots render and the progress line. The last-execution projection keeps requested.effort as the ROW's effort and carries a one-off declaration as declared_effort, so an agent's order never shows as saved configuration (ReviewSlot.declared_effort). Docs: ARCHITECTURE plan review and settings table, subagents.py module docstring and legacy-field reason. (cherry picked from commit 64f1743b867a60c58a1c93440a6ccbaf9458907c)
This commit is contained in:
parent
aecff3cee2
commit
1005b6d01c
13 changed files with 221 additions and 37 deletions
|
|
@ -1714,7 +1714,7 @@ Plan review, task acceptance, commit review, and deep self-review answer differe
|
|||
|
||||
`plan_task` reviews an INTENTION before the work starts — the same organ for code, research, deliverables, and actions (BIBLE P3). The obligation is constitutional and the finalization gate is structural (`owner_hurry.force_plan_decision`); the order in which a task asks, explores and plans is the mind's judgment, and no prompt choreographs it. The submitted envelope carries the goal, the plan prose, and a typed domain-neutral SPEC (`in_scope`, `non_goals`, `acceptance_claims`, `invariants`, `decisions` with rejected alternatives, `deferred`, `affected_resources`, `evidence`); `ouroboros/tools/plan_spec.py` validates and normalizes its full operative content without shortening strings or dropping excess items, then mints the ids that are the only valid `breaks` targets (`goal`, `claim_N`, `invariant_N`, `decision_N`, `deferred_N`) — host-minted positionally, because a caller-chosen id could shadow another target and corrupt what a blocking finding `breaks` — and hashes it. Governance documents always come from the system repository; declared targets and evidence resolve against `active_repo_dir_for(ctx)`, and a path escaping the active subject or an unreadable root is a named omission, never a silent gap.
|
||||
|
||||
ONE structural fact tiers the governance pack: `constitutional` is true iff a declared `affected_resources`/`evidence` PATH locator resolves under the Ouroboros system repository (an `evidence` path only if it exists). A constitutional plan carries BIBLE.md and ARCHITECTURE.md in full (an `api_chat` row inline; a retrieving `agent_session` row as mandatory full reads) — the constitutional packet is not tiered, so assembling it without either is a typed failure, never a disclosure — and every other plan carries the runtime heading-derived navigation maps (`context_layout.generate_doc_nav_map`, never a copy) plus resolvable pointers. There is no plan-kind taxonomy, no agent-declared `plan_class`, no planning scouts, and no plan Atlas.
|
||||
ONE structural fact tiers the governance pack: `constitutional` is true iff a declared `affected_resources`/`evidence` PATH locator resolves under the Ouroboros system repository (an `evidence` path only if it exists). A constitutional plan carries BIBLE.md and ARCHITECTURE.md in full (an `api_chat` row inline; a retrieving `agent_session` row as mandatory full reads) — the constitutional packet is not tiered, so assembling it without either is a typed failure, never a disclosure — and every other plan carries the runtime heading-derived navigation maps (`context_layout.generate_doc_nav_map`, never a copy) plus resolvable pointers. There is no plan-kind taxonomy, no agent-declared `plan_class`, no planning scouts, and no plan Atlas. The ONE caller-facing strength axis is the envelope's optional `reviewer_effort`: the panel's effort for THIS order, placed on the default rung of each reviewer row's ladder (`reviewer_slot_config.row_effort`: an explicit per-row effort, then a compound Cursor/Agy route slug, then the declaration, then the owner's `OUROBOROS_EFFORT_REVIEW` setting), passed as an ARGUMENT of `plan_review_slots` only, so the commit gate, scope, acceptance and skill review keep reading the untouched rows. Effort is roster identity: a different declaration re-dispatches a paid panel within `OUROBOROS_REVIEW_MAX_CYCLES`, the same one replays free, and a wave records the declaration (`reviewer_effort`) so a collection rebuilds the roster it was dispatched with. Scoping against BIBLE P1 ("Cognitive horizon is part of continuity"): this is the review panel's strength for one order, ordered by the mind under the owner's setting as the default (the owner chose this scoping: the setting is the default and Ouroboros may order stronger or weaker per envelope), never the core's own model, effort or horizon; the disclosed residual is that a cheap panel's closed GREEN is earned authority for that envelope and a later stronger order does not reopen it. The last-execution projection keeps a declared effort apart from the row's saved effort (`declared_effort`).
|
||||
|
||||
Declared evidence is resolved by `ouroboros/tools/plan_evidence.py` against exactly two allowed roots — the active workspace and the system repository — with the shared sensitive-name policy applied to locator and target; every refused, missing, truncated, oversized, binary, or URL locator becomes a typed omission row in the manifest (the host never fetches a URL), and a `need_evidence` locator a reviewer names is attached by the host on the next cycle through the same policy. What that policy cannot attach is a named `[reviewer-requested]` omission row the panel is dispatched with and judges with, never a reason to run no reviewer; the evidence continuation uses a fresh full-packet dispatch only when no exact artifact reference exists, disclosed per slot as a `capability_delta`. An unreadable referenced artifact instead fails closed with `plan_review_exact_artifact_unavailable` and never mints replacement authority. A locator may carry an exact range — `::lines=A-B`, `::bytes=A-B`, `::tail=N`, or `::symbol=Name` for `.py` sources — and a source above the per-item byte bound is attached head-first with the cut named. The manifest hash joins the spec hash and `constitutional` in the wave fingerprint, so changing what the reviewers can see changes the identity of the review.
|
||||
|
||||
|
|
@ -1982,7 +1982,7 @@ A registry of `config.SETTINGS_DEFAULTS` (exact defaults stay canonical in `conf
|
|||
| OUROBOROS_PROMPT_CACHE_TTL | 1h | Prompt-cache tier (default/5m/1h). The policy acts at the final send-time wire boundary so it can legalize provider ordering without prompt builders creating provider-specific TTL policy; `review_helpers.cached_prompt_blocks` and `usage_accounting._reservation_cost` also consult it; usage records the applied tier |
|
||||
| OUROBOROS_EFFORT_TASK | medium | Task reasoning effort (scale none/minimal/low/medium/high/xhigh/max/ultra; Settings exposes all but `minimal`); provider adaptation is exact-route, success-confirmed, disclosed in `request_wire` |
|
||||
| OUROBOROS_EFFORT_EVOLUTION | high | Evolution effort |
|
||||
| OUROBOROS_EFFORT_REVIEW | high | Review effort |
|
||||
| OUROBOROS_EFFORT_REVIEW | high | Review effort; reaches plan review as every row's default rung unless the envelope declares `reviewer_effort` |
|
||||
| OUROBOROS_EFFORT_SCOPE_REVIEW | high | Scope-review effort |
|
||||
| OUROBOROS_EFFORT_DEEP_SELF_REVIEW | high | Deep-self-review effort — the surface default; a saved `deep_review` row's own effort outranks it |
|
||||
| OUROBOROS_EFFORT_CONSCIOUSNESS | high | Consciousness effort |
|
||||
|
|
|
|||
|
|
@ -148,6 +148,11 @@ class ReviewSlot:
|
|||
subagent_id: str = ""
|
||||
# Host sampling hint, resolved at dispatch; an explicit temperature wins.
|
||||
default_temperature: float | None = None
|
||||
# The effort this row runs at because the CALLER declared it for one order
|
||||
# (plan review's ``reviewer_effort``): '' when the row's own effort, a
|
||||
# compound route slug or the surface setting applied. Disclosure for the
|
||||
# last-execution projection; identity already rides ``effort``.
|
||||
declared_effort: str = ""
|
||||
|
||||
@property
|
||||
def native_retrieval(self) -> bool:
|
||||
|
|
|
|||
|
|
@ -729,10 +729,12 @@ def _delivery_slot(
|
|||
|
||||
# ABI-4: the local-route fact is read off the typed target constructed at
|
||||
# the review seam, not re-derived per model string here.
|
||||
own_effort = _row_own_effort(row)
|
||||
return ReviewSlot(
|
||||
slot_id=row.slot_id,
|
||||
model=row.target_id,
|
||||
effort=row_effort(row, effort_surface, default=default_effort),
|
||||
declared_effort=default_effort if default_effort and not own_effort else "",
|
||||
role_hint=role_hint,
|
||||
use_local=(row.use_local if row.use_local is not None else resolved_review_model_target(row.target_id).provider_route == "local"),
|
||||
route=(ReviewRouteKind.AGENT_SESSION if row.is_session
|
||||
|
|
@ -854,6 +856,20 @@ def commit_triad_delivery() -> Dict[str, Any]:
|
|||
}
|
||||
|
||||
|
||||
def _row_own_effort(row: ConfiguredReviewerSlot) -> str:
|
||||
"""The effort the ROW itself carries: its explicit field, else a Cursor/Agy
|
||||
compound slug's encoded effort; '' when the row leaves it to its caller."""
|
||||
if row.effort:
|
||||
return row.effort
|
||||
if row.is_session:
|
||||
return compound_session_effort(RouteSpec(
|
||||
kind=SHARED_ROUTE_KIND_SESSION,
|
||||
target_id=row.session_target or row.target_id,
|
||||
credential_profile_id=row.profile_id,
|
||||
)) or ""
|
||||
return ""
|
||||
|
||||
|
||||
def row_effort(
|
||||
row: ConfiguredReviewerSlot,
|
||||
surface: str,
|
||||
|
|
@ -865,18 +881,11 @@ def row_effort(
|
|||
An explicit row field wins. When it is absent, a Cursor/Agy compound model
|
||||
slug already carries the requested effort and therefore wins over the
|
||||
surface default. Ordinary rows retain the existing surface default (or a
|
||||
caller's established local default, as Plan Review does).
|
||||
caller's declared default, as a plan review order may carry).
|
||||
"""
|
||||
if row.effort:
|
||||
return row.effort
|
||||
if row.is_session:
|
||||
encoded = compound_session_effort(RouteSpec(
|
||||
kind=SHARED_ROUTE_KIND_SESSION,
|
||||
target_id=row.session_target or row.target_id,
|
||||
credential_profile_id=row.profile_id,
|
||||
))
|
||||
if encoded:
|
||||
return encoded
|
||||
own = _row_own_effort(row)
|
||||
if own:
|
||||
return own
|
||||
if default:
|
||||
return default
|
||||
from ouroboros.config import resolve_effort
|
||||
|
|
@ -1103,7 +1112,11 @@ def record_reviewer_slot_executions(surface: str, actors: Any, slots_by_id: Dict
|
|||
"requested": {
|
||||
"route_kind": route_kind,
|
||||
"model": str(getattr(slot, "model", "") or ""),
|
||||
"effort": str(getattr(slot, "effort", "") or ""),
|
||||
# The ROW's effort. A caller-declared one-off (plan review's
|
||||
# reviewer_effort) is disclosed separately, never shown as the
|
||||
# row's saved configuration.
|
||||
"effort": "" if getattr(slot, "declared_effort", "") else str(getattr(slot, "effort", "") or ""),
|
||||
**({"declared_effort": str(slot.declared_effort)} if getattr(slot, "declared_effort", "") else {}),
|
||||
"session_target": str(getattr(slot, "session_target", "") or ""),
|
||||
"profile_id": str(getattr(slot, "session_profile", "") or ""),
|
||||
# Actor binding, when the row is a configured-subagent
|
||||
|
|
|
|||
|
|
@ -17,7 +17,11 @@ cheapest model to the strongest reasoning and no rule reconciled them). And a
|
|||
harness route carries its OWN effort, so a parent asking ``low`` against a route
|
||||
pinned to ``xhigh`` had no rule for who wins. It is removed rather than ranked
|
||||
against the lane (BIBLE P2: remove the class). The owner still controls effort
|
||||
exactly as before, through ``config.resolve_effort(task_type)``.
|
||||
exactly as before, through ``config.resolve_effort(task_type)``. The ONE
|
||||
caller-facing strength axis lives elsewhere: a plan review order may declare
|
||||
its reviewer panel's effort as the default rung of each row's ladder
|
||||
(``plan_task.reviewer_effort`` → ``plan_review_runtime.plan_review_slots``);
|
||||
that is a review panel, not a subagent.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -885,9 +889,10 @@ SUBAGENT_INTENT_FIELDS: tuple[str, ...] = (
|
|||
# a load never fails over one (BIBLE P1: no silent loss, and no crash either).
|
||||
LEGACY_SUBAGENT_FIELDS: Dict[str, str] = {
|
||||
"reasoning_effort": (
|
||||
"effort is no longer an owner-facing axis: it is derived from the owner's "
|
||||
"configured effort for this task type, because a public effort was a second "
|
||||
"knob for the question model_lane already answers"
|
||||
"effort is not a subagent axis: it is derived from the owner's configured "
|
||||
"effort for this task type, because a public effort was a second knob for "
|
||||
"the question model_lane already answers (a plan review order declares its "
|
||||
"panel's strength through plan_task.reviewer_effort instead)"
|
||||
),
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -282,6 +282,10 @@ def _render_wave(
|
|||
findings = list(wave.get("findings") or [])
|
||||
findings_total = int(wave.get("findings_total") or len(findings))
|
||||
finding_page = findings[:MAX_FINDINGS_PER_SLOT]
|
||||
if wave.get("reviewer_effort"):
|
||||
actor_lines.append(
|
||||
f"- declared reviewer effort: {wave['reviewer_effort']} (this envelope's order; an explicit "
|
||||
"per-row effort or a compound route slug outranks it)")
|
||||
lines += [
|
||||
"", "### Reviewer slots", "", *actor_lines,
|
||||
"", "### Findings (per slot; finding_id = slot:id)", "", "```json",
|
||||
|
|
|
|||
|
|
@ -57,6 +57,7 @@ from ouroboros.tools.plan_review_runtime import (
|
|||
plan_fanout_inputs as _plan_fanout_inputs,
|
||||
plan_in_flight_custody_error as _plan_in_flight_custody_error,
|
||||
plan_deadline_skip as _plan_deadline_skip,
|
||||
REVIEWER_EFFORT_SCHEMA as _REVIEWER_EFFORT_SCHEMA,
|
||||
publish_plan_review_projection as _publish_plan_review_projection,
|
||||
publish_rendered_wave as _publish_rendered_wave,
|
||||
plan_payload_roots as _plan_payload_roots,
|
||||
|
|
@ -118,6 +119,7 @@ class _PlanRequest:
|
|||
goal: str
|
||||
plan: str
|
||||
spec: Any
|
||||
reviewer_effort: str = "" # the envelope's declared panel strength ('' = the owner's setting)
|
||||
|
||||
_SPEC_SCHEMA = {
|
||||
"type": "object",
|
||||
|
|
@ -242,7 +244,7 @@ def get_tools():
|
|||
"when another paid cycle is available. Cycles are bounded by the owner's Max review cycles; an unchanged "
|
||||
"envelope replays the recorded result for free (a locator a reviewer asked for "
|
||||
"with need_evidence is attached by the host next time and makes the envelope "
|
||||
"new). Under blocking enforcement an "
|
||||
"new; a different reviewer_effort re-dispatches a paid panel). Under blocking enforcement an "
|
||||
"open review holds finalization; under advisory you may proceed with the "
|
||||
"review open and the host discloses it. Declare evidence reviewers need; "
|
||||
"declare affected_resources so a self-modification gets the constitutional pack."
|
||||
|
|
@ -253,9 +255,10 @@ def get_tools():
|
|||
"goal": {"type": "string", "description": "Why — the outcome the work serves."},
|
||||
"plan": {"type": "string", "description": "Accompanying prose: how you intend to do it (context for reviewers; the spec is what is judged)."},
|
||||
"spec": _SPEC_SCHEMA,
|
||||
"reviewer_effort": _REVIEWER_EFFORT_SCHEMA,
|
||||
"review_disposition": _DISPOSITION_SCHEMA,
|
||||
},
|
||||
# Two exclusive modes: goal+plan+spec (review) or review_disposition alone.
|
||||
# Two exclusive modes: goal+plan+spec (+ optional reviewer_effort) or review_disposition alone.
|
||||
"required": [],
|
||||
},
|
||||
},
|
||||
|
|
@ -294,7 +297,7 @@ def _typed_refusal(ctx: ToolContext, code: str, text: str) -> str:
|
|||
def _handle_plan_task(ctx: ToolContext, **params) -> str:
|
||||
raw_disposition = params.get("review_disposition")
|
||||
# The registry refuses unknown params; a vacuous envelope field carries no plan.
|
||||
envelope_fields = [k for k in ("goal", "plan", "spec") if not _vacuous(k, params.get(k))]
|
||||
envelope_fields = [k for k in ("goal", "plan", "spec", "reviewer_effort") if not _vacuous(k, params.get(k))]
|
||||
if raw_disposition is not None and not _vacuous_disposition(raw_disposition):
|
||||
if envelope_fields:
|
||||
return _typed_refusal(
|
||||
|
|
@ -317,6 +320,7 @@ def _handle_plan_task(ctx: ToolContext, **params) -> str:
|
|||
)
|
||||
request = _PlanRequest(
|
||||
goal=str(params.get("goal") or ""), plan=str(params.get("plan") or ""), spec=params.get("spec"),
|
||||
reviewer_effort=str(params.get("reviewer_effort") or "").strip().lower(),
|
||||
)
|
||||
try: # the ToolEntry envelope is the outer settlement bound (plan_review_collect.run_plan_coroutine)
|
||||
return _collect.run_plan_coroutine(_run_plan_review_async(ctx, request))
|
||||
|
|
@ -437,6 +441,8 @@ def _prepare_plan_inputs(ctx: ToolContext, request: "_PlanRequest", state_root:
|
|||
spec, errors = plan_spec.normalize_spec(raw_spec if isinstance(raw_spec, dict) else None)
|
||||
if not request.plan.strip():
|
||||
errors = ["plan: required non-empty prose", *errors]
|
||||
if request.reviewer_effort and request.reviewer_effort not in _REVIEWER_EFFORT_SCHEMA["enum"]:
|
||||
errors.append(f"reviewer_effort: not on the effort scale {list(_REVIEWER_EFFORT_SCHEMA['enum'])}")
|
||||
if errors:
|
||||
return {"error": "ERROR: PLAN_SPEC_INVALID: " + "; ".join(errors) + ". No reviewer was called.",
|
||||
"code": "TOOL_ARG_ERROR"}
|
||||
|
|
@ -531,6 +537,8 @@ async def _run_plan_review_async(ctx: ToolContext, request: _PlanRequest, *, col
|
|||
previous_override: Optional[dict] = None
|
||||
replay_snapshot: Any = _PLAN_NO_SNAPSHOT
|
||||
resume_in_flight = False
|
||||
# The declaration wraps the builder ONLY when non-empty: zero-arg stubs of the builder stay valid.
|
||||
slots_fn = (lambda: _plan_review_slots(default_effort=request.reviewer_effort)) if request.reviewer_effort else _plan_review_slots
|
||||
existing = plan_review_wave(state, fingerprint)
|
||||
if existing is not None and not isinstance(existing.get("spec"), dict):
|
||||
existing = None # C-09: a COMPACTED row (no frozen spec) is never authority
|
||||
|
|
@ -563,7 +571,7 @@ async def _run_plan_review_async(ctx: ToolContext, request: _PlanRequest, *, col
|
|||
return _publish_rendered_wave(ctx, existing, cap=cap, cycles_paid=cycles_paid,
|
||||
enforcement=enforcement, cached=True, reminder=reminder)
|
||||
elif not resume_in_flight: # stale ⇒ identical envelope re-dispatches fresh
|
||||
stale, replay_snapshot = _plan_wave_replay_decision(_plan_review_slots, existing)
|
||||
stale, replay_snapshot = _plan_wave_replay_decision(slots_fn, existing)
|
||||
if not stale:
|
||||
if enforcement == "advisory":
|
||||
# Still-OPEN wave: re-invoke the emitter so a durable append that FAILED
|
||||
|
|
@ -596,7 +604,7 @@ async def _run_plan_review_async(ctx: ToolContext, request: _PlanRequest, *, col
|
|||
return _plan_unavailable(
|
||||
ctx, f"ERROR: Invalid reviewer-slot configuration blocks plan review — {err}. "
|
||||
"Fix Review lanes on the Agents tab in Settings.", "reviewer_slot_config_invalid")
|
||||
slots = _plan_review_slots()
|
||||
slots = slots_fn()
|
||||
if not slots:
|
||||
return _plan_unavailable(
|
||||
ctx, "ERROR: No review models configured. Configure Review lanes "
|
||||
|
|
@ -702,7 +710,7 @@ async def _run_plan_review_async(ctx: ToolContext, request: _PlanRequest, *, col
|
|||
constitutional=constitutional, constitutional_note=constitutional_note,
|
||||
cycle_index=cycle_index, retry_key=retry_key, enforcement=enforcement, cap=cap,
|
||||
quorum=quorum, configured_slots=configured_slots,
|
||||
health_evidence=health_evidence,
|
||||
health_evidence=health_evidence, reviewer_effort=request.reviewer_effort,
|
||||
)
|
||||
aggregate = str(wave["aggregate"])
|
||||
exact_wave = _exact_wave(
|
||||
|
|
|
|||
|
|
@ -93,6 +93,7 @@ def prepared_from_wave(ctx: Any, exact: Dict[str, Any]) -> tuple[Any, Dict[str,
|
|||
spec = dict(exact.get("spec") or {})
|
||||
request = _PlanRequest(
|
||||
goal=str(spec.get("goal") or ""), plan=str(exact.get("plan_prose") or ""), spec=spec,
|
||||
reviewer_effort=str(exact.get("reviewer_effort") or ""), # the same roster the wave dispatched with
|
||||
)
|
||||
prepared = {
|
||||
"spec": spec, "system_root": system_root, "active_root": active_root,
|
||||
|
|
|
|||
|
|
@ -22,6 +22,7 @@ from typing import Any, Dict, List, Optional
|
|||
|
||||
from ouroboros.tools.tool_result import ToolResult, _publish_tool_result
|
||||
|
||||
from ouroboros.config import EFFORT_SCALE
|
||||
from ouroboros.deadline_utils import parse_deadline_ts, utc_now
|
||||
from ouroboros.llm import LLMClient
|
||||
from ouroboros.review_execution_projection import review_executions_from_actor_usage
|
||||
|
|
@ -39,7 +40,23 @@ from ouroboros.tools.plan_review_artifacts import ( # noqa: E402, F401 - compat
|
|||
persist_wave as persist_plan_review_wave_artifact,
|
||||
read_wave as read_plan_review_wave_artifact,
|
||||
)
|
||||
PLAN_REVIEW_EFFORT = "high"
|
||||
# The one caller-facing strength axis of a review panel: the plan envelope may
|
||||
# declare the panel's effort for THIS order as the default rung of each row's
|
||||
# ladder (explicit per-row effort and compound route slugs still outrank it; the
|
||||
# owner's review-effort setting applies when nothing is declared). Owner
|
||||
# decision 2026-09-11 (batch 1/Q6, batch 2/Q3=A): the setting is the default,
|
||||
# Ouroboros may order stronger or weaker; a different strength is a different
|
||||
# envelope and re-dispatches a paid panel within OUROBOROS_REVIEW_MAX_CYCLES.
|
||||
REVIEWER_EFFORT_SCHEMA = {
|
||||
"type": "string", "enum": list(EFFORT_SCALE),
|
||||
"description": (
|
||||
"Optional reviewer-panel strength for THIS order (the default rung of each "
|
||||
"reviewer row's effort ladder; an explicit per-row effort or a compound route "
|
||||
"slug still wins; omitted = the owner's review-effort setting). A different "
|
||||
"strength is a different envelope: it re-dispatches a paid panel within "
|
||||
"OUROBOROS_REVIEW_MAX_CYCLES, while the same strength replays free."
|
||||
),
|
||||
}
|
||||
# ``None`` means no plan-local cognition cutoff. The substrate settles against
|
||||
# the owner deadline or shared transport bound, keeping the historical 560s
|
||||
# number from being reused as an HTTP timeout.
|
||||
|
|
@ -262,17 +279,21 @@ def record_raw_plan_request_attempt(
|
|||
return fingerprint
|
||||
|
||||
|
||||
def plan_review_slots() -> list:
|
||||
def plan_review_slots(default_effort: str = "") -> list:
|
||||
"""The configured commit-triad rows as plan-review ``ReviewSlot`` objects:
|
||||
the shared ``triad_delivery_slots`` builder (one reader of the triad rows
|
||||
for plan, skill and acceptance review) with plan review's own slot
|
||||
properties — timeout, output budget, temperature, and ``PLAN_REVIEW_EFFORT``
|
||||
as the effort default. Both delivery kinds ride; slot ids are the rows' own.
|
||||
properties — timeout, output budget, temperature — and the envelope's
|
||||
declared ``reviewer_effort`` as the rows' default rung (``''`` = the owner's
|
||||
review-effort setting, exactly like the commit triad). The declaration is an
|
||||
ARGUMENT of this builder only, never a contextvar: the commit gate, scope,
|
||||
acceptance and skill review keep reading the untouched rows. Both delivery
|
||||
kinds ride; slot ids are the rows' own.
|
||||
"""
|
||||
from ouroboros.reviewer_slot_config import triad_delivery_slots
|
||||
|
||||
return triad_delivery_slots(
|
||||
role_hint="plan reviewer", default_effort=PLAN_REVIEW_EFFORT,
|
||||
role_hint="plan reviewer", default_effort=str(default_effort or ""),
|
||||
timeout_sec=PLAN_REVIEW_SLOT_TIMEOUT_SEC, max_tokens=PLAN_REVIEW_MAX_TOKENS,
|
||||
default_temperature=0.2,
|
||||
)
|
||||
|
|
@ -499,7 +520,7 @@ def synthesize_plan_review_wave(
|
|||
fingerprint: str, previous: Optional[dict], manifest: dict, manifest_hash: str,
|
||||
constitutional: bool, constitutional_note: str, cycle_index: int, retry_key: str,
|
||||
enforcement: str, cap: Any, quorum: int, configured_slots: list,
|
||||
health_evidence: Any,
|
||||
health_evidence: Any, reviewer_effort: str = "",
|
||||
) -> tuple[dict, set[str], dict]:
|
||||
"""Validate raw actor rows and build one durable plan-review wave."""
|
||||
from ouroboros.tools import plan_spec
|
||||
|
|
@ -559,6 +580,7 @@ def synthesize_plan_review_wave(
|
|||
"paid": any(_row_has_physical_dispatch(row) for row in slot_records),
|
||||
"health_epoch": plan_health_epoch(health_evidence),
|
||||
"reviewer_config_fingerprint": plan_reviewer_config_fingerprint(configured_slots),
|
||||
"reviewer_effort": str(reviewer_effort or ""), # the envelope's declared panel strength ('' = setting)
|
||||
**plan_quorum_unreachable_facts(slot_records, quorum=quorum), "reviewed_at": utc_now_iso(),
|
||||
}
|
||||
# A partially-settled paid cycle is not yet allowed to mutate the next
|
||||
|
|
@ -676,6 +698,8 @@ def plan_wave_progress_line(
|
|||
line += f"; slot reasons: {reasons}"
|
||||
if (wave or {}).get("custody_pending"):
|
||||
line += "; late result pending (reviewer slots still in flight, not yet collected)"
|
||||
if (wave or {}).get("reviewer_effort"):
|
||||
line += f"; declared reviewer effort {wave['reviewer_effort']}"
|
||||
return line
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -114,7 +114,6 @@ def _capture_panel(monkeypatch):
|
|||
def test_triad_delivery_slots_is_the_one_builder_shared_by_plan_and_commit_vectors(structured_env):
|
||||
from ouroboros.reviewer_slot_config import commit_triad_delivery
|
||||
from ouroboros.tools.plan_review_runtime import (
|
||||
PLAN_REVIEW_EFFORT,
|
||||
PLAN_REVIEW_MAX_TOKENS,
|
||||
plan_review_slots,
|
||||
)
|
||||
|
|
@ -129,8 +128,7 @@ def test_triad_delivery_slots_is_the_one_builder_shared_by_plan_and_commit_vecto
|
|||
assert all(s.role_hint == "task acceptance" for s in acceptance)
|
||||
# Effort: explicit row → row; compound/none → the caller's default (plan) or the
|
||||
# roster row's own effort (actor row).
|
||||
assert [s.effort for s in plan] == ["high", "xhigh", "medium"]
|
||||
assert plan[0].effort != PLAN_REVIEW_EFFORT or PLAN_REVIEW_EFFORT == "high"
|
||||
assert [s.effort for s in plan] == ["high", "xhigh", "medium"] # the review setting, as for every triad surface
|
||||
# The commit/skill vectors are a projection of the same slots.
|
||||
vectors = commit_triad_delivery()
|
||||
assert vectors["slot_ids"] == [s.slot_id for s in acceptance]
|
||||
|
|
|
|||
|
|
@ -508,7 +508,7 @@ class TestPlanReviewToolRegistration(unittest.TestCase):
|
|||
from ouroboros.tools.plan_review import get_tools
|
||||
tool = next(t for t in get_tools() if t.name == "plan_task")
|
||||
params = tool.schema["parameters"]["properties"]
|
||||
self.assertEqual(set(params), {"plan", "goal", "spec", "review_disposition"})
|
||||
self.assertEqual(set(params), {"plan", "goal", "spec", "reviewer_effort", "review_disposition"})
|
||||
spec = params["spec"]["properties"]
|
||||
self.assertEqual(set(spec), {
|
||||
"in_scope", "non_goals", "acceptance_claims", "invariants", "decisions",
|
||||
|
|
|
|||
|
|
@ -494,3 +494,55 @@ def test_loop_reminder_does_not_repromise_a_spent_panel(monkeypatch):
|
|||
assert "a changed spec starts the next paid cycle" not in text
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "unlimited")
|
||||
assert "re-dispatches a fresh panel" in plan_review_reminder({"outcome": "DEGRADED", "reviewer_slots_degraded": True, "cycles_paid": 2})
|
||||
|
||||
|
||||
def _effort_aware_builder(harness, monkeypatch):
|
||||
"""A builder stub that honours ``default_effort`` (the engine wraps the builder
|
||||
only when the envelope declares an effort; zero-arg stubs stay valid)."""
|
||||
from ouroboros.tools import plan_review as pr
|
||||
|
||||
def build(default_effort=""):
|
||||
return [dataclasses.replace(slot, effort=default_effort or slot.effort,
|
||||
declared_effort=default_effort)
|
||||
for slot in harness.state["slots"]]
|
||||
|
||||
monkeypatch.setattr(pr, "_plan_review_slots", build)
|
||||
|
||||
|
||||
def test_declared_reviewer_effort_re_dispatches_and_the_same_declaration_replays_free(harness, monkeypatch):
|
||||
"""Owner batch 2 Q3=A: the envelope declares the panel's strength; effort is
|
||||
roster identity, so a changed declaration is a new paid panel while repeating
|
||||
the same declaration replays the recorded wave for free."""
|
||||
_patch_health(monkeypatch, lambda slots: {})
|
||||
_effort_aware_builder(harness, monkeypatch)
|
||||
open_finding = json.dumps([_finding("n1", "blocking", breaks="claim_1")])
|
||||
sub = harness.install({"s1": open_finding, "s2": CLEAN, "s3": CLEAN})
|
||||
ctx = harness.make_ctx()
|
||||
first = _call(ctx, reviewer_effort="low")
|
||||
assert _control(first) == {"outcome": "REVIEW_REQUIRED", "closed": False}
|
||||
assert len(sub.calls) == 1 and [s.effort for s in sub.calls[0]["slots"]] == ["low", "low", "low"]
|
||||
wave = _state(harness)["waves"][-1]
|
||||
assert wave["reviewer_effort"] == "low" and "declared reviewer effort: low" in first
|
||||
assert _call(ctx, reviewer_effort="low").count("cached exact review") == 1 and len(sub.calls) == 1
|
||||
stronger = _call(ctx, reviewer_effort="max")
|
||||
assert "cached exact review" not in stronger and len(sub.calls) == 2
|
||||
assert [s.effort for s in sub.calls[1]["slots"]] == ["max", "max", "max"]
|
||||
assert _state(harness)["cycles_paid"] == 2
|
||||
# Off the scale: a typed argument refusal, no reviewer called, no attempt recorded.
|
||||
refused = _call(ctx, reviewer_effort="turbo")
|
||||
assert refused.startswith("ERROR: PLAN_SPEC_INVALID") and "reviewer_effort" in refused
|
||||
assert len(sub.calls) == 2
|
||||
|
||||
|
||||
def test_a_closed_verdict_from_a_cheap_panel_is_not_reopened_by_a_stronger_declaration(harness, monkeypatch):
|
||||
"""The disclosed residual named to the owner (batch 2, Q3): a closed verdict is
|
||||
earned authority for this envelope; ordering a stronger panel afterwards
|
||||
replays the closed wave for free instead of re-dispatching."""
|
||||
_patch_health(monkeypatch, lambda slots: {})
|
||||
_effort_aware_builder(harness, monkeypatch)
|
||||
sub = harness.install({"s1": CLEAN, "s2": CLEAN, "s3": CLEAN})
|
||||
ctx = harness.make_ctx(task_id="task-cheap")
|
||||
assert _control(_call(ctx, reviewer_effort="none")) == {"outcome": "GREEN", "closed": True}
|
||||
again = _call(ctx, reviewer_effort="max")
|
||||
assert _control(again) == {"outcome": "GREEN", "closed": True}
|
||||
assert "cached exact review" in again and len(sub.calls) == 1
|
||||
|
|
|
|||
|
|
@ -623,3 +623,26 @@ def test_two_step_wave_emits_one_advisory_open_event_and_keeps_the_paid_identity
|
|||
trace = {"tool_calls": [{"plan_review_outcome": "DEGRADED"}, {"plan_review_outcome": "DEGRADED"}],
|
||||
"acceptance_obligations": []}
|
||||
assert acceptance_paid_identity("cand", trace) == acceptance_paid_identity("cand", {"tool_calls": [], "acceptance_obligations": []})
|
||||
|
||||
|
||||
def test_collection_rebuilds_the_roster_the_wave_was_dispatched_with(harness, monkeypatch):
|
||||
"""A declared effort is roster identity; the collection reuses the wave's
|
||||
recorded declaration, so it finds the same roster instead of refusing."""
|
||||
import dataclasses
|
||||
|
||||
from ouroboros.tools import plan_review as pr
|
||||
|
||||
def build(default_effort=""):
|
||||
return [dataclasses.replace(slot, effort=default_effort or slot.effort, declared_effort=default_effort)
|
||||
for slot in harness.state["slots"]]
|
||||
|
||||
monkeypatch.setattr(pr, "_plan_review_slots", build)
|
||||
calls = []
|
||||
_install_barrier_substrate(monkeypatch, calls)
|
||||
ctx = harness.make_ctx()
|
||||
_call(ctx, reviewer_effort="xhigh")
|
||||
wave = _state(harness)["waves"][-1]
|
||||
assert wave["custody_pending"] is True and wave["reviewer_effort"] == "xhigh"
|
||||
collected = _collect(ctx, wave["request_fingerprint"])
|
||||
assert _control(collected) == {"outcome": "GREEN", "closed": True}
|
||||
assert calls[1]["retry_key"] == calls[0]["retry_key"]
|
||||
|
|
|
|||
|
|
@ -137,10 +137,61 @@ def test_compound_session_effort_precedes_surface_defaults(monkeypatch):
|
|||
assert [row.effort for row in config.triad] == ["", ""]
|
||||
assert commit_triad_delivery()["efforts"] == ["xhigh", "low"]
|
||||
assert [slot.effort for slot in structured_scope_review_slots()] == ["max"]
|
||||
assert [slot.effort for slot in plan_review_runtime.plan_review_slots()] == [
|
||||
"xhigh",
|
||||
plan_review_runtime.PLAN_REVIEW_EFFORT,
|
||||
# The owner's review-effort setting reaches the plan panel exactly like the
|
||||
# commit triad (no plan-local constant overrides it any more).
|
||||
assert [slot.effort for slot in plan_review_runtime.plan_review_slots()] == ["xhigh", "low"]
|
||||
assert [slot.declared_effort for slot in plan_review_runtime.plan_review_slots()] == ["", ""]
|
||||
|
||||
|
||||
def test_declared_plan_effort_is_the_default_rung_and_touches_no_other_surface(monkeypatch):
|
||||
"""The envelope's reviewer_effort fills only the rows that leave effort to the
|
||||
caller: a compound slug and an explicit per-row effort still win. It travels as
|
||||
an argument of the plan builder alone, so the commit gate, scope, acceptance and
|
||||
skill-review identities are byte-identical before and after a declaration."""
|
||||
from ouroboros.skill_review_cycles import skill_review_contract_fingerprint
|
||||
from ouroboros.tools import plan_review_runtime
|
||||
from ouroboros.tools.commit_gate import commit_review_contract_fingerprint
|
||||
|
||||
payload = _payload()
|
||||
payload["triad"] = [
|
||||
{"slot_id": "cursor-row", "route": {"kind": "agent_session", "target_id": "cursor=cursor-grok-4.6-xhigh-fast"}},
|
||||
{"slot_id": "plain-row", "route": {"kind": "agent_session", "target_id": "codex=gpt-5.6-sol"}},
|
||||
{"slot_id": "pinned-row", "route": {"kind": "api_chat", "target_id": "openai/gpt-5.6-sol"}, "effort": "low"},
|
||||
]
|
||||
monkeypatch.setenv(REVIEWER_SLOTS_ENV, json.dumps(payload))
|
||||
monkeypatch.setenv("OUROBOROS_EFFORT_REVIEW", "medium")
|
||||
before = (commit_triad_delivery(), [s.effort for s in structured_scope_review_slots()],
|
||||
commit_review_contract_fingerprint(),
|
||||
skill_review_contract_fingerprint(["m"], delivery=commit_triad_delivery()))
|
||||
declared = plan_review_runtime.plan_review_slots(default_effort="max")
|
||||
assert [s.effort for s in declared] == ["xhigh", "max", "low"]
|
||||
assert [s.declared_effort for s in declared] == ["", "max", ""]
|
||||
assert [s.effort for s in plan_review_runtime.plan_review_slots()] == ["xhigh", "medium", "low"]
|
||||
after = (commit_triad_delivery(), [s.effort for s in structured_scope_review_slots()],
|
||||
commit_review_contract_fingerprint(),
|
||||
skill_review_contract_fingerprint(["m"], delivery=commit_triad_delivery()))
|
||||
assert before == after and before[0]["efforts"] == ["xhigh", "medium", "low"]
|
||||
|
||||
|
||||
def test_last_execution_projection_keeps_a_declared_effort_apart_from_the_row(tmp_path, monkeypatch):
|
||||
"""«Выполняется как» must not show the agent's one-off panel strength as the
|
||||
row's saved configuration: requested.effort is the ROW's effort ('' when the
|
||||
declaration filled it) and the declaration rides its own field."""
|
||||
from types import SimpleNamespace
|
||||
|
||||
from ouroboros import reviewer_slot_config
|
||||
from ouroboros.review_substrate import ReviewSlot
|
||||
|
||||
monkeypatch.setattr(reviewer_slot_config, "_last_execution_path", lambda: tmp_path / "last.json")
|
||||
slots = {
|
||||
"declared": ReviewSlot(slot_id="declared", model="m/a", effort="max", declared_effort="max"),
|
||||
"own": ReviewSlot(slot_id="own", model="m/b", effort="low"),
|
||||
}
|
||||
actors = [SimpleNamespace(slot_id=sid, status="ok", usage={}) for sid in slots]
|
||||
reviewer_slot_config.record_reviewer_slot_executions("plan_review", actors, slots)
|
||||
last = reviewer_slot_config.reviewer_slot_last_executions()
|
||||
assert last["declared"]["requested"]["effort"] == "" and last["declared"]["requested"]["declared_effort"] == "max"
|
||||
assert last["own"]["requested"]["effort"] == "low" and "declared_effort" not in last["own"]["requested"]
|
||||
|
||||
|
||||
def test_compound_effort_stabilizes_replay_identity_against_global_drift(monkeypatch):
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue